{
 "S1sh1::signage": {
  "fp": "f0142979c86dcc17",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S1sh1": {
  "input_fingerprint": "b8503bc5a18f8e77",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바닷속 쓰레기 더미 위로 튀어나온 낡은 로봇 손가락에 입을 맞추듯 주둥이를 댄 물고기의 측면.\n\nLOCATION (lock): Underwater at the seabed, beside a submerged rubbish heap with a robot finger protruding from it. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Protruding robot finger (Old and protruding from submerged rubbish, touching the fish's snout) — Seen from the side, with its contact surface level with the fish's snout; used as Small focal anchor to the right of the contact point; Submerged rubbish pile (Accumulated on the seabed around the protruding finger); used as Layered lower-frame context that establishes where the finger emerges.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Blue underwater ambient light and restrained contrast preserve the quiet, tactile contact without exaggerating it into a luminous effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): In blue daytime seawater, an old scrap robot lies buried in seabed rubbish with a finger protruding; its body is entangled in netting, and its chest bears a worn Ubik logo. Unidentified fish swim among the submerged debris.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바닷속 쓰레기 더미 위로 튀어나온 낡은 로봇 손가락에 입을 맞추듯 주둥이를 댄 물고기의 측면.\n\nLOCATION (lock): Underwater at the seabed, beside a submerged rubbish heap with a robot finger protruding from it. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Protruding robot finger (Old and protruding from submerged rubbish, touching the fish's snout) — Seen from the side, with its contact surface level with the fish's snout; used as Small focal anchor to the right of the contact point; Submerged rubbish pile (Accumulated on the seabed around the protruding finger); used as Layered lower-frame context that establishes where the finger emerges.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Blue underwater ambient light and restrained contrast preserve the quiet, tactile contact without exaggerating it into a luminous effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): In blue daytime seawater, an old scrap robot lies buried in seabed rubbish with a finger protruding; its body is entangled in netting, and its chest bears a worn Ubik logo. Unidentified fish swim among the submerged debris.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바닷속 쓰레기 더미 위로 튀어나온 낡은 로봇 손가락에 입을 맞추듯 주둥이를 댄 물고기의 측면.\n\nLOCATION (lock): Underwater at the seabed, beside a submerged rubbish heap with a robot finger protruding from it. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Protruding robot finger (Old and protruding from submerged rubbish, touching the fish's snout) — Seen from the side, with its contact surface level with the fish's snout; used as Small focal anchor to the right of the contact point; Submerged rubbish pile (Accumulated on the seabed around the protruding finger); used as Layered lower-frame context that establishes where the finger emerges.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Blue underwater ambient light and restrained contrast preserve the quiet, tactile contact without exaggerating it into a luminous effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): In blue daytime seawater, an old scrap robot lies buried in seabed rubbish with a finger protruding; its body is entangled in netting, and its chest bears a worn Ubik logo. Unidentified fish swim among the submerged debris.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "수평으로 헤엄치는 물고기의 주둥이가 화면 우측에서 뻗어 나온 로봇 손가락 끝을 향해 정확히 맞닿아 있습니다.",
    "built_space": "푸른 빛이 도는 수중 해저면에 플라스틱, 그물 등 다양한 질감의 쓰레기 더미가 층을 이루어 쌓여 있습니다.",
    "entities": "측면 구도의 물고기, 부식된 질감의 낡은 로봇 손과 손가락, 바닷속 수중 쓰레기 더미가 모두 명확히 확인됩니다.",
    "hard_violations": [],
    "physics": "물고기는 수중에 자연스럽게 떠 있으며, 로봇 손은 쓰레기 더미 속에 물리적으로 안정되게 박혀 지지받고 있습니다."
   },
   {
    "label": "B",
    "direction": "물고기가 사선으로 뻗어 나온 로봇 손가락 끝을 향해 입을 대고 있습니다.",
    "built_space": "바닷속 해저면에 타이어와 플라스틱 등 여러 형태의 폐기물들이 흩어져 쌓여 있습니다.",
    "entities": "물고기의 측면, 금속 질감의 로봇 손과 위로 솟은 손가락, 해저 쓰레기 더미가 식별됩니다.",
    "hard_violations": [],
    "physics": "물고기는 물의 부력으로 떠 있고, 로봇 손은 바닥의 쓰레기 더미 위에 얹혀 지탱되고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "프롬프트가 요구한 대로 물고기의 주둥이와 로봇 손가락의 접촉면이 수평을 이루는 측면 구도를 매우 정확하고 자연스럽게 구현했습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "물고기와 로봇 손가락의 접촉은 잘 표현되었으나, 손가락이 위로 솟구친 사선 형태여서 접촉면이 수평을 이룬다는 지시사항의 정확도가 A에 비해 약간 떨어집니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "수평으로 헤엄치는 물고기의 주둥이가 화면 우측에서 뻗어 나온 로봇 손가락 끝을 향해 정확히 맞닿아 있습니다.",
        "built_space": "푸른 빛이 도는 수중 해저면에 플라스틱, 그물 등 다양한 질감의 쓰레기 더미가 층을 이루어 쌓여 있습니다.",
        "entities": "측면 구도의 물고기, 부식된 질감의 낡은 로봇 손과 손가락, 바닷속 수중 쓰레기 더미가 모두 명확히 확인됩니다.",
        "hard_violations": [],
        "physics": "물고기는 수중에 자연스럽게 떠 있으며, 로봇 손은 쓰레기 더미 속에 물리적으로 안정되게 박혀 지지받고 있습니다."
       },
       {
        "label": "B",
        "direction": "물고기가 사선으로 뻗어 나온 로봇 손가락 끝을 향해 입을 대고 있습니다.",
        "built_space": "바닷속 해저면에 타이어와 플라스틱 등 여러 형태의 폐기물들이 흩어져 쌓여 있습니다.",
        "entities": "물고기의 측면, 금속 질감의 로봇 손과 위로 솟은 손가락, 해저 쓰레기 더미가 식별됩니다.",
        "hard_violations": [],
        "physics": "물고기는 물의 부력으로 떠 있고, 로봇 손은 바닥의 쓰레기 더미 위에 얹혀 지탱되고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "프롬프트가 요구한 대로 물고기의 주둥이와 로봇 손가락의 접촉면이 수평을 이루는 측면 구도를 매우 정확하고 자연스럽게 구현했습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "물고기와 로봇 손가락의 접촉은 잘 표현되었으나, 손가락이 위로 솟구친 사선 형태여서 접촉면이 수평을 이룬다는 지시사항의 정확도가 A에 비해 약간 떨어집니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "수평으로 헤엄치는 물고기의 주둥이가 화면 우측에서 뻗어 나온 로봇 손가락 끝을 향해 정확히 맞닿아 있습니다.",
        "built_space": "푸른 빛이 도는 수중 해저면에 플라스틱, 그물 등 다양한 질감의 쓰레기 더미가 층을 이루어 쌓여 있습니다.",
        "entities": "측면 구도의 물고기, 부식된 질감의 낡은 로봇 손과 손가락, 바닷속 수중 쓰레기 더미가 모두 명확히 확인됩니다.",
        "hard_violations": [],
        "physics": "물고기는 수중에 자연스럽게 떠 있으며, 로봇 손은 쓰레기 더미 속에 물리적으로 안정되게 박혀 지지받고 있습니다."
       },
       {
        "label": "B",
        "direction": "물고기가 사선으로 뻗어 나온 로봇 손가락 끝을 향해 입을 대고 있습니다.",
        "built_space": "바닷속 해저면에 타이어와 플라스틱 등 여러 형태의 폐기물들이 흩어져 쌓여 있습니다.",
        "entities": "물고기의 측면, 금속 질감의 로봇 손과 위로 솟은 손가락, 해저 쓰레기 더미가 식별됩니다.",
        "hard_violations": [],
        "physics": "물고기는 물의 부력으로 떠 있고, 로봇 손은 바닥의 쓰레기 더미 위에 얹혀 지탱되고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "물고기 측면을 더 크게 담고 접촉점 오른쪽의 손가락을 비교적 작게 유지하여, 주둥이 접촉을 중심으로 한 클로즈업 지시에 더 충실하다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "수평으로 마주 닿은 주둥이와 손끝은 정확하지만, 물고기보다 로봇 손과 넓은 해저 맥락의 비중이 커 작은 손가락을 보조 초점으로 삼으라는 구도에서 다소 벗어난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "주 물고기는 왼쪽에서 오른쪽을 향한 측면이며, 주둥이가 오른쪽 로봇 손끝에 직접 닿는다. 손가락은 오른쪽 아래에서 왼쪽 위로 뻗어 물고기의 입 높이에서 만난다. 배경 물고기들은 여러 방향으로 놓여 있으나 흐려서 개별 시선이나 이동 방향은 확정하기 어렵다.",
        "built_space": "건축물이나 고정 설비는 없다. 하단에 겹겹이 쌓인 폐기물 더미가 있고, 그 오른쪽에서 로봇 손 하나가 드러난다. 손 뒤에는 타이어 하나, 왼쪽 아래에는 밧줄과 그물성 섬유, 주변에는 파손된 용기와 부품들이 보인다. 손의 밑부분은 잔해 속에 묻혀 있어 지정된 해저 장소와 맞는다.",
        "entities": "큰 비늘과 지느러미를 가진 물고기 한 마리가 주 피사체이며 배경에도 여러 물고기가 있다. 접촉 대상은 부식과 관절 구조가 보이는 낡은 금속 로봇 손가락이다. 푸른 낮 수중광, 해저 쓰레기, 얽힌 섬유가 표현되어 있다. 로봇 몸통과 가슴은 가려져 그물의 몸통 결박 상태와 가슴 표식은 확인할 수 없지만, 클로즈업에서 드러낼 필요는 없다. 사람이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "물고기는 물속에서 부력과 지느러미로 유영 자세를 유지하므로 공중 부유 문제가 없다. 로봇 손가락은 관절을 통해 손에 연결되고, 손의 밑부분은 해저 잔해에 지지된다. 타이어와 다른 폐기물도 바닥이나 서로 겹친 잔해에 놓여 있다. 주둥이와 손끝의 가벼운 접촉은 물리적으로 자연스럽다."
       },
       {
        "label": "B",
        "direction": "주 물고기는 왼쪽에서 오른쪽을 바라보는 측면이며, 입술이 왼쪽으로 뻗은 로봇 손끝에 닿는다. 손가락은 거의 수평으로 뻗어 접촉면과 주둥이 높이가 일치한다. 배경의 작은 물고기들은 흐려서 정확한 시선과 진행 방향을 판별하기 어렵다.",
        "built_space": "건축물이나 고정 설비는 없다. 오른쪽 하단의 쓰레기 더미에서 로봇 손 하나가 드러나고, 왼쪽에는 모래 해저가 넓게 보인다. 손 주변에 그물과 굵은 밧줄, 기울어진 폐용기와 판재가 층을 이룬다. 지정된 장소에는 맞지만 손 전체와 주변 잔해가 차지하는 화면 비중이 크다.",
        "entities": "주 피사체는 어두운 비늘의 물고기 한 마리이며 배경에도 작은 물고기들이 보인다. 상대 물체는 녹슬고 마모된 관절식 로봇 손으로, 펼친 손가락 하나와 접힌 다른 손가락들이 식별된다. 푸른 수중광과 해저 폐기물, 그물이 있다. 묻힌 몸통과 가슴 표식은 보이지 않아 확인 대상 밖이다. 사람, 읽을 수 있는 글자나 그래픽은 없다.",
        "hard_violations": [],
        "physics": "물고기는 물속에서 지느러미를 펼친 정상적인 유영 자세다. 뻗은 손가락은 로봇 손과 연결되어 있고, 손과 손목은 아래쪽 잔해에 받쳐져 있다. 그물과 밧줄은 폐기물 위에 걸쳐 있으며 주변 용기들도 잔해에 기대어 있다. 주둥이 접촉이나 물체의 지지에서 물리적 모순은 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "물고기 측면을 더 크게 담고 접촉점 오른쪽의 손가락을 비교적 작게 유지하여, 주둥이 접촉을 중심으로 한 클로즈업 지시에 더 충실하다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "수평으로 마주 닿은 주둥이와 손끝은 정확하지만, 물고기보다 로봇 손과 넓은 해저 맥락의 비중이 커 작은 손가락을 보조 초점으로 삼으라는 구도에서 다소 벗어난다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "주 물고기는 왼쪽에서 오른쪽을 향한 측면이며, 주둥이가 오른쪽 로봇 손끝에 직접 닿는다. 손가락은 오른쪽 아래에서 왼쪽 위로 뻗어 물고기의 입 높이에서 만난다. 배경 물고기들은 여러 방향으로 놓여 있으나 흐려서 개별 시선이나 이동 방향은 확정하기 어렵다.",
        "built_space": "건축물이나 고정 설비는 없다. 하단에 겹겹이 쌓인 폐기물 더미가 있고, 그 오른쪽에서 로봇 손 하나가 드러난다. 손 뒤에는 타이어 하나, 왼쪽 아래에는 밧줄과 그물성 섬유, 주변에는 파손된 용기와 부품들이 보인다. 손의 밑부분은 잔해 속에 묻혀 있어 지정된 해저 장소와 맞는다.",
        "entities": "큰 비늘과 지느러미를 가진 물고기 한 마리가 주 피사체이며 배경에도 여러 물고기가 있다. 접촉 대상은 부식과 관절 구조가 보이는 낡은 금속 로봇 손가락이다. 푸른 낮 수중광, 해저 쓰레기, 얽힌 섬유가 표현되어 있다. 로봇 몸통과 가슴은 가려져 그물의 몸통 결박 상태와 가슴 표식은 확인할 수 없지만, 클로즈업에서 드러낼 필요는 없다. 사람이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "물고기는 물속에서 부력과 지느러미로 유영 자세를 유지하므로 공중 부유 문제가 없다. 로봇 손가락은 관절을 통해 손에 연결되고, 손의 밑부분은 해저 잔해에 지지된다. 타이어와 다른 폐기물도 바닥이나 서로 겹친 잔해에 놓여 있다. 주둥이와 손끝의 가벼운 접촉은 물리적으로 자연스럽다."
       },
       {
        "label": "A",
        "direction": "주 물고기는 왼쪽에서 오른쪽을 바라보는 측면이며, 입술이 왼쪽으로 뻗은 로봇 손끝에 닿는다. 손가락은 거의 수평으로 뻗어 접촉면과 주둥이 높이가 일치한다. 배경의 작은 물고기들은 흐려서 정확한 시선과 진행 방향을 판별하기 어렵다.",
        "built_space": "건축물이나 고정 설비는 없다. 오른쪽 하단의 쓰레기 더미에서 로봇 손 하나가 드러나고, 왼쪽에는 모래 해저가 넓게 보인다. 손 주변에 그물과 굵은 밧줄, 기울어진 폐용기와 판재가 층을 이룬다. 지정된 장소에는 맞지만 손 전체와 주변 잔해가 차지하는 화면 비중이 크다.",
        "entities": "주 피사체는 어두운 비늘의 물고기 한 마리이며 배경에도 작은 물고기들이 보인다. 상대 물체는 녹슬고 마모된 관절식 로봇 손으로, 펼친 손가락 하나와 접힌 다른 손가락들이 식별된다. 푸른 수중광과 해저 폐기물, 그물이 있다. 묻힌 몸통과 가슴 표식은 보이지 않아 확인 대상 밖이다. 사람, 읽을 수 있는 글자나 그래픽은 없다.",
        "hard_violations": [],
        "physics": "물고기는 물속에서 지느러미를 펼친 정상적인 유영 자세다. 뻗은 손가락은 로봇 손과 연결되어 있고, 손과 손목은 아래쪽 잔해에 받쳐져 있다. 그물과 밧줄은 폐기물 위에 걸쳐 있으며 주변 용기들도 잔해에 기대어 있다. 주둥이 접촉이나 물체의 지지에서 물리적 모순은 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.889,
    "B": 1.875
   },
   "adjusted": {
    "A": 1.889,
    "B": 1.875
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1889,
   "B": 1875
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1889,
    "verdict_ko": "프롬프트가 요구한 대로 물고기의 주둥이와 로봇 손가락의 접촉면이 수평을 이루는 측면 구도를 매우 정확하고 자연스럽게 구현했습니다."
   },
   {
    "label": "B",
    "score": 1875,
    "verdict_ko": "물고기와 로봇 손가락의 접촉은 잘 표현되었으나, 손가락이 위로 솟구친 사선 형태여서 접촉면이 수평을 이룬다는 지시사항의 정확도가 A에 비해 약간 떨어집니다."
   }
  ],
  "refs": [],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-89cf-7206-bd09-c68069044000",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S1sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T03:45:43.737998+00:00",
  "fingerprint": "b0c147d7b0e90dbf70411ddea65fd3242b49e40f54a6210ee03502f955f07e1a",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S1sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S1sh1_sel.png",
  "source_sha256": "31585cd1caef41d122e806cf99793db1e6651e0ddf18c54ee9489a9c91764d02",
  "file": "S1sh1_cine.png",
  "staged_sha256": "d12bac9d46b562171b8c7dd50ff590d8e816301f84130d95dcca37c98f8f9d7c",
  "latency_ms": 11212
 },
 "S2sh3::signage": {
  "fp": "f33ea41b6728da5d",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S2sh3": {
  "input_fingerprint": "f25512556d71a15d",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쏟아진 쓰레기 더미 사이, 그물에 몸이 감긴 채 널브러져 있는 낡은 고철 로봇 찰리의 전신.\n\nLOCATION (lock): On the open deck of a garbage collection ship, among freshly dumped rubbish and tangled fishing net. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Net around 찰리 (Wrapped around his body among the dumped rubbish); used as Crossing lines that reveal confinement while leaving the full-body silhouette readable; Dumped rubbish (Deposited on the collection ship's deck around 찰리); used as Uneven foreground and background layers surrounding the revealed body; Collection ship deck (Receiving the collected rubbish) — Its upper surface is seen obliquely beneath gaps in the rubbish; used as A spatial base establishing that the body is now aboard the ship.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light with controlled contrast keeps the net and aged robot body legible without romanticizing the discarded surroundings.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is sprawled motionless among the rubbish deposited on the collection ship's deck, with a net wrapped around his body. The source does not establish which side of his body faces upward or the individual positions of his head, arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Rubbish has accumulated on the collection ship's deck, with the old scrap robot Charlie lying among it, still entangled in netting and bearing a worn Ubik chest logo. Floating waste around the ship includes shattered helicopter wreckage and a ship broken in half.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쏟아진 쓰레기 더미 사이, 그물에 몸이 감긴 채 널브러져 있는 낡은 고철 로봇 찰리의 전신.\n\nLOCATION (lock): On the open deck of a garbage collection ship, among freshly dumped rubbish and tangled fishing net. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Net around 찰리 (Wrapped around his body among the dumped rubbish); used as Crossing lines that reveal confinement while leaving the full-body silhouette readable; Dumped rubbish (Deposited on the collection ship's deck around 찰리); used as Uneven foreground and background layers surrounding the revealed body; Collection ship deck (Receiving the collected rubbish) — Its upper surface is seen obliquely beneath gaps in the rubbish; used as A spatial base establishing that the body is now aboard the ship.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light with controlled contrast keeps the net and aged robot body legible without romanticizing the discarded surroundings.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is sprawled motionless among the rubbish deposited on the collection ship's deck, with a net wrapped around his body. The source does not establish which side of his body faces upward or the individual positions of his head, arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Rubbish has accumulated on the collection ship's deck, with the old scrap robot Charlie lying among it, still entangled in netting and bearing a worn Ubik chest logo. Floating waste around the ship includes shattered helicopter wreckage and a ship broken in half.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쏟아진 쓰레기 더미 사이, 그물에 몸이 감긴 채 널브러져 있는 낡은 고철 로봇 찰리의 전신.\n\nLOCATION (lock): On the open deck of a garbage collection ship, among freshly dumped rubbish and tangled fishing net. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Net around 찰리 (Wrapped around his body among the dumped rubbish); used as Crossing lines that reveal confinement while leaving the full-body silhouette readable; Dumped rubbish (Deposited on the collection ship's deck around 찰리); used as Uneven foreground and background layers surrounding the revealed body; Collection ship deck (Receiving the collected rubbish) — Its upper surface is seen obliquely beneath gaps in the rubbish; used as A spatial base establishing that the body is now aboard the ship.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light with controlled contrast keeps the net and aged robot body legible without romanticizing the discarded surroundings.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is sprawled motionless among the rubbish deposited on the collection ship's deck, with a net wrapped around his body. The source does not establish which side of his body faces upward or the individual positions of his head, arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Rubbish has accumulated on the collection ship's deck, with the old scrap robot Charlie lying among it, still entangled in netting and bearing a worn Ubik chest logo. Floating waste around the ship includes shattered helicopter wreckage and a ship broken in half.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S2sh3__bgfirst_bg.png",
     "asset_id": "22001d0d-2139-4c65-801f-8fe7b7978313",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S2sh3.png",
     "asset_id": "b077dc93-9217-44ee-9e06-e7c3ca0dcd21",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_collection_deck_b9e7b9.png",
     "asset_id": "ad655f79-61dd-4341-a8cb-31e490c3f150",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "로봇의 머리는 하늘을 정면으로 향해 있음.",
    "built_space": "레퍼런스와 동일한 갑판 환경 및 바다 위 난파선 배경이 보임.",
    "entities": "찰리의 얼굴 형태가 레퍼런스와 다르고 Ubik 로고가 없으며, 로봇과 밧줄이 평면적인 2D 그래픽으로 묘사됨.",
    "hard_violations": [
     "[gemini-pro] 로봇과 밧줄이 실사 배경 위에 평면 그래픽으로 삽입된 콜라주 형태 (사실적 재질감 위반)"
    ],
    "physics": "로봇의 몸체가 쓰레기 더미 위에 평평하게 얹혀 있음."
   },
   {
    "label": "B",
    "direction": "로봇의 얼굴은 약간 오른쪽 위 허공을 향해 누워 있음.",
    "built_space": "쓰레기 수거선 갑판 구조와 배경의 난파선들이 레퍼런스 위치에 정확히 구현됨.",
    "entities": "찰리의 샌드 베이지 장갑, 흰색 마스크, 가슴의 Ubik 로고가 레퍼런스와 일치하며 입체적이고 사실적임.",
    "hard_violations": [
     "[gpt-high] 가슴에 ‘Ubik’라는 판독 가능한 로고가 노출되어, 읽을 수 있는 글자와 로고를 금지한 지시를 위반한다."
    ],
    "physics": "로봇은 갑판의 쓰레기 더미 위에 무게감을 지닌 채 자연스럽게 널브러져 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "캐릭터의 외형, 가슴의 Ubik 로고, 그물에 감긴 채 널브러진 포즈, 그리고 주변 배경과 질감을 사실적이고 정확하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "사진과 같은 사실성 지시를 위반하고 로봇을 2D 그래픽 형태로 렌더링했으며, 캐릭터 디자인과 로고 등 주요 디테일이 누락되었습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "로봇의 얼굴은 약간 오른쪽 위 허공을 향해 누워 있음.",
        "built_space": "쓰레기 수거선 갑판 구조와 배경의 난파선들이 레퍼런스 위치에 정확히 구현됨.",
        "entities": "찰리의 샌드 베이지 장갑, 흰색 마스크, 가슴의 Ubik 로고가 레퍼런스와 일치하며 입체적이고 사실적임.",
        "hard_violations": [],
        "physics": "로봇은 갑판의 쓰레기 더미 위에 무게감을 지닌 채 자연스럽게 널브러져 있음."
       },
       {
        "label": "A",
        "direction": "로봇의 머리는 하늘을 정면으로 향해 있음.",
        "built_space": "레퍼런스와 동일한 갑판 환경 및 바다 위 난파선 배경이 보임.",
        "entities": "찰리의 얼굴 형태가 레퍼런스와 다르고 Ubik 로고가 없으며, 로봇과 밧줄이 평면적인 2D 그래픽으로 묘사됨.",
        "hard_violations": [
         "로봇과 밧줄이 실사 배경 위에 평면 그래픽으로 삽입된 콜라주 형태 (사실적 재질감 위반)"
        ],
        "physics": "로봇의 몸체가 쓰레기 더미 위에 평평하게 얹혀 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "캐릭터의 외형, 가슴의 Ubik 로고, 그물에 감긴 채 널브러진 포즈, 그리고 주변 배경과 질감을 사실적이고 정확하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "사진과 같은 사실성 지시를 위반하고 로봇을 2D 그래픽 형태로 렌더링했으며, 캐릭터 디자인과 로고 등 주요 디테일이 누락되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "로봇의 얼굴은 약간 오른쪽 위 허공을 향해 누워 있음.",
        "built_space": "쓰레기 수거선 갑판 구조와 배경의 난파선들이 레퍼런스 위치에 정확히 구현됨.",
        "entities": "찰리의 샌드 베이지 장갑, 흰색 마스크, 가슴의 Ubik 로고가 레퍼런스와 일치하며 입체적이고 사실적임.",
        "hard_violations": [],
        "physics": "로봇은 갑판의 쓰레기 더미 위에 무게감을 지닌 채 자연스럽게 널브러져 있음."
       },
       {
        "label": "A",
        "direction": "로봇의 머리는 하늘을 정면으로 향해 있음.",
        "built_space": "레퍼런스와 동일한 갑판 환경 및 바다 위 난파선 배경이 보임.",
        "entities": "찰리의 얼굴 형태가 레퍼런스와 다르고 Ubik 로고가 없으며, 로봇과 밧줄이 평면적인 2D 그래픽으로 묘사됨.",
        "hard_violations": [
         "로봇과 밧줄이 실사 배경 위에 평면 그래픽으로 삽입된 콜라주 형태 (사실적 재질감 위반)"
        ],
        "physics": "로봇의 몸체가 쓰레기 더미 위에 평평하게 얹혀 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "갑판과 그물 속 전신 배치는 부합하지만, 가슴의 읽을 수 있는 ‘Ubik’ 표기가 명시적인 문자·로고 금지를 위반하며 다리도 참조보다 지나치게 길다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "쓰레기에 받쳐져 누운 전신과 몸을 가로질러 감싼 그물이 요구한 순간을 잘 구현하고 판독 가능한 문자도 없지만, 긴 다리와 다소 삽화적인 장갑판 윤곽은 참조 및 실사 질감에서 아쉽다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 머리는 화면 오른쪽 뒤에, 발은 왼쪽 및 중앙 전경에 놓여 있다. 흰 얼굴은 위쪽과 카메라 쪽으로 기울어 있으며 특정 대상을 응시하는 행동은 없다. 양팔은 몸 옆 쓰레기 쪽으로 내려가 있다. 조준하거나 이동하는 물체는 없다.",
        "built_space": "녹슨 금속 현측이 갑판의 왼쪽과 뒤쪽을 둘러싼다. 둘레의 계선주 다섯 개와 왼쪽 전경의 수직 관 하나, 뒤쪽 중앙의 윈치 한 대, 오른쪽 크레인 한 대와 직사각형 금속 함 하나가 보인다. 참조의 주요 설비와 배치가 유지되며, 쓰레기 사이와 몸 뒤쪽으로 젖은 갑판 윗면이 비스듬히 드러난다. 찰리는 난간 안쪽 쓰레기 더미에 놓여 있다.",
        "entities": "등장 개체는 로봇 찰리 한 대뿐이다. 흰 마스크형 얼굴, 점 형태 눈과 선 형태 입, 베이지색 각진 상체 장갑과 육중한 팔은 참조에 가깝다. 그러나 길게 뻗은 원통형 하퇴와 발은 참조의 짧은 다리 및 발 형태와 다르다. 몸통과 다리에 녹색 어망이 걸쳐져 있고 주변에는 부표, 상자, 타이어와 폐기물이 있다. 바다에는 헬리콥터 잔해와 두 동강 난 선박이 보인다. 가슴의 ‘Ubik’는 읽을 수 있어 최종 문자 금지와 충돌한다. 낮의 자연광과 녹슨 금속·젖은 쓰레기 질감은 확인된다.",
        "hard_violations": [
         "가슴에 ‘Ubik’라는 판독 가능한 로고가 노출되어, 읽을 수 있는 글자와 로고를 금지한 지시를 위반한다."
        ],
        "physics": "몸통은 뒤쪽 쓰레기 더미에 기대어 있고 골반과 다리는 아래의 폐기물에 놓여 있다. 양손도 주변 잔해 위에 내려앉아 있다. 머리는 목과 뒤쪽 잔해에 이어져 있으며, 명백히 공중에 떠 있는 신체 부위는 보이지 않는다. 그물은 몸과 쓰레기 위에 걸쳐져 지지된다. 다만 상체와 머리가 비교적 세워져 있어 완전히 축 늘어진 느낌은 약하다."
       },
       {
        "label": "B",
        "direction": "머리는 오른쪽 후경, 두 발은 왼쪽과 중앙 전경을 향한다. 얼굴은 하늘과 카메라 쪽으로 비스듬히 향하고 특정 표적을 바라보지 않는다. 양팔은 몸 옆으로 떨어져 있으며 오른쪽 화면의 손은 흰 폐기물 위를 향해 내려앉아 있다. 이동이나 조준 행동은 없다.",
        "built_space": "참조와 같은 녹슨 현측, 둘레의 계선주 다섯 개와 왼쪽 수직 관 하나가 보인다. 뒤쪽 중앙에 윈치 한 대, 오른쪽에 크레인 한 대와 금속 함 하나가 있으며 설비의 중복은 없다. 찰리는 갑판 안쪽에 누워 있고 몸 주변의 틈과 후경에서 갑판 표면이 드러난다. 전경과 후경의 쓰레기가 몸을 둘러싸면서도 전신의 배치를 읽을 수 있다.",
        "entities": "찰리 한 대가 있으며 추가 인물은 없다. 베이지색 각진 장갑, 흰 마스크형 얼굴, 작은 눈과 선 형태 입, 큰 어깨와 굵은 팔이 확인된다. 다만 참조의 고릴라형 비례보다 다리가 길고 장갑판의 세부 형태도 다르다. 굵은 망줄이 몸통과 사지를 교차하며 감싸고, 주변에 녹색 어망과 상자·부표·타이어 등 쓰레기가 쌓여 있다. 배 밖에는 헬리콥터 잔해와 분리된 선체 두 부분이 있다. 가슴 표식은 마모되어 읽히지 않는다. 낮의 배경은 실사적이나 로봇의 짙은 윤곽과 균일한 표면 음영은 다소 삽화처럼 보인다.",
        "hard_violations": [],
        "physics": "등과 머리는 뒤쪽 쓰레기층에 기대어 있고 골반·다리·발은 갑판 위 잔해에 받쳐져 있다. 화면 오른쪽 손은 흰색 폐기물 표면에 직접 놓여 있으며 반대쪽 팔도 몸 옆 쓰레기에 내려앉아 있다. 몸을 능동적으로 들어 올린 자세나 지지 없이 떠 있는 부위는 보이지 않는다. 그물도 몸의 굴곡과 주변 폐기물에 걸려 무게를 받는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "갑판과 그물 속 전신 배치는 부합하지만, 가슴의 읽을 수 있는 ‘Ubik’ 표기가 명시적인 문자·로고 금지를 위반하며 다리도 참조보다 지나치게 길다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "쓰레기에 받쳐져 누운 전신과 몸을 가로질러 감싼 그물이 요구한 순간을 잘 구현하고 판독 가능한 문자도 없지만, 긴 다리와 다소 삽화적인 장갑판 윤곽은 참조 및 실사 질감에서 아쉽다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 머리는 화면 오른쪽 뒤에, 발은 왼쪽 및 중앙 전경에 놓여 있다. 흰 얼굴은 위쪽과 카메라 쪽으로 기울어 있으며 특정 대상을 응시하는 행동은 없다. 양팔은 몸 옆 쓰레기 쪽으로 내려가 있다. 조준하거나 이동하는 물체는 없다.",
        "built_space": "녹슨 금속 현측이 갑판의 왼쪽과 뒤쪽을 둘러싼다. 둘레의 계선주 다섯 개와 왼쪽 전경의 수직 관 하나, 뒤쪽 중앙의 윈치 한 대, 오른쪽 크레인 한 대와 직사각형 금속 함 하나가 보인다. 참조의 주요 설비와 배치가 유지되며, 쓰레기 사이와 몸 뒤쪽으로 젖은 갑판 윗면이 비스듬히 드러난다. 찰리는 난간 안쪽 쓰레기 더미에 놓여 있다.",
        "entities": "등장 개체는 로봇 찰리 한 대뿐이다. 흰 마스크형 얼굴, 점 형태 눈과 선 형태 입, 베이지색 각진 상체 장갑과 육중한 팔은 참조에 가깝다. 그러나 길게 뻗은 원통형 하퇴와 발은 참조의 짧은 다리 및 발 형태와 다르다. 몸통과 다리에 녹색 어망이 걸쳐져 있고 주변에는 부표, 상자, 타이어와 폐기물이 있다. 바다에는 헬리콥터 잔해와 두 동강 난 선박이 보인다. 가슴의 ‘Ubik’는 읽을 수 있어 최종 문자 금지와 충돌한다. 낮의 자연광과 녹슨 금속·젖은 쓰레기 질감은 확인된다.",
        "hard_violations": [
         "가슴에 ‘Ubik’라는 판독 가능한 로고가 노출되어, 읽을 수 있는 글자와 로고를 금지한 지시를 위반한다."
        ],
        "physics": "몸통은 뒤쪽 쓰레기 더미에 기대어 있고 골반과 다리는 아래의 폐기물에 놓여 있다. 양손도 주변 잔해 위에 내려앉아 있다. 머리는 목과 뒤쪽 잔해에 이어져 있으며, 명백히 공중에 떠 있는 신체 부위는 보이지 않는다. 그물은 몸과 쓰레기 위에 걸쳐져 지지된다. 다만 상체와 머리가 비교적 세워져 있어 완전히 축 늘어진 느낌은 약하다."
       },
       {
        "label": "A",
        "direction": "머리는 오른쪽 후경, 두 발은 왼쪽과 중앙 전경을 향한다. 얼굴은 하늘과 카메라 쪽으로 비스듬히 향하고 특정 표적을 바라보지 않는다. 양팔은 몸 옆으로 떨어져 있으며 오른쪽 화면의 손은 흰 폐기물 위를 향해 내려앉아 있다. 이동이나 조준 행동은 없다.",
        "built_space": "참조와 같은 녹슨 현측, 둘레의 계선주 다섯 개와 왼쪽 수직 관 하나가 보인다. 뒤쪽 중앙에 윈치 한 대, 오른쪽에 크레인 한 대와 금속 함 하나가 있으며 설비의 중복은 없다. 찰리는 갑판 안쪽에 누워 있고 몸 주변의 틈과 후경에서 갑판 표면이 드러난다. 전경과 후경의 쓰레기가 몸을 둘러싸면서도 전신의 배치를 읽을 수 있다.",
        "entities": "찰리 한 대가 있으며 추가 인물은 없다. 베이지색 각진 장갑, 흰 마스크형 얼굴, 작은 눈과 선 형태 입, 큰 어깨와 굵은 팔이 확인된다. 다만 참조의 고릴라형 비례보다 다리가 길고 장갑판의 세부 형태도 다르다. 굵은 망줄이 몸통과 사지를 교차하며 감싸고, 주변에 녹색 어망과 상자·부표·타이어 등 쓰레기가 쌓여 있다. 배 밖에는 헬리콥터 잔해와 분리된 선체 두 부분이 있다. 가슴 표식은 마모되어 읽히지 않는다. 낮의 배경은 실사적이나 로봇의 짙은 윤곽과 균일한 표면 음영은 다소 삽화처럼 보인다.",
        "hard_violations": [],
        "physics": "등과 머리는 뒤쪽 쓰레기층에 기대어 있고 골반·다리·발은 갑판 위 잔해에 받쳐져 있다. 화면 오른쪽 손은 흰색 폐기물 표면에 직접 놓여 있으며 반대쪽 팔도 몸 옆 쓰레기에 내려앉아 있다. 몸을 능동적으로 들어 올린 자세나 지지 없이 떠 있는 부위는 보이지 않는다. 그물도 몸의 굴곡과 주변 폐기물에 걸려 무게를 받는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.429,
    "B": 1.375
   },
   "adjusted": {
    "A": 1.179,
    "B": 1.125
   },
   "violations": {
    "A": [
     "[gemini-pro] 로봇과 밧줄이 실사 배경 위에 평면 그래픽으로 삽입된 콜라주 형태 (사실적 재질감 위반)"
    ],
    "B": [
     "[gpt-high] 가슴에 ‘Ubik’라는 판독 가능한 로고가 노출되어, 읽을 수 있는 글자와 로고를 금지한 지시를 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1125,
   "A": 1179
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1125,
    "verdict_ko": "캐릭터의 외형, 가슴의 Ubik 로고, 그물에 감긴 채 널브러진 포즈, 그리고 주변 배경과 질감을 사실적이고 정확하게 구현했습니다.  ★위반: [gpt-high] 가슴에 ‘Ubik’라는 판독 가능한 로고가 노출되어, 읽을 수 있는 글자와 로고를 금지한 지시를 위반한다."
   },
   {
    "label": "A",
    "score": 1179,
    "verdict_ko": "사진과 같은 사실성 지시를 위반하고 로봇을 2D 그래픽 형태로 렌더링했으며, 캐릭터 디자인과 로고 등 주요 디테일이 누락되었습니다.  ★위반: [gemini-pro] 로봇과 밧줄이 실사 배경 위에 평면 그래픽으로 삽입된 콜라주 형태 (사실적 재질감 위반)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_collection_deck_b9e7b9.png",
    "asset_id": "ad655f79-61dd-4341-a8cb-31e490c3f150",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-8ba4-77fd-8259-9e76fc782f25",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S2sh3__bgfirst_bg.png",
   "bg_asset_id": "22001d0d-2139-4c65-801f-8fe7b7978313",
   "bg_record_key": "S2sh3::bgfirst_bg",
   "chain_winner": true,
   "authority": "groupbg",
   "group_key": "collection_deck",
   "groupbg_asset_id": "ad655f79-61dd-4341-a8cb-31e490c3f150"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S2sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T10:54:04.157839+00:00",
  "fingerprint": "228ac898204fb1324a7f0fa973a95bd362218ec8e9bfe50592ec388cceb73d51",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S2sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S2sh3_sel.png",
  "source_sha256": "55d8ddc5ed65f374e7921fdb56dc2d95b8d601f6d7bda5a58c2d332e73642ec9",
  "file": "S2sh3_cine.png",
  "staged_sha256": "37a969ff02aca21d6ddc5647ed51a01fd82ebbd3bfbb842945c8cbe7b43f336f",
  "latency_ms": 9703
 },
 "S3sh2::confined_fp_apt": {
  "applies": true,
  "reason_ko": "덤프트럭 운전석 내부라는 제한된 통제 공간에서 운전대와 운전자의 위치 및 방향 관계가 정확하게 묘사되어야 하므로 레이아웃 가이드가 필요합니다.",
  "input_fingerprint": "5886dc4de21d8287"
 },
 "S3sh2::signage": {
  "fp": "7203ed35b954bc2c",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "confinedfp::a1ed7fb485b1": {
  "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/confinedfp_base_a1ed7fb485b1.png",
  "place_text": "Inside the moving dump truck's compact driver's cab, at the steering position with daylight beyond the windshield.",
  "input_fingerprint": "c8d955a7a74b3395"
 },
 "S3sh2::confined_fp": {
  "reads": {
   "controls": "A steering wheel is located at the left-side seat.",
   "mirrors": "No mirrors are depicted in the diagram.",
   "camera": "Positioned behind and between the two seats, angled diagonally forward and left to point directly at the left seat.",
   "occupants": "A person occupies the left seat (driver's seat)."
  },
  "mismatches": [],
  "scene_description_en": "The camera views the cab interior diagonally, looking forward and left from a point between the seats. The driver sits in the left-hand seat, appearing in the center-left of the frame from a rear-oblique angle. A steering wheel sits directly in front of the driver on the left side of the screen. The windshield stretches across the background space ahead. No mirrors or reflective surfaces are visible from this camera position.",
  "fixed": false,
  "input_fingerprint": "a5d7961103c50693"
 },
 "era_assess::96792db9ad07ec93": {
  "subjects": [],
  "subject_text": "인천 난민촌 외곽 도로와 무인점포 앞\n임시 주거지 외곽에서 도심 방향으로 이어지는 도로. 길 건너편에 무인점포의 전면과 출입구가 보인다.",
  "identity": "canonical",
  "scope_id": "L194",
  "scope_role": "location_exterior",
  "scope_sha": "194604be5aa0a228"
 },
 "S3sh2": {
  "input_fingerprint": "5ca986d0e5e7ceb3",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 덤프트럭 운전석에서 고개를 떨군 채 조는 운전사의 모습.\n\nLOCATION (lock): Inside the moving dump truck's compact driver's cab, at the steering position with daylight beyond the windshield. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Dump truck cab interior (Occupied by the dozing driver while the truck travels autonomously) — Viewed diagonally from the passenger side toward the driving position; used as Close spatial enclosure around the driver's lowered head and shoulders.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light within the cab maintains natural facial detail and restrained tonal contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): The dump-truck driver is seated asleep in the driver's seat with his head drooping forward. The source does not specify the placement of his arms and legs or whether his torso rests against the seat back.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The self-driving dump truck carries a full load of scrap, including the old robot Charlie, still entangled in netting with a worn Ubik logo on its chest.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera views the cab interior diagonally, looking forward and left from a point between the seats. The driver sits in the left-hand seat, appearing in the center-left of the frame from a rear-oblique angle. A steering wheel sits directly in front of the driver on the left side of the screen. The windshield stretches across the background space ahead. No mirrors or reflective surfaces are visible from this camera position.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 덤프트럭 운전석에서 고개를 떨군 채 조는 운전사의 모습.\n\nLOCATION (lock): Inside the moving dump truck's compact driver's cab, at the steering position with daylight beyond the windshield. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light within the cab maintains natural facial detail and restrained tonal contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): The dump-truck driver is seated asleep in the driver's seat with his head drooping forward. The source does not specify the placement of his arms and legs or whether his torso rests against the seat back.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The self-driving dump truck carries a full load of scrap, including the old robot Charlie, still entangled in netting with a worn Ubik logo on its chest.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera views the cab interior diagonally, looking forward and left from a point between the seats. The driver sits in the left-hand seat, appearing in the center-left of the frame from a rear-oblique angle. A steering wheel sits directly in front of the driver on the left side of the screen. The windshield stretches across the background space ahead. No mirrors or reflective surfaces are visible from this camera position.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 덤프트럭 운전석에서 고개를 떨군 채 조는 운전사의 모습.\n\nLOCATION (lock): Inside the moving dump truck's compact driver's cab, at the steering position with daylight beyond the windshield. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light within the cab maintains natural facial detail and restrained tonal contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): The dump-truck driver is seated asleep in the driver's seat with his head drooping forward. The source does not specify the placement of his arms and legs or whether his torso rests against the seat back.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The self-driving dump truck carries a full load of scrap, including the old robot Charlie, still entangled in netting with a worn Ubik logo on its chest.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S3sh2_confinedfp.png",
     "asset_id": null,
     "role": null
    }
   ],
   "B": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S3sh2_confinedfp.png",
     "asset_id": null,
     "role": null
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라가 조수석에서 운전석을 대각선으로 바라보며, 운전자의 고개는 무릎을 향해 푹 수그러져 있음.",
    "built_space": "덤프트럭 내부. 좌측에 온전한 형태의 스티어링 휠과 대시보드가 정상적으로 배치되어 있으며 창밖으로 낮의 도로가 보임.",
    "entities": "회색 셔츠를 입고 졸고 있는 중년의 한국인 남성 운전자.",
    "hard_violations": [],
    "physics": "운전자의 몸에 힘이 빠져 있으며, 오른손은 무릎 위에 완전히 지탱되어 중력 지침(근육 노력 없음)을 정확히 따름."
   },
   {
    "label": "B",
    "direction": "조수석 측에서 운전석을 대각선으로 촬영함. 운전자의 고개가 아래로 향해 있음.",
    "built_space": "트럭 운전석 내부. 창밖으로 고속도로와 밝은 낮 풍경이 보임.",
    "entities": "작업복 재킷을 입은 중년의 한국인 남성.",
    "hard_violations": [
     "[gemini-pro] 스티어링 휠의 좌측 테두리가 끊어지고 대시보드와 융합되어 사라지는 물리적으로 불가능한 구조."
    ],
    "physics": "운전자가 의식을 잃은 듯 졸고 있으나, 오른손이 운전대 위에 얹혀 있어 완전히 이완된 상태로 보기에는 약간의 어색함이 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "요구된 대각선 구도와 수면 중인 운전자의 무기력한 자세(중력 지침)를 훌륭하게 구현했으며 공간의 물리적 구조가 안정적입니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "스티어링 휠의 형태가 부서져 대시보드와 융합된 물리적 불가능성(Hard Violation)으로 인해 실격 수준의 치명적 결함이 있습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라가 조수석에서 운전석을 대각선으로 바라보며, 운전자의 고개는 무릎을 향해 푹 수그러져 있음.",
        "built_space": "덤프트럭 내부. 좌측에 온전한 형태의 스티어링 휠과 대시보드가 정상적으로 배치되어 있으며 창밖으로 낮의 도로가 보임.",
        "entities": "회색 셔츠를 입고 졸고 있는 중년의 한국인 남성 운전자.",
        "hard_violations": [],
        "physics": "운전자의 몸에 힘이 빠져 있으며, 오른손은 무릎 위에 완전히 지탱되어 중력 지침(근육 노력 없음)을 정확히 따름."
       },
       {
        "label": "B",
        "direction": "조수석 측에서 운전석을 대각선으로 촬영함. 운전자의 고개가 아래로 향해 있음.",
        "built_space": "트럭 운전석 내부. 창밖으로 고속도로와 밝은 낮 풍경이 보임.",
        "entities": "작업복 재킷을 입은 중년의 한국인 남성.",
        "hard_violations": [
         "스티어링 휠의 좌측 테두리가 끊어지고 대시보드와 융합되어 사라지는 물리적으로 불가능한 구조."
        ],
        "physics": "운전자가 의식을 잃은 듯 졸고 있으나, 오른손이 운전대 위에 얹혀 있어 완전히 이완된 상태로 보기에는 약간의 어색함이 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "요구된 대각선 구도와 수면 중인 운전자의 무기력한 자세(중력 지침)를 훌륭하게 구현했으며 공간의 물리적 구조가 안정적입니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "스티어링 휠의 형태가 부서져 대시보드와 융합된 물리적 불가능성(Hard Violation)으로 인해 실격 수준의 치명적 결함이 있습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라가 조수석에서 운전석을 대각선으로 바라보며, 운전자의 고개는 무릎을 향해 푹 수그러져 있음.",
        "built_space": "덤프트럭 내부. 좌측에 온전한 형태의 스티어링 휠과 대시보드가 정상적으로 배치되어 있으며 창밖으로 낮의 도로가 보임.",
        "entities": "회색 셔츠를 입고 졸고 있는 중년의 한국인 남성 운전자.",
        "hard_violations": [],
        "physics": "운전자의 몸에 힘이 빠져 있으며, 오른손은 무릎 위에 완전히 지탱되어 중력 지침(근육 노력 없음)을 정확히 따름."
       },
       {
        "label": "B",
        "direction": "조수석 측에서 운전석을 대각선으로 촬영함. 운전자의 고개가 아래로 향해 있음.",
        "built_space": "트럭 운전석 내부. 창밖으로 고속도로와 밝은 낮 풍경이 보임.",
        "entities": "작업복 재킷을 입은 중년의 한국인 남성.",
        "hard_violations": [
         "스티어링 휠의 좌측 테두리가 끊어지고 대시보드와 융합되어 사라지는 물리적으로 불가능한 구조."
        ],
        "physics": "운전자가 의식을 잃은 듯 졸고 있으나, 오른손이 운전대 위에 얹혀 있어 완전히 이완된 상태로 보기에는 약간의 어색함이 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "낮의 운전실에서 고개를 떨군 운전사는 잘 구현했지만, 운전대를 쥔 손의 긴장이 남아 있어 완전히 힘을 빼고 조는 자세는 B보다 덜 충실하다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "조수석 쪽 대각 구도와 앞으로 처진 머리, 허벅지에 내려놓은 양손이 운전석에서 잠든 순간과 중력에 순응하는 자세를 가장 충실하게 구현한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "운전사는 눈을 감고 턱을 가슴 쪽으로 떨어뜨렸으며 얼굴은 무릎 방향을 향한다. 전방 도로나 카메라를 응시하지 않는다. 운전대와 계기판은 운전석 쪽을 향하고, 카메라는 조수석 쪽에서 운전사를 비스듬히 바라본다.",
        "built_space": "운전대 1개, 계기판과 대시보드 1세트, 운전석 등받이와 머리받침 각 1개, 운전석 측면 문과 창, 뒤쪽 창이 보인다. 측면에는 상하로 나뉜 외부 거울 2개가 있으며 서로 다른 높이의 도로를 비추는 배치는 가능하다. 운전사는 운전대 뒤 좌석에 앉아 있다. 조수석 자체는 대부분 프레임 밖이므로 도면의 두 좌석 전체 배치를 확인할 수는 없다. 머리부터 허벅지까지 보이지만 대시보드와 도로도 넓게 들어와 요청한 밀착감은 다소 약하다.",
        "entities": "중년의 동아시아계 남성 운전사 1명이며 한국인 설정과 외형상 충돌하지 않는다. 회색 작업복과 어두운 바지, 안전벨트를 착용했다. 낡은 대형 트럭 운전실과 주간 도로가 보이며 추가 인물은 없다. 적재함의 고철과 로봇 찰리는 이 구도에서 보이지 않는 것이 자연스럽다. 운전대 중심에 작은 엠블럼 형태가 남아 있어 로고 금지 조건에는 아쉬움이 있다. 판독 가능한 문장은 보이지 않는다.",
        "hard_violations": [],
        "physics": "엉덩이와 허벅지는 좌석에 지지되고 몸통은 등받이와 안전벨트에 기대어 있다. 머리는 앞으로 처져 있다. 양손은 운전대 아래쪽 테두리에 접촉하므로 허공에 떠 있지는 않지만, 손가락이 테두리를 잡는 모양이라 완전히 힘이 빠져 내려앉은 수면 자세로는 덜 명확하다. 발은 프레임 밖이어서 지지 상태를 판단하지 않는다."
       },
       {
        "label": "B",
        "direction": "운전사는 눈을 감고 고개를 앞으로 숙여 얼굴이 자기 허벅지 쪽을 향한다. 손은 운전대를 조작하지 않는다. 운전대와 계기판은 운전사를 향하며, 카메라는 오른쪽 조수석 부근에서 왼쪽 운전석을 대각으로 바라본다.",
        "built_space": "운전대 1개, 운전석 1개, 오른쪽 전경의 빈 조수석 1개가 보여 도면의 좌우 좌석 관계가 명확하다. 두 좌석은 모두 앞쪽 대시보드를 향한다. 중앙 하단에 변속 레버 1개, 대시보드 위에 소형 화면 1개가 있으며 운전석 옆 문과 측면 창, 뒤쪽 창도 보인다. 운전사는 운전대 바로 뒤 자기 좌석에 정상적으로 앉아 있다. 머리부터 허벅지까지 담은 미디엄 계열 구도지만 조수석과 대시보드가 차지하는 면적은 조금 크다.",
        "entities": "중년의 동아시아계 남성 운전사 1명이며 회색 긴팔 상의와 어두운 바지를 입었다. 별도 인물 정체성이나 복장을 고정하는 참조는 없으므로 충돌하지 않는다. 사용감 있는 대형 트럭 운전실과 낮의 도로가 보인다. 추가 사람은 없으며 적재함과 고철, 그물에 얽힌 찰리는 구도 밖이다. 작은 경고 스티커와 계기 표시는 있으나 문구를 확실하게 판독할 수는 없다. 운전대 중심의 작은 표식은 로고 배제 요구에 다소 미흡하다.",
        "hard_violations": [],
        "physics": "엉덩이와 허벅지는 좌석 쿠션에 놓이고 몸통은 등받이 쪽에 기대어 있다. 양손과 팔은 허벅지와 무릎 위로 내려앉아 지지되며, 머리는 가슴 방향으로 처져 있다. 운전대를 잡거나 물건을 들어 올린 부위가 없어 수면 중 무긴장 자세가 자연스럽다. 화면과 변속 레버도 각각 대시보드와 바닥 장착부에 고정되어 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "낮의 운전실에서 고개를 떨군 운전사는 잘 구현했지만, 운전대를 쥔 손의 긴장이 남아 있어 완전히 힘을 빼고 조는 자세는 B보다 덜 충실하다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "조수석 쪽 대각 구도와 앞으로 처진 머리, 허벅지에 내려놓은 양손이 운전석에서 잠든 순간과 중력에 순응하는 자세를 가장 충실하게 구현한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "운전사는 눈을 감고 턱을 가슴 쪽으로 떨어뜨렸으며 얼굴은 무릎 방향을 향한다. 전방 도로나 카메라를 응시하지 않는다. 운전대와 계기판은 운전석 쪽을 향하고, 카메라는 조수석 쪽에서 운전사를 비스듬히 바라본다.",
        "built_space": "운전대 1개, 계기판과 대시보드 1세트, 운전석 등받이와 머리받침 각 1개, 운전석 측면 문과 창, 뒤쪽 창이 보인다. 측면에는 상하로 나뉜 외부 거울 2개가 있으며 서로 다른 높이의 도로를 비추는 배치는 가능하다. 운전사는 운전대 뒤 좌석에 앉아 있다. 조수석 자체는 대부분 프레임 밖이므로 도면의 두 좌석 전체 배치를 확인할 수는 없다. 머리부터 허벅지까지 보이지만 대시보드와 도로도 넓게 들어와 요청한 밀착감은 다소 약하다.",
        "entities": "중년의 동아시아계 남성 운전사 1명이며 한국인 설정과 외형상 충돌하지 않는다. 회색 작업복과 어두운 바지, 안전벨트를 착용했다. 낡은 대형 트럭 운전실과 주간 도로가 보이며 추가 인물은 없다. 적재함의 고철과 로봇 찰리는 이 구도에서 보이지 않는 것이 자연스럽다. 운전대 중심에 작은 엠블럼 형태가 남아 있어 로고 금지 조건에는 아쉬움이 있다. 판독 가능한 문장은 보이지 않는다.",
        "hard_violations": [],
        "physics": "엉덩이와 허벅지는 좌석에 지지되고 몸통은 등받이와 안전벨트에 기대어 있다. 머리는 앞으로 처져 있다. 양손은 운전대 아래쪽 테두리에 접촉하므로 허공에 떠 있지는 않지만, 손가락이 테두리를 잡는 모양이라 완전히 힘이 빠져 내려앉은 수면 자세로는 덜 명확하다. 발은 프레임 밖이어서 지지 상태를 판단하지 않는다."
       },
       {
        "label": "A",
        "direction": "운전사는 눈을 감고 고개를 앞으로 숙여 얼굴이 자기 허벅지 쪽을 향한다. 손은 운전대를 조작하지 않는다. 운전대와 계기판은 운전사를 향하며, 카메라는 오른쪽 조수석 부근에서 왼쪽 운전석을 대각으로 바라본다.",
        "built_space": "운전대 1개, 운전석 1개, 오른쪽 전경의 빈 조수석 1개가 보여 도면의 좌우 좌석 관계가 명확하다. 두 좌석은 모두 앞쪽 대시보드를 향한다. 중앙 하단에 변속 레버 1개, 대시보드 위에 소형 화면 1개가 있으며 운전석 옆 문과 측면 창, 뒤쪽 창도 보인다. 운전사는 운전대 바로 뒤 자기 좌석에 정상적으로 앉아 있다. 머리부터 허벅지까지 담은 미디엄 계열 구도지만 조수석과 대시보드가 차지하는 면적은 조금 크다.",
        "entities": "중년의 동아시아계 남성 운전사 1명이며 회색 긴팔 상의와 어두운 바지를 입었다. 별도 인물 정체성이나 복장을 고정하는 참조는 없으므로 충돌하지 않는다. 사용감 있는 대형 트럭 운전실과 낮의 도로가 보인다. 추가 사람은 없으며 적재함과 고철, 그물에 얽힌 찰리는 구도 밖이다. 작은 경고 스티커와 계기 표시는 있으나 문구를 확실하게 판독할 수는 없다. 운전대 중심의 작은 표식은 로고 배제 요구에 다소 미흡하다.",
        "hard_violations": [],
        "physics": "엉덩이와 허벅지는 좌석 쿠션에 놓이고 몸통은 등받이 쪽에 기대어 있다. 양손과 팔은 허벅지와 무릎 위로 내려앉아 지지되며, 머리는 가슴 방향으로 처져 있다. 운전대를 잡거나 물건을 들어 올린 부위가 없어 수면 중 무긴장 자세가 자연스럽다. 화면과 변속 레버도 각각 대시보드와 바닥 장착부에 고정되어 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.206
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.956
   },
   "violations": {
    "B": [
     "[gemini-pro] 스티어링 휠의 좌측 테두리가 끊어지고 대시보드와 융합되어 사라지는 물리적으로 불가능한 구조."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 956
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "요구된 대각선 구도와 수면 중인 운전자의 무기력한 자세(중력 지침)를 훌륭하게 구현했으며 공간의 물리적 구조가 안정적입니다."
   },
   {
    "label": "B",
    "score": 956,
    "verdict_ko": "스티어링 휠의 형태가 부서져 대시보드와 융합된 물리적 불가능성(Hard Violation)으로 인해 실격 수준의 치명적 결함이 있습니다.  ★위반: [gemini-pro] 스티어링 휠의 좌측 테두리가 끊어지고 대시보드와 융합되어 사라지는 물리적으로 불가능한 구조."
   }
  ],
  "refs": [
   {
    "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S3sh2_confinedfp.png",
    "asset_id": null,
    "role": null
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-90a6-785d-9219-32ec30a961a3",
  "confined_fp": {
   "base_key": "confinedfp::a1ed7fb485b1",
   "apt_reason": "덤프트럭 운전석 내부라는 제한된 통제 공간에서 운전대와 운전자의 위치 및 방향 관계가 정확하게 묘사되어야 하므로 레이아웃 가이드가 필요합니다.",
   "fixed": false,
   "mismatches": []
  },
  "ref_mode": "confined_fp: 도면+장면설명+엔티티",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S3sh2::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T03:48:54.780163+00:00",
  "fingerprint": "01f3c3b47e62731b477c69eae1f038d3f864b2f3b10226c0a8b826f2825666a9",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S3sh2_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S3sh2_sel.png",
  "source_sha256": "4cfc692cf2b9568283faa0584da19cb23ea11a1c3d669a5b950ad8af932eaf26",
  "file": "S3sh2_cine.png",
  "staged_sha256": "4e6da132ed00aadbb98647034f695a87fa6d8867d475047c4092a06f1ad75c79",
  "latency_ms": 10027
 },
 "S4sh1::confined_fp_apt": {
  "applies": true,
  "reason_ko": "차량 내부의 앞좌석이라는 통제 장치가 있는 제한된 공간을 배경으로 하며, 페드로가 운전석에 앉고 현우가 조수석에 앉아 상호작용하는 정확한 위치 관계가 이야기 전달에 필수적이므로 평면도 레이아웃 보조가 필요합니다.",
  "input_fingerprint": "70ce91fadbf89495"
 },
 "S4sh1::signage": {
  "fp": "f07dec417db507b8",
  "inscriptions": [],
  "cues": [
   {
    "text_native": "",
    "source": "scene_text_implied",
    "source_quote": "구겨진 도면"
   }
  ],
  "dropped": []
 },
 "confinedfp::510c33afc573": {
  "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/confinedfp_base_510c33afc573.png",
  "place_text": "Inside the front seating compartment of a car parked in an affluent residential neighborhood, with daylight entering through the windows.",
  "input_fingerprint": "f9b6d15cd2a96a4b"
 },
 "S4sh1::confined_fp": {
  "reads": {
   "controls": "The steering wheel is in front of the left seat. A center console sits between the left and right seats.",
   "mirrors": "A rearview mirror is mounted at the top center, facing the rear of the cabin.",
   "camera": "The camera is located behind the left seat, angled diagonally forward and to the right.",
   "occupants": "페드로 (Pedro) occupies the left seat (driver's side). 현우 (Hyunwoo) occupies the right seat (passenger's side)."
  },
  "mismatches": [],
  "scene_description_en": "The camera is positioned in the rear left of the car's interior, looking diagonally forward and right from behind the driver's seat. In the left foreground, the back of Pedro's head and shoulder are visible as he occupies the driver's seat. To his right, in the center midground, sits the center console where Pedro's hand reaches toward a crumpled floor plan. In the right midground, Hyunwoo occupies the passenger seat, angled slightly to look down at the console. The steering wheel is located in front of Pedro on the left. High in the upper center background, the rearview mirror faces rearward, positioned to physically reflect the rear of the cabin and the occupants' faces from the camera's perspective.",
  "fixed": true,
  "input_fingerprint": "df64102d5ba7dda3"
 },
 "era_assess::947f776320e3b609": {
  "subjects": [],
  "subject_text": "페드로의 자동차 내부\n운전석과 조수석이 나란한 자동차 앞좌석 공간. 운전대와 대시보드, 노트북과 구겨진 도면이 있으며 창으로 낮빛이 들어온다.",
  "identity": "canonical",
  "scope_id": "L150",
  "scope_role": "location_interior",
  "scope_sha": "5d50a17e313ca4db"
 },
 "S4sh1": {
  "input_fingerprint": "b30a218b3b5f14a0",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 정차해 있는 차 안, 구겨진 도면을 가리키는 페드로의 손과 이를 미간을 찌푸린 채 내려다보는 현우의 모습.\n\nLOCATION (lock): Inside the front seating compartment of a car parked in an affluent residential neighborhood, with daylight entering through the windows. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Crumpled drawing (Crumpled and poorly drawn, being examined and indicated) — The marked face tilts upward toward the occupants and remains visible to the camera from above; used as Shared lower-center attention anchor linking the hand and 현우's reaction; Stationary car interior (Occupied by 페드로 in the driving position and 현우 beside him) — Viewed diagonally forward from behind the driving position; used as Maintains the compressed relationship between the two occupants.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light inside the car provides clear facial and paper detail with understated contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The car is parked in the upscale residential district, and the house plan is crumpled and badly drawn. 페드로: Pedro occupies the driver's seat and wears an earpiece. 현우: Hyunwoo is seated in the passenger position, holding and examining the crumpled plan.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera is positioned in the rear left of the car's interior, looking diagonally forward and right from behind the driver's seat. In the left foreground, the back of Pedro's head and shoulder are visible as he occupies the driver's seat. To his right, in the center midground, sits the center console where Pedro's hand reaches toward a crumpled floor plan. In the right midground, Hyunwoo occupies the passenger seat, angled slightly to look down at the console. The steering wheel is located in front of Pedro on the left. High in the upper center background, the rearview mirror faces rearward, positioned to physically reflect the rear of the cabin and the occupants' faces from the camera's perspective.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 정차해 있는 차 안, 구겨진 도면을 가리키는 페드로의 손과 이를 미간을 찌푸린 채 내려다보는 현우의 모습.\n\nLOCATION (lock): Inside the front seating compartment of a car parked in an affluent residential neighborhood, with daylight entering through the windows. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light inside the car provides clear facial and paper detail with understated contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The car is parked in the upscale residential district, and the house plan is crumpled and badly drawn. 페드로: Pedro occupies the driver's seat and wears an earpiece. 현우: Hyunwoo is seated in the passenger position, holding and examining the crumpled plan.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera is positioned in the rear left of the car's interior, looking diagonally forward and right from behind the driver's seat. In the left foreground, the back of Pedro's head and shoulder are visible as he occupies the driver's seat. To his right, in the center midground, sits the center console where Pedro's hand reaches toward a crumpled floor plan. In the right midground, Hyunwoo occupies the passenger seat, angled slightly to look down at the console. The steering wheel is located in front of Pedro on the left. High in the upper center background, the rearview mirror faces rearward, positioned to physically reflect the rear of the cabin and the occupants' faces from the camera's perspective.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 정차해 있는 차 안, 구겨진 도면을 가리키는 페드로의 손과 이를 미간을 찌푸린 채 내려다보는 현우의 모습.\n\nLOCATION (lock): Inside the front seating compartment of a car parked in an affluent residential neighborhood, with daylight entering through the windows. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light inside the car provides clear facial and paper detail with understated contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The car is parked in the upscale residential district, and the house plan is crumpled and badly drawn. 페드로: Pedro occupies the driver's seat and wears an earpiece. 현우: Hyunwoo is seated in the passenger position, holding and examining the crumpled plan.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S4sh1_confinedfp.png",
     "asset_id": null,
     "role": null
    },
    {
     "label": "현우",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "페드로",
     "path": "<bytes:1278830>",
     "asset_id": "b09df655-64d4-4db4-a1b3-2f0bb5d29c95",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S4sh1_confinedfp.png",
     "asset_id": null,
     "role": null
    },
    {
     "label": "현우",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "페드로",
     "path": "<bytes:1278830>",
     "asset_id": "b09df655-64d4-4db4-a1b3-2f0bb5d29c95",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "페드로의 손과 정체불명의 손이 도면을 가리키며, 현우는 도면을 응시함.",
    "built_space": "차 안 앞좌석. 룸미러에 카메라 위치상 불가능한 각도로 엉뚱한 인물들이 반사됨.",
    "entities": "페드로가 참조보다 훨씬 나이 들어 보임. 룸미러 속 인물은 프롬프트에 없는 발명된 사람임.",
    "hard_violations": [
     "[gemini-pro] 룸미러에 프롬프트에 없는 두 인물이 반사됨",
     "[gemini-pro] 불가능한 각도의 룸미러 반사",
     "[gemini-pro] 하단에 연결된 신체가 없는 카키색 소매의 여분 손 존재",
     "[gpt-high] 룸미러에 기준 인물과 대응하지 않는 성인 남성과 긴 머리 인물이 나타나, 허용되지 않은 추가 인물을 도입한다."
    ],
    "physics": "도면을 가리키는 하단의 카키색 소매 손은 지지하는 신체 없이 허공에 나타난 불가능한 구조임."
   },
   {
    "label": "B",
    "direction": "페드로의 오른손이 도면을 향하고, 현우는 미간을 찌푸린 채 이를 내려다봄.",
    "built_space": "차 안 앞좌석. 룸미러에 운전자의 얼굴이 반사되며 광학적으로 자연스러움.",
    "entities": "두 인물 모두 18세 참조 이미지와 정확히 일치하며 페드로의 이어피스도 구현됨.",
    "hard_violations": [
     "[gpt-high] 운전대에 명확한 현대 로고가 노출되어 로고 금지 조건을 위반한다.",
     "[gpt-high] 선바이저에 읽을 수 있는 영문 경고 표제가 노출되어 가독성 있는 글자 금지 조건을 위반한다."
    ],
    "physics": "현우의 양손이 도면을 안정적으로 쥐고 있으며 물체와 신체의 지지 구조가 자연스러움."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "룸미러에 엉뚱한 인물이 반사되고 출처를 알 수 없는 세 번째 손이 등장하는 등 치명적인 물리적 오류가 있습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "인물들의 외모, 이어피스, 차량 내부 구도 및 동작을 물리적 오류 없이 지시문과 참조에 맞게 구현했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "페드로의 손과 정체불명의 손이 도면을 가리키며, 현우는 도면을 응시함.",
        "built_space": "차 안 앞좌석. 룸미러에 카메라 위치상 불가능한 각도로 엉뚱한 인물들이 반사됨.",
        "entities": "페드로가 참조보다 훨씬 나이 들어 보임. 룸미러 속 인물은 프롬프트에 없는 발명된 사람임.",
        "hard_violations": [
         "룸미러에 프롬프트에 없는 두 인물이 반사됨",
         "불가능한 각도의 룸미러 반사",
         "하단에 연결된 신체가 없는 카키색 소매의 여분 손 존재"
        ],
        "physics": "도면을 가리키는 하단의 카키색 소매 손은 지지하는 신체 없이 허공에 나타난 불가능한 구조임."
       },
       {
        "label": "B",
        "direction": "페드로의 오른손이 도면을 향하고, 현우는 미간을 찌푸린 채 이를 내려다봄.",
        "built_space": "차 안 앞좌석. 룸미러에 운전자의 얼굴이 반사되며 광학적으로 자연스러움.",
        "entities": "두 인물 모두 18세 참조 이미지와 정확히 일치하며 페드로의 이어피스도 구현됨.",
        "hard_violations": [],
        "physics": "현우의 양손이 도면을 안정적으로 쥐고 있으며 물체와 신체의 지지 구조가 자연스러움."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "룸미러에 엉뚱한 인물이 반사되고 출처를 알 수 없는 세 번째 손이 등장하는 등 치명적인 물리적 오류가 있습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "인물들의 외모, 이어피스, 차량 내부 구도 및 동작을 물리적 오류 없이 지시문과 참조에 맞게 구현했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "페드로의 손과 정체불명의 손이 도면을 가리키며, 현우는 도면을 응시함.",
        "built_space": "차 안 앞좌석. 룸미러에 카메라 위치상 불가능한 각도로 엉뚱한 인물들이 반사됨.",
        "entities": "페드로가 참조보다 훨씬 나이 들어 보임. 룸미러 속 인물은 프롬프트에 없는 발명된 사람임.",
        "hard_violations": [
         "룸미러에 프롬프트에 없는 두 인물이 반사됨",
         "불가능한 각도의 룸미러 반사",
         "하단에 연결된 신체가 없는 카키색 소매의 여분 손 존재"
        ],
        "physics": "도면을 가리키는 하단의 카키색 소매 손은 지지하는 신체 없이 허공에 나타난 불가능한 구조임."
       },
       {
        "label": "B",
        "direction": "페드로의 오른손이 도면을 향하고, 현우는 미간을 찌푸린 채 이를 내려다봄.",
        "built_space": "차 안 앞좌석. 룸미러에 운전자의 얼굴이 반사되며 광학적으로 자연스러움.",
        "entities": "두 인물 모두 18세 참조 이미지와 정확히 일치하며 페드로의 이어피스도 구현됨.",
        "hard_violations": [],
        "physics": "현우의 양손이 도면을 안정적으로 쥐고 있으며 물체와 신체의 지지 구조가 자연스러움."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "도면을 가리키는 손과 미간을 찌푸린 현우의 반응, 좌석 배치는 더 충실하지만, 명확한 자동차 로고와 읽히는 경고 문구가 금지 조건을 위반한다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "가리키기와 도면을 내려다보는 동작은 맞지만, 룸미러에 허용되지 않은 인물들이 나타나고 페드로도 기준보다 상당히 나이 들어 보인다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "페드로의 검지 끝은 하단 중앙 도면의 방 구획을 가리킨다. 현우는 미간을 찌푸리고 같은 종이를 내려다보며, 페드로의 고개도 종이 쪽으로 숙여져 있다. 도면의 표시 면은 위쪽과 카메라 쪽으로 기울어져 있고, 몸을 안쪽으로 돌린 두 사람이 함께 볼 수 있는 위치다.",
        "built_space": "앞좌석 등받이 두 개, 왼쪽 운전대 한 개, 중앙 대시보드와 콘솔 한 벌, 룸미러 한 개, 선바이저 두 개가 보인다. 페드로는 운전대 뒤 왼쪽 좌석, 현우는 오른쪽 조수석에 있다. 운전석 뒤에서 전방을 보는 미디엄 구도이며, 지정된 사선보다 다소 중앙에 가깝다. 룸미러 속 한 얼굴은 페드로로 읽히며, 뒤쪽 카메라와 운전자 사이의 반사 경로로 성립할 수 있다. 창밖에는 낮의 고급 주택가가 보인다.",
        "entities": "직접 보이는 인물은 두 명이다. 현우의 앳된 동아시아계 남성 외모, 헝클어진 검은 머리와 남색 티셔츠는 기준에 가깝다. 페드로는 얼굴이 부분적으로만 보이지만 짙은 머리, 젊은 남성의 체격, 남색 티셔츠와 이어피스가 확인된다. 종이는 심하게 구겨져 있고 방 배치가 거칠게 그려져 있어 요청한 도면에 부합한다. 다만 운전대에 현대 로고가 선명하고 선바이저 경고 표제도 읽힌다.",
        "hard_violations": [
         "운전대에 명확한 현대 로고가 노출되어 로고 금지 조건을 위반한다.",
         "선바이저에 읽을 수 있는 영문 경고 표제가 노출되어 가독성 있는 글자 금지 조건을 위반한다."
        ],
        "physics": "두 사람의 몸은 각각 앞좌석에 지지되어 있다. 현우의 손이 종이 오른쪽 가장자리와 아래쪽을 잡아 지탱하고, 페드로의 팔과 손가락은 몸에서 자연스럽게 이어져 종이를 가리킨다. 종이의 접힘과 처짐도 손으로 든 상태에 맞으며, 지지 없이 떠 있는 물체나 신체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "페드로의 뻗은 검지는 도면 중앙 왼쪽 구획을 향하고, 아래쪽의 다른 검지도 인접한 지점을 가리킨다. 현우의 시선은 종이에 내려가 있지만 미간을 찌푸린 반응은 A보다 약하다. 도면은 상당히 세워져 카메라에 넓게 노출되어 있어, 탑승자들을 향해 위로 기울어진 배치보다는 카메라에 보여주는 성격이 강하다.",
        "built_space": "앞좌석 등받이 두 개, 왼쪽 운전대 일부, 중앙 대시보드와 콘솔, 룸미러 한 개, 선바이저 두 개와 오른쪽 사이드미러 한 개가 보인다. 두 주인공은 운전석과 조수석을 각각 차지하며, 운전석 뒤에서 대각선으로 보는 미디엄 구도는 대체로 맞는다. 그러나 룸미러에는 직접 보이는 두 사람과 대응하지 않는 성인 남성과 긴 머리 인물의 얼굴이 함께 나타난다. 창밖은 낮의 넓은 잔디와 고급 주택가다.",
        "entities": "현우는 젊은 동아시아계 남성, 검은 머리, 남색 티셔츠라는 기준을 대체로 따른다. 페드로는 이어피스와 남색 티셔츠를 착용했지만 수염과 얼굴 윤곽 때문에 앳된 18세 기준보다 나이 들어 보인다. 구겨진 집 도면은 있으나 선과 구획이 비교적 정돈되어 있어 서투르게 그린 도면이라는 특성은 약하다. 룸미러 속 두 얼굴은 허용된 인물 구성과 맞지 않는다.",
        "hard_violations": [
         "룸미러에 기준 인물과 대응하지 않는 성인 남성과 긴 머리 인물이 나타나, 허용되지 않은 추가 인물을 도입한다."
        ],
        "physics": "두 사람은 각자의 좌석에 앉아 있고 현우의 손이 종이 오른쪽 위 가장자리를 잡고 있다. 종이는 그 손에 지지되며, 가리키는 손들도 팔에 연결되어 있다. 상체를 안쪽으로 돌린 자세는 정차 중 가능한 동작이다. 지지 없이 떠 있는 신체나 소품은 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "도면을 가리키는 손과 미간을 찌푸린 현우의 반응, 좌석 배치는 더 충실하지만, 명확한 자동차 로고와 읽히는 경고 문구가 금지 조건을 위반한다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "가리키기와 도면을 내려다보는 동작은 맞지만, 룸미러에 허용되지 않은 인물들이 나타나고 페드로도 기준보다 상당히 나이 들어 보인다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "페드로의 검지 끝은 하단 중앙 도면의 방 구획을 가리킨다. 현우는 미간을 찌푸리고 같은 종이를 내려다보며, 페드로의 고개도 종이 쪽으로 숙여져 있다. 도면의 표시 면은 위쪽과 카메라 쪽으로 기울어져 있고, 몸을 안쪽으로 돌린 두 사람이 함께 볼 수 있는 위치다.",
        "built_space": "앞좌석 등받이 두 개, 왼쪽 운전대 한 개, 중앙 대시보드와 콘솔 한 벌, 룸미러 한 개, 선바이저 두 개가 보인다. 페드로는 운전대 뒤 왼쪽 좌석, 현우는 오른쪽 조수석에 있다. 운전석 뒤에서 전방을 보는 미디엄 구도이며, 지정된 사선보다 다소 중앙에 가깝다. 룸미러 속 한 얼굴은 페드로로 읽히며, 뒤쪽 카메라와 운전자 사이의 반사 경로로 성립할 수 있다. 창밖에는 낮의 고급 주택가가 보인다.",
        "entities": "직접 보이는 인물은 두 명이다. 현우의 앳된 동아시아계 남성 외모, 헝클어진 검은 머리와 남색 티셔츠는 기준에 가깝다. 페드로는 얼굴이 부분적으로만 보이지만 짙은 머리, 젊은 남성의 체격, 남색 티셔츠와 이어피스가 확인된다. 종이는 심하게 구겨져 있고 방 배치가 거칠게 그려져 있어 요청한 도면에 부합한다. 다만 운전대에 현대 로고가 선명하고 선바이저 경고 표제도 읽힌다.",
        "hard_violations": [
         "운전대에 명확한 현대 로고가 노출되어 로고 금지 조건을 위반한다.",
         "선바이저에 읽을 수 있는 영문 경고 표제가 노출되어 가독성 있는 글자 금지 조건을 위반한다."
        ],
        "physics": "두 사람의 몸은 각각 앞좌석에 지지되어 있다. 현우의 손이 종이 오른쪽 가장자리와 아래쪽을 잡아 지탱하고, 페드로의 팔과 손가락은 몸에서 자연스럽게 이어져 종이를 가리킨다. 종이의 접힘과 처짐도 손으로 든 상태에 맞으며, 지지 없이 떠 있는 물체나 신체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "페드로의 뻗은 검지는 도면 중앙 왼쪽 구획을 향하고, 아래쪽의 다른 검지도 인접한 지점을 가리킨다. 현우의 시선은 종이에 내려가 있지만 미간을 찌푸린 반응은 A보다 약하다. 도면은 상당히 세워져 카메라에 넓게 노출되어 있어, 탑승자들을 향해 위로 기울어진 배치보다는 카메라에 보여주는 성격이 강하다.",
        "built_space": "앞좌석 등받이 두 개, 왼쪽 운전대 일부, 중앙 대시보드와 콘솔, 룸미러 한 개, 선바이저 두 개와 오른쪽 사이드미러 한 개가 보인다. 두 주인공은 운전석과 조수석을 각각 차지하며, 운전석 뒤에서 대각선으로 보는 미디엄 구도는 대체로 맞는다. 그러나 룸미러에는 직접 보이는 두 사람과 대응하지 않는 성인 남성과 긴 머리 인물의 얼굴이 함께 나타난다. 창밖은 낮의 넓은 잔디와 고급 주택가다.",
        "entities": "현우는 젊은 동아시아계 남성, 검은 머리, 남색 티셔츠라는 기준을 대체로 따른다. 페드로는 이어피스와 남색 티셔츠를 착용했지만 수염과 얼굴 윤곽 때문에 앳된 18세 기준보다 나이 들어 보인다. 구겨진 집 도면은 있으나 선과 구획이 비교적 정돈되어 있어 서투르게 그린 도면이라는 특성은 약하다. 룸미러 속 두 얼굴은 허용된 인물 구성과 맞지 않는다.",
        "hard_violations": [
         "룸미러에 기준 인물과 대응하지 않는 성인 남성과 긴 머리 인물이 나타나, 허용되지 않은 추가 인물을 도입한다."
        ],
        "physics": "두 사람은 각자의 좌석에 앉아 있고 현우의 손이 종이 오른쪽 위 가장자리를 잡고 있다. 종이는 그 손에 지지되며, 가리키는 손들도 팔에 연결되어 있다. 상체를 안쪽으로 돌린 자세는 정차 중 가능한 동작이다. 지지 없이 떠 있는 신체나 소품은 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.829,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.579,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 룸미러에 프롬프트에 없는 두 인물이 반사됨",
     "[gemini-pro] 불가능한 각도의 룸미러 반사",
     "[gemini-pro] 하단에 연결된 신체가 없는 카키색 소매의 여분 손 존재",
     "[gpt-high] 룸미러에 기준 인물과 대응하지 않는 성인 남성과 긴 머리 인물이 나타나, 허용되지 않은 추가 인물을 도입한다."
    ],
    "B": [
     "[gpt-high] 운전대에 명확한 현대 로고가 노출되어 로고 금지 조건을 위반한다.",
     "[gpt-high] 선바이저에 읽을 수 있는 영문 경고 표제가 노출되어 가독성 있는 글자 금지 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "A": 579,
   "B": 1750
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 579,
    "verdict_ko": "룸미러에 엉뚱한 인물이 반사되고 출처를 알 수 없는 세 번째 손이 등장하는 등 치명적인 물리적 오류가 있습니다.  ★위반: [gemini-pro] 룸미러에 프롬프트에 없는 두 인물이 반사됨 / [gemini-pro] 불가능한 각도의 룸미러 반사 / [gemini-pro] 하단에 연결된 신체가 없는 카키색 소매의 여분 손 존재 / [gpt-high] 룸미러에 기준 인물과 대응하지 않는 성인 남성과 긴 머리 인물이 나타나, 허용되지 않은 추가 인물을 도입한다."
   },
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "인물들의 외모, 이어피스, 차량 내부 구도 및 동작을 물리적 오류 없이 지시문과 참조에 맞게 구현했습니다.  ★위반: [gpt-high] 운전대에 명확한 현대 로고가 노출되어 로고 금지 조건을 위반한다. / [gpt-high] 선바이저에 읽을 수 있는 영문 경고 표제가 노출되어 가독성 있는 글자 금지 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S4sh1_confinedfp.png",
    "asset_id": null,
    "role": null
   },
   {
    "label": "현우",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "페드로",
    "path": "<bytes:1278830>",
    "asset_id": "b09df655-64d4-4db4-a1b3-2f0bb5d29c95",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-93e8-7a82-83b1-474b70f58709",
  "confined_fp": {
   "base_key": "confinedfp::510c33afc573",
   "apt_reason": "차량 내부의 앞좌석이라는 통제 장치가 있는 제한된 공간을 배경으로 하며, 페드로가 운전석에 앉고 현우가 조수석에 앉아 상호작용하는 정확한 위치 관계가 이야기 전달에 필수적이므로 평면도 레이아웃 보조가 필요합니다.",
   "fixed": true,
   "mismatches": []
  },
  "ref_mode": "confined_fp: 도면+장면설명+엔티티",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S4sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T03:52:29.102059+00:00",
  "fingerprint": "493baa6047273b48d43706e5f0406bfaf189bd4b143626bf01b9399f5f33102f",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S4sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S4sh1_sel.png",
  "source_sha256": "a2472b6c70eca7a8857d1b4a520af32b0d06e9be60fe3194b32ca0fb67957503",
  "file": "S4sh1_cine.png",
  "staged_sha256": "ea0e53cd1c45824a42063ccefbca60683a62fc0288a0d9b60edbc19278b6e7e2",
  "latency_ms": 10798
 },
 "S5sh1::signage": {
  "fp": "8d5506ac21c64cce",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::683a2594bb638b15": {
  "subjects": [],
  "subject_text": "고급주택 거실\n큰 통유리창으로 외부가 보이는 고급주택 거실. 열리는 창과 현관 쪽 동선, 2층으로 올라가는 계단 입구가 연결된다.",
  "identity": "canonical",
  "scope_id": "L152",
  "scope_role": "location_interior",
  "scope_sha": "468a1a21e5d74ad9"
 },
 "S5sh1::bgfirst_bg": {
  "input_fingerprint": "35c60ca0064c7215",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 고급주택 거실 통유리창 너머로 먼지를 터는 젊은 여자를 응시하며 현관문 앞에 바짝 붙어선 현우의 모습.\n\nLOCATION (lock): Outside the front door of an upscale house, beside the large living-room window through which the occupant is visible.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Living-room window with the dusting woman visible beyond in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Living-room window (Closed at this moment, with the room and dusting woman visible through it) — Seen obliquely from outside; the woman is visible beyond the pane; used as Separates the concealed observer from the unaware interior figure; Entrance door (Not yet opened, with an old-style lock) — Its exterior side is beside 현우 in the left foreground; used as Explains his compressed posture and establishes the entrance-to-window route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Balanced daytime ambient light preserves visibility through the window without introducing a dominant reflection or stylized interior glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 고급주택 거실 통유리창 너머로 먼지를 터는 젊은 여자를 응시하며 현관문 앞에 바짝 붙어선 현우의 모습.\n\nLOCATION (lock): Outside the front door of an upscale house, beside the large living-room window through which the occupant is visible.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Living-room window with the dusting woman visible beyond in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Living-room window (Closed at this moment, with the room and dusting woman visible through it) — Seen obliquely from outside; the woman is visible beyond the pane; used as Separates the concealed observer from the unaware interior figure; Entrance door (Not yet opened, with an old-style lock) — Its exterior side is beside 현우 in the left foreground; used as Explains his compressed posture and establishes the entrance-to-window route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Balanced daytime ambient light preserves visibility through the window without introducing a dominant reflection or stylized interior glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S5sh1__bgfirst_bg.png",
  "asset_id": "33fddc9a-f643-4e6e-9f6c-5eeb879142cd",
  "input_asset_ids": [
   "dde54bdc-8795-4460-ac4c-7a559b6d6b14",
   "1ced485f-4388-40e8-9db1-c1e4432bbd3c"
  ]
 },
 "S5sh1": {
  "input_fingerprint": "d63ea52509a078cc",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 고급주택 거실 통유리창 너머로 먼지를 터는 젊은 여자를 응시하며 현관문 앞에 바짝 붙어선 현우의 모습.\n\nLOCATION (lock): Outside the front door of an upscale house, beside the large living-room window through which the occupant is visible. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Living-room window with the dusting woman visible beyond in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Living-room window (Closed at this moment, with the room and dusting woman visible through it) — Seen obliquely from outside; the woman is visible beyond the pane; used as Separates the concealed observer from the unaware interior figure; Entrance door (Not yet opened, with an old-style lock) — Its exterior side is beside 현우 in the left foreground; used as Explains his compressed posture and establishes the entrance-to-window route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Balanced daytime ambient light preserves visibility through the window without introducing a dominant reflection or stylized interior glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The house has a full-height living-room window and an old-fashioned keyed front-door lock. The young female housekeeping robot is intact at this point. 현우: Hyunwoo has crossed the boundary wall and is keeping close to the front entrance.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 젊은 아시아계 여자 로봇 (젊은 여성형 얼굴, 한국인 외모, 사람과 같은 얼굴 표면, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 고급주택 거실 통유리창 너머로 먼지를 터는 젊은 여자를 응시하며 현관문 앞에 바짝 붙어선 현우의 모습.\n\nLOCATION (lock): Outside the front door of an upscale house, beside the large living-room window through which the occupant is visible. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Living-room window with the dusting woman visible beyond in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Living-room window (Closed at this moment, with the room and dusting woman visible through it) — Seen obliquely from outside; the woman is visible beyond the pane; used as Separates the concealed observer from the unaware interior figure; Entrance door (Not yet opened, with an old-style lock) — Its exterior side is beside 현우 in the left foreground; used as Explains his compressed posture and establishes the entrance-to-window route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Balanced daytime ambient light preserves visibility through the window without introducing a dominant reflection or stylized interior glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The house has a full-height living-room window and an old-fashioned keyed front-door lock. The young female housekeeping robot is intact at this point. 현우: Hyunwoo has crossed the boundary wall and is keeping close to the front entrance.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 젊은 아시아계 여자 로봇 (젊은 여성형 얼굴, 한국인 외모, 사람과 같은 얼굴 표면, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 고급주택 거실 통유리창 너머로 먼지를 터는 젊은 여자를 응시하며 현관문 앞에 바짝 붙어선 현우의 모습.\n\nLOCATION (lock): Outside the front door of an upscale house, beside the large living-room window through which the occupant is visible. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Living-room window with the dusting woman visible beyond in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Living-room window (Closed at this moment, with the room and dusting woman visible through it) — Seen obliquely from outside; the woman is visible beyond the pane; used as Separates the concealed observer from the unaware interior figure; Entrance door (Not yet opened, with an old-style lock) — Its exterior side is beside 현우 in the left foreground; used as Explains his compressed posture and establishes the entrance-to-window route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Balanced daytime ambient light preserves visibility through the window without introducing a dominant reflection or stylized interior glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The house has a full-height living-room window and an old-fashioned keyed front-door lock. The young female housekeeping robot is intact at this point. 현우: Hyunwoo has crossed the boundary wall and is keeping close to the front entrance.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 젊은 아시아계 여자 로봇 (젊은 여성형 얼굴, 한국인 외모, 사람과 같은 얼굴 표면, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S5sh1__bgfirst_bg.png",
     "asset_id": "33fddc9a-f643-4e6e-9f6c-5eeb879142cd",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S5sh1.png",
     "asset_id": "dde54bdc-8795-4460-ac4c-7a559b6d6b14",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 젊은 아시아계 여자 로봇: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:922551>",
     "asset_id": "99a76981-3435-4649-9a90-459804a78b99",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L152B03.png",
     "asset_id": "1ced485f-4388-40e8-9db1-c1e4432bbd3c",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 젊은 아시아계 여자 로봇: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:922551>",
     "asset_id": "99a76981-3435-4649-9a90-459804a78b99",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "남자는 유리창 안쪽을 향해 고개를 돌리고 있고, 로봇은 먼지떨이를 아래 가구 쪽으로 향하고 있음.",
    "built_space": "왼쪽의 어두운 현관문과 오른쪽의 거실 통유리창이 있는 주택 외부 구조가 레퍼런스와 일치함. 남자는 문에서 떨어진 통로 한가운데 서 있음.",
    "entities": "남자는 검은 머리를 가졌으나 뒷모습만 보이며 하늘색 긴팔 옷을 입음. 창문 너머 로봇은 검은 머리와 기계 팔다리를 가짐.",
    "hard_violations": [
     "[gpt-high] 현우가 현관문에 밀착해 숨은 위치가 아니라 현관 계단 앞 열린 보행 공간에 떨어져 서 있다."
    ],
    "physics": "남자는 야외 통로 바닥에 두 발을 딛고 안정적으로 서 있으며, 로봇도 실내 바닥에 서서 체중을 지탱하고 있음."
   },
   {
    "label": "B",
    "direction": "남자는 웅크린 채 유리창 너머 실내의 로봇을 주시하고 있으며, 로봇은 먼지떨이를 배경의 선반 쪽으로 향하고 있음.",
    "built_space": "레퍼런스와 동일한 현관문과 통유리창 구조이며, 남자는 현관문 옆 벽면에 바짝 붙어 올바른 위치에 자리잡고 있음.",
    "entities": "남자는 레퍼런스와 일치하는 앳된 얼굴, 헝클어진 머리, 남색 반팔 티셔츠를 착용함. 로봇은 검은 머리와 기계 신체 구조를 보여줌.",
    "hard_violations": [
     "[gpt-high] 현우가 현관문 앞에 바짝 붙어 있어야 하는데, 실제로는 문에서 떨어진 창 옆 벽 끝에서 몸을 내밀고 있다."
    ],
    "physics": "남자는 무릎을 굽히고 다리에 손을 얹어 웅크린 자세로 체중을 자연스럽게 지탱하고 있으며, 로봇은 실내 바닥에 딛고 서 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "현우가 문에 바짝 붙어 몸을 웅크린 자세가 아니라 통로 한가운데 멀리 서 있으며, 얼굴이 보이지 않아 핵심 연출 지시에서 벗어났습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "현우가 현관문 옆 벽면에 바짝 붙어 웅크린 자세와 시선, 그리고 인물 레퍼런스의 외모와 의상까지 지시문을 매우 충실하게 구현했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "남자는 유리창 안쪽을 향해 고개를 돌리고 있고, 로봇은 먼지떨이를 아래 가구 쪽으로 향하고 있음.",
        "built_space": "왼쪽의 어두운 현관문과 오른쪽의 거실 통유리창이 있는 주택 외부 구조가 레퍼런스와 일치함. 남자는 문에서 떨어진 통로 한가운데 서 있음.",
        "entities": "남자는 검은 머리를 가졌으나 뒷모습만 보이며 하늘색 긴팔 옷을 입음. 창문 너머 로봇은 검은 머리와 기계 팔다리를 가짐.",
        "hard_violations": [],
        "physics": "남자는 야외 통로 바닥에 두 발을 딛고 안정적으로 서 있으며, 로봇도 실내 바닥에 서서 체중을 지탱하고 있음."
       },
       {
        "label": "B",
        "direction": "남자는 웅크린 채 유리창 너머 실내의 로봇을 주시하고 있으며, 로봇은 먼지떨이를 배경의 선반 쪽으로 향하고 있음.",
        "built_space": "레퍼런스와 동일한 현관문과 통유리창 구조이며, 남자는 현관문 옆 벽면에 바짝 붙어 올바른 위치에 자리잡고 있음.",
        "entities": "남자는 레퍼런스와 일치하는 앳된 얼굴, 헝클어진 머리, 남색 반팔 티셔츠를 착용함. 로봇은 검은 머리와 기계 신체 구조를 보여줌.",
        "hard_violations": [],
        "physics": "남자는 무릎을 굽히고 다리에 손을 얹어 웅크린 자세로 체중을 자연스럽게 지탱하고 있으며, 로봇은 실내 바닥에 딛고 서 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "현우가 문에 바짝 붙어 몸을 웅크린 자세가 아니라 통로 한가운데 멀리 서 있으며, 얼굴이 보이지 않아 핵심 연출 지시에서 벗어났습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "현우가 현관문 옆 벽면에 바짝 붙어 웅크린 자세와 시선, 그리고 인물 레퍼런스의 외모와 의상까지 지시문을 매우 충실하게 구현했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "남자는 유리창 안쪽을 향해 고개를 돌리고 있고, 로봇은 먼지떨이를 아래 가구 쪽으로 향하고 있음.",
        "built_space": "왼쪽의 어두운 현관문과 오른쪽의 거실 통유리창이 있는 주택 외부 구조가 레퍼런스와 일치함. 남자는 문에서 떨어진 통로 한가운데 서 있음.",
        "entities": "남자는 검은 머리를 가졌으나 뒷모습만 보이며 하늘색 긴팔 옷을 입음. 창문 너머 로봇은 검은 머리와 기계 팔다리를 가짐.",
        "hard_violations": [],
        "physics": "남자는 야외 통로 바닥에 두 발을 딛고 안정적으로 서 있으며, 로봇도 실내 바닥에 서서 체중을 지탱하고 있음."
       },
       {
        "label": "B",
        "direction": "남자는 웅크린 채 유리창 너머 실내의 로봇을 주시하고 있으며, 로봇은 먼지떨이를 배경의 선반 쪽으로 향하고 있음.",
        "built_space": "레퍼런스와 동일한 현관문과 통유리창 구조이며, 남자는 현관문 옆 벽면에 바짝 붙어 올바른 위치에 자리잡고 있음.",
        "entities": "남자는 레퍼런스와 일치하는 앳된 얼굴, 헝클어진 머리, 남색 반팔 티셔츠를 착용함. 로봇은 검은 머리와 기계 신체 구조를 보여줌.",
        "hard_violations": [],
        "physics": "남자는 무릎을 굽히고 다리에 손을 얹어 웅크린 자세로 체중을 자연스럽게 지탱하고 있으며, 로봇은 실내 바닥에 딛고 서 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "넓은 구도와 몸을 낮춘 관찰 자세, 인물 의상은 더 충실하지만, 현우가 현관문에 바짝 붙지 않고 창가 쪽으로 이동해 있다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "주택 구조와 청소 동작은 명확하지만, 현우가 현관에서 떨어진 전경에 크게 배치되어 핵심 은신 위치와 인물 크기를 어기며 두 인물의 의상도 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 고개를 오른쪽으로 내밀어 창 너머 여성 쪽을 보고 있다. 여성은 현우를 돌아보지 않고 왼쪽 위로 들어 올린 먼지떨이와 선반 쪽을 향한다. 먼지떨이는 여성의 손에서 선반 윗부분을 향해 뻗어 있어 청소 대상이 읽힌다.",
        "built_space": "왼쪽에 닫힌 현관문 하나와 세로형 잠금장치 하나, 위쪽에 매입등 하나가 보인다. 오른쪽에는 중앙 세로 프레임으로 나뉜 전면 높이의 큰 창과 흰 커튼이 있으며, 실내에는 소파와 낮은 탁자, 책장이 보인다. 회색 외벽과 검은 창틀은 참조와 유사하지만 실내 가구 배치는 달라졌다. 현우는 현관문 바로 앞이 아니라 문과 창 사이 벽의 창 쪽 끝에 있다. 창은 다소 정면에 가깝고, 나무와 맞은편 건물의 반사는 외부 촬영에서 가능한 방향이다.",
        "entities": "인물은 현우와 젊은 여성형 로봇 두 명뿐이다. 현우의 앳된 동아시아계 외모, 헝클어진 검은 머리와 남색 반소매 상의는 참조에 가깝다. 여성은 검은 머리와 사람 같은 얼굴, 검정·은색 몸체를 갖추어 참조의 로봇으로 읽히지만 작은 얼굴의 정확한 일치까지 확인하기는 어렵다. 먼지떨이는 보이며 몸체 손상은 보이지 않는다. 문 잠금장치는 현대적인 세로형 장치로 보여 구식 열쇠식이라는 조건은 명확하지 않다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "현우가 현관문 앞에 바짝 붙어 있어야 하는데, 실제로는 문에서 떨어진 창 옆 벽 끝에서 몸을 내밀고 있다."
        ],
        "physics": "현우는 무릎과 허리를 굽힌 자세이며 하퇴 아래는 관목에 가려져 있다. 발 접지는 직접 확인되지 않지만 지면 위에 웅크린 자세로 연결되며 공중에 떠 있다는 증거는 없다. 여성의 하체는 실내 가구와 반사에 일부 가려져 있고, 먼지떨이는 뻗은 손에 잡혀 있다. 팔을 들어 선반 먼지를 터는 동작은 가능한 자세다."
       },
       {
        "label": "B",
        "direction": "현우의 뒤통수와 옆얼굴이 오른쪽 창 속 여성을 향한다. 여성은 아래쪽 탁자와 먼지떨이를 내려다보고 있어 외부 관찰자를 모르는 상태로 읽힌다. 손에 든 먼지떨이는 오른쪽 아래 탁자 쪽으로 향한다.",
        "built_space": "왼쪽에 닫힌 현관문 하나, 현관 단차와 상부 매입등 하나가 보이고 잠금장치 부분은 현우에게 가려져 있다. 오른쪽에는 중앙 세로 프레임이 있는 큰 닫힌 창과 양옆 흰 커튼이 있다. 실내의 소파 하나, 낮은 탁자, 텔레비전 하나와 긴 목재 장식장은 위치 참조에 가깝다. 그러나 현우는 현관 안쪽이 아니라 계단 앞 외부 보행 공간에 서 있다. 창을 비스듬히 보는 방향은 맞지만 나무와 건물 반사가 상당히 강하다. 그 반사 자체는 외부 카메라 위치에서 가능하다.",
        "entities": "현우와 여성형 로봇 두 명만 보인다. 현우의 검은 머리와 젊은 동아시아계 인상은 부합하지만 얼굴 대부분이 숨겨져 정확한 동일성은 판단하기 어렵다. 연한 청회색 긴소매 상의는 참조의 남색 반소매와 다르다. 여성은 검은 머리, 사람 같은 젊은 동아시아계 얼굴과 온전한 기계 팔다리를 갖췄지만, 밝은 반소매 상의와 짧은 하의는 참조의 검정 몸체 의상과 다르다. 손에 먼지떨이가 있고 구식 열쇠식 잠금장치는 확인되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "현우가 현관문에 밀착해 숨은 위치가 아니라 현관 계단 앞 열린 보행 공간에 떨어져 서 있다."
        ],
        "physics": "현우는 상체를 조금 앞으로 기울였으며 다리는 화면 아래로 이어진다. 발이 잘렸지만 공중에 떠 있는 형태는 아니다. 여성의 신발은 실내 바닥에 놓여 있고, 허리를 숙인 상태에서 손으로 먼지떨이 손잡이를 잡고 있다. 몸과 도구의 지지는 자연스럽고 청소 동작도 물리적으로 가능하다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "넓은 구도와 몸을 낮춘 관찰 자세, 인물 의상은 더 충실하지만, 현우가 현관문에 바짝 붙지 않고 창가 쪽으로 이동해 있다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "주택 구조와 청소 동작은 명확하지만, 현우가 현관에서 떨어진 전경에 크게 배치되어 핵심 은신 위치와 인물 크기를 어기며 두 인물의 의상도 다르다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 고개를 오른쪽으로 내밀어 창 너머 여성 쪽을 보고 있다. 여성은 현우를 돌아보지 않고 왼쪽 위로 들어 올린 먼지떨이와 선반 쪽을 향한다. 먼지떨이는 여성의 손에서 선반 윗부분을 향해 뻗어 있어 청소 대상이 읽힌다.",
        "built_space": "왼쪽에 닫힌 현관문 하나와 세로형 잠금장치 하나, 위쪽에 매입등 하나가 보인다. 오른쪽에는 중앙 세로 프레임으로 나뉜 전면 높이의 큰 창과 흰 커튼이 있으며, 실내에는 소파와 낮은 탁자, 책장이 보인다. 회색 외벽과 검은 창틀은 참조와 유사하지만 실내 가구 배치는 달라졌다. 현우는 현관문 바로 앞이 아니라 문과 창 사이 벽의 창 쪽 끝에 있다. 창은 다소 정면에 가깝고, 나무와 맞은편 건물의 반사는 외부 촬영에서 가능한 방향이다.",
        "entities": "인물은 현우와 젊은 여성형 로봇 두 명뿐이다. 현우의 앳된 동아시아계 외모, 헝클어진 검은 머리와 남색 반소매 상의는 참조에 가깝다. 여성은 검은 머리와 사람 같은 얼굴, 검정·은색 몸체를 갖추어 참조의 로봇으로 읽히지만 작은 얼굴의 정확한 일치까지 확인하기는 어렵다. 먼지떨이는 보이며 몸체 손상은 보이지 않는다. 문 잠금장치는 현대적인 세로형 장치로 보여 구식 열쇠식이라는 조건은 명확하지 않다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "현우가 현관문 앞에 바짝 붙어 있어야 하는데, 실제로는 문에서 떨어진 창 옆 벽 끝에서 몸을 내밀고 있다."
        ],
        "physics": "현우는 무릎과 허리를 굽힌 자세이며 하퇴 아래는 관목에 가려져 있다. 발 접지는 직접 확인되지 않지만 지면 위에 웅크린 자세로 연결되며 공중에 떠 있다는 증거는 없다. 여성의 하체는 실내 가구와 반사에 일부 가려져 있고, 먼지떨이는 뻗은 손에 잡혀 있다. 팔을 들어 선반 먼지를 터는 동작은 가능한 자세다."
       },
       {
        "label": "A",
        "direction": "현우의 뒤통수와 옆얼굴이 오른쪽 창 속 여성을 향한다. 여성은 아래쪽 탁자와 먼지떨이를 내려다보고 있어 외부 관찰자를 모르는 상태로 읽힌다. 손에 든 먼지떨이는 오른쪽 아래 탁자 쪽으로 향한다.",
        "built_space": "왼쪽에 닫힌 현관문 하나, 현관 단차와 상부 매입등 하나가 보이고 잠금장치 부분은 현우에게 가려져 있다. 오른쪽에는 중앙 세로 프레임이 있는 큰 닫힌 창과 양옆 흰 커튼이 있다. 실내의 소파 하나, 낮은 탁자, 텔레비전 하나와 긴 목재 장식장은 위치 참조에 가깝다. 그러나 현우는 현관 안쪽이 아니라 계단 앞 외부 보행 공간에 서 있다. 창을 비스듬히 보는 방향은 맞지만 나무와 건물 반사가 상당히 강하다. 그 반사 자체는 외부 카메라 위치에서 가능하다.",
        "entities": "현우와 여성형 로봇 두 명만 보인다. 현우의 검은 머리와 젊은 동아시아계 인상은 부합하지만 얼굴 대부분이 숨겨져 정확한 동일성은 판단하기 어렵다. 연한 청회색 긴소매 상의는 참조의 남색 반소매와 다르다. 여성은 검은 머리, 사람 같은 젊은 동아시아계 얼굴과 온전한 기계 팔다리를 갖췄지만, 밝은 반소매 상의와 짧은 하의는 참조의 검정 몸체 의상과 다르다. 손에 먼지떨이가 있고 구식 열쇠식 잠금장치는 확인되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "현우가 현관문에 밀착해 숨은 위치가 아니라 현관 계단 앞 열린 보행 공간에 떨어져 서 있다."
        ],
        "physics": "현우는 상체를 조금 앞으로 기울였으며 다리는 화면 아래로 이어진다. 발이 잘렸지만 공중에 떠 있는 형태는 아니다. 여성의 신발은 실내 바닥에 놓여 있고, 허리를 숙인 상태에서 손으로 먼지떨이 손잡이를 잡고 있다. 몸과 도구의 지지는 자연스럽고 청소 동작도 물리적으로 가능하다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.171,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.921,
    "B": 1.75
   },
   "violations": {
    "B": [
     "[gpt-high] 현우가 현관문 앞에 바짝 붙어 있어야 하는데, 실제로는 문에서 떨어진 창 옆 벽 끝에서 몸을 내밀고 있다."
    ],
    "A": [
     "[gpt-high] 현우가 현관문에 밀착해 숨은 위치가 아니라 현관 계단 앞 열린 보행 공간에 떨어져 서 있다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "A": 921,
   "B": 1750
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 921,
    "verdict_ko": "현우가 문에 바짝 붙어 몸을 웅크린 자세가 아니라 통로 한가운데 멀리 서 있으며, 얼굴이 보이지 않아 핵심 연출 지시에서 벗어났습니다.  ★위반: [gpt-high] 현우가 현관문에 밀착해 숨은 위치가 아니라 현관 계단 앞 열린 보행 공간에 떨어져 서 있다."
   },
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "현우가 현관문 옆 벽면에 바짝 붙어 웅크린 자세와 시선, 그리고 인물 레퍼런스의 외모와 의상까지 지시문을 매우 충실하게 구현했습니다.  ★위반: [gpt-high] 현우가 현관문 앞에 바짝 붙어 있어야 하는데, 실제로는 문에서 떨어진 창 옆 벽 끝에서 몸을 내밀고 있다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L152B03.png",
    "asset_id": "1ced485f-4388-40e8-9db1-c1e4432bbd3c",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 젊은 아시아계 여자 로봇: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:922551>",
    "asset_id": "99a76981-3435-4649-9a90-459804a78b99",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-974c-7fd0-87f6-3cb810c29499",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S5sh1__bgfirst_bg.png",
   "bg_asset_id": "33fddc9a-f643-4e6e-9f6c-5eeb879142cd",
   "bg_record_key": "S5sh1::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S5sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T03:54:10.357693+00:00",
  "fingerprint": "26d5110e66e4679945e38f22bf8509093400e86d7dbc884304d5bb7305bd2831",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S5sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S5sh1_sel.png",
  "source_sha256": "d53f32a10ddf4039fbecb542b67ca7ddf45f97e8ec1cc0d6304cd84306fc43e6",
  "file": "S5sh1_cine.png",
  "staged_sha256": "574f42be232916faa1cbcae2ff32ba95b7e59ed1d32e83430ea9b9671319e09c",
  "latency_ms": 10364
 },
 "S5sh9::signage": {
  "fp": "b3c980b8f47dc17b",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S5sh9": {
  "input_fingerprint": "67093e27d473da03",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 셰퍼드에게 물어뜯긴 여자의 머리 가죽 아래로 은빛 금속 재질이 드러난 상태.\n\nLOCATION (lock): On the ground immediately outside the opened living-room window of an upscale house, where the fallen household robot is attacked. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Living-room floor (Supporting the fallen robot woman after impact); used as A narrow contextual plane around the head and shoulder; Open living-room window (Opened by the woman before her collapse) — Only an edge of the opening remains visible behind the low foreground action; used as Maintains continuity with the exterior approach and the next movement into the house.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light makes the exposed silver metal distinct from the torn outer covering without adding a supernatural glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 젊은 아시아계 여자 로봇 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): The disabled female robot is collapsed on the floor after being dropped by Hyunwoo, with her head at floor level as the shepherd tears its covering and exposes the metal underneath. The source does not specify her torso's orientation or the positions of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The living-room window remains open, and the female housekeeping robot lies on the floor with torn head skin exposing metal underneath. The phone used for the call remains at the house, and the German shepherd is biting the robot's head.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 젊은 아시아계 여자 로봇 (젊은 여성형 얼굴, 한국인 외모, 사람과 같은 얼굴 표면, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 셰퍼드에게 물어뜯긴 여자의 머리 가죽 아래로 은빛 금속 재질이 드러난 상태.\n\nLOCATION (lock): On the ground immediately outside the opened living-room window of an upscale house, where the fallen household robot is attacked. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Living-room floor (Supporting the fallen robot woman after impact); used as A narrow contextual plane around the head and shoulder; Open living-room window (Opened by the woman before her collapse) — Only an edge of the opening remains visible behind the low foreground action; used as Maintains continuity with the exterior approach and the next movement into the house.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light makes the exposed silver metal distinct from the torn outer covering without adding a supernatural glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 젊은 아시아계 여자 로봇 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): The disabled female robot is collapsed on the floor after being dropped by Hyunwoo, with her head at floor level as the shepherd tears its covering and exposes the metal underneath. The source does not specify her torso's orientation or the positions of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The living-room window remains open, and the female housekeeping robot lies on the floor with torn head skin exposing metal underneath. The phone used for the call remains at the house, and the German shepherd is biting the robot's head.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 젊은 아시아계 여자 로봇 (젊은 여성형 얼굴, 한국인 외모, 사람과 같은 얼굴 표면, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 셰퍼드에게 물어뜯긴 여자의 머리 가죽 아래로 은빛 금속 재질이 드러난 상태.\n\nLOCATION (lock): On the ground immediately outside the opened living-room window of an upscale house, where the fallen household robot is attacked. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Living-room floor (Supporting the fallen robot woman after impact); used as A narrow contextual plane around the head and shoulder; Open living-room window (Opened by the woman before her collapse) — Only an edge of the opening remains visible behind the low foreground action; used as Maintains continuity with the exterior approach and the next movement into the house.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light makes the exposed silver metal distinct from the torn outer covering without adding a supernatural glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 젊은 아시아계 여자 로봇 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): The disabled female robot is collapsed on the floor after being dropped by Hyunwoo, with her head at floor level as the shepherd tears its covering and exposes the metal underneath. The source does not specify her torso's orientation or the positions of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The living-room window remains open, and the female housekeeping robot lies on the floor with torn head skin exposing metal underneath. The phone used for the call remains at the house, and the German shepherd is biting the robot's head.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 젊은 아시아계 여자 로봇 (젊은 여성형 얼굴, 한국인 외모, 사람과 같은 얼굴 표면, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "셰퍼드의 시선과 입이 바닥에 쓰러진 로봇의 머리를 향해 있으며, 드러난 금속 부위를 물고 있음.",
    "built_space": "실내 거실의 나무 바닥이 배경으로 쓰였으며, 왼쪽 화면 가장자리에 야외로 이어지는 열린 창문 프레임이 올바르게 배치됨.",
    "entities": "셰퍼드, 머리 가죽이 벗겨져 내부의 은빛 금속 골조가 드러난 젊은 아시아계 여자 로봇(검은색과 은색이 섞인 의상)이 명시된 외형에 맞게 등장함.",
    "hard_violations": [
     "[gpt-high] 창 바로 바깥에 있어야 하는 쓰러진 로봇과 공격 중인 셰퍼드를 창 레일 안쪽의 거실 목재 바닥에 배치했다."
    ],
    "physics": "로봇은 중력에 따라 거실 바닥에 완전히 기대어 쓰러져 있고, 셰퍼드는 바닥을 딛고 서서 물어뜯는 동작을 물리적으로 자연스럽게 취하고 있음."
   },
   {
    "label": "B",
    "direction": "셰퍼드가 로봇의 머리에서 찢어진 검은 피부 조각을 입으로 물고 당기고 있음.",
    "built_space": "캐릭터가 거실 내부가 아닌 열린 창문 바깥의 자갈밭 위에 위치해 있으며, 뒤쪽으로 거실 바닥이 보임.",
    "entities": "셰퍼드, 금속이 드러난 여자 로봇이 등장하나, 프롬프트나 레퍼런스에 없는 금속 막대기 소품이 추가됨.",
    "hard_violations": [
     "[gemini-pro] 지정된 '거실 바닥(Living-room floor)'이 아닌 야외 자갈밭에 캐릭터가 배치된 위치 오류 (a person placed where the staging does not put them)",
     "[gemini-pro] 로봇의 오른팔 아래에 팔과 연결되지 않은 여분의 로봇 손이 금속 막대를 쥐고 있는 신체 구조 및 중력 위반 (duplicated or extra bodies / floating object)",
     "[gemini-pro] 프롬프트에 지시되지 않은 금속 막대기가 등장함 (invented objects)",
     "[gpt-high] 로봇의 두 손에 검은 손잡이가 달린 은색 막대 소품을 추가했다. 이 소품은 해당 샷의 지시나 참고에서 같은 물건으로 확인되지 않는다."
    ],
    "physics": "로봇 본체와 셰퍼드는 바닥에 지탱되어 있으나, 막대를 쥐고 있는 두 번째 로봇 손은 팔이 연결되지 않은 채 물리적으로 불가능하게 놓여 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "지정된 거실 바닥 배경에서 셰퍼드가 로봇의 머리를 물어뜯어 은빛 금속이 드러난 장면을 프롬프트에 맞게 사실적으로 구현했습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "거실 바닥이 아닌 야외 자갈밭에 캐릭터를 배치했으며, 프롬프트에 없는 막대기와 여분의 잘린 로봇 손이 등장하여 치명적인 오류가 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "셰퍼드의 시선과 입이 바닥에 쓰러진 로봇의 머리를 향해 있으며, 드러난 금속 부위를 물고 있음.",
        "built_space": "실내 거실의 나무 바닥이 배경으로 쓰였으며, 왼쪽 화면 가장자리에 야외로 이어지는 열린 창문 프레임이 올바르게 배치됨.",
        "entities": "셰퍼드, 머리 가죽이 벗겨져 내부의 은빛 금속 골조가 드러난 젊은 아시아계 여자 로봇(검은색과 은색이 섞인 의상)이 명시된 외형에 맞게 등장함.",
        "hard_violations": [],
        "physics": "로봇은 중력에 따라 거실 바닥에 완전히 기대어 쓰러져 있고, 셰퍼드는 바닥을 딛고 서서 물어뜯는 동작을 물리적으로 자연스럽게 취하고 있음."
       },
       {
        "label": "B",
        "direction": "셰퍼드가 로봇의 머리에서 찢어진 검은 피부 조각을 입으로 물고 당기고 있음.",
        "built_space": "캐릭터가 거실 내부가 아닌 열린 창문 바깥의 자갈밭 위에 위치해 있으며, 뒤쪽으로 거실 바닥이 보임.",
        "entities": "셰퍼드, 금속이 드러난 여자 로봇이 등장하나, 프롬프트나 레퍼런스에 없는 금속 막대기 소품이 추가됨.",
        "hard_violations": [
         "지정된 '거실 바닥(Living-room floor)'이 아닌 야외 자갈밭에 캐릭터가 배치된 위치 오류 (a person placed where the staging does not put them)",
         "로봇의 오른팔 아래에 팔과 연결되지 않은 여분의 로봇 손이 금속 막대를 쥐고 있는 신체 구조 및 중력 위반 (duplicated or extra bodies / floating object)",
         "프롬프트에 지시되지 않은 금속 막대기가 등장함 (invented objects)"
        ],
        "physics": "로봇 본체와 셰퍼드는 바닥에 지탱되어 있으나, 막대를 쥐고 있는 두 번째 로봇 손은 팔이 연결되지 않은 채 물리적으로 불가능하게 놓여 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "지정된 거실 바닥 배경에서 셰퍼드가 로봇의 머리를 물어뜯어 은빛 금속이 드러난 장면을 프롬프트에 맞게 사실적으로 구현했습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "거실 바닥이 아닌 야외 자갈밭에 캐릭터를 배치했으며, 프롬프트에 없는 막대기와 여분의 잘린 로봇 손이 등장하여 치명적인 오류가 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "셰퍼드의 시선과 입이 바닥에 쓰러진 로봇의 머리를 향해 있으며, 드러난 금속 부위를 물고 있음.",
        "built_space": "실내 거실의 나무 바닥이 배경으로 쓰였으며, 왼쪽 화면 가장자리에 야외로 이어지는 열린 창문 프레임이 올바르게 배치됨.",
        "entities": "셰퍼드, 머리 가죽이 벗겨져 내부의 은빛 금속 골조가 드러난 젊은 아시아계 여자 로봇(검은색과 은색이 섞인 의상)이 명시된 외형에 맞게 등장함.",
        "hard_violations": [],
        "physics": "로봇은 중력에 따라 거실 바닥에 완전히 기대어 쓰러져 있고, 셰퍼드는 바닥을 딛고 서서 물어뜯는 동작을 물리적으로 자연스럽게 취하고 있음."
       },
       {
        "label": "B",
        "direction": "셰퍼드가 로봇의 머리에서 찢어진 검은 피부 조각을 입으로 물고 당기고 있음.",
        "built_space": "캐릭터가 거실 내부가 아닌 열린 창문 바깥의 자갈밭 위에 위치해 있으며, 뒤쪽으로 거실 바닥이 보임.",
        "entities": "셰퍼드, 금속이 드러난 여자 로봇이 등장하나, 프롬프트나 레퍼런스에 없는 금속 막대기 소품이 추가됨.",
        "hard_violations": [
         "지정된 '거실 바닥(Living-room floor)'이 아닌 야외 자갈밭에 캐릭터가 배치된 위치 오류 (a person placed where the staging does not put them)",
         "로봇의 오른팔 아래에 팔과 연결되지 않은 여분의 로봇 손이 금속 막대를 쥐고 있는 신체 구조 및 중력 위반 (duplicated or extra bodies / floating object)",
         "프롬프트에 지시되지 않은 금속 막대기가 등장함 (invented objects)"
        ],
        "physics": "로봇 본체와 셰퍼드는 바닥에 지탱되어 있으나, 막대를 쥐고 있는 두 번째 로봇 손은 팔이 연결되지 않은 채 물리적으로 불가능하게 놓여 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "창 바로 바깥에 쓰러진 로봇과 두피 아래 은빛 금속은 맞지만, 불필요한 막대 소품이 추가되고 머리·어깨 중심의 클로즈업보다 넓게 잡혔다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "두피를 무는 순간과 밀착 클로즈업은 더 정확하지만, 로봇과 개를 창밖이 아닌 거실 안쪽에 배치하여 지정 장소를 위반했다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "셰퍼드는 오른쪽에서 왼쪽 아래로 주둥이를 내려 로봇 정수리의 벗겨진 외피를 물고 있다. 공격 대상은 지정된 로봇의 머리다. 로봇의 눈은 바닥 쪽으로 풀려 있으며 카메라를 응시하지 않는다.",
        "built_space": "머리와 팔은 열린 창 바로 바깥의 회색 바닥과 자갈 경계에 놓여 있다. 창 개구부 하나의 하부 레일과 세로 프레임, 왼쪽 흰 커튼이 보이고, 안쪽에는 목재 바닥과 소파 일부, 탁자 다리와 러그가 보인다. 참고 장소의 재료와 안팎 관계는 대체로 맞지만, 배경 바닥과 가구가 넓게 보여 창 가장자리만 남기는 좁은 구도에는 미달한다.",
        "entities": "젊은 동아시아계 여성형 로봇 한 명과 검정·갈색 저먼 셰퍼드 한 마리가 보이며 다른 사람은 없다. 로봇의 검은 머리, 검은 상의와 은색 배색, 금속 팔은 참고와 대체로 맞는다. 정수리의 찢긴 덮개 아래 은색 기계 구조가 드러나지만 덮개의 뒷면이 피부보다 회색 고무처럼 보인다. 두 손 아래에는 요청에 없는 검은 손잡이의 은색 막대가 추가되어 있다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "로봇의 두 손에 검은 손잡이가 달린 은색 막대 소품을 추가했다. 이 소품은 해당 샷의 지시나 참고에서 같은 물건으로 확인되지 않는다."
        ],
        "physics": "로봇의 옆머리는 바닥에 펼쳐진 머리카락 위에 내려앉아 있고 어깨와 팔, 손도 바닥 또는 자갈에 지지된다. 막대는 손과 바닥에 걸쳐 있어 떠 있지 않다. 셰퍼드의 하체는 대부분 잘렸지만 오른쪽 아래 다리 일부가 지면으로 이어진다. 들어 올린 외피 조각은 개의 입에 붙잡혀 있어 당기는 동작으로 설명된다."
       },
       {
        "label": "B",
        "direction": "셰퍼드의 주둥이는 왼쪽 아래를 향하며 벌어진 턱과 이빨이 로봇의 찢긴 이마·두피 부위에 직접 닿는다. 공격 방향과 대상이 명확하다. 로봇의 얼굴은 위로 돌아가 있고 눈 부위는 손상과 개의 입에 가려져 시선을 판독하기 어렵다.",
        "built_space": "왼쪽에 열린 창 하나의 프레임과 하부 레일이 있고, 그 바깥으로 회색 외부 바닥과 자갈이 보인다. 반면 로봇의 머리와 어깨, 셰퍼드의 발은 레일 안쪽의 목재 거실 바닥에 놓여 있다. 머리 주변의 좁은 바닥과 창 가장자리만 담은 클로즈업은 잘 맞지만, 안팎 배치가 지정된 창 바로 바깥 장소와 반대다.",
        "entities": "젊은 동아시아계 여성형 로봇 한 명과 저먼 셰퍼드 한 마리가 보인다. 검은 머리와 은색 선이 있는 검은 상의는 참고와 부합한다. 얼굴이 크게 훼손되고 가려져 정확한 얼굴 동일성은 확인하기 어렵다. 찢긴 피부 가장자리와 그 아래 반사되는 은빛 금속은 분명하며 손상이 두피에서 볼까지 넓게 이어진다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "창 바로 바깥에 있어야 하는 쓰러진 로봇과 공격 중인 셰퍼드를 창 레일 안쪽의 거실 목재 바닥에 배치했다."
        ],
        "physics": "로봇의 뒤통수와 머리카락은 목재 바닥에 닿고, 목과 어깨도 누운 몸에 연결되어 지지된다. 셰퍼드의 보이는 앞발은 목재 바닥을 딛고 있으며 몸을 숙여 머리를 무는 자세가 가능하다. 찢어진 외피는 머리에 붙어 있거나 개의 이빨에 걸려 있고, 지지 없이 떠 있는 신체나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "창 바로 바깥에 쓰러진 로봇과 두피 아래 은빛 금속은 맞지만, 불필요한 막대 소품이 추가되고 머리·어깨 중심의 클로즈업보다 넓게 잡혔다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "두피를 무는 순간과 밀착 클로즈업은 더 정확하지만, 로봇과 개를 창밖이 아닌 거실 안쪽에 배치하여 지정 장소를 위반했다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "셰퍼드는 오른쪽에서 왼쪽 아래로 주둥이를 내려 로봇 정수리의 벗겨진 외피를 물고 있다. 공격 대상은 지정된 로봇의 머리다. 로봇의 눈은 바닥 쪽으로 풀려 있으며 카메라를 응시하지 않는다.",
        "built_space": "머리와 팔은 열린 창 바로 바깥의 회색 바닥과 자갈 경계에 놓여 있다. 창 개구부 하나의 하부 레일과 세로 프레임, 왼쪽 흰 커튼이 보이고, 안쪽에는 목재 바닥과 소파 일부, 탁자 다리와 러그가 보인다. 참고 장소의 재료와 안팎 관계는 대체로 맞지만, 배경 바닥과 가구가 넓게 보여 창 가장자리만 남기는 좁은 구도에는 미달한다.",
        "entities": "젊은 동아시아계 여성형 로봇 한 명과 검정·갈색 저먼 셰퍼드 한 마리가 보이며 다른 사람은 없다. 로봇의 검은 머리, 검은 상의와 은색 배색, 금속 팔은 참고와 대체로 맞는다. 정수리의 찢긴 덮개 아래 은색 기계 구조가 드러나지만 덮개의 뒷면이 피부보다 회색 고무처럼 보인다. 두 손 아래에는 요청에 없는 검은 손잡이의 은색 막대가 추가되어 있다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "로봇의 두 손에 검은 손잡이가 달린 은색 막대 소품을 추가했다. 이 소품은 해당 샷의 지시나 참고에서 같은 물건으로 확인되지 않는다."
        ],
        "physics": "로봇의 옆머리는 바닥에 펼쳐진 머리카락 위에 내려앉아 있고 어깨와 팔, 손도 바닥 또는 자갈에 지지된다. 막대는 손과 바닥에 걸쳐 있어 떠 있지 않다. 셰퍼드의 하체는 대부분 잘렸지만 오른쪽 아래 다리 일부가 지면으로 이어진다. 들어 올린 외피 조각은 개의 입에 붙잡혀 있어 당기는 동작으로 설명된다."
       },
       {
        "label": "A",
        "direction": "셰퍼드의 주둥이는 왼쪽 아래를 향하며 벌어진 턱과 이빨이 로봇의 찢긴 이마·두피 부위에 직접 닿는다. 공격 방향과 대상이 명확하다. 로봇의 얼굴은 위로 돌아가 있고 눈 부위는 손상과 개의 입에 가려져 시선을 판독하기 어렵다.",
        "built_space": "왼쪽에 열린 창 하나의 프레임과 하부 레일이 있고, 그 바깥으로 회색 외부 바닥과 자갈이 보인다. 반면 로봇의 머리와 어깨, 셰퍼드의 발은 레일 안쪽의 목재 거실 바닥에 놓여 있다. 머리 주변의 좁은 바닥과 창 가장자리만 담은 클로즈업은 잘 맞지만, 안팎 배치가 지정된 창 바로 바깥 장소와 반대다.",
        "entities": "젊은 동아시아계 여성형 로봇 한 명과 저먼 셰퍼드 한 마리가 보인다. 검은 머리와 은색 선이 있는 검은 상의는 참고와 부합한다. 얼굴이 크게 훼손되고 가려져 정확한 얼굴 동일성은 확인하기 어렵다. 찢긴 피부 가장자리와 그 아래 반사되는 은빛 금속은 분명하며 손상이 두피에서 볼까지 넓게 이어진다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "창 바로 바깥에 있어야 하는 쓰러진 로봇과 공격 중인 셰퍼드를 창 레일 안쪽의 거실 목재 바닥에 배치했다."
        ],
        "physics": "로봇의 뒤통수와 머리카락은 목재 바닥에 닿고, 목과 어깨도 누운 몸에 연결되어 지지된다. 셰퍼드의 보이는 앞발은 목재 바닥을 딛고 있으며 몸을 숙여 머리를 무는 자세가 가능하다. 찢어진 외피는 머리에 붙어 있거나 개의 이빨에 걸려 있고, 지지 없이 떠 있는 신체나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.75,
    "B": 1.25
   },
   "adjusted": {
    "A": 1.5,
    "B": 1.0
   },
   "violations": {
    "B": [
     "[gemini-pro] 지정된 '거실 바닥(Living-room floor)'이 아닌 야외 자갈밭에 캐릭터가 배치된 위치 오류 (a person placed where the staging does not put them)",
     "[gemini-pro] 로봇의 오른팔 아래에 팔과 연결되지 않은 여분의 로봇 손이 금속 막대를 쥐고 있는 신체 구조 및 중력 위반 (duplicated or extra bodies / floating object)",
     "[gemini-pro] 프롬프트에 지시되지 않은 금속 막대기가 등장함 (invented objects)",
     "[gpt-high] 로봇의 두 손에 검은 손잡이가 달린 은색 막대 소품을 추가했다. 이 소품은 해당 샷의 지시나 참고에서 같은 물건으로 확인되지 않는다."
    ],
    "A": [
     "[gpt-high] 창 바로 바깥에 있어야 하는 쓰러진 로봇과 공격 중인 셰퍼드를 창 레일 안쪽의 거실 목재 바닥에 배치했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1500,
   "B": 1000
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1500,
    "verdict_ko": "지정된 거실 바닥 배경에서 셰퍼드가 로봇의 머리를 물어뜯어 은빛 금속이 드러난 장면을 프롬프트에 맞게 사실적으로 구현했습니다.  ★위반: [gpt-high] 창 바로 바깥에 있어야 하는 쓰러진 로봇과 공격 중인 셰퍼드를 창 레일 안쪽의 거실 목재 바닥에 배치했다."
   },
   {
    "label": "B",
    "score": 1000,
    "verdict_ko": "거실 바닥이 아닌 야외 자갈밭에 캐릭터를 배치했으며, 프롬프트에 없는 막대기와 여분의 잘린 로봇 손이 등장하여 치명적인 오류가 발생했습니다.  ★위반: [gemini-pro] 지정된 '거실 바닥(Living-room floor)'이 아닌 야외 자갈밭에 캐릭터가 배치된 위치 오류 (a person placed where the staging does not put them) / [gemini-pro] 로봇의 오른팔 아래에 팔과 연결되지 않은 여분의 로봇 손이 금속 막대를 쥐고 있는 신체 구조 및 중력 위반 (duplicated or extra bodies / floating object) / [gemini-pro] 프롬프트에 지시되지 않은 금속 막대기가 등장함 (invented objects) / [gpt-high] 로봇의 두 손에 검은 손잡이가 달린 은색 막대 소품을 추가했다. 이 소품은 해당 샷의 지시나 참고에서 같은 물건으로 확인되지 않는다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 젊은 아시아계 여자 로봇 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S5sh1_sel.png",
    "asset_id": "a55d4fae-40d0-49ba-8c8c-947e911e0b7c",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 젊은 아시아계 여자 로봇: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:922551>",
    "asset_id": "99a76981-3435-4649-9a90-459804a78b99",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-9aca-745d-aba3-5b6536fa6693",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S5sh1"
  }
 },
 "S5sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T10:55:23.510851+00:00",
  "fingerprint": "7a3cf5a1f9a848b5e2fb45c0142859c1c142760a880a54126e12fb8e2d317b9f",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S5sh9_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S5sh9_sel.png",
  "source_sha256": "a15a7d8290428f184ab9893dd66a0410ffb4b5afce07b0fb7c8e1225b88eb810",
  "file": "S5sh9_cine.png",
  "staged_sha256": "617e55290eab768dae735553b450cce3b28f0259205f1deffc5f83482adbe1fc",
  "latency_ms": 9579
 },
 "S5sh16::signage": {
  "fp": "3ff843c6248d527a",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S5sh16::bgfirst_bg": {
  "input_fingerprint": "415cd6aa16584117",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 경찰에게 양팔을 붙잡힌 채 페드로의 차가 사라진 텅 빈 길가 쪽을 원망스럽게 돌아보는 현우의 찡그린 얼굴.\n\nLOCATION (lock): Outside the upscale house near the street, facing the empty roadside where the getaway car had been parked.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Empty roadside formerly occupied by 페드로's car in the middle-left of the frame, background.\n- KEY BACKGROUND ELEMENTS: Roadside where 페드로's car had been (Empty; 페드로 and the car are no longer present); used as A visible left-background gap that gives 현우's backward look its meaning.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light and controlled facial contrast keep the resentment intimate and the empty roadside plainly readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 경찰에게 양팔을 붙잡힌 채 페드로의 차가 사라진 텅 빈 길가 쪽을 원망스럽게 돌아보는 현우의 찡그린 얼굴.\n\nLOCATION (lock): Outside the upscale house near the street, facing the empty roadside where the getaway car had been parked.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Empty roadside formerly occupied by 페드로's car in the middle-left of the frame, background.\n- KEY BACKGROUND ELEMENTS: Roadside where 페드로's car had been (Empty; 페드로 and the car are no longer present); used as A visible left-background gap that gives 현우's backward look its meaning.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light and controlled facial contrast keep the resentment intimate and the empty roadside plainly readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S5sh16__bgfirst_bg.png",
  "asset_id": "87ad0c21-6893-428f-8870-cce81e611fed",
  "input_asset_ids": [
   "32035306-c502-4408-8bfc-f7eedf49376d",
   "e2cccbcb-d0f6-4b30-ac00-d35ae2ddd10b"
  ]
 },
 "S5sh16": {
  "input_fingerprint": "153494638d80443f",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 경찰에게 양팔을 붙잡힌 채 페드로의 차가 사라진 텅 빈 길가 쪽을 원망스럽게 돌아보는 현우의 찡그린 얼굴.\n\nLOCATION (lock): Outside the upscale house near the street, facing the empty roadside where the getaway car had been parked. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Empty roadside formerly occupied by 페드로's car in the middle-left of the frame, background.\n- KEY BACKGROUND ELEMENTS: Roadside where 페드로's car had been (Empty; 페드로 and the car are no longer present); used as A visible left-background gap that gives 현우's backward look its meaning.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light and controlled facial contrast keep the resentment intimate and the empty roadside plainly readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The car has disappeared from the roadside; the living-room window and the second-floor escape window remain open. The female housekeeping robot remains on the floor with torn head skin and exposed metal. 현우: Hyunwoo has a fresh dog-bite wound on his leg and is being led away under arrest.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 경찰에게 양팔을 붙잡힌 채 페드로의 차가 사라진 텅 빈 길가 쪽을 원망스럽게 돌아보는 현우의 찡그린 얼굴.\n\nLOCATION (lock): Outside the upscale house near the street, facing the empty roadside where the getaway car had been parked. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Empty roadside formerly occupied by 페드로's car in the middle-left of the frame, background.\n- KEY BACKGROUND ELEMENTS: Roadside where 페드로's car had been (Empty; 페드로 and the car are no longer present); used as A visible left-background gap that gives 현우's backward look its meaning.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light and controlled facial contrast keep the resentment intimate and the empty roadside plainly readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The car has disappeared from the roadside; the living-room window and the second-floor escape window remain open. The female housekeeping robot remains on the floor with torn head skin and exposed metal. 현우: Hyunwoo has a fresh dog-bite wound on his leg and is being led away under arrest.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 경찰에게 양팔을 붙잡힌 채 페드로의 차가 사라진 텅 빈 길가 쪽을 원망스럽게 돌아보는 현우의 찡그린 얼굴.\n\nLOCATION (lock): Outside the upscale house near the street, facing the empty roadside where the getaway car had been parked. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Empty roadside formerly occupied by 페드로's car in the middle-left of the frame, background.\n- KEY BACKGROUND ELEMENTS: Roadside where 페드로's car had been (Empty; 페드로 and the car are no longer present); used as A visible left-background gap that gives 현우's backward look its meaning.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light and controlled facial contrast keep the resentment intimate and the empty roadside plainly readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The car has disappeared from the roadside; the living-room window and the second-floor escape window remain open. The female housekeeping robot remains on the floor with torn head skin and exposed metal. 현우: Hyunwoo has a fresh dog-bite wound on his leg and is being led away under arrest.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S5sh16__bgfirst_bg.png",
     "asset_id": "87ad0c21-6893-428f-8870-cce81e611fed",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S5sh16.png",
     "asset_id": "32035306-c502-4408-8bfc-f7eedf49376d",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L152B01.png",
     "asset_id": "e2cccbcb-d0f6-4b30-ac00-d35ae2ddd10b",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우가 왼쪽의 텅 빈 도로 방향을 바라보고 있음.",
    "built_space": "제공된 로케이션 레퍼런스와 무관한 매끄러운 담장과 다른 형태의 주택들이 배치됨.",
    "entities": "현우의 외형은 레퍼런스와 비슷하나, 프롬프트 상 다리에 있어야 할 상처가 팔에 나타남. 두 명의 경찰관이 보임.",
    "hard_violations": [
     "[gemini-pro] 로케이션 레퍼런스 불일치 (완전히 다른 공간 및 건축물)"
    ],
    "physics": "두 경찰관에게 양팔이 붙잡힌 상태로 지면에 서 있음."
   },
   {
    "label": "B",
    "direction": "현우가 어깨 너머로 왼쪽 배경에 있는 텅 빈 도로 쪽을 시선으로 향하고 있음.",
    "built_space": "로케이션 레퍼런스의 낡은 콘크리트 담장, 녹슨 철문, 건물 창문 및 도로 구조가 정확한 위치에 구현됨.",
    "entities": "현우의 얼굴과 체형이 레퍼런스와 일치하며, 경찰 제복을 입은 사람들의 팔이 현우를 잡고 있음.",
    "hard_violations": [
     "[gpt-high] 전경 경찰의 소매 패치에 'POLICE'와 한글이 읽혀, 읽을 수 있는 글자와 로고를 금지한 조건을 위반한다."
    ],
    "physics": "땅에 발을 딛고 서 있으며, 옆과 뒤에서 경찰들에게 몸과 팔이 제압된 상태로 안정적임."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지정된 로케이션 레퍼런스의 디테일을 정확히 재현했으며, 프롬프트가 요구한 인물의 표정과 배경 구도를 충실히 묘사했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "필수인 로케이션 레퍼런스를 무시하고 전혀 다른 장소를 생성하여 가장 중요한 공간 제약을 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "현우가 어깨 너머로 왼쪽 배경에 있는 텅 빈 도로 쪽을 시선으로 향하고 있음.",
        "built_space": "로케이션 레퍼런스의 낡은 콘크리트 담장, 녹슨 철문, 건물 창문 및 도로 구조가 정확한 위치에 구현됨.",
        "entities": "현우의 얼굴과 체형이 레퍼런스와 일치하며, 경찰 제복을 입은 사람들의 팔이 현우를 잡고 있음.",
        "hard_violations": [],
        "physics": "땅에 발을 딛고 서 있으며, 옆과 뒤에서 경찰들에게 몸과 팔이 제압된 상태로 안정적임."
       },
       {
        "label": "A",
        "direction": "현우가 왼쪽의 텅 빈 도로 방향을 바라보고 있음.",
        "built_space": "제공된 로케이션 레퍼런스와 무관한 매끄러운 담장과 다른 형태의 주택들이 배치됨.",
        "entities": "현우의 외형은 레퍼런스와 비슷하나, 프롬프트 상 다리에 있어야 할 상처가 팔에 나타남. 두 명의 경찰관이 보임.",
        "hard_violations": [
         "로케이션 레퍼런스 불일치 (완전히 다른 공간 및 건축물)"
        ],
        "physics": "두 경찰관에게 양팔이 붙잡힌 상태로 지면에 서 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지정된 로케이션 레퍼런스의 디테일을 정확히 재현했으며, 프롬프트가 요구한 인물의 표정과 배경 구도를 충실히 묘사했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "필수인 로케이션 레퍼런스를 무시하고 전혀 다른 장소를 생성하여 가장 중요한 공간 제약을 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우가 어깨 너머로 왼쪽 배경에 있는 텅 빈 도로 쪽을 시선으로 향하고 있음.",
        "built_space": "로케이션 레퍼런스의 낡은 콘크리트 담장, 녹슨 철문, 건물 창문 및 도로 구조가 정확한 위치에 구현됨.",
        "entities": "현우의 얼굴과 체형이 레퍼런스와 일치하며, 경찰 제복을 입은 사람들의 팔이 현우를 잡고 있음.",
        "hard_violations": [],
        "physics": "땅에 발을 딛고 서 있으며, 옆과 뒤에서 경찰들에게 몸과 팔이 제압된 상태로 안정적임."
       },
       {
        "label": "A",
        "direction": "현우가 왼쪽의 텅 빈 도로 방향을 바라보고 있음.",
        "built_space": "제공된 로케이션 레퍼런스와 무관한 매끄러운 담장과 다른 형태의 주택들이 배치됨.",
        "entities": "현우의 외형은 레퍼런스와 비슷하나, 프롬프트 상 다리에 있어야 할 상처가 팔에 나타남. 두 명의 경찰관이 보임.",
        "hard_violations": [
         "로케이션 레퍼런스 불일치 (완전히 다른 공간 및 건축물)"
        ],
        "physics": "두 경찰관에게 양팔이 붙잡힌 상태로 지면에 서 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "장소와 현우의 외형은 가깝지만, 읽히는 경찰 표기가 금지 조건을 위반하며 시선도 왼쪽 빈 길가가 아닌 오른쪽으로 향하고 얼굴 클로즈업보다 넓다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "양팔을 붙잡힌 동작과 왼쪽의 빈 도로는 명확하지만, 시선이 그 도로를 향하지 않고 구도가 너무 넓으며 주택 외관도 지정 장소와 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 몸에 비해 얼굴을 뒤로 돌리고 미간을 찌푸렸지만, 눈동자는 화면 오른쪽 경찰 쪽으로 치우쳐 있다. 왼쪽 배경의 빈 길가를 바라보는 시선은 아니다. 오른쪽 경찰은 현우를 내려다본다. 무기나 이동 중인 차량은 없다.",
        "built_space": "왼쪽에 큰 미닫이창 한 벌, 녹슨 금속 대문 하나, 우편함 하나, 얼룩진 콘크리트 담장과 뒤쪽 기와지붕이 보여 장소 참조와 상당히 가깝다. 빈 길가는 왼쪽 중하단에 보이지만 도로의 깊은 원근은 오른쪽 배경에 놓인다. 큰 창은 닫힌 것으로 보이고, 2층 탈출 창은 프레임에 없다. 현우와 경찰은 담장 밖 도로에 있다.",
        "entities": "현우는 앳된 동아시아계 남성으로 보이며 헝클어진 검은 머리, 얼굴 윤곽과 남색 티셔츠가 참조에 가깝다. 국적은 외형만으로 확인할 수 없다. 경찰은 오른쪽의 얼굴과 상체 일부가 보이는 한 명, 회색 소매와 손이 보이는 한 명이다. 경찰은 장면 문구에 명시된 인물이다. 페드로와 차는 없다. 다리 상처와 실내 로봇은 이 구도에서 보이지 않아 평가 대상이 아니다. 전경 경찰 패치에는 'POLICE'와 한글 표기가 읽힌다.",
        "hard_violations": [
         "전경 경찰의 소매 패치에 'POLICE'와 한글이 읽혀, 읽을 수 있는 글자와 로고를 금지한 조건을 위반한다."
        ],
        "physics": "회색 소매의 손이 현우의 가까운 위팔을 실제로 움켜쥐고, 다른 손들이 반대쪽 팔과 어깨 부근에 접촉한다. 양팔을 제압하는 상황은 대체로 성립하지만 한 손은 팔보다 어깨를 누른다. 몸통과 목의 비틀림은 가능한 자세다. 발은 프레임 밖이며 공중에 뜬 몸이나 지지 없는 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "현우는 어깨 너머로 얼굴을 돌려 찡그리지만 눈동자는 화면 오른쪽, 오른쪽 경찰이 있는 방향으로 향한다. 왼쪽 중경의 빈 도로가 시선의 목표로 연결되지 않는다. 양옆 경찰은 현우 쪽으로 머리를 향하고 있으나 눈은 명확히 보이지 않는다. 무기나 움직이는 차량은 없다.",
        "built_space": "왼쪽 중경에 차 없는 도로가 확실히 확보된다. 오른쪽 주택에는 열린 아래층 창 한 벌과 열린 위층 창 한 벌, 출입문 하나와 매끈한 담장이 보인다. 그러나 참조의 낡은 콘크리트 벽, 녹슨 철문, 우편함과 큰 미닫이창 대신 다른 외관의 주택을 보여 정확한 장소 일치는 약하다. 현우는 도로 가장자리에서 양옆 경찰 사이에 서 있다. 얼굴뿐 아니라 허리 부근까지 보여 지정된 얼굴 클로즈업보다 넓다.",
        "entities": "현우의 앳된 얼굴, 검은 헝클어진 머리와 체격은 참조에 가깝고, 티셔츠는 참조보다 검정에 가까운 짙은 색이다. 양옆에는 제복을 입은 성인 남성 경찰 두 명의 부분 신체가 보인다. 페드로와 차량, 추가 행인은 없다. 아래쪽에 드러난 팔에는 붉은 상처 같은 자국이 있지만, 지정된 다리의 개물림 상처는 프레임 밖이라 확인할 수 없다. 실내 로봇 역시 보이지 않는다.",
        "hard_violations": [],
        "physics": "왼쪽 경찰의 손은 현우의 한쪽 위팔을 감싸고, 오른쪽 경찰의 손은 반대쪽 위팔을 잡아 양팔 구속이 명확하다. 두 손은 각각 경찰의 팔과 연결되고 옷과 피부에 접촉한다. 몸통은 진행 방향을 유지한 채 목과 어깨를 돌린 자세로 물리적으로 가능하다. 발은 잘려 있지만 부유나 불가능한 지지는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "장소와 현우의 외형은 가깝지만, 읽히는 경찰 표기가 금지 조건을 위반하며 시선도 왼쪽 빈 길가가 아닌 오른쪽으로 향하고 얼굴 클로즈업보다 넓다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "양팔을 붙잡힌 동작과 왼쪽의 빈 도로는 명확하지만, 시선이 그 도로를 향하지 않고 구도가 너무 넓으며 주택 외관도 지정 장소와 다르다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 몸에 비해 얼굴을 뒤로 돌리고 미간을 찌푸렸지만, 눈동자는 화면 오른쪽 경찰 쪽으로 치우쳐 있다. 왼쪽 배경의 빈 길가를 바라보는 시선은 아니다. 오른쪽 경찰은 현우를 내려다본다. 무기나 이동 중인 차량은 없다.",
        "built_space": "왼쪽에 큰 미닫이창 한 벌, 녹슨 금속 대문 하나, 우편함 하나, 얼룩진 콘크리트 담장과 뒤쪽 기와지붕이 보여 장소 참조와 상당히 가깝다. 빈 길가는 왼쪽 중하단에 보이지만 도로의 깊은 원근은 오른쪽 배경에 놓인다. 큰 창은 닫힌 것으로 보이고, 2층 탈출 창은 프레임에 없다. 현우와 경찰은 담장 밖 도로에 있다.",
        "entities": "현우는 앳된 동아시아계 남성으로 보이며 헝클어진 검은 머리, 얼굴 윤곽과 남색 티셔츠가 참조에 가깝다. 국적은 외형만으로 확인할 수 없다. 경찰은 오른쪽의 얼굴과 상체 일부가 보이는 한 명, 회색 소매와 손이 보이는 한 명이다. 경찰은 장면 문구에 명시된 인물이다. 페드로와 차는 없다. 다리 상처와 실내 로봇은 이 구도에서 보이지 않아 평가 대상이 아니다. 전경 경찰 패치에는 'POLICE'와 한글 표기가 읽힌다.",
        "hard_violations": [
         "전경 경찰의 소매 패치에 'POLICE'와 한글이 읽혀, 읽을 수 있는 글자와 로고를 금지한 조건을 위반한다."
        ],
        "physics": "회색 소매의 손이 현우의 가까운 위팔을 실제로 움켜쥐고, 다른 손들이 반대쪽 팔과 어깨 부근에 접촉한다. 양팔을 제압하는 상황은 대체로 성립하지만 한 손은 팔보다 어깨를 누른다. 몸통과 목의 비틀림은 가능한 자세다. 발은 프레임 밖이며 공중에 뜬 몸이나 지지 없는 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "현우는 어깨 너머로 얼굴을 돌려 찡그리지만 눈동자는 화면 오른쪽, 오른쪽 경찰이 있는 방향으로 향한다. 왼쪽 중경의 빈 도로가 시선의 목표로 연결되지 않는다. 양옆 경찰은 현우 쪽으로 머리를 향하고 있으나 눈은 명확히 보이지 않는다. 무기나 움직이는 차량은 없다.",
        "built_space": "왼쪽 중경에 차 없는 도로가 확실히 확보된다. 오른쪽 주택에는 열린 아래층 창 한 벌과 열린 위층 창 한 벌, 출입문 하나와 매끈한 담장이 보인다. 그러나 참조의 낡은 콘크리트 벽, 녹슨 철문, 우편함과 큰 미닫이창 대신 다른 외관의 주택을 보여 정확한 장소 일치는 약하다. 현우는 도로 가장자리에서 양옆 경찰 사이에 서 있다. 얼굴뿐 아니라 허리 부근까지 보여 지정된 얼굴 클로즈업보다 넓다.",
        "entities": "현우의 앳된 얼굴, 검은 헝클어진 머리와 체격은 참조에 가깝고, 티셔츠는 참조보다 검정에 가까운 짙은 색이다. 양옆에는 제복을 입은 성인 남성 경찰 두 명의 부분 신체가 보인다. 페드로와 차량, 추가 행인은 없다. 아래쪽에 드러난 팔에는 붉은 상처 같은 자국이 있지만, 지정된 다리의 개물림 상처는 프레임 밖이라 확인할 수 없다. 실내 로봇 역시 보이지 않는다.",
        "hard_violations": [],
        "physics": "왼쪽 경찰의 손은 현우의 한쪽 위팔을 감싸고, 오른쪽 경찰의 손은 반대쪽 위팔을 잡아 양팔 구속이 명확하다. 두 손은 각각 경찰의 팔과 연결되고 옷과 피부에 접촉한다. 몸통은 진행 방향을 유지한 채 목과 어깨를 돌린 자세로 물리적으로 가능하다. 발은 잘려 있지만 부유나 불가능한 지지는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.429,
    "B": 1.5
   },
   "adjusted": {
    "A": 1.179,
    "B": 1.25
   },
   "violations": {
    "A": [
     "[gemini-pro] 로케이션 레퍼런스 불일치 (완전히 다른 공간 및 건축물)"
    ],
    "B": [
     "[gpt-high] 전경 경찰의 소매 패치에 'POLICE'와 한글이 읽혀, 읽을 수 있는 글자와 로고를 금지한 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1250,
   "A": 1179
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1250,
    "verdict_ko": "지정된 로케이션 레퍼런스의 디테일을 정확히 재현했으며, 프롬프트가 요구한 인물의 표정과 배경 구도를 충실히 묘사했습니다.  ★위반: [gpt-high] 전경 경찰의 소매 패치에 'POLICE'와 한글이 읽혀, 읽을 수 있는 글자와 로고를 금지한 조건을 위반한다."
   },
   {
    "label": "A",
    "score": 1179,
    "verdict_ko": "필수인 로케이션 레퍼런스를 무시하고 전혀 다른 장소를 생성하여 가장 중요한 공간 제약을 위반했습니다.  ★위반: [gemini-pro] 로케이션 레퍼런스 불일치 (완전히 다른 공간 및 건축물)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L152B01.png",
    "asset_id": "e2cccbcb-d0f6-4b30-ac00-d35ae2ddd10b",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-9c86-75e8-b580-1abbc86b8fc4",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S5sh16__bgfirst_bg.png",
   "bg_asset_id": "87ad0c21-6893-428f-8870-cce81e611fed",
   "bg_record_key": "S5sh16::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S5sh16::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T03:57:07.117965+00:00",
  "fingerprint": "ab1e61464228cf5de73ebc12cff266750efb2d8d9bfd11a33cd748e9f544a12f",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S5sh16_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S5sh16_sel.png",
  "source_sha256": "c1fc7f5f50a357c35bb36404585a3b4a60d5e26f85887338b7ffac2cae6cdf7b",
  "file": "S5sh16_cine.png",
  "staged_sha256": "e680ef142c37eed59f0ed698211d69891305615c969e091465597163c76872a0",
  "latency_ms": 10966
 },
 "S6sh15::signage": {
  "fp": "caa95ef3c360673b",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S6sh15": {
  "input_fingerprint": "b58f44179b56cde0",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쓰레기를 머리에 뒤집어쓴 채 상체를 반쯤 일으킨 지탱 자세로 고릴라 형태의 로봇 찰리의 두 눈에 파란 불빛이 켜져 있는 순간.\n\nLOCATION (lock): Within an exposed rubbish mound at the refugee settlement's dump, at the spot where a large robot emerges. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Refuse surrounding 찰리 (Disturbed by his emergence, with rubbish still resting on his head and surrounding his lower body); used as Frames the supporting arms and makes the effort of emergence physically readable; Larger rubbish heap (Extending behind the emergence point); used as Provides layered background context without obscuring the head outline.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light remains steady while the blue illumination in 찰리's eyes becomes the restrained focal accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The repaired jukebox has its horn-shaped speaker connected and an LP fitted, with tools and additional records scattered around the blanket. Charlie is emerging from the rubbish as a dirt-covered, net-entangled gorilla-shaped robot with blue-lit eyes, stiff joints, and a worn Ubik chest logo.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쓰레기를 머리에 뒤집어쓴 채 상체를 반쯤 일으킨 지탱 자세로 고릴라 형태의 로봇 찰리의 두 눈에 파란 불빛이 켜져 있는 순간.\n\nLOCATION (lock): Within an exposed rubbish mound at the refugee settlement's dump, at the spot where a large robot emerges. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Refuse surrounding 찰리 (Disturbed by his emergence, with rubbish still resting on his head and surrounding his lower body); used as Frames the supporting arms and makes the effort of emergence physically readable; Larger rubbish heap (Extending behind the emergence point); used as Provides layered background context without obscuring the head outline.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light remains steady while the blue illumination in 찰리's eyes becomes the restrained focal accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The repaired jukebox has its horn-shaped speaker connected and an LP fitted, with tools and additional records scattered around the blanket. Charlie is emerging from the rubbish as a dirt-covered, net-entangled gorilla-shaped robot with blue-lit eyes, stiff joints, and a worn Ubik chest logo.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쓰레기를 머리에 뒤집어쓴 채 상체를 반쯤 일으킨 지탱 자세로 고릴라 형태의 로봇 찰리의 두 눈에 파란 불빛이 켜져 있는 순간.\n\nLOCATION (lock): Within an exposed rubbish mound at the refugee settlement's dump, at the spot where a large robot emerges. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Refuse surrounding 찰리 (Disturbed by his emergence, with rubbish still resting on his head and surrounding his lower body); used as Frames the supporting arms and makes the effort of emergence physically readable; Larger rubbish heap (Extending behind the emergence point); used as Provides layered background context without obscuring the head outline.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light remains steady while the blue illumination in 찰리's eyes becomes the restrained focal accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The repaired jukebox has its horn-shaped speaker connected and an LP fitted, with tools and additional records scattered around the blanket. Charlie is emerging from the rubbish as a dirt-covered, net-entangled gorilla-shaped robot with blue-lit eyes, stiff joints, and a worn Ubik chest logo.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S6sh15__bgfirst_bg.png",
     "asset_id": "8d75d911-3d18-4135-b6ab-fc26e495c901",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S6sh15.png",
     "asset_id": "c554fa18-115e-4e6b-9462-3d5f65b3a88e",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_dump_emergence_3bffab.png",
     "asset_id": "27c70e61-a2a2-4e06-b8da-d8a7b9325903",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "시선은 정면을 향하며 두 눈에 파란 불빛이 정확히 켜져 있음.",
    "built_space": "쓰레기장 배경. 좌측의 파란 컨테이너와 전경의 타이어, 우측의 굴착기 등 레퍼런스 이미지의 공간 배치를 매우 정확하게 재현함.",
    "entities": "고릴라 형태의 찰리 로봇이 머리에 그물과 쓰레기를 뒤집어쓴 채 등장함. 하체는 쓰레기 더미에 안정적으로 파묻혀 있음.",
    "hard_violations": [
     "[gpt-high] 가슴에 'Ubik'이라는 판독 가능한 글자와 로고가 남아 있어, 모든 글자를 읽히지 않게 하라는 조건을 위반한다."
    ],
    "physics": "두 팔로 바닥의 쓰레기 더미를 짚고 체중을 실어 상체를 들어올리는 지탱 자세가 물리적으로 자연스러움."
   },
   {
    "label": "B",
    "direction": "시선은 정면을 향하며 두 눈에 파란 불빛이 켜져 있음.",
    "built_space": "쓰레기장 배경. 우측에 굴착기, 좌측에 주크박스가 배치되어 있으나 전체적인 공간 구성이 레퍼런스 위치와 다소 차이가 있음.",
    "entities": "찰리 로봇. 머리 위에 쓰레기가 얹혀 있고 그물망은 왼쪽 어깨에 치우쳐 있음. 가슴에 불필요한 문자가 뒤섞여 있음.",
    "hard_violations": [
     "[gpt-high] 가슴에 'Ubik'이라는 읽을 수 있는 글자와 로고가 노출되어, 판독 가능한 글자와 로고를 금지한 조건을 위반한다."
    ],
    "physics": "팔로 바닥을 짚고 지탱하고 있으나, 무릎과 하체가 밖으로 크게 노출되어 쓰레기를 뚫고 나오는 물리적 상황과 다소 맞지 않음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "프롬프트가 요구한 반쯤 일으킨 지탱 자세와 하체를 둘러싼 쓰레기 더미를 정확히 표현했으며, 레퍼런스의 배경 요소(컨테이너, 타이어, 굴착기)를 훌륭하게 재현했습니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "로봇의 외형과 파란 눈빛은 잘 구현되었으나, 하체가 과하게 노출되어 쓰레기 더미에 파묻힌 느낌이 덜하고 머리에 쓰레기를 뒤집어쓴 형태가 다소 어색합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 정면을 향하며 두 눈에 파란 불빛이 정확히 켜져 있음.",
        "built_space": "쓰레기장 배경. 좌측의 파란 컨테이너와 전경의 타이어, 우측의 굴착기 등 레퍼런스 이미지의 공간 배치를 매우 정확하게 재현함.",
        "entities": "고릴라 형태의 찰리 로봇이 머리에 그물과 쓰레기를 뒤집어쓴 채 등장함. 하체는 쓰레기 더미에 안정적으로 파묻혀 있음.",
        "hard_violations": [],
        "physics": "두 팔로 바닥의 쓰레기 더미를 짚고 체중을 실어 상체를 들어올리는 지탱 자세가 물리적으로 자연스러움."
       },
       {
        "label": "B",
        "direction": "시선은 정면을 향하며 두 눈에 파란 불빛이 켜져 있음.",
        "built_space": "쓰레기장 배경. 우측에 굴착기, 좌측에 주크박스가 배치되어 있으나 전체적인 공간 구성이 레퍼런스 위치와 다소 차이가 있음.",
        "entities": "찰리 로봇. 머리 위에 쓰레기가 얹혀 있고 그물망은 왼쪽 어깨에 치우쳐 있음. 가슴에 불필요한 문자가 뒤섞여 있음.",
        "hard_violations": [],
        "physics": "팔로 바닥을 짚고 지탱하고 있으나, 무릎과 하체가 밖으로 크게 노출되어 쓰레기를 뚫고 나오는 물리적 상황과 다소 맞지 않음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "프롬프트가 요구한 반쯤 일으킨 지탱 자세와 하체를 둘러싼 쓰레기 더미를 정확히 표현했으며, 레퍼런스의 배경 요소(컨테이너, 타이어, 굴착기)를 훌륭하게 재현했습니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "로봇의 외형과 파란 눈빛은 잘 구현되었으나, 하체가 과하게 노출되어 쓰레기 더미에 파묻힌 느낌이 덜하고 머리에 쓰레기를 뒤집어쓴 형태가 다소 어색합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 정면을 향하며 두 눈에 파란 불빛이 정확히 켜져 있음.",
        "built_space": "쓰레기장 배경. 좌측의 파란 컨테이너와 전경의 타이어, 우측의 굴착기 등 레퍼런스 이미지의 공간 배치를 매우 정확하게 재현함.",
        "entities": "고릴라 형태의 찰리 로봇이 머리에 그물과 쓰레기를 뒤집어쓴 채 등장함. 하체는 쓰레기 더미에 안정적으로 파묻혀 있음.",
        "hard_violations": [],
        "physics": "두 팔로 바닥의 쓰레기 더미를 짚고 체중을 실어 상체를 들어올리는 지탱 자세가 물리적으로 자연스러움."
       },
       {
        "label": "B",
        "direction": "시선은 정면을 향하며 두 눈에 파란 불빛이 켜져 있음.",
        "built_space": "쓰레기장 배경. 우측에 굴착기, 좌측에 주크박스가 배치되어 있으나 전체적인 공간 구성이 레퍼런스 위치와 다소 차이가 있음.",
        "entities": "찰리 로봇. 머리 위에 쓰레기가 얹혀 있고 그물망은 왼쪽 어깨에 치우쳐 있음. 가슴에 불필요한 문자가 뒤섞여 있음.",
        "hard_violations": [],
        "physics": "팔로 바닥을 짚고 지탱하고 있으나, 무릎과 하체가 밖으로 크게 노출되어 쓰레기를 뚫고 나오는 물리적 상황과 다소 맞지 않음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "머리의 쓰레기와 양팔 지지는 구현했지만 상체가 상당히 세워져 있고 배경 주크박스가 과도하게 부각되며, 읽히는 가슴 글자가 명시적 금지 조건을 위반한다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "잔해에 양팔을 짚고 상체를 반쯤 일으키는 미디엄 숏을 더 정확히 구현했지만, 가슴의 읽히는 글자 때문에 최종 사용 조건에는 미달한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴과 두 파란 눈은 거의 카메라 정면을 향한다. 지정된 시선 대상은 없으므로 방향 위반은 아니다. 양팔은 몸 앞쪽 아래의 쓰레기 더미를 향해 내려가 있다.",
        "built_space": "야외 폐기물 더미 안에 로봇 한 대가 있고, 앞에는 금속 잔해와 타이어 하나, 뒤에는 큰 쓰레기 산이 보인다. 오른쪽 원경에는 집게 굴착기 한 대와 산업 구조물이 있다. 왼쪽 주크박스 한 대는 장소 참조보다 훨씬 크게 드러나 배경 소품의 존재감이 강하다. 별도의 실내 설비나 반사는 없다.",
        "entities": "찰리 한 대만 등장하며 추가 인물은 없다. 흰 각진 마스크, 모래색 장갑판, 긴 중량감 있는 팔과 노출된 기계 관절은 캐릭터 참조에 가깝다. 두 눈은 요청대로 파랗게 빛난다. 머리에 천과 폐기물이 얹혀 있고 그물이 어깨와 팔을 감싼다. 하체는 잔해에 가려져 있다. 가슴에는 'Ubik'이 선명하게 읽혀 마지막의 글자·로고 금지와 충돌한다.",
        "hard_violations": [
         "가슴에 'Ubik'이라는 읽을 수 있는 글자와 로고가 노출되어, 판독 가능한 글자와 로고를 금지한 조건을 위반한다."
        ],
        "physics": "화면 오른쪽 손은 전경 잔해에 닿아 있고 왼쪽 손의 접점은 일부 가려져 있으나 팔이 잔해 속까지 이어진다. 양팔과 매몰된 하체가 몸을 지지하는 자세로, 공중에 뜬 상태는 아니다. 다만 몸통이 상당히 수직으로 올라와 반쯤 일으키는 중간 동작은 상대적으로 약하다. 머리의 쓰레기는 머리와 어깨에 걸쳐 지지되고 그물은 장갑판을 따라 늘어진다."
       },
       {
        "label": "B",
        "direction": "얼굴은 화면 오른쪽으로 약간 돌아가며 두 파란 눈도 오른쪽 전방을 향한다. 특정 목표를 보라는 지시는 없어 허용되는 방향이다. 몸통은 앞쪽으로 기울고 양팔은 전방 아래의 잔해를 향해 뻗어 있다.",
        "built_space": "폐기물 더미 속 로봇의 상체와 지지하는 양팔을 미디엄 숏으로 담았다. 왼쪽 가장자리에 파란 금속 컨테이너 일부와 전경 타이어 하나가 있고, 뒤쪽으로 넓은 쓰레기 산이 이어진다. 오른쪽 원경의 집게 굴착기 한 대와 산업 구조물은 작게 유지된다. 장소 참조의 야외 고철 폐기장 재질과 공간 관계를 잘 이어가며 불가능한 반사나 중복 설비는 보이지 않는다.",
        "entities": "등장 주체는 찰리 한 대뿐이다. 흰 마스크형 얼굴, 모래색 각진 장갑, 육중한 긴 팔과 기계 관절이 참조의 정체성을 유지한다. 두 눈에는 파란 불빛이 켜져 있다. 머리 위에는 파이프와 잔해가 걸려 있고 그물이 머리와 어깨를 덮는다. 하체 주변은 폐기물에 묻혀 있다. 가슴의 'Ubik' 글자는 비스듬하지만 읽을 수 있어 글자 금지 조건을 충족하지 못한다.",
        "hard_violations": [
         "가슴에 'Ubik'이라는 판독 가능한 글자와 로고가 남아 있어, 모든 글자를 읽히지 않게 하라는 조건을 위반한다."
        ],
        "physics": "화면 오른쪽 손가락이 잔해 표면에 닿아 있고 반대쪽 팔도 전경 폐기물 속에 박혀 지지한다. 몸통은 앞으로 기울어진 채 양팔 사이에서 올라와, 쓰레기를 밀며 상체를 반쯤 일으키는 하중 전달이 읽힌다. 하체는 더미에 지지되며 떠 있지 않다. 머리 위 파이프와 잔해는 머리·등 쪽의 적층물과 그물에 걸쳐 있어 지지 없이 떠 있는 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "머리의 쓰레기와 양팔 지지는 구현했지만 상체가 상당히 세워져 있고 배경 주크박스가 과도하게 부각되며, 읽히는 가슴 글자가 명시적 금지 조건을 위반한다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "잔해에 양팔을 짚고 상체를 반쯤 일으키는 미디엄 숏을 더 정확히 구현했지만, 가슴의 읽히는 글자 때문에 최종 사용 조건에는 미달한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴과 두 파란 눈은 거의 카메라 정면을 향한다. 지정된 시선 대상은 없으므로 방향 위반은 아니다. 양팔은 몸 앞쪽 아래의 쓰레기 더미를 향해 내려가 있다.",
        "built_space": "야외 폐기물 더미 안에 로봇 한 대가 있고, 앞에는 금속 잔해와 타이어 하나, 뒤에는 큰 쓰레기 산이 보인다. 오른쪽 원경에는 집게 굴착기 한 대와 산업 구조물이 있다. 왼쪽 주크박스 한 대는 장소 참조보다 훨씬 크게 드러나 배경 소품의 존재감이 강하다. 별도의 실내 설비나 반사는 없다.",
        "entities": "찰리 한 대만 등장하며 추가 인물은 없다. 흰 각진 마스크, 모래색 장갑판, 긴 중량감 있는 팔과 노출된 기계 관절은 캐릭터 참조에 가깝다. 두 눈은 요청대로 파랗게 빛난다. 머리에 천과 폐기물이 얹혀 있고 그물이 어깨와 팔을 감싼다. 하체는 잔해에 가려져 있다. 가슴에는 'Ubik'이 선명하게 읽혀 마지막의 글자·로고 금지와 충돌한다.",
        "hard_violations": [
         "가슴에 'Ubik'이라는 읽을 수 있는 글자와 로고가 노출되어, 판독 가능한 글자와 로고를 금지한 조건을 위반한다."
        ],
        "physics": "화면 오른쪽 손은 전경 잔해에 닿아 있고 왼쪽 손의 접점은 일부 가려져 있으나 팔이 잔해 속까지 이어진다. 양팔과 매몰된 하체가 몸을 지지하는 자세로, 공중에 뜬 상태는 아니다. 다만 몸통이 상당히 수직으로 올라와 반쯤 일으키는 중간 동작은 상대적으로 약하다. 머리의 쓰레기는 머리와 어깨에 걸쳐 지지되고 그물은 장갑판을 따라 늘어진다."
       },
       {
        "label": "A",
        "direction": "얼굴은 화면 오른쪽으로 약간 돌아가며 두 파란 눈도 오른쪽 전방을 향한다. 특정 목표를 보라는 지시는 없어 허용되는 방향이다. 몸통은 앞쪽으로 기울고 양팔은 전방 아래의 잔해를 향해 뻗어 있다.",
        "built_space": "폐기물 더미 속 로봇의 상체와 지지하는 양팔을 미디엄 숏으로 담았다. 왼쪽 가장자리에 파란 금속 컨테이너 일부와 전경 타이어 하나가 있고, 뒤쪽으로 넓은 쓰레기 산이 이어진다. 오른쪽 원경의 집게 굴착기 한 대와 산업 구조물은 작게 유지된다. 장소 참조의 야외 고철 폐기장 재질과 공간 관계를 잘 이어가며 불가능한 반사나 중복 설비는 보이지 않는다.",
        "entities": "등장 주체는 찰리 한 대뿐이다. 흰 마스크형 얼굴, 모래색 각진 장갑, 육중한 긴 팔과 기계 관절이 참조의 정체성을 유지한다. 두 눈에는 파란 불빛이 켜져 있다. 머리 위에는 파이프와 잔해가 걸려 있고 그물이 머리와 어깨를 덮는다. 하체 주변은 폐기물에 묻혀 있다. 가슴의 'Ubik' 글자는 비스듬하지만 읽을 수 있어 글자 금지 조건을 충족하지 못한다.",
        "hard_violations": [
         "가슴에 'Ubik'이라는 판독 가능한 글자와 로고가 남아 있어, 모든 글자를 읽히지 않게 하라는 조건을 위반한다."
        ],
        "physics": "화면 오른쪽 손가락이 잔해 표면에 닿아 있고 반대쪽 팔도 전경 폐기물 속에 박혀 지지한다. 몸통은 앞으로 기울어진 채 양팔 사이에서 올라와, 쓰레기를 밀며 상체를 반쯤 일으키는 하중 전달이 읽힌다. 하체는 더미에 지지되며 떠 있지 않다. 머리 위 파이프와 잔해는 머리·등 쪽의 적층물과 그물에 걸쳐 있어 지지 없이 떠 있는 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.464
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.214
   },
   "violations": {
    "B": [
     "[gpt-high] 가슴에 'Ubik'이라는 읽을 수 있는 글자와 로고가 노출되어, 판독 가능한 글자와 로고를 금지한 조건을 위반한다."
    ],
    "A": [
     "[gpt-high] 가슴에 'Ubik'이라는 판독 가능한 글자와 로고가 남아 있어, 모든 글자를 읽히지 않게 하라는 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 1214
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "프롬프트가 요구한 반쯤 일으킨 지탱 자세와 하체를 둘러싼 쓰레기 더미를 정확히 표현했으며, 레퍼런스의 배경 요소(컨테이너, 타이어, 굴착기)를 훌륭하게 재현했습니다.  ★위반: [gpt-high] 가슴에 'Ubik'이라는 판독 가능한 글자와 로고가 남아 있어, 모든 글자를 읽히지 않게 하라는 조건을 위반한다."
   },
   {
    "label": "B",
    "score": 1214,
    "verdict_ko": "로봇의 외형과 파란 눈빛은 잘 구현되었으나, 하체가 과하게 노출되어 쓰레기 더미에 파묻힌 느낌이 덜하고 머리에 쓰레기를 뒤집어쓴 형태가 다소 어색합니다.  ★위반: [gpt-high] 가슴에 'Ubik'이라는 읽을 수 있는 글자와 로고가 노출되어, 판독 가능한 글자와 로고를 금지한 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_dump_emergence_3bffab.png",
    "asset_id": "27c70e61-a2a2-4e06-b8da-d8a7b9325903",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-a002-755b-90fd-cd527be09269",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S6sh15__bgfirst_bg.png",
   "bg_asset_id": "8d75d911-3d18-4135-b6ab-fc26e495c901",
   "bg_record_key": "S6sh15::bgfirst_bg",
   "chain_winner": true,
   "authority": "groupbg",
   "group_key": "dump_emergence",
   "groupbg_asset_id": "27c70e61-a2a2-4e06-b8da-d8a7b9325903"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S6sh15::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T10:57:47.581200+00:00",
  "fingerprint": "f7e6fe768dc02a3d54e30d24a59fe751419ac9a6217b1c051dfd1c9638b9dbf7",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S6sh15_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S6sh15_sel.png",
  "source_sha256": "259fd7edf56f47e5ceaf09a5f2fc8a7b12503357773eb90f813cd8b485a08b7e",
  "file": "S6sh15_cine.png",
  "staged_sha256": "a7ff82f1fd4b5b2faa966960ee619c6aa6079cbcedc49df8101b0ed4d25bd677",
  "latency_ms": 11478
 },
 "S6sh17::signage": {
  "fp": "456993398ebdc7bf",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S6sh17": {
  "input_fingerprint": "89ca3ca87d311b41",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 고릴라 형태의 찰리가 거대한 두 팔로 앰버의 작은 몸을 빈틈없이 감싸 안은 상태.\n\nLOCATION (lock): On the rubbish-strewn ground beside the robot's emergence point in the refugee settlement dump. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Refuse around the pair (Surrounding the place where 찰리 emerged); used as A subdued lower and rear surround that preserves the harsh setting around the intimate contact.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Steady daytime ambient light and restrained blue eye illumination preserve tenderness without changing the established tonal treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same daylight, exposed refuse, and surrounding scrap heaps at the robot's emergence site. Exclude airborne rubbish and any debris still falling from the emergence.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie is upright with his arms closed in an embrace, retaining his blue-lit eyes, dirty aged casing, netting, and worn Ubik chest logo. The repaired jukebox, fitted horn speaker, LP, blanket, and scattered tools remain at the rubbish heap. 앰버: Amber has approached the robot's position and is caught in a tight embrace, looking bewildered; her mask and waist tool pouch remain in place.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 고릴라 형태의 찰리가 거대한 두 팔로 앰버의 작은 몸을 빈틈없이 감싸 안은 상태.\n\nLOCATION (lock): On the rubbish-strewn ground beside the robot's emergence point in the refugee settlement dump. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Refuse around the pair (Surrounding the place where 찰리 emerged); used as A subdued lower and rear surround that preserves the harsh setting around the intimate contact.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Steady daytime ambient light and restrained blue eye illumination preserve tenderness without changing the established tonal treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same daylight, exposed refuse, and surrounding scrap heaps at the robot's emergence site. Exclude airborne rubbish and any debris still falling from the emergence.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie is upright with his arms closed in an embrace, retaining his blue-lit eyes, dirty aged casing, netting, and worn Ubik chest logo. The repaired jukebox, fitted horn speaker, LP, blanket, and scattered tools remain at the rubbish heap. 앰버: Amber has approached the robot's position and is caught in a tight embrace, looking bewildered; her mask and waist tool pouch remain in place.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 고릴라 형태의 찰리가 거대한 두 팔로 앰버의 작은 몸을 빈틈없이 감싸 안은 상태.\n\nLOCATION (lock): On the rubbish-strewn ground beside the robot's emergence point in the refugee settlement dump. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Refuse around the pair (Surrounding the place where 찰리 emerged); used as A subdued lower and rear surround that preserves the harsh setting around the intimate contact.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Steady daytime ambient light and restrained blue eye illumination preserve tenderness without changing the established tonal treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same daylight, exposed refuse, and surrounding scrap heaps at the robot's emergence site. Exclude airborne rubbish and any debris still falling from the emergence.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie is upright with his arms closed in an embrace, retaining his blue-lit eyes, dirty aged casing, netting, and worn Ubik chest logo. The repaired jukebox, fitted horn speaker, LP, blanket, and scattered tools remain at the rubbish heap. 앰버: Amber has approached the robot's position and is caught in a tight embrace, looking bewildered; her mask and waist tool pouch remain in place.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "B",
    "direction": "찰리는 머리를 숙여 앰버의 머리 쪽을 향한다. 앰버는 찰리를 올려다보지 않고 카메라 왼쪽 가까운 화면 밖을 바라보며 입을 조금 벌리고 있다. 두 팔과 손은 앰버의 몸통을 향해 안쪽으로 닫혀 있다. 무기나 이동 중인 물체는 없다.",
    "built_space": "야외 폐기물 더미 안의 두 인물을 허리 부근까지 담은 미디엄 숏이다. 왼쪽에 기울어진 파란 금속 용기 하나, 왼쪽 아래에 타이어 하나, 오른쪽 뒤에 녹슨 가구형 폐기물이 보인다. 이전 장면의 노출된 고철과 콘크리트 잔해는 이어지지만 뒤쪽 더미의 윤곽과 구성은 달라졌다. 두 인물은 잔해 사이 같은 자리에 밀착해 있으며 좌석이나 반사면은 없다.",
    "entities": "찰리 한 개체와 앰버 한 명만 보인다. 찰리의 거대한 긴 팔, 낡은 샌드 베이지 장갑, 흰 각진 얼굴, 파란 눈, 그물과 머리 위 관은 이전 장면과 대체로 일치한다. 앰버는 금발의 어린 여자아이로 둥근 얼굴과 남색 상의가 참고 이미지에 부합하며 혼혈 설정과 모순되는 뚜렷한 특징은 없다. 허리 공구 주머니는 보이지만 마스크는 코와 입을 덮지 않고 턱 아래에 있다. 가슴에는 희미해도 판독 가능한 글자가 있다. 주크박스, 혼 스피커, LP, 담요와 개별 공구는 이 프레임에서 식별되지 않는다.",
    "hard_violations": [
     "찰리 가슴의 상표 글자가 판독 가능하여 읽을 수 있는 글자와 로고를 금지한 조건을 위반한다."
    ],
    "physics": "찰리의 두 전완은 앰버의 몸 앞에서 겹치고 손은 어깨와 몸통에 접촉해 포옹을 지탱한다. 앰버의 하체는 아래로 이어지다가 잔해와 화면 경계에 가려진다. 두 인물의 발은 보이지 않지만 공중에 떠 있는 자세는 아니며, 앰버에게는 로봇 팔의 지지가 보인다. 주변 잔해는 지면이나 다른 잔해 위에 놓여 있고 낙하 중인 쓰레기는 없다."
   },
   {
    "label": "A",
    "direction": "찰리의 얼굴은 아래로 기울어 앰버의 머리와 상체 쪽을 향한다. 앰버는 화면 왼쪽 바깥을 긴장한 눈으로 바라본다. 위쪽 팔은 앰버의 어깨와 등 쪽을, 아래쪽 팔은 허리를 감싸며 모두 아이를 중심으로 닫혀 있다. 무기나 이동 물체는 없다.",
    "built_space": "폐기물 사이의 두 인물을 중심으로 한 미디엄 숏이다. 왼쪽에는 기울어진 파란 금속 용기 하나와 아래쪽 타이어 하나가 있고, 오른쪽 뒤에는 큰 녹슨 상자형 폐기물 두 개가 서로 다른 높이로 놓여 있다. 배경 전체가 고철과 잔해로 채워져 이전 장면의 재료와 낮의 환경을 유지한다. 더미의 개별 배치는 달라졌지만 중복된 고정 설비나 불가능한 반사는 보이지 않는다.",
    "entities": "찰리와 앰버 외의 인물은 없다. 찰리는 이전 장면의 낡은 베이지 장갑, 그물, 흰 마스크형 얼굴과 파란 눈을 유지하며, 눈의 발광은 A보다 조금 강하다. 앰버는 참고 이미지와 부합하는 금발, 어린 여자아이의 체격, 남색 상의를 갖췄다. 코와 입을 마스크가 덮어 얼굴 전체의 일치 여부는 확인할 수 없으며, 눈과 눈썹에는 경계하고 당황한 기색이 있다. 허리 주머니는 팔과 잔해에 가려 확인되지 않는다. 가슴의 상표 글자는 읽힌다. 주크박스, 혼 스피커, LP, 담요와 개별 공구는 프레임에서 식별되지 않는다.",
    "hard_violations": [
     "찰리 가슴의 상표 글자가 판독 가능하여 읽을 수 있는 글자와 로고를 금지한 조건을 위반한다."
    ],
    "physics": "찰리의 한 손은 앰버의 어깨와 등 옆에, 다른 손은 허리 아래쪽에 밀착해 있어 작은 몸을 감싸고 지지하는 관계가 명확하다. 팔의 관절과 전완은 몸통으로 자연스럽게 이어진다. 앰버의 하체는 아래로 이어지며 잔해에 가려지고, 찰리의 하체도 화면 밖에 있어 접지점 자체는 확인되지 않는다. 지지 없이 떠 있는 몸이나 물체는 보이지 않으며 잔해도 모두 쌓인 상태다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": null,
     "normalized": null,
     "ok": false
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "미디엄 숏과 두 팔의 포옹은 구현했지만 마스크가 턱 아래로 내려갔으며, 읽을 수 있는 가슴 글자가 명시적 금지 조건을 위반한다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "앰버의 마스크를 유지하고 어깨와 허리를 두 팔로 밀착해 감싼 순간이 더 충실하지만, 가슴 글자가 읽혀 최종 사용 조건은 충족하지 못한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 머리를 숙여 앰버의 머리 쪽을 향한다. 앰버는 찰리를 올려다보지 않고 카메라 왼쪽 가까운 화면 밖을 바라보며 입을 조금 벌리고 있다. 두 팔과 손은 앰버의 몸통을 향해 안쪽으로 닫혀 있다. 무기나 이동 중인 물체는 없다.",
        "built_space": "야외 폐기물 더미 안의 두 인물을 허리 부근까지 담은 미디엄 숏이다. 왼쪽에 기울어진 파란 금속 용기 하나, 왼쪽 아래에 타이어 하나, 오른쪽 뒤에 녹슨 가구형 폐기물이 보인다. 이전 장면의 노출된 고철과 콘크리트 잔해는 이어지지만 뒤쪽 더미의 윤곽과 구성은 달라졌다. 두 인물은 잔해 사이 같은 자리에 밀착해 있으며 좌석이나 반사면은 없다.",
        "entities": "찰리 한 개체와 앰버 한 명만 보인다. 찰리의 거대한 긴 팔, 낡은 샌드 베이지 장갑, 흰 각진 얼굴, 파란 눈, 그물과 머리 위 관은 이전 장면과 대체로 일치한다. 앰버는 금발의 어린 여자아이로 둥근 얼굴과 남색 상의가 참고 이미지에 부합하며 혼혈 설정과 모순되는 뚜렷한 특징은 없다. 허리 공구 주머니는 보이지만 마스크는 코와 입을 덮지 않고 턱 아래에 있다. 가슴에는 희미해도 판독 가능한 글자가 있다. 주크박스, 혼 스피커, LP, 담요와 개별 공구는 이 프레임에서 식별되지 않는다.",
        "hard_violations": [
         "찰리 가슴의 상표 글자가 판독 가능하여 읽을 수 있는 글자와 로고를 금지한 조건을 위반한다."
        ],
        "physics": "찰리의 두 전완은 앰버의 몸 앞에서 겹치고 손은 어깨와 몸통에 접촉해 포옹을 지탱한다. 앰버의 하체는 아래로 이어지다가 잔해와 화면 경계에 가려진다. 두 인물의 발은 보이지 않지만 공중에 떠 있는 자세는 아니며, 앰버에게는 로봇 팔의 지지가 보인다. 주변 잔해는 지면이나 다른 잔해 위에 놓여 있고 낙하 중인 쓰레기는 없다."
       },
       {
        "label": "B",
        "direction": "찰리의 얼굴은 아래로 기울어 앰버의 머리와 상체 쪽을 향한다. 앰버는 화면 왼쪽 바깥을 긴장한 눈으로 바라본다. 위쪽 팔은 앰버의 어깨와 등 쪽을, 아래쪽 팔은 허리를 감싸며 모두 아이를 중심으로 닫혀 있다. 무기나 이동 물체는 없다.",
        "built_space": "폐기물 사이의 두 인물을 중심으로 한 미디엄 숏이다. 왼쪽에는 기울어진 파란 금속 용기 하나와 아래쪽 타이어 하나가 있고, 오른쪽 뒤에는 큰 녹슨 상자형 폐기물 두 개가 서로 다른 높이로 놓여 있다. 배경 전체가 고철과 잔해로 채워져 이전 장면의 재료와 낮의 환경을 유지한다. 더미의 개별 배치는 달라졌지만 중복된 고정 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "찰리와 앰버 외의 인물은 없다. 찰리는 이전 장면의 낡은 베이지 장갑, 그물, 흰 마스크형 얼굴과 파란 눈을 유지하며, 눈의 발광은 A보다 조금 강하다. 앰버는 참고 이미지와 부합하는 금발, 어린 여자아이의 체격, 남색 상의를 갖췄다. 코와 입을 마스크가 덮어 얼굴 전체의 일치 여부는 확인할 수 없으며, 눈과 눈썹에는 경계하고 당황한 기색이 있다. 허리 주머니는 팔과 잔해에 가려 확인되지 않는다. 가슴의 상표 글자는 읽힌다. 주크박스, 혼 스피커, LP, 담요와 개별 공구는 프레임에서 식별되지 않는다.",
        "hard_violations": [
         "찰리 가슴의 상표 글자가 판독 가능하여 읽을 수 있는 글자와 로고를 금지한 조건을 위반한다."
        ],
        "physics": "찰리의 한 손은 앰버의 어깨와 등 옆에, 다른 손은 허리 아래쪽에 밀착해 있어 작은 몸을 감싸고 지지하는 관계가 명확하다. 팔의 관절과 전완은 몸통으로 자연스럽게 이어진다. 앰버의 하체는 아래로 이어지며 잔해에 가려지고, 찰리의 하체도 화면 밖에 있어 접지점 자체는 확인되지 않는다. 지지 없이 떠 있는 몸이나 물체는 보이지 않으며 잔해도 모두 쌓인 상태다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "미디엄 숏과 두 팔의 포옹은 구현했지만 마스크가 턱 아래로 내려갔으며, 읽을 수 있는 가슴 글자가 명시적 금지 조건을 위반한다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "앰버의 마스크를 유지하고 어깨와 허리를 두 팔로 밀착해 감싼 순간이 더 충실하지만, 가슴 글자가 읽혀 최종 사용 조건은 충족하지 못한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "찰리는 머리를 숙여 앰버의 머리 쪽을 향한다. 앰버는 찰리를 올려다보지 않고 카메라 왼쪽 가까운 화면 밖을 바라보며 입을 조금 벌리고 있다. 두 팔과 손은 앰버의 몸통을 향해 안쪽으로 닫혀 있다. 무기나 이동 중인 물체는 없다.",
        "built_space": "야외 폐기물 더미 안의 두 인물을 허리 부근까지 담은 미디엄 숏이다. 왼쪽에 기울어진 파란 금속 용기 하나, 왼쪽 아래에 타이어 하나, 오른쪽 뒤에 녹슨 가구형 폐기물이 보인다. 이전 장면의 노출된 고철과 콘크리트 잔해는 이어지지만 뒤쪽 더미의 윤곽과 구성은 달라졌다. 두 인물은 잔해 사이 같은 자리에 밀착해 있으며 좌석이나 반사면은 없다.",
        "entities": "찰리 한 개체와 앰버 한 명만 보인다. 찰리의 거대한 긴 팔, 낡은 샌드 베이지 장갑, 흰 각진 얼굴, 파란 눈, 그물과 머리 위 관은 이전 장면과 대체로 일치한다. 앰버는 금발의 어린 여자아이로 둥근 얼굴과 남색 상의가 참고 이미지에 부합하며 혼혈 설정과 모순되는 뚜렷한 특징은 없다. 허리 공구 주머니는 보이지만 마스크는 코와 입을 덮지 않고 턱 아래에 있다. 가슴에는 희미해도 판독 가능한 글자가 있다. 주크박스, 혼 스피커, LP, 담요와 개별 공구는 이 프레임에서 식별되지 않는다.",
        "hard_violations": [
         "찰리 가슴의 상표 글자가 판독 가능하여 읽을 수 있는 글자와 로고를 금지한 조건을 위반한다."
        ],
        "physics": "찰리의 두 전완은 앰버의 몸 앞에서 겹치고 손은 어깨와 몸통에 접촉해 포옹을 지탱한다. 앰버의 하체는 아래로 이어지다가 잔해와 화면 경계에 가려진다. 두 인물의 발은 보이지 않지만 공중에 떠 있는 자세는 아니며, 앰버에게는 로봇 팔의 지지가 보인다. 주변 잔해는 지면이나 다른 잔해 위에 놓여 있고 낙하 중인 쓰레기는 없다."
       },
       {
        "label": "A",
        "direction": "찰리의 얼굴은 아래로 기울어 앰버의 머리와 상체 쪽을 향한다. 앰버는 화면 왼쪽 바깥을 긴장한 눈으로 바라본다. 위쪽 팔은 앰버의 어깨와 등 쪽을, 아래쪽 팔은 허리를 감싸며 모두 아이를 중심으로 닫혀 있다. 무기나 이동 물체는 없다.",
        "built_space": "폐기물 사이의 두 인물을 중심으로 한 미디엄 숏이다. 왼쪽에는 기울어진 파란 금속 용기 하나와 아래쪽 타이어 하나가 있고, 오른쪽 뒤에는 큰 녹슨 상자형 폐기물 두 개가 서로 다른 높이로 놓여 있다. 배경 전체가 고철과 잔해로 채워져 이전 장면의 재료와 낮의 환경을 유지한다. 더미의 개별 배치는 달라졌지만 중복된 고정 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "찰리와 앰버 외의 인물은 없다. 찰리는 이전 장면의 낡은 베이지 장갑, 그물, 흰 마스크형 얼굴과 파란 눈을 유지하며, 눈의 발광은 A보다 조금 강하다. 앰버는 참고 이미지와 부합하는 금발, 어린 여자아이의 체격, 남색 상의를 갖췄다. 코와 입을 마스크가 덮어 얼굴 전체의 일치 여부는 확인할 수 없으며, 눈과 눈썹에는 경계하고 당황한 기색이 있다. 허리 주머니는 팔과 잔해에 가려 확인되지 않는다. 가슴의 상표 글자는 읽힌다. 주크박스, 혼 스피커, LP, 담요와 개별 공구는 프레임에서 식별되지 않는다.",
        "hard_violations": [
         "찰리 가슴의 상표 글자가 판독 가능하여 읽을 수 있는 글자와 로고를 금지한 조건을 위반한다."
        ],
        "physics": "찰리의 한 손은 앰버의 어깨와 등 옆에, 다른 손은 허리 아래쪽에 밀착해 있어 작은 몸을 감싸고 지지하는 관계가 명확하다. 팔의 관절과 전완은 몸통으로 자연스럽게 이어진다. 앰버의 하체는 아래로 이어지며 잔해에 가려지고, 찰리의 하체도 화면 밖에 있어 접지점 자체는 확인되지 않는다. 지지 없이 떠 있는 몸이나 물체는 보이지 않으며 잔해도 모두 쌓인 상태다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gemini-pro"
   ],
   "route": "single_reverse"
  },
  "totals": {
   "B": 3,
   "A": 4
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 3,
    "verdict_ko": "미디엄 숏과 두 팔의 포옹은 구현했지만 마스크가 턱 아래로 내려갔으며, 읽을 수 있는 가슴 글자가 명시적 금지 조건을 위반한다."
   },
   {
    "label": "A",
    "score": 4,
    "verdict_ko": "앰버의 마스크를 유지하고 어깨와 허리를 두 팔로 밀착해 감싼 순간이 더 충실하지만, 가슴 글자가 읽혀 최종 사용 조건은 충족하지 못한다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S6sh15_sel.png",
    "asset_id": "47c38513-8294-4e92-85e0-3e6a6ea306f3",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-a4f4-7fe7-b965-9d9ccb65b705",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S6sh15"
  }
 },
 "S6sh17::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T10:59:10.567103+00:00",
  "fingerprint": "342e91a0b3275fcf4494716a11ff37d50332e3a918e48aa0fcb88e13d3b1f6fc",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S6sh17_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S6sh17_sel.png",
  "source_sha256": "2c887b1f2eed1612239989c99ed6cacbdedb53a131ae77424fbe8ba2c936c4f5",
  "file": "S6sh17_cine.png",
  "staged_sha256": "1bd8dacc58c202091865ce1ba3fb537b07d9942670911a7720f91b8e8fcdddcb",
  "latency_ms": 9628
 },
 "S6sh26::signage": {
  "fp": "4f5353d86025d5cd",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::351d6ebea8b99a6a": {
  "subjects": [],
  "subject_text": "인천 난민촌 쓰레기장\n거대한 쓰레기 산이 솟은 야외 폐기물 지대. 고철과 폐가전, 로봇 부품이 뒤섞여 있고 입구에는 ‘난민 거주지역’ 표지판이 걸려 있다.",
  "identity": "canonical",
  "scope_id": "L154",
  "scope_role": "location_exterior",
  "scope_sha": "476e0d3d155a1022"
 },
 "S6sh26::bgfirst_bg": {
  "input_fingerprint": "2b0b5a907c39d06c",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 쓰레기로 덧대어 지어진 거대한 난민촌 판잣집들을 배경으로 흙먼지 날리는 언덕을 걷고 있는 mid-stride 상태의 앰버, 라울, 위장한 찰리의 뒷모습 풀샷.\n\nLOCATION (lock): On a dusty hillside path leaving the refugee settlement dump, with extensive makeshift dwellings spread behind it.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Refugee settlement dwellings (An extensive settlement of makeshift homes patched with discarded materials) — Overlapping sides and fronts extend beyond the departing figures; used as Broad background scale that will become the focus of the following upward tilt; Route toward 페드로's home (Being traversed by the three companions) — Recedes from the lower foreground into the settlement; used as Connects the full-body walking figures with their destination.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light with restrained contrast keeps the departing figures distinct against the extensive settlement.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 쓰레기로 덧대어 지어진 거대한 난민촌 판잣집들을 배경으로 흙먼지 날리는 언덕을 걷고 있는 mid-stride 상태의 앰버, 라울, 위장한 찰리의 뒷모습 풀샷.\n\nLOCATION (lock): On a dusty hillside path leaving the refugee settlement dump, with extensive makeshift dwellings spread behind it.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Refugee settlement dwellings (An extensive settlement of makeshift homes patched with discarded materials) — Overlapping sides and fronts extend beyond the departing figures; used as Broad background scale that will become the focus of the following upward tilt; Route toward 페드로's home (Being traversed by the three companions) — Recedes from the lower foreground into the settlement; used as Connects the full-body walking figures with their destination.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light with restrained contrast keeps the departing figures distinct against the extensive settlement.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S6sh26__bgfirst_bg.png",
  "asset_id": "dc27e10d-eecc-46d8-bfa1-0d822e7b99c4",
  "input_asset_ids": [
   "e7cb1336-68ab-42f8-8caa-a8db9daa34a5",
   "c7293dea-4a59-4086-b333-889f46e571b8"
  ]
 },
 "S6sh26": {
  "input_fingerprint": "69fae035120fc26e",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쓰레기로 덧대어 지어진 거대한 난민촌 판잣집들을 배경으로 흙먼지 날리는 언덕을 걷고 있는 mid-stride 상태의 앰버, 라울, 위장한 찰리의 뒷모습 풀샷.\n\nLOCATION (lock): On a dusty hillside path leaving the refugee settlement dump, with extensive makeshift dwellings spread behind it. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Refugee settlement dwellings (An extensive settlement of makeshift homes patched with discarded materials) — Overlapping sides and fronts extend beyond the departing figures; used as Broad background scale that will become the focus of the following upward tilt; Route toward 페드로's home (Being traversed by the three companions) — Recedes from the lower foreground into the settlement; used as Connects the full-body walking figures with their destination.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light with restrained contrast keeps the departing figures distinct against the extensive settlement.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie is disguised in an old overcoat and hat over his dirty, aged gorilla-shaped body, with the netting and worn chest logo beneath the disguise. His blue-lit eyes remain unchanged, and the repaired jukebox was last left switched off. 앰버: Amber is walking away from the rubbish dump toward Pedro's home, still wearing her mask and waist tool pouch. 라울: Raul is walking toward Pedro's home after escaping notice.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쓰레기로 덧대어 지어진 거대한 난민촌 판잣집들을 배경으로 흙먼지 날리는 언덕을 걷고 있는 mid-stride 상태의 앰버, 라울, 위장한 찰리의 뒷모습 풀샷.\n\nLOCATION (lock): On a dusty hillside path leaving the refugee settlement dump, with extensive makeshift dwellings spread behind it. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Refugee settlement dwellings (An extensive settlement of makeshift homes patched with discarded materials) — Overlapping sides and fronts extend beyond the departing figures; used as Broad background scale that will become the focus of the following upward tilt; Route toward 페드로's home (Being traversed by the three companions) — Recedes from the lower foreground into the settlement; used as Connects the full-body walking figures with their destination.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light with restrained contrast keeps the departing figures distinct against the extensive settlement.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie is disguised in an old overcoat and hat over his dirty, aged gorilla-shaped body, with the netting and worn chest logo beneath the disguise. His blue-lit eyes remain unchanged, and the repaired jukebox was last left switched off. 앰버: Amber is walking away from the rubbish dump toward Pedro's home, still wearing her mask and waist tool pouch. 라울: Raul is walking toward Pedro's home after escaping notice.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쓰레기로 덧대어 지어진 거대한 난민촌 판잣집들을 배경으로 흙먼지 날리는 언덕을 걷고 있는 mid-stride 상태의 앰버, 라울, 위장한 찰리의 뒷모습 풀샷.\n\nLOCATION (lock): On a dusty hillside path leaving the refugee settlement dump, with extensive makeshift dwellings spread behind it. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Refugee settlement dwellings (An extensive settlement of makeshift homes patched with discarded materials) — Overlapping sides and fronts extend beyond the departing figures; used as Broad background scale that will become the focus of the following upward tilt; Route toward 페드로's home (Being traversed by the three companions) — Recedes from the lower foreground into the settlement; used as Connects the full-body walking figures with their destination.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light with restrained contrast keeps the departing figures distinct against the extensive settlement.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie is disguised in an old overcoat and hat over his dirty, aged gorilla-shaped body, with the netting and worn chest logo beneath the disguise. His blue-lit eyes remain unchanged, and the repaired jukebox was last left switched off. 앰버: Amber is walking away from the rubbish dump toward Pedro's home, still wearing her mask and waist tool pouch. 라울: Raul is walking toward Pedro's home after escaping notice.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S6sh26__bgfirst_bg.png",
     "asset_id": "dc27e10d-eecc-46d8-bfa1-0d822e7b99c4",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S6sh26.png",
     "asset_id": "e7cb1336-68ab-42f8-8caa-a8db9daa34a5",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163202>",
     "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_refugee_dump_sel.png",
     "asset_id": "c7293dea-4a59-4086-b333-889f46e571b8",
     "role": "location_seed_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163202>",
     "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "B",
    "direction": "앰버는 왼쪽, 찰리는 중앙, 라울은 오른쪽에서 카메라 반대편 난민촌 내부로 걷는다. 앰버의 고개는 오른쪽 찰리 쪽으로 조금 돌아가 있고, 찰리도 오른쪽으로 돌아 얼굴 일부가 드러난다. 라울은 앞쪽 길을 향한다. 페드로의 집 자체는 식별되지 않지만 이동 방향은 주거지로 연결된다.",
    "built_space": "전경 흙길 하나가 아래쪽 난민촌 골목으로 이어지고, 양옆과 먼 배경에 수십 채의 철판 판잣집이 겹쳐 있다. 큰 쓰레기 더미는 왼쪽 위에 있으며 오른쪽에는 집들이 밀집한다. 참조의 쓰레기 사면과 금속 주거지라는 재료 관계는 남아 있지만, 참조에서 보이는 넓은 도로와 거대한 연속 사면의 정확한 배치는 확인되지 않는다. 출입문 구조물은 보이지 않으나 다른 촬영 위치에서 제외될 수 있다. 세 인물은 모두 길 위에 있다.",
    "entities": "인물은 정확히 셋이다. 앰버는 어린 금발 소녀로 남색 반소매, 마스크와 허리 도구 주머니가 보여 부합한다. 라울은 갈색 피부의 어린 소년이며 뒤로 묶은 곱슬머리와 남색 반소매가 참조에 가깝다. 두 아이의 얼굴과 세부 혈통은 뒷모습으로 확인할 수 없다. 찰리는 낡은 외투와 모자를 착용하고 기계 손과 청색 발광부가 보이지만, 어깨와 팔이 상대적으로 가늘고 다리가 길어 지정된 고릴라형 몸체와 다르다. 가슴 그물과 문양은 외투에 가려져 평가하지 않는다. 주크박스는 장면에 없으며 추가할 필요가 없다. 읽을 수 있는 글씨는 없다.",
    "hard_violations": [],
    "physics": "앰버는 오른발을 땅에 두고 왼발을 들어 올렸고, 라울과 찰리는 왼발로 지지하면서 오른발을 옮기는 보행 순간이다. 몸의 지지와 발밑 먼지가 자연스럽게 연결된다. 모자는 머리에 얹혀 있고 외투는 어깨에, 도구 주머니는 허리띠에 지지된다. 공중에 근거 없이 떠 있는 몸이나 물건은 보이지 않는다."
   },
   {
    "label": "A",
    "direction": "왼쪽 앰버, 중앙 라울, 오른쪽 찰리가 모두 카메라에서 멀어지며 화면 중앙의 오르막 길을 따라 난민촌으로 향한다. 라울은 정면의 길을 보고 앰버는 약간 오른쪽을 향한다. 찰리의 얼굴 방향은 모자와 깃에 가려 불명확하며, 뒤쪽 목 부근에 보이는 청색 발광부는 눈의 위치로 읽기에 다소 모호하다. 무기나 겨냥하는 소품은 없다.",
    "built_space": "전경에서 시작한 흙길 하나가 화면 중앙을 지나 언덕 위로 굽어 올라간다. 양쪽에는 수십 채의 낮은 철판 주택, 왼쪽에는 연속 철판 울타리 한 줄, 길 주변에는 여러 전신주와 전선이 있다. 세 인물은 모두 같은 길 위에 서 있고 건축물과 인물의 크기 관계도 자연스럽다. 다만 쓰레기가 집들 사이 사면 전반에 흩어져 있어, 참조 장소의 도로 한쪽을 압도하는 거대한 연속 쓰레기 산과 동일한 배치로 보이지 않는다. 두 후보 모두 정확한 장소 잠금은 충분히 재현하지 못했다.",
    "entities": "인물은 지정된 셋뿐이다. 앰버는 어린 금발 소녀이며 남색 반소매, 옆으로 보이는 마스크와 허리 도구 주머니가 맞는다. 라울은 어린 체격과 뒤로 묶은 머리가 부합하지만 참조의 남색 반소매 대신 회청색 후드 상의를 입었다. 얼굴이 가려져 두 아이의 세부 얼굴 정체성은 판단할 수 없다. 찰리는 낡은 모자와 외투 아래 넓은 등, 육중하고 긴 팔, 큰 기계 손과 짧고 굵은 다리가 보여 A보다 지정 체격에 가깝다. 노출된 장갑은 참조보다 회색에 가깝다. 가슴 문양과 그물은 외투 안에 가려져 있다. 읽을 수 있는 글씨나 별도 인물은 없다.",
    "hard_violations": [],
    "physics": "앰버와 라울은 한쪽 발로 체중을 지지하고 반대쪽 뒤꿈치를 들어 보행하고 있다. 찰리도 왼발 쪽에 체중을 두고 오른발을 옮기는 자세이며 발밑 먼지가 지면 접촉을 뒷받침한다. 외투는 넓은 어깨에서 내려오고 허리띠에 묶여 있으며 모자와 도구 주머니도 각각 머리와 허리에 지지된다. 지지 없는 부유나 불가능한 도약은 없다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": null,
     "normalized": null,
     "ok": false
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "세 인물의 전신 보행과 먼지는 잘 보이지만, 인물이 배경보다 크게 강조되고 찰리가 긴 다리의 보통 성인 체형에 가까워 지정된 와이드 구성과 고릴라형 체격에서 B보다 뒤진다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "난민촌으로 이어지는 길과 세 인물의 뒷모습 전신을 넓게 담고 찰리의 육중한 체격도 더 충실하지만, 참조 장소의 거대한 연속 쓰레기 사면과 라울의 의상은 일치하지 않는다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버는 왼쪽, 찰리는 중앙, 라울은 오른쪽에서 카메라 반대편 난민촌 내부로 걷는다. 앰버의 고개는 오른쪽 찰리 쪽으로 조금 돌아가 있고, 찰리도 오른쪽으로 돌아 얼굴 일부가 드러난다. 라울은 앞쪽 길을 향한다. 페드로의 집 자체는 식별되지 않지만 이동 방향은 주거지로 연결된다.",
        "built_space": "전경 흙길 하나가 아래쪽 난민촌 골목으로 이어지고, 양옆과 먼 배경에 수십 채의 철판 판잣집이 겹쳐 있다. 큰 쓰레기 더미는 왼쪽 위에 있으며 오른쪽에는 집들이 밀집한다. 참조의 쓰레기 사면과 금속 주거지라는 재료 관계는 남아 있지만, 참조에서 보이는 넓은 도로와 거대한 연속 사면의 정확한 배치는 확인되지 않는다. 출입문 구조물은 보이지 않으나 다른 촬영 위치에서 제외될 수 있다. 세 인물은 모두 길 위에 있다.",
        "entities": "인물은 정확히 셋이다. 앰버는 어린 금발 소녀로 남색 반소매, 마스크와 허리 도구 주머니가 보여 부합한다. 라울은 갈색 피부의 어린 소년이며 뒤로 묶은 곱슬머리와 남색 반소매가 참조에 가깝다. 두 아이의 얼굴과 세부 혈통은 뒷모습으로 확인할 수 없다. 찰리는 낡은 외투와 모자를 착용하고 기계 손과 청색 발광부가 보이지만, 어깨와 팔이 상대적으로 가늘고 다리가 길어 지정된 고릴라형 몸체와 다르다. 가슴 그물과 문양은 외투에 가려져 평가하지 않는다. 주크박스는 장면에 없으며 추가할 필요가 없다. 읽을 수 있는 글씨는 없다.",
        "hard_violations": [],
        "physics": "앰버는 오른발을 땅에 두고 왼발을 들어 올렸고, 라울과 찰리는 왼발로 지지하면서 오른발을 옮기는 보행 순간이다. 몸의 지지와 발밑 먼지가 자연스럽게 연결된다. 모자는 머리에 얹혀 있고 외투는 어깨에, 도구 주머니는 허리띠에 지지된다. 공중에 근거 없이 떠 있는 몸이나 물건은 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "왼쪽 앰버, 중앙 라울, 오른쪽 찰리가 모두 카메라에서 멀어지며 화면 중앙의 오르막 길을 따라 난민촌으로 향한다. 라울은 정면의 길을 보고 앰버는 약간 오른쪽을 향한다. 찰리의 얼굴 방향은 모자와 깃에 가려 불명확하며, 뒤쪽 목 부근에 보이는 청색 발광부는 눈의 위치로 읽기에 다소 모호하다. 무기나 겨냥하는 소품은 없다.",
        "built_space": "전경에서 시작한 흙길 하나가 화면 중앙을 지나 언덕 위로 굽어 올라간다. 양쪽에는 수십 채의 낮은 철판 주택, 왼쪽에는 연속 철판 울타리 한 줄, 길 주변에는 여러 전신주와 전선이 있다. 세 인물은 모두 같은 길 위에 서 있고 건축물과 인물의 크기 관계도 자연스럽다. 다만 쓰레기가 집들 사이 사면 전반에 흩어져 있어, 참조 장소의 도로 한쪽을 압도하는 거대한 연속 쓰레기 산과 동일한 배치로 보이지 않는다. 두 후보 모두 정확한 장소 잠금은 충분히 재현하지 못했다.",
        "entities": "인물은 지정된 셋뿐이다. 앰버는 어린 금발 소녀이며 남색 반소매, 옆으로 보이는 마스크와 허리 도구 주머니가 맞는다. 라울은 어린 체격과 뒤로 묶은 머리가 부합하지만 참조의 남색 반소매 대신 회청색 후드 상의를 입었다. 얼굴이 가려져 두 아이의 세부 얼굴 정체성은 판단할 수 없다. 찰리는 낡은 모자와 외투 아래 넓은 등, 육중하고 긴 팔, 큰 기계 손과 짧고 굵은 다리가 보여 A보다 지정 체격에 가깝다. 노출된 장갑은 참조보다 회색에 가깝다. 가슴 문양과 그물은 외투 안에 가려져 있다. 읽을 수 있는 글씨나 별도 인물은 없다.",
        "hard_violations": [],
        "physics": "앰버와 라울은 한쪽 발로 체중을 지지하고 반대쪽 뒤꿈치를 들어 보행하고 있다. 찰리도 왼발 쪽에 체중을 두고 오른발을 옮기는 자세이며 발밑 먼지가 지면 접촉을 뒷받침한다. 외투는 넓은 어깨에서 내려오고 허리띠에 묶여 있으며 모자와 도구 주머니도 각각 머리와 허리에 지지된다. 지지 없는 부유나 불가능한 도약은 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "세 인물의 전신 보행과 먼지는 잘 보이지만, 인물이 배경보다 크게 강조되고 찰리가 긴 다리의 보통 성인 체형에 가까워 지정된 와이드 구성과 고릴라형 체격에서 B보다 뒤진다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "난민촌으로 이어지는 길과 세 인물의 뒷모습 전신을 넓게 담고 찰리의 육중한 체격도 더 충실하지만, 참조 장소의 거대한 연속 쓰레기 사면과 라울의 의상은 일치하지 않는다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "앰버는 왼쪽, 찰리는 중앙, 라울은 오른쪽에서 카메라 반대편 난민촌 내부로 걷는다. 앰버의 고개는 오른쪽 찰리 쪽으로 조금 돌아가 있고, 찰리도 오른쪽으로 돌아 얼굴 일부가 드러난다. 라울은 앞쪽 길을 향한다. 페드로의 집 자체는 식별되지 않지만 이동 방향은 주거지로 연결된다.",
        "built_space": "전경 흙길 하나가 아래쪽 난민촌 골목으로 이어지고, 양옆과 먼 배경에 수십 채의 철판 판잣집이 겹쳐 있다. 큰 쓰레기 더미는 왼쪽 위에 있으며 오른쪽에는 집들이 밀집한다. 참조의 쓰레기 사면과 금속 주거지라는 재료 관계는 남아 있지만, 참조에서 보이는 넓은 도로와 거대한 연속 사면의 정확한 배치는 확인되지 않는다. 출입문 구조물은 보이지 않으나 다른 촬영 위치에서 제외될 수 있다. 세 인물은 모두 길 위에 있다.",
        "entities": "인물은 정확히 셋이다. 앰버는 어린 금발 소녀로 남색 반소매, 마스크와 허리 도구 주머니가 보여 부합한다. 라울은 갈색 피부의 어린 소년이며 뒤로 묶은 곱슬머리와 남색 반소매가 참조에 가깝다. 두 아이의 얼굴과 세부 혈통은 뒷모습으로 확인할 수 없다. 찰리는 낡은 외투와 모자를 착용하고 기계 손과 청색 발광부가 보이지만, 어깨와 팔이 상대적으로 가늘고 다리가 길어 지정된 고릴라형 몸체와 다르다. 가슴 그물과 문양은 외투에 가려져 평가하지 않는다. 주크박스는 장면에 없으며 추가할 필요가 없다. 읽을 수 있는 글씨는 없다.",
        "hard_violations": [],
        "physics": "앰버는 오른발을 땅에 두고 왼발을 들어 올렸고, 라울과 찰리는 왼발로 지지하면서 오른발을 옮기는 보행 순간이다. 몸의 지지와 발밑 먼지가 자연스럽게 연결된다. 모자는 머리에 얹혀 있고 외투는 어깨에, 도구 주머니는 허리띠에 지지된다. 공중에 근거 없이 떠 있는 몸이나 물건은 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "왼쪽 앰버, 중앙 라울, 오른쪽 찰리가 모두 카메라에서 멀어지며 화면 중앙의 오르막 길을 따라 난민촌으로 향한다. 라울은 정면의 길을 보고 앰버는 약간 오른쪽을 향한다. 찰리의 얼굴 방향은 모자와 깃에 가려 불명확하며, 뒤쪽 목 부근에 보이는 청색 발광부는 눈의 위치로 읽기에 다소 모호하다. 무기나 겨냥하는 소품은 없다.",
        "built_space": "전경에서 시작한 흙길 하나가 화면 중앙을 지나 언덕 위로 굽어 올라간다. 양쪽에는 수십 채의 낮은 철판 주택, 왼쪽에는 연속 철판 울타리 한 줄, 길 주변에는 여러 전신주와 전선이 있다. 세 인물은 모두 같은 길 위에 서 있고 건축물과 인물의 크기 관계도 자연스럽다. 다만 쓰레기가 집들 사이 사면 전반에 흩어져 있어, 참조 장소의 도로 한쪽을 압도하는 거대한 연속 쓰레기 산과 동일한 배치로 보이지 않는다. 두 후보 모두 정확한 장소 잠금은 충분히 재현하지 못했다.",
        "entities": "인물은 지정된 셋뿐이다. 앰버는 어린 금발 소녀이며 남색 반소매, 옆으로 보이는 마스크와 허리 도구 주머니가 맞는다. 라울은 어린 체격과 뒤로 묶은 머리가 부합하지만 참조의 남색 반소매 대신 회청색 후드 상의를 입었다. 얼굴이 가려져 두 아이의 세부 얼굴 정체성은 판단할 수 없다. 찰리는 낡은 모자와 외투 아래 넓은 등, 육중하고 긴 팔, 큰 기계 손과 짧고 굵은 다리가 보여 A보다 지정 체격에 가깝다. 노출된 장갑은 참조보다 회색에 가깝다. 가슴 문양과 그물은 외투 안에 가려져 있다. 읽을 수 있는 글씨나 별도 인물은 없다.",
        "hard_violations": [],
        "physics": "앰버와 라울은 한쪽 발로 체중을 지지하고 반대쪽 뒤꿈치를 들어 보행하고 있다. 찰리도 왼발 쪽에 체중을 두고 오른발을 옮기는 자세이며 발밑 먼지가 지면 접촉을 뒷받침한다. 외투는 넓은 어깨에서 내려오고 허리띠에 묶여 있으며 모자와 도구 주머니도 각각 머리와 허리에 지지된다. 지지 없는 부유나 불가능한 도약은 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gemini-pro"
   ],
   "route": "single_reverse"
  },
  "totals": {
   "B": 6,
   "A": 7
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 6,
    "verdict_ko": "세 인물의 전신 보행과 먼지는 잘 보이지만, 인물이 배경보다 크게 강조되고 찰리가 긴 다리의 보통 성인 체형에 가까워 지정된 와이드 구성과 고릴라형 체격에서 B보다 뒤진다."
   },
   {
    "label": "A",
    "score": 7,
    "verdict_ko": "난민촌으로 이어지는 길과 세 인물의 뒷모습 전신을 넓게 담고 찰리의 육중한 체격도 더 충실하지만, 참조 장소의 거대한 연속 쓰레기 사면과 라울의 의상은 일치하지 않는다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_refugee_dump_sel.png",
    "asset_id": "c7293dea-4a59-4086-b333-889f46e571b8",
    "role": "location_seed_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163202>",
    "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-a6cf-7755-afdb-6e81a944f955",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S6sh26__bgfirst_bg.png",
   "bg_asset_id": "dc27e10d-eecc-46d8-bfa1-0d822e7b99c4",
   "bg_record_key": "S6sh26::bgfirst_bg",
   "chain_winner": true,
   "authority": "seed_bg"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  },
  "lane_policy": "ab_select_ready"
 },
 "S6sh26::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:01:24.789757+00:00",
  "fingerprint": "5dff3cf211fd591d611ecdd6f3839d6497a95852d43fdc7a3cefc0fefe191f94",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S6sh26_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S6sh26_sel.png",
  "source_sha256": "c7a2d05bf78813f0b2a3e7488f6805cac961650d8d27b158bb8f756ac3a13328",
  "file": "S6sh26_cine.png",
  "staged_sha256": "651d9f71525fa96873693ea76f7a4bee6b6857d71eedd87ad970d1d984f9b514",
  "latency_ms": 16346
 },
 "S7sh11::signage": {
  "fp": "0b0e716578a800e7",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::bb81f5972953eb12": {
  "subjects": [],
  "subject_text": "인천 난민촌 인공제방과 공사장\n바다를 막아선 거대한 콘크리트 제방과 보수 공사장. 벽 곳곳의 깊은 균열과 젖은 지지대, 적재된 보수 자재와 돌무더기가 보인다.",
  "identity": "canonical",
  "scope_id": "L156",
  "scope_role": "location_exterior",
  "scope_sha": "2bf570d14c3dfff2"
 },
 "groupbg::seawall_worksite": {
  "input_fingerprint": "aff63496c4b484a8",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "seawall_worksite",
    "tags": [
     "S13sh13",
     "S13sh20",
     "S33sh14",
     "S7sh11",
     "S7sh13",
     "S7sh15"
    ]
   },
   "context_sig": "f22b22c1cb13823e"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On the ground beside the cracked seawall at the refugee settlement's suspended repair site.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n박철진이 탑승한 전투 헬기 내부: 아래 지상을 지휘 통제할 수 있는 장비가 갖춰진 군용 헬기 조종석 옆자리. (특징: 복잡한 계기판과 유리창; 정보가 표시된 휴대용 태블릿 모니터; 아래로 넓게 펼쳐진 지상 시야) / 인천 난민촌 인공제방과 공사장: 바다를 막고 있으나 곳곳에 심한 균열이 간 거대한 콘크리트 장벽과 그 앞의 작업 구역. (특징: 표면에 물이 스며들고 굵은 금이 간 거대한 제방 벽; 자재들이 쌓여 있는 공사 현장; 물웅덩이가 파인 질척이는 흙바닥; 지게차와 주차된 군용 트럭)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 제방 벽 곳곳이 심하게 금이 가 있다.\n- 어둠 속이지만 여기저기 제방 갈라진 틈이 보인다. 지지대에도 물이 흐르고.\n- 제방 근처에 있던 보수작업을 위한 중장비 기계들과 건설 장비들이 한 순간에 바닷물에 휩쓸려 간다.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On the ground beside the cracked seawall at the refugee settlement's suspended repair site.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n박철진이 탑승한 전투 헬기 내부: 아래 지상을 지휘 통제할 수 있는 장비가 갖춰진 군용 헬기 조종석 옆자리. (특징: 복잡한 계기판과 유리창; 정보가 표시된 휴대용 태블릿 모니터; 아래로 넓게 펼쳐진 지상 시야) / 인천 난민촌 인공제방과 공사장: 바다를 막고 있으나 곳곳에 심한 균열이 간 거대한 콘크리트 장벽과 그 앞의 작업 구역. (특징: 표면에 물이 스며들고 굵은 금이 간 거대한 제방 벽; 자재들이 쌓여 있는 공사 현장; 물웅덩이가 파인 질척이는 흙바닥; 지게차와 주차된 군용 트럭)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 제방 벽 곳곳이 심하게 금이 가 있다.\n- 어둠 속이지만 여기저기 제방 갈라진 틈이 보인다. 지지대에도 물이 흐르고.\n- 제방 근처에 있던 보수작업을 위한 중장비 기계들과 건설 장비들이 한 순간에 바닷물에 휩쓸려 간다.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_seawall_worksite_44c24f.png",
  "asset_id": "13d1bb52-f02a-4d90-ae80-f1d788351b63",
  "input_asset_ids": [
   "d68199d9-a355-482b-9cd4-df9585548d5e"
  ],
  "origin_tag": "S7sh11",
  "place_text": "On the ground beside the cracked seawall at the refugee settlement's suspended repair site.",
  "origin_inputs": {
   "place_text": "On the ground beside the cracked seawall at the refugee settlement's suspended repair site.",
   "time_of_day_en": "day",
   "conti_asset_id": "d68199d9-a355-482b-9cd4-df9585548d5e"
  }
 },
 "S7sh11::bgfirst_bg": {
  "input_fingerprint": "da74c208a9ecfc8e",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 심하게 금이 간 방벽 쪽을 손가락으로 가리키며 눈을 부릅뜬 미연의 절박한 얼굴.\n\nLOCATION (lock): On the ground beside the cracked seawall at the refugee settlement's suspended repair site.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Cracked embankment wall indicated by the pointing arm in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Cracked embankment wall (Severely cracked, with water leaking through the wall) — A narrow section of the damaged face is visible beyond the pointing arm on screen right; used as Provides visible evidence for the appeal without competing with the face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light preserves facial detail and restrained contrast without embellishing the urgency with a new lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 심하게 금이 간 방벽 쪽을 손가락으로 가리키며 눈을 부릅뜬 미연의 절박한 얼굴.\n\nLOCATION (lock): On the ground beside the cracked seawall at the refugee settlement's suspended repair site.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Cracked embankment wall indicated by the pointing arm in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Cracked embankment wall (Severely cracked, with water leaking through the wall) — A narrow section of the damaged face is visible beyond the pointing arm on screen right; used as Provides visible evidence for the appeal without competing with the face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light preserves facial detail and restrained contrast without embellishing the urgency with a new lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S7sh11__bgfirst_bg.png",
  "asset_id": "0e42b760-1b39-40e2-a02b-2089485b8bc4",
  "input_asset_ids": [
   "d68199d9-a355-482b-9cd4-df9585548d5e",
   "13d1bb52-f02a-4d90-ae80-f1d788351b63"
  ]
 },
 "S7sh11": {
  "input_fingerprint": "b8334088f47f0164",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 심하게 금이 간 방벽 쪽을 손가락으로 가리키며 눈을 부릅뜬 미연의 절박한 얼굴.\n\nLOCATION (lock): On the ground beside the cracked seawall at the refugee settlement's suspended repair site. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Cracked embankment wall indicated by the pointing arm in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Cracked embankment wall (Severely cracked, with water leaking through the wall) — A narrow section of the damaged face is visible beyond the pointing arm on screen right; used as Provides visible evidence for the appeal without competing with the face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light preserves facial detail and restrained contrast without embellishing the urgency with a new lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall has severe cracks and visible water leakage, with puddles on the worksite ground. The military truck remains stopped in front of the forklift amid the halted repair equipment and materials. 미연: Miyeon is out of the forklift and standing among the gathered workers, urgently indicating the leaking wall.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 심하게 금이 간 방벽 쪽을 손가락으로 가리키며 눈을 부릅뜬 미연의 절박한 얼굴.\n\nLOCATION (lock): On the ground beside the cracked seawall at the refugee settlement's suspended repair site. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Cracked embankment wall indicated by the pointing arm in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Cracked embankment wall (Severely cracked, with water leaking through the wall) — A narrow section of the damaged face is visible beyond the pointing arm on screen right; used as Provides visible evidence for the appeal without competing with the face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light preserves facial detail and restrained contrast without embellishing the urgency with a new lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall has severe cracks and visible water leakage, with puddles on the worksite ground. The military truck remains stopped in front of the forklift amid the halted repair equipment and materials. 미연: Miyeon is out of the forklift and standing among the gathered workers, urgently indicating the leaking wall.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 심하게 금이 간 방벽 쪽을 손가락으로 가리키며 눈을 부릅뜬 미연의 절박한 얼굴.\n\nLOCATION (lock): On the ground beside the cracked seawall at the refugee settlement's suspended repair site. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Cracked embankment wall indicated by the pointing arm in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Cracked embankment wall (Severely cracked, with water leaking through the wall) — A narrow section of the damaged face is visible beyond the pointing arm on screen right; used as Provides visible evidence for the appeal without competing with the face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light preserves facial detail and restrained contrast without embellishing the urgency with a new lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall has severe cracks and visible water leakage, with puddles on the worksite ground. The military truck remains stopped in front of the forklift amid the halted repair equipment and materials. 미연: Miyeon is out of the forklift and standing among the gathered workers, urgently indicating the leaking wall.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S7sh11__bgfirst_bg.png",
     "asset_id": "0e42b760-1b39-40e2-a02b-2089485b8bc4",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S7sh11.png",
     "asset_id": "d68199d9-a355-482b-9cd4-df9585548d5e",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1113064>",
     "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_seawall_worksite_44c24f.png",
     "asset_id": "13d1bb52-f02a-4d90-ae80-f1d788351b63",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1113064>",
     "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "미연은 우측 방벽을 손가락으로 가리키고 있으나 시선은 화면 좌측 밖을 향함.",
    "built_space": "우측에 금이 가고 물이 새는 방벽, 바닥의 웅덩이, 뒤편의 트럭 및 지게차가 레퍼런스에 맞게 배치됨.",
    "entities": "미연은 레퍼런스와 일치하나, 숏 텍스트에 없는 인물 4명(전경 1명, 배경 3명)이 추가됨. 미연의 머리 왼쪽 부근에 검은색 화살표 마커가 그려져 있음.",
    "hard_violations": [
     "[gemini-pro] invented people (숏 텍스트에 없는 인물 4명 추가)",
     "[gemini-pro] leaked markers (화면 좌측 중앙에 화살표 마커 누출)",
     "[gpt-high] 미연만 등장해야 하는 장면에 전경 남성과 배경 작업자 등 추가 인물 다섯 명이 등장합니다.",
     "[gpt-high] 왼쪽 배경 작업자 얼굴 옆에 장면 내 물체로 설명되지 않는 검은 화살표 표식이 노출되어 있습니다."
    ],
    "physics": "인물들은 지면에 안정적으로 서 있으며, 물은 중력에 맞게 흘러내림."
   },
   {
    "label": "B",
    "direction": "미연은 우측 방벽을 가리키며 시선도 그 방향을 향하고 있음.",
    "built_space": "우측에 균열이 생겨 물이 쏟아지는 방벽과 진흙 바닥이 프레임에 알맞게 구성됨.",
    "entities": "미연은 레퍼런스와 일치하며, 지시대로 다른 인물은 등장하지 않음.",
    "hard_violations": [],
    "physics": "인물의 자세와 팔의 무게 중심이 자연스럽고, 방벽에서 새어나오는 물줄기도 물리적으로 타당함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "프롬프트가 요구한 클로즈업 숏 크기, 피사체의 배치, 배경의 균열된 방벽 묘사를 충실히 구현했으며 엑스트라를 배제하라는 지시를 잘 따랐습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정되지 않은 인물들이 다수 등장하고, 화면 좌측에 마커 기호가 누출되어 치명적인 오류가 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "미연은 우측 방벽을 손가락으로 가리키고 있으나 시선은 화면 좌측 밖을 향함.",
        "built_space": "우측에 금이 가고 물이 새는 방벽, 바닥의 웅덩이, 뒤편의 트럭 및 지게차가 레퍼런스에 맞게 배치됨.",
        "entities": "미연은 레퍼런스와 일치하나, 숏 텍스트에 없는 인물 4명(전경 1명, 배경 3명)이 추가됨. 미연의 머리 왼쪽 부근에 검은색 화살표 마커가 그려져 있음.",
        "hard_violations": [
         "invented people (숏 텍스트에 없는 인물 4명 추가)",
         "leaked markers (화면 좌측 중앙에 화살표 마커 누출)"
        ],
        "physics": "인물들은 지면에 안정적으로 서 있으며, 물은 중력에 맞게 흘러내림."
       },
       {
        "label": "B",
        "direction": "미연은 우측 방벽을 가리키며 시선도 그 방향을 향하고 있음.",
        "built_space": "우측에 균열이 생겨 물이 쏟아지는 방벽과 진흙 바닥이 프레임에 알맞게 구성됨.",
        "entities": "미연은 레퍼런스와 일치하며, 지시대로 다른 인물은 등장하지 않음.",
        "hard_violations": [],
        "physics": "인물의 자세와 팔의 무게 중심이 자연스럽고, 방벽에서 새어나오는 물줄기도 물리적으로 타당함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "프롬프트가 요구한 클로즈업 숏 크기, 피사체의 배치, 배경의 균열된 방벽 묘사를 충실히 구현했으며 엑스트라를 배제하라는 지시를 잘 따랐습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정되지 않은 인물들이 다수 등장하고, 화면 좌측에 마커 기호가 누출되어 치명적인 오류가 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "미연은 우측 방벽을 손가락으로 가리키고 있으나 시선은 화면 좌측 밖을 향함.",
        "built_space": "우측에 금이 가고 물이 새는 방벽, 바닥의 웅덩이, 뒤편의 트럭 및 지게차가 레퍼런스에 맞게 배치됨.",
        "entities": "미연은 레퍼런스와 일치하나, 숏 텍스트에 없는 인물 4명(전경 1명, 배경 3명)이 추가됨. 미연의 머리 왼쪽 부근에 검은색 화살표 마커가 그려져 있음.",
        "hard_violations": [
         "invented people (숏 텍스트에 없는 인물 4명 추가)",
         "leaked markers (화면 좌측 중앙에 화살표 마커 누출)"
        ],
        "physics": "인물들은 지면에 안정적으로 서 있으며, 물은 중력에 맞게 흘러내림."
       },
       {
        "label": "B",
        "direction": "미연은 우측 방벽을 가리키며 시선도 그 방향을 향하고 있음.",
        "built_space": "우측에 균열이 생겨 물이 쏟아지는 방벽과 진흙 바닥이 프레임에 알맞게 구성됨.",
        "entities": "미연은 레퍼런스와 일치하며, 지시대로 다른 인물은 등장하지 않음.",
        "hard_violations": [],
        "physics": "인물의 자세와 팔의 무게 중심이 자연스럽고, 방벽에서 새어나오는 물줄기도 물리적으로 타당함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "미연만 등장하는 근접 구도와 균열을 가리키는 동작은 충실하지만, 방벽이 요구한 좁은 배경보다 넓고 눈을 부릅뜬 절박함은 다소 약합니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "누수 지시와 절박한 표정은 명확하지만, 금지된 추가 인물들과 화살표가 등장하고 구도도 얼굴 클로즈업에서 넓어진 중대한 위반입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "미연의 검지는 화면 오른쪽 위로 뻗어 손끝 너머 방벽의 굵은 세로 균열을 가리킵니다. 눈은 카메라 오른쪽 화면 밖을 향하며, 시선이 닿는 상대는 보이지 않습니다. 무기나 이동 중인 물체는 없습니다.",
        "built_space": "미연은 방벽 바로 옆 전경에 있고, 오른쪽에는 균열과 누수가 있는 콘크리트 방벽 한 면이 이어집니다. 상단 난간 일부와 벽 아래 잔해, 왼쪽의 젖은 작업장 바닥이 보입니다. 지상 시점과 장소의 재질은 참고에 부합하지만, 손 너머 좁은 부분만 보여야 할 방벽이 화면 오른쪽의 상당 부분을 차지합니다. 중복 설비나 불가능한 반사는 보이지 않습니다.",
        "entities": "보이는 사람은 미연 한 명입니다. 중년 동아시아 여성의 얼굴 윤곽, 검은 단발머리, 남색 둥근목 상의가 인물 참고와 잘 맞습니다. 눈과 입을 약간 벌린 걱정스러운 표정이지만, 요구된 강한 절박함과 부릅뜬 눈은 비교적 약합니다. 심한 균열, 흐르는 물, 바닥의 고인 물은 확인됩니다. 군용 트럭과 지게차는 이 근접 구도에서 명확하게 식별되지 않으며, 이를 누락으로 판단하지 않습니다. 읽을 수 있는 글자나 추가 인물은 없습니다.",
        "hard_violations": [],
        "physics": "뻗은 팔은 어깨에서 손목과 손까지 자연스럽게 연결되어 있고 검지로 가리키는 자세가 가능합니다. 발은 프레임 밖이지만 상체가 떠 있다는 징후는 없습니다. 물은 균열에서 벽을 따라 아래로 흐르고 잔해는 벽 아래 지면에 놓여 있습니다. 지지 없이 떠 있는 물체는 보이지 않습니다."
       },
       {
        "label": "B",
        "direction": "미연의 팔과 검지는 오른쪽 방벽의 누수 균열을 향합니다. 미연은 눈을 크게 뜨고 왼쪽 전경 남성의 얼굴을 바라보며 호소합니다. 뒤쪽 작업자들은 대체로 미연과 전경의 대화 상대 쪽을 봅니다. 왼쪽 배경 작업자 얼굴 옆에는 왼쪽을 향하는 작은 검은 화살표가 보입니다.",
        "built_space": "오른쪽의 방벽 한 면과 상단 난간 한 줄, 그 아래 잔해와 파이프 적재물, 젖은 작업장 바닥이 보입니다. 중앙 배경에는 군용 트럭 한 대와 그 오른쪽 뒤 지게차 한 대가 있어 참고 장소의 배치와 잘 맞습니다. 다만 전경 남성의 어깨 너머로 미연의 상체와 작업자들의 다리까지 담아, 지정된 얼굴 클로즈업보다 훨씬 넓습니다. 방벽 역시 좁은 배경 단면이 아니라 긴 작업장 전경으로 제시됩니다.",
        "entities": "미연의 중년 얼굴, 검은 머리와 남색 상의는 참고와 대체로 맞고, 크게 뜬 눈과 열린 입은 절박함을 전달합니다. 그러나 왼쪽 전경 남성 한 명, 왼쪽 배경의 안전모 인물 두 명, 중앙 배경의 안전모 남성 두 명으로 미연 외에 다섯 명이 추가되어 있습니다. 균열, 누수, 물웅덩이, 정차한 트럭과 지게차는 확인됩니다. 읽을 수 있는 문자는 없지만 작은 화살표 표식이 있습니다.",
        "hard_violations": [
         "미연만 등장해야 하는 장면에 전경 남성과 배경 작업자 등 추가 인물 다섯 명이 등장합니다.",
         "왼쪽 배경 작업자 얼굴 옆에 장면 내 물체로 설명되지 않는 검은 화살표 표식이 노출되어 있습니다."
        ],
        "physics": "미연은 상체를 앞으로 기울인 채 어깨와 연결된 팔을 옆으로 뻗고 있어 가능한 동작입니다. 미연의 발은 프레임 밖이며 부유의 증거는 없습니다. 뒤쪽 작업자들의 보이는 신발은 지면에 닿고, 차량은 바퀴로 지면에 지지됩니다. 파이프는 적재대와 지면에 놓여 있으며 누수는 아래로 흐릅니다. 지지 없이 떠 있는 신체나 물체는 없습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "미연만 등장하는 근접 구도와 균열을 가리키는 동작은 충실하지만, 방벽이 요구한 좁은 배경보다 넓고 눈을 부릅뜬 절박함은 다소 약합니다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "누수 지시와 절박한 표정은 명확하지만, 금지된 추가 인물들과 화살표가 등장하고 구도도 얼굴 클로즈업에서 넓어진 중대한 위반입니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "미연의 검지는 화면 오른쪽 위로 뻗어 손끝 너머 방벽의 굵은 세로 균열을 가리킵니다. 눈은 카메라 오른쪽 화면 밖을 향하며, 시선이 닿는 상대는 보이지 않습니다. 무기나 이동 중인 물체는 없습니다.",
        "built_space": "미연은 방벽 바로 옆 전경에 있고, 오른쪽에는 균열과 누수가 있는 콘크리트 방벽 한 면이 이어집니다. 상단 난간 일부와 벽 아래 잔해, 왼쪽의 젖은 작업장 바닥이 보입니다. 지상 시점과 장소의 재질은 참고에 부합하지만, 손 너머 좁은 부분만 보여야 할 방벽이 화면 오른쪽의 상당 부분을 차지합니다. 중복 설비나 불가능한 반사는 보이지 않습니다.",
        "entities": "보이는 사람은 미연 한 명입니다. 중년 동아시아 여성의 얼굴 윤곽, 검은 단발머리, 남색 둥근목 상의가 인물 참고와 잘 맞습니다. 눈과 입을 약간 벌린 걱정스러운 표정이지만, 요구된 강한 절박함과 부릅뜬 눈은 비교적 약합니다. 심한 균열, 흐르는 물, 바닥의 고인 물은 확인됩니다. 군용 트럭과 지게차는 이 근접 구도에서 명확하게 식별되지 않으며, 이를 누락으로 판단하지 않습니다. 읽을 수 있는 글자나 추가 인물은 없습니다.",
        "hard_violations": [],
        "physics": "뻗은 팔은 어깨에서 손목과 손까지 자연스럽게 연결되어 있고 검지로 가리키는 자세가 가능합니다. 발은 프레임 밖이지만 상체가 떠 있다는 징후는 없습니다. 물은 균열에서 벽을 따라 아래로 흐르고 잔해는 벽 아래 지면에 놓여 있습니다. 지지 없이 떠 있는 물체는 보이지 않습니다."
       },
       {
        "label": "A",
        "direction": "미연의 팔과 검지는 오른쪽 방벽의 누수 균열을 향합니다. 미연은 눈을 크게 뜨고 왼쪽 전경 남성의 얼굴을 바라보며 호소합니다. 뒤쪽 작업자들은 대체로 미연과 전경의 대화 상대 쪽을 봅니다. 왼쪽 배경 작업자 얼굴 옆에는 왼쪽을 향하는 작은 검은 화살표가 보입니다.",
        "built_space": "오른쪽의 방벽 한 면과 상단 난간 한 줄, 그 아래 잔해와 파이프 적재물, 젖은 작업장 바닥이 보입니다. 중앙 배경에는 군용 트럭 한 대와 그 오른쪽 뒤 지게차 한 대가 있어 참고 장소의 배치와 잘 맞습니다. 다만 전경 남성의 어깨 너머로 미연의 상체와 작업자들의 다리까지 담아, 지정된 얼굴 클로즈업보다 훨씬 넓습니다. 방벽 역시 좁은 배경 단면이 아니라 긴 작업장 전경으로 제시됩니다.",
        "entities": "미연의 중년 얼굴, 검은 머리와 남색 상의는 참고와 대체로 맞고, 크게 뜬 눈과 열린 입은 절박함을 전달합니다. 그러나 왼쪽 전경 남성 한 명, 왼쪽 배경의 안전모 인물 두 명, 중앙 배경의 안전모 남성 두 명으로 미연 외에 다섯 명이 추가되어 있습니다. 균열, 누수, 물웅덩이, 정차한 트럭과 지게차는 확인됩니다. 읽을 수 있는 문자는 없지만 작은 화살표 표식이 있습니다.",
        "hard_violations": [
         "미연만 등장해야 하는 장면에 전경 남성과 배경 작업자 등 추가 인물 다섯 명이 등장합니다.",
         "왼쪽 배경 작업자 얼굴 옆에 장면 내 물체로 설명되지 않는 검은 화살표 표식이 노출되어 있습니다."
        ],
        "physics": "미연은 상체를 앞으로 기울인 채 어깨와 연결된 팔을 옆으로 뻗고 있어 가능한 동작입니다. 미연의 발은 프레임 밖이며 부유의 증거는 없습니다. 뒤쪽 작업자들의 보이는 신발은 지면에 닿고, 차량은 바퀴로 지면에 지지됩니다. 파이프는 적재대와 지면에 놓여 있으며 누수는 아래로 흐릅니다. 지지 없이 떠 있는 신체나 물체는 없습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.679,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.429,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] invented people (숏 텍스트에 없는 인물 4명 추가)",
     "[gemini-pro] leaked markers (화면 좌측 중앙에 화살표 마커 누출)",
     "[gpt-high] 미연만 등장해야 하는 장면에 전경 남성과 배경 작업자 등 추가 인물 다섯 명이 등장합니다.",
     "[gpt-high] 왼쪽 배경 작업자 얼굴 옆에 장면 내 물체로 설명되지 않는 검은 화살표 표식이 노출되어 있습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 429
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "프롬프트가 요구한 클로즈업 숏 크기, 피사체의 배치, 배경의 균열된 방벽 묘사를 충실히 구현했으며 엑스트라를 배제하라는 지시를 잘 따랐습니다."
   },
   {
    "label": "A",
    "score": 429,
    "verdict_ko": "지정되지 않은 인물들이 다수 등장하고, 화면 좌측에 마커 기호가 누출되어 치명적인 오류가 발생했습니다.  ★위반: [gemini-pro] invented people (숏 텍스트에 없는 인물 4명 추가) / [gemini-pro] leaked markers (화면 좌측 중앙에 화살표 마커 누출) / [gpt-high] 미연만 등장해야 하는 장면에 전경 남성과 배경 작업자 등 추가 인물 다섯 명이 등장합니다. / [gpt-high] 왼쪽 배경 작업자 얼굴 옆에 장면 내 물체로 설명되지 않는 검은 화살표 표식이 노출되어 있습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_seawall_worksite_44c24f.png",
    "asset_id": "13d1bb52-f02a-4d90-ae80-f1d788351b63",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1113064>",
    "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-aa8a-703c-a506-ebe0e37e0d89",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S7sh11__bgfirst_bg.png",
   "bg_asset_id": "0e42b760-1b39-40e2-a02b-2089485b8bc4",
   "bg_record_key": "S7sh11::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "seawall_worksite",
   "groupbg_asset_id": "13d1bb52-f02a-4d90-ae80-f1d788351b63"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S7sh11::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:03:35.351242+00:00",
  "fingerprint": "720592ce1d3110063c882364e869846ad30e0eb1c1661b345347d90f41ebd265",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S7sh11_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S7sh11_sel.png",
  "source_sha256": "975a237a299ebb0c11001239e4dc755241ff751cae99455915b6230176be0c2e",
  "file": "S7sh11_cine.png",
  "staged_sha256": "d5bf46c43c7c092bdbd46d89681ba117b0a6cdaa1648c8e4c6670963612c8091",
  "latency_ms": 14418
 },
 "S7sh13::signage": {
  "fp": "20f2a3458a91bd2a",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S7sh13": {
  "input_fingerprint": "db7c93285d02bc9d",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 난민 남자의 어깨를 향해 거칠게 뻗은 경비병의 양손과 그 힘에 밀려 몸이 뒤로 크게 기울어진 난민 남자가 포착된 mid-impact의 정점.\n\nLOCATION (lock): In the muddy gathering area beside the seawall repair works, where guards confront the refugee workers. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Embankment repair area (Repair work has been ordered to stop) — Only a recessed portion of the work area remains visible behind the opposing bodies; used as Maintains location continuity while leaving the hand-to-shoulder contact unobstructed; Ground beneath the confrontation (The refugee has not yet fallen onto it) — Visible along the lower edge beneath the displaced bodies; used as Leaves a readable fall direction without prematurely showing the aftermath.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding daylight and controlled contrast so the impact registers through posture rather than a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The severely cracked seawall continues to leak, and puddles remain on the ground. The military truck is still beside the stopped forklift and repair materials.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리); 50대 대머리 한국인 남성 (한국인 남성, 50대, 대머리, 드러난 두피, 중년의 얼굴). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 난민 남자의 어깨를 향해 거칠게 뻗은 경비병의 양손과 그 힘에 밀려 몸이 뒤로 크게 기울어진 난민 남자가 포착된 mid-impact의 정점.\n\nLOCATION (lock): In the muddy gathering area beside the seawall repair works, where guards confront the refugee workers. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Embankment repair area (Repair work has been ordered to stop) — Only a recessed portion of the work area remains visible behind the opposing bodies; used as Maintains location continuity while leaving the hand-to-shoulder contact unobstructed; Ground beneath the confrontation (The refugee has not yet fallen onto it) — Visible along the lower edge beneath the displaced bodies; used as Leaves a readable fall direction without prematurely showing the aftermath.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding daylight and controlled contrast so the impact registers through posture rather than a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The severely cracked seawall continues to leak, and puddles remain on the ground. The military truck is still beside the stopped forklift and repair materials.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리); 50대 대머리 한국인 남성 (한국인 남성, 50대, 대머리, 드러난 두피, 중년의 얼굴). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 난민 남자의 어깨를 향해 거칠게 뻗은 경비병의 양손과 그 힘에 밀려 몸이 뒤로 크게 기울어진 난민 남자가 포착된 mid-impact의 정점.\n\nLOCATION (lock): In the muddy gathering area beside the seawall repair works, where guards confront the refugee workers. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Embankment repair area (Repair work has been ordered to stop) — Only a recessed portion of the work area remains visible behind the opposing bodies; used as Maintains location continuity while leaving the hand-to-shoulder contact unobstructed; Ground beneath the confrontation (The refugee has not yet fallen onto it) — Visible along the lower edge beneath the displaced bodies; used as Leaves a readable fall direction without prematurely showing the aftermath.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding daylight and controlled contrast so the impact registers through posture rather than a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The severely cracked seawall continues to leak, and puddles remain on the ground. The military truck is still beside the stopped forklift and repair materials.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리); 50대 대머리 한국인 남성 (한국인 남성, 50대, 대머리, 드러난 두피, 중년의 얼굴). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "경비병의 양손이 난민의 가슴과 어깨 부위를 향해 곧게 뻗어 있으며, 난민은 그 힘에 밀려 뒤로 기울어짐.",
    "built_space": "금이 간 방파제 벽면의 오목한 부분과 물웅덩이가 있는 진흙 바닥이 프레임 내에 적절히 배치됨.",
    "entities": "한국인 외양의 경비병 1명과 다민족 외양의 난민 남성 1명만 정확하게 등장함.",
    "hard_violations": [
     "[gpt-high] 경비병 조끼의 가슴 표찰에 판독 가능한 한글이 보여 읽을 수 있는 글자를 금지한 조건을 위반한다."
    ],
    "physics": "두 인물 모두 진흙 바닥에 발을 딛고 체중을 지탱하고 있으며, 밀고 밀리는 동작의 물리적 균형이 설득력 있음."
   },
   {
    "label": "B",
    "direction": "경비병이 난민의 팔과 옷깃을 붙잡고 밀고 있으며, 난민의 상체가 뒤로 크게 젖혀짐.",
    "built_space": "방파제 벽면과 진흙 바닥뿐만 아니라, 프레임을 넓혀 뒤쪽의 지게차와 군용 트럭까지 노출됨.",
    "entities": "경비병과 난민 남성 외에, 왼쪽 배경에 지문에 없는 남성 3명이 추가로 존재함.",
    "hard_violations": [
     "[gemini-pro] 지문에 없는 인물들(배경의 세 남자)을 임의로 추가함",
     "[gpt-high] 숏 텍스트가 등장시키지 않은 배경 인물 세 명을 추가했다."
    ],
    "physics": "발은 바닥에 닿아 있으나 난민이 뒤로 넘어가는 각도에 비해 하체의 지지력이 다소 불안정해 보임."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 미디엄 샷 프레임과 배경 제약을 정확히 준수했으며, 지문에 없는 인물을 배제하고 두 인물의 충돌에 집중해 프롬프트를 충실히 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "지문에 명시되지 않은 세 명의 인물을 배경에 임의로 추가하는 치명적 오류를 범했으며, 프레임을 지시사항보다 과도하게 넓게 잡음."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "경비병의 양손이 난민의 가슴과 어깨 부위를 향해 곧게 뻗어 있으며, 난민은 그 힘에 밀려 뒤로 기울어짐.",
        "built_space": "금이 간 방파제 벽면의 오목한 부분과 물웅덩이가 있는 진흙 바닥이 프레임 내에 적절히 배치됨.",
        "entities": "한국인 외양의 경비병 1명과 다민족 외양의 난민 남성 1명만 정확하게 등장함.",
        "hard_violations": [],
        "physics": "두 인물 모두 진흙 바닥에 발을 딛고 체중을 지탱하고 있으며, 밀고 밀리는 동작의 물리적 균형이 설득력 있음."
       },
       {
        "label": "B",
        "direction": "경비병이 난민의 팔과 옷깃을 붙잡고 밀고 있으며, 난민의 상체가 뒤로 크게 젖혀짐.",
        "built_space": "방파제 벽면과 진흙 바닥뿐만 아니라, 프레임을 넓혀 뒤쪽의 지게차와 군용 트럭까지 노출됨.",
        "entities": "경비병과 난민 남성 외에, 왼쪽 배경에 지문에 없는 남성 3명이 추가로 존재함.",
        "hard_violations": [
         "지문에 없는 인물들(배경의 세 남자)을 임의로 추가함"
        ],
        "physics": "발은 바닥에 닿아 있으나 난민이 뒤로 넘어가는 각도에 비해 하체의 지지력이 다소 불안정해 보임."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 미디엄 샷 프레임과 배경 제약을 정확히 준수했으며, 지문에 없는 인물을 배제하고 두 인물의 충돌에 집중해 프롬프트를 충실히 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "지문에 명시되지 않은 세 명의 인물을 배경에 임의로 추가하는 치명적 오류를 범했으며, 프레임을 지시사항보다 과도하게 넓게 잡음."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "경비병의 양손이 난민의 가슴과 어깨 부위를 향해 곧게 뻗어 있으며, 난민은 그 힘에 밀려 뒤로 기울어짐.",
        "built_space": "금이 간 방파제 벽면의 오목한 부분과 물웅덩이가 있는 진흙 바닥이 프레임 내에 적절히 배치됨.",
        "entities": "한국인 외양의 경비병 1명과 다민족 외양의 난민 남성 1명만 정확하게 등장함.",
        "hard_violations": [],
        "physics": "두 인물 모두 진흙 바닥에 발을 딛고 체중을 지탱하고 있으며, 밀고 밀리는 동작의 물리적 균형이 설득력 있음."
       },
       {
        "label": "B",
        "direction": "경비병이 난민의 팔과 옷깃을 붙잡고 밀고 있으며, 난민의 상체가 뒤로 크게 젖혀짐.",
        "built_space": "방파제 벽면과 진흙 바닥뿐만 아니라, 프레임을 넓혀 뒤쪽의 지게차와 군용 트럭까지 노출됨.",
        "entities": "경비병과 난민 남성 외에, 왼쪽 배경에 지문에 없는 남성 3명이 추가로 존재함.",
        "hard_violations": [
         "지문에 없는 인물들(배경의 세 남자)을 임의로 추가함"
        ],
        "physics": "발은 바닥에 닿아 있으나 난민이 뒤로 넘어가는 각도에 비해 하체의 지지력이 다소 불안정해 보임."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "난민이 크게 젖혀지는 충격은 보이지만, 배경에 불필요한 인물 세 명을 추가하고 전신과 작업장을 넓게 보여 지정된 미디엄 숏을 위반한다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "두 사람만 등장하고 양손의 밀기와 발의 지지가 명확해 더 충실하지만, 전신 위주의 구도와 가슴을 미는 접촉 위치, 판독 가능한 표찰이 요구와 어긋난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "오른쪽 경비병과 왼쪽 난민은 서로 얼굴을 바라본다. 경비병의 한 손은 난민의 목 아래 어깨 부근으로, 다른 손은 가슴 쪽으로 뻗어 있어 양손 모두 어깨를 미는 모습은 아니다. 난민의 상체는 경비병에게서 멀어지는 왼쪽으로 크게 기울고, 난민의 손 하나는 경비병의 팔에 닿는다. 배경 세 사람은 대치하는 두 사람 쪽을 향한다.",
        "built_space": "오른쪽에 균열과 누수가 있는 방조제 한 면, 뒤쪽 왼편에 군용 트럭 한 대, 중앙에 지게차 한 대가 보인다. 벽 아래에는 수리용 틀과 잔해가 있고 전경에는 큰 웅덩이가 있다. 재료와 낮의 분위기는 참조와 유사하지만, 작업장이 두 인물 뒤의 좁은 일부로 제한되지 않고 화면 상당 부분을 차지한다. 두 주인공의 발까지 보이는 전신 구도다.",
        "entities": "주요 인물은 검은 제복과 전술 조끼를 입은 젊은 동아시아계 남성 경비병 한 명, 짧은 곱슬머리와 수염이 있고 회갈색 작업복을 입은 성인 남성 난민 한 명이다. 이 두 역할에는 별도 얼굴 지정이 없어 미연이나 대머리 남성으로 바꿀 이유는 없다. 그러나 왼쪽 배경에 작업자 차림 두 명과 제복 차림 한 명이 추가되어 있다. 트럭, 지게차, 수리 자재, 진흙과 웅덩이는 요청된 종류로 보인다.",
        "hard_violations": [
         "숏 텍스트가 등장시키지 않은 배경 인물 세 명을 추가했다."
        ],
        "physics": "경비병은 앞발을 진흙에 딛고 뒤쪽 다리를 뻗어 난민을 밀고 있다. 난민은 상체가 뒤로 젖혀지고 한쪽 부츠가 들려 있으며, 다른 부츠의 접지 부분은 경비병 다리에 상당 부분 가려져 있다. 양손 접촉에 의한 충격과 뒤쪽 진흙 바닥으로의 낙하 방향은 설명 가능하므로, 가려진 발만을 근거로 공중 부양이라고 단정할 수는 없다. 차량과 자재는 지면에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "오른쪽 경비병은 난민의 얼굴을 보고, 난민도 눈을 크게 뜬 채 경비병을 바라본다. 경비병의 두 팔은 왼쪽 난민을 향해 뻗으며, 한 손은 목 아래 쇄골 부근에, 다른 손은 위쪽 가슴에 닿는다. 양손 접촉은 선명하지만 양쪽 어깨를 정확히 겨냥하지는 않는다. 난민은 왼쪽 뒤로 밀리며 무릎을 굽힌다.",
        "built_space": "뒤쪽에는 균열 난 방조제 한 면과 직사각형으로 파인 큰 개구부 한 곳이 보인다. 벽 아래에 누수와 수리용 틀이 있고 두 사람 아래에는 진흙과 웅덩이가 이어진다. 개구부는 참조에서 확인되지 않아 고정 구조의 연속성은 다소 약하다. 트럭과 지게차는 화면 밖이므로 부재 자체는 결함이 아니다. 다만 두 사람의 다리와 발, 넓은 바닥까지 보여 요청한 미디엄 숏보다 넓다.",
        "entities": "젊은 동아시아계 남성 경비병 한 명과 곱슬머리·수염이 있는 성인 남성 난민 한 명만 보인다. 경비병은 회색 상의와 검은 전술 조끼, 남색 바지를 입고 난민은 흙 묻은 회갈색 작업복을 입었다. 미연과 대머리 남성은 없으며 숏 텍스트상 추가할 필요가 없다. 경비병 조끼의 작은 표찰에는 판독 가능한 한글이 노출되어 있다.",
        "hard_violations": [
         "경비병 조끼의 가슴 표찰에 판독 가능한 한글이 보여 읽을 수 있는 글자를 금지한 조건을 위반한다."
        ],
        "physics": "경비병은 앞발로 바닥을 지지하고 뒤쪽 다리를 뻗은 채 양손으로 난민을 민다. 난민의 두 부츠는 진흙 바닥에 닿아 있고 무릎이 굽혀지면서 상체가 뒤로 이동한다. 발 주변의 진흙 튀김도 미끄러지며 균형을 잃는 순간과 맞는다. 아직 넘어져 땅에 누운 상태가 아니며, 지지 없이 떠 있는 몸이나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "난민이 크게 젖혀지는 충격은 보이지만, 배경에 불필요한 인물 세 명을 추가하고 전신과 작업장을 넓게 보여 지정된 미디엄 숏을 위반한다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "두 사람만 등장하고 양손의 밀기와 발의 지지가 명확해 더 충실하지만, 전신 위주의 구도와 가슴을 미는 접촉 위치, 판독 가능한 표찰이 요구와 어긋난다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "오른쪽 경비병과 왼쪽 난민은 서로 얼굴을 바라본다. 경비병의 한 손은 난민의 목 아래 어깨 부근으로, 다른 손은 가슴 쪽으로 뻗어 있어 양손 모두 어깨를 미는 모습은 아니다. 난민의 상체는 경비병에게서 멀어지는 왼쪽으로 크게 기울고, 난민의 손 하나는 경비병의 팔에 닿는다. 배경 세 사람은 대치하는 두 사람 쪽을 향한다.",
        "built_space": "오른쪽에 균열과 누수가 있는 방조제 한 면, 뒤쪽 왼편에 군용 트럭 한 대, 중앙에 지게차 한 대가 보인다. 벽 아래에는 수리용 틀과 잔해가 있고 전경에는 큰 웅덩이가 있다. 재료와 낮의 분위기는 참조와 유사하지만, 작업장이 두 인물 뒤의 좁은 일부로 제한되지 않고 화면 상당 부분을 차지한다. 두 주인공의 발까지 보이는 전신 구도다.",
        "entities": "주요 인물은 검은 제복과 전술 조끼를 입은 젊은 동아시아계 남성 경비병 한 명, 짧은 곱슬머리와 수염이 있고 회갈색 작업복을 입은 성인 남성 난민 한 명이다. 이 두 역할에는 별도 얼굴 지정이 없어 미연이나 대머리 남성으로 바꿀 이유는 없다. 그러나 왼쪽 배경에 작업자 차림 두 명과 제복 차림 한 명이 추가되어 있다. 트럭, 지게차, 수리 자재, 진흙과 웅덩이는 요청된 종류로 보인다.",
        "hard_violations": [
         "숏 텍스트가 등장시키지 않은 배경 인물 세 명을 추가했다."
        ],
        "physics": "경비병은 앞발을 진흙에 딛고 뒤쪽 다리를 뻗어 난민을 밀고 있다. 난민은 상체가 뒤로 젖혀지고 한쪽 부츠가 들려 있으며, 다른 부츠의 접지 부분은 경비병 다리에 상당 부분 가려져 있다. 양손 접촉에 의한 충격과 뒤쪽 진흙 바닥으로의 낙하 방향은 설명 가능하므로, 가려진 발만을 근거로 공중 부양이라고 단정할 수는 없다. 차량과 자재는 지면에 놓여 있다."
       },
       {
        "label": "A",
        "direction": "오른쪽 경비병은 난민의 얼굴을 보고, 난민도 눈을 크게 뜬 채 경비병을 바라본다. 경비병의 두 팔은 왼쪽 난민을 향해 뻗으며, 한 손은 목 아래 쇄골 부근에, 다른 손은 위쪽 가슴에 닿는다. 양손 접촉은 선명하지만 양쪽 어깨를 정확히 겨냥하지는 않는다. 난민은 왼쪽 뒤로 밀리며 무릎을 굽힌다.",
        "built_space": "뒤쪽에는 균열 난 방조제 한 면과 직사각형으로 파인 큰 개구부 한 곳이 보인다. 벽 아래에 누수와 수리용 틀이 있고 두 사람 아래에는 진흙과 웅덩이가 이어진다. 개구부는 참조에서 확인되지 않아 고정 구조의 연속성은 다소 약하다. 트럭과 지게차는 화면 밖이므로 부재 자체는 결함이 아니다. 다만 두 사람의 다리와 발, 넓은 바닥까지 보여 요청한 미디엄 숏보다 넓다.",
        "entities": "젊은 동아시아계 남성 경비병 한 명과 곱슬머리·수염이 있는 성인 남성 난민 한 명만 보인다. 경비병은 회색 상의와 검은 전술 조끼, 남색 바지를 입고 난민은 흙 묻은 회갈색 작업복을 입었다. 미연과 대머리 남성은 없으며 숏 텍스트상 추가할 필요가 없다. 경비병 조끼의 작은 표찰에는 판독 가능한 한글이 노출되어 있다.",
        "hard_violations": [
         "경비병 조끼의 가슴 표찰에 판독 가능한 한글이 보여 읽을 수 있는 글자를 금지한 조건을 위반한다."
        ],
        "physics": "경비병은 앞발로 바닥을 지지하고 뒤쪽 다리를 뻗은 채 양손으로 난민을 민다. 난민의 두 부츠는 진흙 바닥에 닿아 있고 무릎이 굽혀지면서 상체가 뒤로 이동한다. 발 주변의 진흙 튀김도 미끄러지며 균형을 잃는 순간과 맞는다. 아직 넘어져 땅에 누운 상태가 아니며, 지지 없이 떠 있는 몸이나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.829
   },
   "adjusted": {
    "A": 1.75,
    "B": 0.579
   },
   "violations": {
    "B": [
     "[gemini-pro] 지문에 없는 인물들(배경의 세 남자)을 임의로 추가함",
     "[gpt-high] 숏 텍스트가 등장시키지 않은 배경 인물 세 명을 추가했다."
    ],
    "A": [
     "[gpt-high] 경비병 조끼의 가슴 표찰에 판독 가능한 한글이 보여 읽을 수 있는 글자를 금지한 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 579
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "지정된 미디엄 샷 프레임과 배경 제약을 정확히 준수했으며, 지문에 없는 인물을 배제하고 두 인물의 충돌에 집중해 프롬프트를 충실히 구현함.  ★위반: [gpt-high] 경비병 조끼의 가슴 표찰에 판독 가능한 한글이 보여 읽을 수 있는 글자를 금지한 조건을 위반한다."
   },
   {
    "label": "B",
    "score": 579,
    "verdict_ko": "지문에 명시되지 않은 세 명의 인물을 배경에 임의로 추가하는 치명적 오류를 범했으며, 프레임을 지시사항보다 과도하게 넓게 잡음.  ★위반: [gemini-pro] 지문에 없는 인물들(배경의 세 남자)을 임의로 추가함 / [gpt-high] 숏 텍스트가 등장시키지 않은 배경 인물 세 명을 추가했다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features, lighting mood and each person's clothing are LOCKED to this photo; never copy its camera framing. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S7sh11_sel.png",
    "asset_id": "8463927d-a43d-4e9f-a68f-6bce56177ade",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1113064>",
    "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 50대 대머리 한국인 남성: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1399111>",
    "asset_id": "41215473-9904-471e-81f7-2f6c01467eee",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-af7f-7423-a2ac-4f5e7e96ec10",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S7sh11"
  }
 },
 "S7sh13::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T13:24:32.633714+00:00",
  "fingerprint": "d4cb6c9e6fac60e27fc0044050ad667648396b91e4d986405cc8de653cd68d12",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S7sh13_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S7sh13_sel.png",
  "source_sha256": "81582a33a5e2aad14a964e2d756dfa1baf114aff3b59584f0f72241ed8b4c9ab",
  "file": "S7sh13_cine.png",
  "staged_sha256": "5ba94ec5e798aae6bff59eabd8efc0b17dca3765ee9cc31d9f0cba7f9227bc76",
  "latency_ms": 8766
 },
 "S7sh15::signage": {
  "fp": "a5dd6b91acf402d1",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S7sh15": {
  "input_fingerprint": "0b89e75a2ad01377",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 성난 표정으로 몰려드는 사람들을 등진 채, 트럭 조수석 안으로 다급히 상체를 반쯤 밀어 넣은 대머리 남자의 뒷모습.\n\nLOCATION (lock): At the open passenger doorway of a military truck parked on the muddy seawall repair site, viewed from outside. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Open military truck passenger entrance in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Military truck passenger entrance (Open for the administrator to climb inside) — The passenger-side opening is seen diagonally past his back, with the cabin extending beyond the right frame edge; used as Creates the destination of his retreat while remaining a cropped, subordinate portion of the frame; Approach beside the truck (People are converging on the administrator) — The open approach occupies the left side and leads diagonally toward the passenger entrance; used as Keeps the pursuing crowd and the retreat connected within one continuous space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Continue the same daytime ambient illumination with readable separation between the retreating figure, the cabin opening, and the crowd.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The military truck remains at the departure point beside the halted construction site. The seawall's severe cracks, continuing leaks, and ground puddles remain unchanged. 50대 대머리 한국인 남성: The frightened administrator is hurriedly climbing into the truck; his shoes remain puddle-soiled.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 50대 대머리 한국인 남성 (한국인 남성, 50대, 대머리, 드러난 두피, 중년의 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 성난 표정으로 몰려드는 사람들을 등진 채, 트럭 조수석 안으로 다급히 상체를 반쯤 밀어 넣은 대머리 남자의 뒷모습.\n\nLOCATION (lock): At the open passenger doorway of a military truck parked on the muddy seawall repair site, viewed from outside. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Open military truck passenger entrance in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Military truck passenger entrance (Open for the administrator to climb inside) — The passenger-side opening is seen diagonally past his back, with the cabin extending beyond the right frame edge; used as Creates the destination of his retreat while remaining a cropped, subordinate portion of the frame; Approach beside the truck (People are converging on the administrator) — The open approach occupies the left side and leads diagonally toward the passenger entrance; used as Keeps the pursuing crowd and the retreat connected within one continuous space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Continue the same daytime ambient illumination with readable separation between the retreating figure, the cabin opening, and the crowd.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The military truck remains at the departure point beside the halted construction site. The seawall's severe cracks, continuing leaks, and ground puddles remain unchanged. 50대 대머리 한국인 남성: The frightened administrator is hurriedly climbing into the truck; his shoes remain puddle-soiled.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 50대 대머리 한국인 남성 (한국인 남성, 50대, 대머리, 드러난 두피, 중년의 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 성난 표정으로 몰려드는 사람들을 등진 채, 트럭 조수석 안으로 다급히 상체를 반쯤 밀어 넣은 대머리 남자의 뒷모습.\n\nLOCATION (lock): At the open passenger doorway of a military truck parked on the muddy seawall repair site, viewed from outside. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Open military truck passenger entrance in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Military truck passenger entrance (Open for the administrator to climb inside) — The passenger-side opening is seen diagonally past his back, with the cabin extending beyond the right frame edge; used as Creates the destination of his retreat while remaining a cropped, subordinate portion of the frame; Approach beside the truck (People are converging on the administrator) — The open approach occupies the left side and leads diagonally toward the passenger entrance; used as Keeps the pursuing crowd and the retreat connected within one continuous space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Continue the same daytime ambient illumination with readable separation between the retreating figure, the cabin opening, and the crowd.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The military truck remains at the departure point beside the halted construction site. The seawall's severe cracks, continuing leaks, and ground puddles remain unchanged. 50대 대머리 한국인 남성: The frightened administrator is hurriedly climbing into the truck; his shoes remain puddle-soiled.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 50대 대머리 한국인 남성 (한국인 남성, 50대, 대머리, 드러난 두피, 중년의 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "군중은 남자와 트럭을 향해 시선을 고정함. 남자는 트럭 내부로 몸과 시선을 향함.",
    "built_space": "좌측 방조제와 우측 트럭이 정확히 배치됨. 열린 문 백미러에 남자의 정면이 비치나, 카메라 위치와 인물 방향상 광학적으로 불가능한 반사임.",
    "entities": "대머리 남자는 레퍼런스와 일치하는 정장을 착용함. 군중 속에 이전 샷의 경비원과 노동자 복장이 일치하게 등장함.",
    "hard_violations": [
     "[gemini-pro] 카메라 위치 및 인물 방향상 불가능한 거울 반사 (거울에 남자의 얼굴 정면이 비침)",
     "[gpt-high] 조수석으로 지정된 승차 위치가 운전대 바로 앞 좌석으로 이어지는 운전석 출입구로 구현되었다."
    ],
    "physics": "오른발은 계단에, 왼발은 바닥에 딛고 양손으로 차체를 잡아 체중을 자연스럽게 지탱함."
   },
   {
    "label": "B",
    "direction": "군중은 남자를 향해 다가가며 손을 뻗음. 남자는 트럭 안을 향함.",
    "built_space": "방조제와 트럭이 배치됨. 트럭 문의 경첩과 거울 장착 위치 등 구조가 비논리적임.",
    "entities": "대머리 남자가 정장 대신 갈색 재킷을 입어 레퍼런스의 복장 규정을 위반함.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 신체 구조 (남자의 다리가 3개로 렌더링됨)",
     "[gpt-high] 조수석으로 지정된 승차 위치가 운전대 바로 앞 좌석으로 이어지는 운전석 출입구로 구현되었다."
    ],
    "physics": "계단에 두 발을 딛고 있으나, 바닥 쪽에 남자의 바지와 동일한 세 번째 다리가 지면을 딛고 있어 불가능함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "광학적으로 불가능한 거울 반사가 있으나, 지정된 샷 프레이밍, 인물의 정장 복장, 이전 샷 인물들의 배경 연속성을 매우 정확히 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "주인공의 복장 레퍼런스(정장)를 무시했으며, 인물의 다리가 3개로 렌더링되는 치명적인 신체 오류가 발생함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "군중은 남자와 트럭을 향해 시선을 고정함. 남자는 트럭 내부로 몸과 시선을 향함.",
        "built_space": "좌측 방조제와 우측 트럭이 정확히 배치됨. 열린 문 백미러에 남자의 정면이 비치나, 카메라 위치와 인물 방향상 광학적으로 불가능한 반사임.",
        "entities": "대머리 남자는 레퍼런스와 일치하는 정장을 착용함. 군중 속에 이전 샷의 경비원과 노동자 복장이 일치하게 등장함.",
        "hard_violations": [
         "카메라 위치 및 인물 방향상 불가능한 거울 반사 (거울에 남자의 얼굴 정면이 비침)"
        ],
        "physics": "오른발은 계단에, 왼발은 바닥에 딛고 양손으로 차체를 잡아 체중을 자연스럽게 지탱함."
       },
       {
        "label": "B",
        "direction": "군중은 남자를 향해 다가가며 손을 뻗음. 남자는 트럭 안을 향함.",
        "built_space": "방조제와 트럭이 배치됨. 트럭 문의 경첩과 거울 장착 위치 등 구조가 비논리적임.",
        "entities": "대머리 남자가 정장 대신 갈색 재킷을 입어 레퍼런스의 복장 규정을 위반함.",
        "hard_violations": [
         "물리적으로 불가능한 신체 구조 (남자의 다리가 3개로 렌더링됨)"
        ],
        "physics": "계단에 두 발을 딛고 있으나, 바닥 쪽에 남자의 바지와 동일한 세 번째 다리가 지면을 딛고 있어 불가능함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "광학적으로 불가능한 거울 반사가 있으나, 지정된 샷 프레이밍, 인물의 정장 복장, 이전 샷 인물들의 배경 연속성을 매우 정확히 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "주인공의 복장 레퍼런스(정장)를 무시했으며, 인물의 다리가 3개로 렌더링되는 치명적인 신체 오류가 발생함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "군중은 남자와 트럭을 향해 시선을 고정함. 남자는 트럭 내부로 몸과 시선을 향함.",
        "built_space": "좌측 방조제와 우측 트럭이 정확히 배치됨. 열린 문 백미러에 남자의 정면이 비치나, 카메라 위치와 인물 방향상 광학적으로 불가능한 반사임.",
        "entities": "대머리 남자는 레퍼런스와 일치하는 정장을 착용함. 군중 속에 이전 샷의 경비원과 노동자 복장이 일치하게 등장함.",
        "hard_violations": [
         "카메라 위치 및 인물 방향상 불가능한 거울 반사 (거울에 남자의 얼굴 정면이 비침)"
        ],
        "physics": "오른발은 계단에, 왼발은 바닥에 딛고 양손으로 차체를 잡아 체중을 자연스럽게 지탱함."
       },
       {
        "label": "B",
        "direction": "군중은 남자를 향해 다가가며 손을 뻗음. 남자는 트럭 안을 향함.",
        "built_space": "방조제와 트럭이 배치됨. 트럭 문의 경첩과 거울 장착 위치 등 구조가 비논리적임.",
        "entities": "대머리 남자가 정장 대신 갈색 재킷을 입어 레퍼런스의 복장 규정을 위반함.",
        "hard_violations": [
         "물리적으로 불가능한 신체 구조 (남자의 다리가 3개로 렌더링됨)"
        ],
        "physics": "계단에 두 발을 딛고 있으나, 바닥 쪽에 남자의 바지와 동일한 세 번째 다리가 지면을 딛고 있어 불가능함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "군중을 등지고 급히 올라타는 동작은 맞지만, 조수석 대신 운전석으로 진입하며 관리자의 남색 정장이 회갈색 작업복으로 바뀌었다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "남색 정장과 대머리 뒷모습, 왼쪽 군중에서 오른쪽 출입구로 이어지는 구도는 더 충실하지만, 조수석이 아닌 운전석 승차로 보이는 공통 오류가 있다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "관리자는 군중에 등을 돌리고 오른쪽 트럭 내부로 머리와 몸통을 기울인다. 왼쪽 군중은 관리자를 바라보며 다가오고, 앞줄 남성의 뻗은 손도 관리자를 향한다. 무기나 별도의 겨냥 대상은 없다.",
        "built_space": "오른쪽에 열린 문 하나, 출입구 하나, 외부 거울 두 면, 앞바퀴 하나와 승차 발판이 보인다. 왼쪽 진흙 통로가 출입구까지 연결되고 뒤에는 크게 갈라진 콘크리트 방조제가 이어진다. 외부에서 보는 넓은 구도이나 트럭과 관리자가 상당히 크게 잡혔다. 출입구 바로 안쪽의 운전대와 좌석 배치는 조수석보다 운전석 출입구로 읽힌다. 거울에는 주변 현장이 작게 비치며 명백한 반사 모순은 없다.",
        "entities": "관리자 한 명과 약 스무 명의 군중이 보인다. 관리자는 노출된 정수리와 옆머리가 있는 중년 남성으로, 뒷모습이라 얼굴 동일성은 확인할 수 없다. 참고의 남색 정장 대신 회갈색 작업 재킷과 남색 작업 바지를 입었다. 군중에는 동아시아계로 보이는 사람들과 흑인 남성이 있으며 화난 표정이 드러난다. 군용 트럭, 균열 방조제, 진흙과 물웅덩이는 존재하지만 지속적인 누수 흐름은 뚜렷하지 않다.",
        "hard_violations": [
         "조수석으로 지정된 승차 위치가 운전대 바로 앞 좌석으로 이어지는 운전석 출입구로 구현되었다."
        ],
        "physics": "관리자의 왼발은 금속 발판에 놓여 있고 왼손은 문 윗부분을 붙잡는다. 오른쪽 다리를 들어 실내로 옮기는 순간으로, 지지하는 발과 손이 있어 공중에 뜬 자세가 아니다. 군중도 진흙 바닥에 발을 디디거나 보행 중 한 발을 들고 있다. 트럭은 바퀴로 지면에 지지된다."
       },
       {
        "label": "B",
        "direction": "관리자는 등을 보인 채 머리와 어깨를 트럭 내부로 향한다. 왼쪽의 화난 군중은 그와 출입구를 바라보며 전진한다. 관리자의 후퇴와 군중의 접근이 같은 대각선 동선으로 연결된다. 거울에는 관리자 얼굴 일부로 보이는 반사가 있다.",
        "built_space": "오른쪽에 열린 문 하나, 출입구 하나, 직사각형 외부 거울 하나, 앞바퀴 하나와 발판이 보인다. 왼쪽에는 물웅덩이가 있는 접근로, 뒤에는 균열 방조제와 작은 굴착기가 있다. 차량은 오른쪽 가장자리 밖으로 이어지고 군중과 출입구가 한 공간 안에 배치된다. 다만 운전대가 승차 중인 남성 바로 앞에 있어 지정된 조수석이 아니라 운전석으로 읽힌다. 열린 문에 달린 거울이 뒤쪽 관리자의 얼굴 일부를 받는 반사는 이 시점에서 불가능하다고 단정할 근거가 없다.",
        "entities": "관리자 한 명과 열두 명 안팎의 군중이 보인다. 관리자의 대머리와 남색 정장 재킷은 인물 참고에 더 가깝다. 얼굴과 셔츠·넥타이는 대부분 가려져 직접 대조하기 어렵다. 신발에는 진흙이 묻어 있다. 군중에는 동아시아계로 보이는 남성들, 흑인 남성, 조끼 차림 인물이 포함되며 앞줄은 성난 표정이다. 군용 트럭, 갈라진 콘크리트 방조제와 물웅덩이는 맞지만 누수 흐름은 명확히 확인되지 않는다.",
        "hard_violations": [
         "조수석으로 지정된 승차 위치가 운전대 바로 앞 좌석으로 이어지는 운전석 출입구로 구현되었다."
        ],
        "physics": "관리자의 왼발은 발판을 딛고 왼손은 문 윗부분을 잡는다. 오른발을 들어 올린 자세는 지지발과 손으로 체중을 받으며 올라타는 동작으로 가능하다. 상체는 앞으로 기울었지만 완전히 실내에 들어가지는 않았다. 군중의 발은 지면에 닿아 있거나 자연스러운 보행 단계에 있으며, 지지 없이 떠 있는 인물이나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "군중을 등지고 급히 올라타는 동작은 맞지만, 조수석 대신 운전석으로 진입하며 관리자의 남색 정장이 회갈색 작업복으로 바뀌었다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "남색 정장과 대머리 뒷모습, 왼쪽 군중에서 오른쪽 출입구로 이어지는 구도는 더 충실하지만, 조수석이 아닌 운전석 승차로 보이는 공통 오류가 있다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "관리자는 군중에 등을 돌리고 오른쪽 트럭 내부로 머리와 몸통을 기울인다. 왼쪽 군중은 관리자를 바라보며 다가오고, 앞줄 남성의 뻗은 손도 관리자를 향한다. 무기나 별도의 겨냥 대상은 없다.",
        "built_space": "오른쪽에 열린 문 하나, 출입구 하나, 외부 거울 두 면, 앞바퀴 하나와 승차 발판이 보인다. 왼쪽 진흙 통로가 출입구까지 연결되고 뒤에는 크게 갈라진 콘크리트 방조제가 이어진다. 외부에서 보는 넓은 구도이나 트럭과 관리자가 상당히 크게 잡혔다. 출입구 바로 안쪽의 운전대와 좌석 배치는 조수석보다 운전석 출입구로 읽힌다. 거울에는 주변 현장이 작게 비치며 명백한 반사 모순은 없다.",
        "entities": "관리자 한 명과 약 스무 명의 군중이 보인다. 관리자는 노출된 정수리와 옆머리가 있는 중년 남성으로, 뒷모습이라 얼굴 동일성은 확인할 수 없다. 참고의 남색 정장 대신 회갈색 작업 재킷과 남색 작업 바지를 입었다. 군중에는 동아시아계로 보이는 사람들과 흑인 남성이 있으며 화난 표정이 드러난다. 군용 트럭, 균열 방조제, 진흙과 물웅덩이는 존재하지만 지속적인 누수 흐름은 뚜렷하지 않다.",
        "hard_violations": [
         "조수석으로 지정된 승차 위치가 운전대 바로 앞 좌석으로 이어지는 운전석 출입구로 구현되었다."
        ],
        "physics": "관리자의 왼발은 금속 발판에 놓여 있고 왼손은 문 윗부분을 붙잡는다. 오른쪽 다리를 들어 실내로 옮기는 순간으로, 지지하는 발과 손이 있어 공중에 뜬 자세가 아니다. 군중도 진흙 바닥에 발을 디디거나 보행 중 한 발을 들고 있다. 트럭은 바퀴로 지면에 지지된다."
       },
       {
        "label": "A",
        "direction": "관리자는 등을 보인 채 머리와 어깨를 트럭 내부로 향한다. 왼쪽의 화난 군중은 그와 출입구를 바라보며 전진한다. 관리자의 후퇴와 군중의 접근이 같은 대각선 동선으로 연결된다. 거울에는 관리자 얼굴 일부로 보이는 반사가 있다.",
        "built_space": "오른쪽에 열린 문 하나, 출입구 하나, 직사각형 외부 거울 하나, 앞바퀴 하나와 발판이 보인다. 왼쪽에는 물웅덩이가 있는 접근로, 뒤에는 균열 방조제와 작은 굴착기가 있다. 차량은 오른쪽 가장자리 밖으로 이어지고 군중과 출입구가 한 공간 안에 배치된다. 다만 운전대가 승차 중인 남성 바로 앞에 있어 지정된 조수석이 아니라 운전석으로 읽힌다. 열린 문에 달린 거울이 뒤쪽 관리자의 얼굴 일부를 받는 반사는 이 시점에서 불가능하다고 단정할 근거가 없다.",
        "entities": "관리자 한 명과 열두 명 안팎의 군중이 보인다. 관리자의 대머리와 남색 정장 재킷은 인물 참고에 더 가깝다. 얼굴과 셔츠·넥타이는 대부분 가려져 직접 대조하기 어렵다. 신발에는 진흙이 묻어 있다. 군중에는 동아시아계로 보이는 남성들, 흑인 남성, 조끼 차림 인물이 포함되며 앞줄은 성난 표정이다. 군용 트럭, 갈라진 콘크리트 방조제와 물웅덩이는 맞지만 누수 흐름은 명확히 확인되지 않는다.",
        "hard_violations": [
         "조수석으로 지정된 승차 위치가 운전대 바로 앞 좌석으로 이어지는 운전석 출입구로 구현되었다."
        ],
        "physics": "관리자의 왼발은 발판을 딛고 왼손은 문 윗부분을 잡는다. 오른발을 들어 올린 자세는 지지발과 손으로 체중을 받으며 올라타는 동작으로 가능하다. 상체는 앞으로 기울었지만 완전히 실내에 들어가지는 않았다. 군중의 발은 지면에 닿아 있거나 자연스러운 보행 단계에 있으며, 지지 없이 떠 있는 인물이나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.417
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.167
   },
   "violations": {
    "A": [
     "[gemini-pro] 카메라 위치 및 인물 방향상 불가능한 거울 반사 (거울에 남자의 얼굴 정면이 비침)",
     "[gpt-high] 조수석으로 지정된 승차 위치가 운전대 바로 앞 좌석으로 이어지는 운전석 출입구로 구현되었다."
    ],
    "B": [
     "[gemini-pro] 물리적으로 불가능한 신체 구조 (남자의 다리가 3개로 렌더링됨)",
     "[gpt-high] 조수석으로 지정된 승차 위치가 운전대 바로 앞 좌석으로 이어지는 운전석 출입구로 구현되었다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 1167
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "광학적으로 불가능한 거울 반사가 있으나, 지정된 샷 프레이밍, 인물의 정장 복장, 이전 샷 인물들의 배경 연속성을 매우 정확히 구현함.  ★위반: [gemini-pro] 카메라 위치 및 인물 방향상 불가능한 거울 반사 (거울에 남자의 얼굴 정면이 비침) / [gpt-high] 조수석으로 지정된 승차 위치가 운전대 바로 앞 좌석으로 이어지는 운전석 출입구로 구현되었다."
   },
   {
    "label": "B",
    "score": 1167,
    "verdict_ko": "주인공의 복장 레퍼런스(정장)를 무시했으며, 인물의 다리가 3개로 렌더링되는 치명적인 신체 오류가 발생함.  ★위반: [gemini-pro] 물리적으로 불가능한 신체 구조 (남자의 다리가 3개로 렌더링됨) / [gpt-high] 조수석으로 지정된 승차 위치가 운전대 바로 앞 좌석으로 이어지는 운전석 출입구로 구현되었다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features, lighting mood and each person's clothing are LOCKED to this photo; never copy its camera framing. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S7sh13_sel.png",
    "asset_id": "c008521e-d6d6-4425-a36d-c28d1ab58d61",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 50대 대머리 한국인 남성: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1399111>",
    "asset_id": "41215473-9904-471e-81f7-2f6c01467eee",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-b13f-7a4c-bf69-95d0f57eee5d",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S7sh13"
  },
  "lane_policy": "ab_select_bypass:prev"
 },
 "S7sh15::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T13:26:07.348920+00:00",
  "fingerprint": "ace0c0ec3f8a356d8254def2cc38f4bf356c4c9bb3e4977fc5439315246c4237",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S7sh15_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S7sh15_sel.png",
  "source_sha256": "db44e59348feb3ac9311b7eaad231ed494a7990b743ca1a1f3b5edca35f2fc5b",
  "file": "S7sh15_cine.png",
  "staged_sha256": "deee31f22363b21c65c25de58ba986b9def610a47ce5f078377337a66338b4b9",
  "latency_ms": 9781
 },
 "S8sh5::signage": {
  "fp": "fb070707f47a136b",
  "inscriptions": [],
  "cues": [
   {
    "text_native": "",
    "source": "scene_text_implied",
    "source_quote": "붉은 도장"
   },
   {
    "text_native": "",
    "source": "scene_text_implied",
    "source_quote": "서류 종이"
   }
  ],
  "dropped": []
 },
 "era_assess::a338852cede772ba": {
  "subjects": [],
  "subject_text": "난민 재판소 내부\n낡고 허름한 실내 공간. 중앙의 초라한 탁자와 의자, 탁자 위에 높이 쌓인 서류 더미가 두드러진다.",
  "identity": "canonical",
  "scope_id": "L159",
  "scope_role": "location_interior",
  "scope_sha": "cf0a4752f24c4151"
 },
 "S8sh5::bgfirst_bg": {
  "input_fingerprint": "7a76f2066210a064",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 서류 종이 위로 붉은 도장을 거칠게 내리찍어 종이 표면에 막 닿은 재판관의 굵은 손.\n\nLOCATION (lock): At the judge's worn table inside a shabby refugee tribunal, in subdued daytime interior light.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Document receiving the stamp (The stamp has just contacted its surface) — The working face lies upward toward the downward-looking camera; no particular wording or completed seal is asserted; used as Makes the exact contact point visible beneath the hand; Red stamp (Held firmly against the document at first contact) — The gripped upper portion is visible while the stamping face is against the paper; used as Provides the small mechanical endpoint of the forceful hand movement; Shabby table and stacked paperwork (Documents remain piled before the seated judge) — The tabletop is seen obliquely, with its outer edge retained as a spatial reference; used as Grounds the close framing in the ongoing bureaucratic setting.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate interior ambient light keeps the hand and paperwork legible, with the stated red stamp providing the only specified color accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 서류 종이 위로 붉은 도장을 거칠게 내리찍어 종이 표면에 막 닿은 재판관의 굵은 손.\n\nLOCATION (lock): At the judge's worn table inside a shabby refugee tribunal, in subdued daytime interior light.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Document receiving the stamp (The stamp has just contacted its surface) — The working face lies upward toward the downward-looking camera; no particular wording or completed seal is asserted; used as Makes the exact contact point visible beneath the hand; Red stamp (Held firmly against the document at first contact) — The gripped upper portion is visible while the stamping face is against the paper; used as Provides the small mechanical endpoint of the forceful hand movement; Shabby table and stacked paperwork (Documents remain piled before the seated judge) — The tabletop is seen obliquely, with its outer edge retained as a spatial reference; used as Grounds the close framing in the ongoing bureaucratic setting.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate interior ambient light keeps the hand and paperwork legible, with the stated red stamp providing the only specified color accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S8sh5__bgfirst_bg.png",
  "asset_id": "70a6fcd0-a558-4ba0-8445-4a35efe45243",
  "input_asset_ids": [
   "5bd46cd4-f353-4051-a92b-c244f42f383d",
   "8c2a0cd6-3979-4553-8bb3-16258809b12f"
  ]
 },
 "S8sh5": {
  "input_fingerprint": "c40f92405782f6bc",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 서류 종이 위로 붉은 도장을 거칠게 내리찍어 종이 표면에 막 닿은 재판관의 굵은 손.\n\nLOCATION (lock): At the judge's worn table inside a shabby refugee tribunal, in subdued daytime interior light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Document receiving the stamp (The stamp has just contacted its surface) — The working face lies upward toward the downward-looking camera; no particular wording or completed seal is asserted; used as Makes the exact contact point visible beneath the hand; Red stamp (Held firmly against the document at first contact) — The gripped upper portion is visible while the stamping face is against the paper; used as Provides the small mechanical endpoint of the forceful hand movement; Shabby table and stacked paperwork (Documents remain piled before the seated judge) — The tabletop is seen obliquely, with its outer edge retained as a spatial reference; used as Grounds the close framing in the ongoing bureaucratic setting.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate interior ambient light keeps the hand and paperwork legible, with the stated red stamp providing the only specified color accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A shabby table and chair serve as the judge's station, with a backlog of documents piled on the table.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 재판관 right now, so 재판관's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 재판관: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 서류 종이 위로 붉은 도장을 거칠게 내리찍어 종이 표면에 막 닿은 재판관의 굵은 손.\n\nLOCATION (lock): At the judge's worn table inside a shabby refugee tribunal, in subdued daytime interior light. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Document receiving the stamp (The stamp has just contacted its surface) — The working face lies upward toward the downward-looking camera; no particular wording or completed seal is asserted; used as Makes the exact contact point visible beneath the hand; Red stamp (Held firmly against the document at first contact) — The gripped upper portion is visible while the stamping face is against the paper; used as Provides the small mechanical endpoint of the forceful hand movement; Shabby table and stacked paperwork (Documents remain piled before the seated judge) — The tabletop is seen obliquely, with its outer edge retained as a spatial reference; used as Grounds the close framing in the ongoing bureaucratic setting.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate interior ambient light keeps the hand and paperwork legible, with the stated red stamp providing the only specified color accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A shabby table and chair serve as the judge's station, with a backlog of documents piled on the table.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 재판관 right now, so 재판관's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 재판관: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 서류 종이 위로 붉은 도장을 거칠게 내리찍어 종이 표면에 막 닿은 재판관의 굵은 손.\n\nLOCATION (lock): At the judge's worn table inside a shabby refugee tribunal, in subdued daytime interior light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Document receiving the stamp (The stamp has just contacted its surface) — The working face lies upward toward the downward-looking camera; no particular wording or completed seal is asserted; used as Makes the exact contact point visible beneath the hand; Red stamp (Held firmly against the document at first contact) — The gripped upper portion is visible while the stamping face is against the paper; used as Provides the small mechanical endpoint of the forceful hand movement; Shabby table and stacked paperwork (Documents remain piled before the seated judge) — The tabletop is seen obliquely, with its outer edge retained as a spatial reference; used as Grounds the close framing in the ongoing bureaucratic setting.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate interior ambient light keeps the hand and paperwork legible, with the stated red stamp providing the only specified color accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A shabby table and chair serve as the judge's station, with a backlog of documents piled on the table.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 재판관 right now, so 재판관's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 재판관: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S8sh5__bgfirst_bg.png",
     "asset_id": "70a6fcd0-a558-4ba0-8445-4a35efe45243",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S8sh5.png",
     "asset_id": "5bd46cd4-f353-4051-a92b-c244f42f383d",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1113064>",
     "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L159B01.png",
     "asset_id": "8c2a0cd6-3979-4553-8bb3-16258809b12f",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1113064>",
     "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "재판관의 시선은 아래쪽의 도장을 향하고 있으며, 오른손이 도장을 쥐고 책상 위 서류를 향해 누르고 있다.",
    "built_space": "레퍼런스 사진의 법정 내부 요소(창문, 태극기, 캐비닛 등)가 보이지만, 지시된 클로즈업 구도가 아닌 미디엄 샷으로 넓게 촬영되어 카메라 앵글 지시를 위반했다.",
    "entities": "법복을 입은 재판관의 상체와 얼굴, 붉은색 사각형 도장, 서류 더미. 도장 밑으로 붉은 액체가 과도하게 번져 있다.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 연출 (도장에서 피처럼 흘러나와 고여 있는 비현실적인 붉은 액체 현상)"
    ],
    "physics": "손이 도장을 잡고 서류 표면을 누르고 있으나, 도장에서 뿜어져 나온 듯한 액체의 번짐 현상은 인주의 물리적 반응으로 볼 수 없다."
   },
   {
    "label": "B",
    "direction": "손이 붉은 도장을 단단히 쥐고 서류 위를 향해 수직에 가깝게 누르고 있다.",
    "built_space": "낡은 나무 책상의 상판과 가장자리, 쌓여 있는 서류 더미가 하향 카메라 앵글을 통해 클로즈업으로 정확하게 묘사되었다.",
    "entities": "옷소매가 보이는 굵은 손, 둥근 형태의 붉은 도장, 텍스트가 읽히지 않도록 흐릿하게 처리된 서류 용지들.",
    "hard_violations": [],
    "physics": "손아귀가 도장의 윗부분을 단단하게 지지하며 쥐고 있고, 도장의 밑면이 종이 표면에 자연스럽고 안정적으로 맞닿아 있다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "요구된 클로즈업 샷과 하향 카메라 앵글을 정확히 구현하여, 서류에 맞닿은 붉은 도장과 굵은 손의 물리적 접촉을 완벽하게 묘사한 결과물입니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "클로즈업 프레이밍 지시를 어기고 인물의 상체와 배경 전체를 넓게 보여주었으며, 도장에서 피처럼 번지는 비현실적인 붉은 액체가 생성되어 감점되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "재판관의 시선은 아래쪽의 도장을 향하고 있으며, 오른손이 도장을 쥐고 책상 위 서류를 향해 누르고 있다.",
        "built_space": "레퍼런스 사진의 법정 내부 요소(창문, 태극기, 캐비닛 등)가 보이지만, 지시된 클로즈업 구도가 아닌 미디엄 샷으로 넓게 촬영되어 카메라 앵글 지시를 위반했다.",
        "entities": "법복을 입은 재판관의 상체와 얼굴, 붉은색 사각형 도장, 서류 더미. 도장 밑으로 붉은 액체가 과도하게 번져 있다.",
        "hard_violations": [
         "물리적으로 불가능한 연출 (도장에서 피처럼 흘러나와 고여 있는 비현실적인 붉은 액체 현상)"
        ],
        "physics": "손이 도장을 잡고 서류 표면을 누르고 있으나, 도장에서 뿜어져 나온 듯한 액체의 번짐 현상은 인주의 물리적 반응으로 볼 수 없다."
       },
       {
        "label": "B",
        "direction": "손이 붉은 도장을 단단히 쥐고 서류 위를 향해 수직에 가깝게 누르고 있다.",
        "built_space": "낡은 나무 책상의 상판과 가장자리, 쌓여 있는 서류 더미가 하향 카메라 앵글을 통해 클로즈업으로 정확하게 묘사되었다.",
        "entities": "옷소매가 보이는 굵은 손, 둥근 형태의 붉은 도장, 텍스트가 읽히지 않도록 흐릿하게 처리된 서류 용지들.",
        "hard_violations": [],
        "physics": "손아귀가 도장의 윗부분을 단단하게 지지하며 쥐고 있고, 도장의 밑면이 종이 표면에 자연스럽고 안정적으로 맞닿아 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "요구된 클로즈업 샷과 하향 카메라 앵글을 정확히 구현하여, 서류에 맞닿은 붉은 도장과 굵은 손의 물리적 접촉을 완벽하게 묘사한 결과물입니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "클로즈업 프레이밍 지시를 어기고 인물의 상체와 배경 전체를 넓게 보여주었으며, 도장에서 피처럼 번지는 비현실적인 붉은 액체가 생성되어 감점되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "재판관의 시선은 아래쪽의 도장을 향하고 있으며, 오른손이 도장을 쥐고 책상 위 서류를 향해 누르고 있다.",
        "built_space": "레퍼런스 사진의 법정 내부 요소(창문, 태극기, 캐비닛 등)가 보이지만, 지시된 클로즈업 구도가 아닌 미디엄 샷으로 넓게 촬영되어 카메라 앵글 지시를 위반했다.",
        "entities": "법복을 입은 재판관의 상체와 얼굴, 붉은색 사각형 도장, 서류 더미. 도장 밑으로 붉은 액체가 과도하게 번져 있다.",
        "hard_violations": [
         "물리적으로 불가능한 연출 (도장에서 피처럼 흘러나와 고여 있는 비현실적인 붉은 액체 현상)"
        ],
        "physics": "손이 도장을 잡고 서류 표면을 누르고 있으나, 도장에서 뿜어져 나온 듯한 액체의 번짐 현상은 인주의 물리적 반응으로 볼 수 없다."
       },
       {
        "label": "B",
        "direction": "손이 붉은 도장을 단단히 쥐고 서류 위를 향해 수직에 가깝게 누르고 있다.",
        "built_space": "낡은 나무 책상의 상판과 가장자리, 쌓여 있는 서류 더미가 하향 카메라 앵글을 통해 클로즈업으로 정확하게 묘사되었다.",
        "entities": "옷소매가 보이는 굵은 손, 둥근 형태의 붉은 도장, 텍스트가 읽히지 않도록 흐릿하게 처리된 서류 용지들.",
        "hard_violations": [],
        "physics": "손아귀가 도장의 윗부분을 단단하게 지지하며 쥐고 있고, 도장의 밑면이 종이 표면에 자연스럽고 안정적으로 맞닿아 있다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "굵은 손과 도장의 종이 접촉점을 중심으로 한 클로즈업이 정확하며, 낡은 책상 가장자리와 서류 더미도 유지한다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "도장의 접촉 방향과 장소는 맞지만, 얼굴·상체·방까지 넓힌 구도가 손 중심 클로즈업에서 벗어나고 첫 접촉치고 잉크 번짐이 과하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "오른쪽 위에서 들어온 손이 도장 손잡이를 감싸 쥐고 아래쪽 서류를 누른다. 도장의 날인면은 종이를 향하고 실제로 맞닿아 있으며, 서류의 작업면은 비스듬히 내려다보는 카메라 쪽으로 열려 있다. 얼굴과 시선은 프레임 밖이다.",
        "built_space": "낡은 목제 책상 하나의 상판과 오른쪽 바깥 가장자리가 보인다. 서류는 전경, 왼쪽 중경, 왼쪽 후경, 중앙 후경, 오른쪽 후경에 대략 다섯 무더기로 놓여 있다. 배경에는 흐릿한 창과 수납장 일부가 보이며 참고 장소의 목재와 주간 채광에 부합한다. 의자와 재판관의 착석 위치는 이 클로즈업으로 확인할 수 없고, 중복 설비나 반사는 없다.",
        "entities": "보이는 인물 부분은 검은 소매에서 이어지는 굵고 주름진 성인 손과 손목·팔뚝 하나뿐이다. 재판관의 굵은 손이라는 지시에 부합하며, 손만으로 정확한 나이나 한국인 정체성을 확정할 수는 없다. 붉은 갈색 목제 도장 하나, 접촉 대상 서류, 쌓인 문서와 닳은 책상이 모두 있다. 도장은 선명한 빨강보다는 어두운 적갈색이고, 문서의 문자 흔적은 판독되지 않는다. 현우·앰버·미연은 등장하지 않는다.",
        "hard_violations": [],
        "physics": "손가락이 손잡이를 확실히 잡고 있으며 도장 밑면은 책상이 받치는 종이에 밀착한다. 손목에서 도장으로 내려가는 힘의 연결이 자연스럽다. 문서 더미도 상판 위에 놓여 있어 지지 없는 물체는 없다. 강한 타격의 속도감은 절제되어 있지만 접촉 순간의 자세는 가능하다."
       },
       {
        "label": "B",
        "direction": "재판관은 앞으로 숙여 도장과 서류 쪽을 내려다본다. 전경의 손이 붉은 사각 도장을 아래로 누르고, 날인면은 위로 펼쳐진 문서에 닿아 있다. 시선과 도장의 작동 방향 모두 대상 서류로 향한다.",
        "built_space": "책상 하나의 상판과 전면 가장자리, 왼쪽의 큰 서류 더미와 그 뒤의 작은 문서 묶음들, 오른쪽 필기구통 하나와 받침이 보인다. 배경에는 왼쪽 창 하나와 수납장 하나, 태극기 하나, 벽 액자 하나, 오른쪽 문 하나가 보여 참고 장소의 배치와 대체로 맞는다. 다만 손뿐 아니라 재판관의 얼굴·상체와 실내 대부분을 포함하여 지정된 클로즈업보다 넓다. 의자는 가려져 착석 상태를 확인할 수 없다.",
        "entities": "검은 법복과 흰 옷깃을 입은 중년 이상 동아시아계로 보이는 남성 재판관 한 명이 있으며, 굵은 손과 얼굴이 함께 보인다. 붉은 도장 하나, 날인 대상 문서, 낡은 목제 책상과 밀린 서류가 있다. 장면에 필요 없는 다른 인물은 없다. 문서 오른쪽에 작은 문자 형태가 남아 있으나 확실히 읽히는 문구는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "손이 도장 윗부분을 잡고 밑면을 종이에 누르며, 종이와 서류 더미는 책상이 받친다. 몸은 책상 뒤에서 앞으로 기울어져 있고 팔과 손의 연결에 명백한 물리적 불가능은 없다. 다만 도장 둘레로 넓게 퍼진 붉은 잉크는 통상적인 도장의 첫 접촉보다 과장되어, 막 닿은 순간의 재료 반응으로는 설득력이 떨어진다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "굵은 손과 도장의 종이 접촉점을 중심으로 한 클로즈업이 정확하며, 낡은 책상 가장자리와 서류 더미도 유지한다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "도장의 접촉 방향과 장소는 맞지만, 얼굴·상체·방까지 넓힌 구도가 손 중심 클로즈업에서 벗어나고 첫 접촉치고 잉크 번짐이 과하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "오른쪽 위에서 들어온 손이 도장 손잡이를 감싸 쥐고 아래쪽 서류를 누른다. 도장의 날인면은 종이를 향하고 실제로 맞닿아 있으며, 서류의 작업면은 비스듬히 내려다보는 카메라 쪽으로 열려 있다. 얼굴과 시선은 프레임 밖이다.",
        "built_space": "낡은 목제 책상 하나의 상판과 오른쪽 바깥 가장자리가 보인다. 서류는 전경, 왼쪽 중경, 왼쪽 후경, 중앙 후경, 오른쪽 후경에 대략 다섯 무더기로 놓여 있다. 배경에는 흐릿한 창과 수납장 일부가 보이며 참고 장소의 목재와 주간 채광에 부합한다. 의자와 재판관의 착석 위치는 이 클로즈업으로 확인할 수 없고, 중복 설비나 반사는 없다.",
        "entities": "보이는 인물 부분은 검은 소매에서 이어지는 굵고 주름진 성인 손과 손목·팔뚝 하나뿐이다. 재판관의 굵은 손이라는 지시에 부합하며, 손만으로 정확한 나이나 한국인 정체성을 확정할 수는 없다. 붉은 갈색 목제 도장 하나, 접촉 대상 서류, 쌓인 문서와 닳은 책상이 모두 있다. 도장은 선명한 빨강보다는 어두운 적갈색이고, 문서의 문자 흔적은 판독되지 않는다. 현우·앰버·미연은 등장하지 않는다.",
        "hard_violations": [],
        "physics": "손가락이 손잡이를 확실히 잡고 있으며 도장 밑면은 책상이 받치는 종이에 밀착한다. 손목에서 도장으로 내려가는 힘의 연결이 자연스럽다. 문서 더미도 상판 위에 놓여 있어 지지 없는 물체는 없다. 강한 타격의 속도감은 절제되어 있지만 접촉 순간의 자세는 가능하다."
       },
       {
        "label": "A",
        "direction": "재판관은 앞으로 숙여 도장과 서류 쪽을 내려다본다. 전경의 손이 붉은 사각 도장을 아래로 누르고, 날인면은 위로 펼쳐진 문서에 닿아 있다. 시선과 도장의 작동 방향 모두 대상 서류로 향한다.",
        "built_space": "책상 하나의 상판과 전면 가장자리, 왼쪽의 큰 서류 더미와 그 뒤의 작은 문서 묶음들, 오른쪽 필기구통 하나와 받침이 보인다. 배경에는 왼쪽 창 하나와 수납장 하나, 태극기 하나, 벽 액자 하나, 오른쪽 문 하나가 보여 참고 장소의 배치와 대체로 맞는다. 다만 손뿐 아니라 재판관의 얼굴·상체와 실내 대부분을 포함하여 지정된 클로즈업보다 넓다. 의자는 가려져 착석 상태를 확인할 수 없다.",
        "entities": "검은 법복과 흰 옷깃을 입은 중년 이상 동아시아계로 보이는 남성 재판관 한 명이 있으며, 굵은 손과 얼굴이 함께 보인다. 붉은 도장 하나, 날인 대상 문서, 낡은 목제 책상과 밀린 서류가 있다. 장면에 필요 없는 다른 인물은 없다. 문서 오른쪽에 작은 문자 형태가 남아 있으나 확실히 읽히는 문구는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "손이 도장 윗부분을 잡고 밑면을 종이에 누르며, 종이와 서류 더미는 책상이 받친다. 몸은 책상 뒤에서 앞으로 기울어져 있고 팔과 손의 연결에 명백한 물리적 불가능은 없다. 다만 도장 둘레로 넓게 퍼진 붉은 잉크는 통상적인 도장의 첫 접촉보다 과장되어, 막 닿은 순간의 재료 반응으로는 설득력이 떨어진다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.0,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.75,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 물리적으로 불가능한 연출 (도장에서 피처럼 흘러나와 고여 있는 비현실적인 붉은 액체 현상)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 750
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "요구된 클로즈업 샷과 하향 카메라 앵글을 정확히 구현하여, 서류에 맞닿은 붉은 도장과 굵은 손의 물리적 접촉을 완벽하게 묘사한 결과물입니다."
   },
   {
    "label": "A",
    "score": 750,
    "verdict_ko": "클로즈업 프레이밍 지시를 어기고 인물의 상체와 배경 전체를 넓게 보여주었으며, 도장에서 피처럼 번지는 비현실적인 붉은 액체가 생성되어 감점되었습니다.  ★위반: [gemini-pro] 물리적으로 불가능한 연출 (도장에서 피처럼 흘러나와 고여 있는 비현실적인 붉은 액체 현상)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L159B01.png",
    "asset_id": "8c2a0cd6-3979-4553-8bb3-16258809b12f",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1113064>",
    "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-b30f-7c2e-86cf-c1a72219e533",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S8sh5__bgfirst_bg.png",
   "bg_asset_id": "70a6fcd0-a558-4ba0-8445-4a35efe45243",
   "bg_record_key": "S8sh5::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S8sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:07:47.044483+00:00",
  "fingerprint": "64f8d6a5d9b9d016ffe436413ad1cf793ae464664936b8e476f54cb3ecfefba0",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S8sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S8sh5_sel.png",
  "source_sha256": "8a695eac10e1cd607fd3a0e6a8032c0dcedbb44e8a4ac45969db9565945bf7fa",
  "file": "S8sh5_cine.png",
  "staged_sha256": "d3eb702a4d68819e8fb8f9471c5f868c76710931c252e7930140c9bd40b6f05a",
  "latency_ms": 11305
 },
 "S8sh8::signage": {
  "fp": "06c16cf5d24a5c12",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S8sh8": {
  "input_fingerprint": "47d56c964f11cb1d",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바닥에 쓰러진 재판관의 멱살을 양손으로 꽉 틀어쥔 채 핏발 선 눈으로 노려보는 현우의 상체.\n\nLOCATION (lock): On the floor beside the judge's table and fallen chair inside the shabby refugee tribunal, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Courtroom floor (The judge has fallen onto it) — A narrow strip remains visible around the judge's shoulder; used as Confirms the low physical relationship between the figures; Shabby table (Remains beside the confrontation) — Its outer end is cropped at the background edge, viewed from below tabletop height; used as Connects the floor confrontation to the preceding document shot without obstructing either hand.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the courtroom's ambient illumination and controlled contrast, keeping this physical confrontation visually distinct from the later black-and-white nightmare.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The shabby courtroom table, chair, and piled case documents remain in place. 현우: Hyunwoo is bent over at floor level with both hands clenched in a collar-gripping posture; his dog-bitten leg remains injured and he is still feverish.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바닥에 쓰러진 재판관의 멱살을 양손으로 꽉 틀어쥔 채 핏발 선 눈으로 노려보는 현우의 상체.\n\nLOCATION (lock): On the floor beside the judge's table and fallen chair inside the shabby refugee tribunal, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Courtroom floor (The judge has fallen onto it) — A narrow strip remains visible around the judge's shoulder; used as Confirms the low physical relationship between the figures; Shabby table (Remains beside the confrontation) — Its outer end is cropped at the background edge, viewed from below tabletop height; used as Connects the floor confrontation to the preceding document shot without obstructing either hand.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the courtroom's ambient illumination and controlled contrast, keeping this physical confrontation visually distinct from the later black-and-white nightmare.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The shabby courtroom table, chair, and piled case documents remain in place. 현우: Hyunwoo is bent over at floor level with both hands clenched in a collar-gripping posture; his dog-bitten leg remains injured and he is still feverish.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바닥에 쓰러진 재판관의 멱살을 양손으로 꽉 틀어쥔 채 핏발 선 눈으로 노려보는 현우의 상체.\n\nLOCATION (lock): On the floor beside the judge's table and fallen chair inside the shabby refugee tribunal, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Courtroom floor (The judge has fallen onto it) — A narrow strip remains visible around the judge's shoulder; used as Confirms the low physical relationship between the figures; Shabby table (Remains beside the confrontation) — Its outer end is cropped at the background edge, viewed from below tabletop height; used as Connects the floor confrontation to the preceding document shot without obstructing either hand.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the courtroom's ambient illumination and controlled contrast, keeping this physical confrontation visually distinct from the later black-and-white nightmare.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The shabby courtroom table, chair, and piled case documents remain in place. 현우: Hyunwoo is bent over at floor level with both hands clenched in a collar-gripping posture; his dog-bitten leg remains injured and he is still feverish.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선이 바닥에 쓰러진 재판관의 얼굴을 명확하게 향하고 있다.",
    "built_space": "법정 바닥. 오른쪽에 화면 가장자리에 잘린 낡은 테이블이 지정된 대로 낮게 배치되어 있다. 뒤쪽에 쓰러진 의자가 있으나, 화면 왼쪽 배경에 지시되지 않은 책상과 의자가 추가로 존재한다.",
    "entities": "현우(핏발 선 눈, 헝클어진 머리, 양손으로 멱살을 단단히 쥠), 바닥에 쓰러진 재판관. 그러나 왼쪽 배경에 프롬프트에 없는 신원 미상의 인물의 다리와 발이 등장한다.",
    "hard_violations": [
     "[gemini-pro] 프롬프트에 명시되지 않은 인물 추가 (왼쪽 배경에 앉아 있는 사람의 발과 다리)",
     "[gemini-pro] 물리적으로 불가능한 구조 (쓰러진 재판관의 목 부위에 나무 의자 등받이가 관통하듯 융합됨)",
     "[gpt-high] 화면 왼쪽 끝에 장면이 요구하지 않은 제삼자의 검은 바지 차림 다리 일부가 보인다.",
     "[gpt-high] 뒤쪽의 넘어진 의자 외에 재판관을 받치는 전경 의자가 추가되어 의자가 중복된다.",
     "[gpt-high] 바닥에 쓰러져 있어야 할 재판관의 머리와 상체가 전경 의자에 기대어 지지되는 배치로 바뀌었다."
    ],
    "physics": "현우는 쓰러진 재판관 위에 몸을 숙이고 바닥과 상대방의 몸을 통해 체중을 지지하고 있으나, 재판관의 머리를 받치고 있는 의자의 형태가 신체와 기형적으로 결합되어 있어 물리적으로 성립하지 않는다."
   },
   {
    "label": "B",
    "direction": "현우의 시선이 재판관의 얼굴을 향해 내리꽂히고 있다.",
    "built_space": "법정 바닥. 뒤편 배경에 낡은 테이블 전체와 쓰러진 의자가 온전히 보인다. 테이블의 바깥쪽 끝이 화면 가장자리에 잘려야 한다는 카메라 프레이밍 지시를 전혀 따르지 않았다.",
    "entities": "현우(핏발 선 눈, 헝클어진 검은 머리), 쓰러진 재판관. 개에 물려 붕대를 감은 다리가 화면 오른쪽에 나타난다.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 해부학 (재판관의 어깨와 몸통 사이에서 붕대를 감은 맨다리가 기형적으로 뻗어 나옴, 현우의 다리라고 가정해도 위치와 각도가 불가능함)"
    ],
    "physics": "현우가 재판관의 옷깃을 틀어쥐고 몸을 지탱하고 있으나, 화면 우측에 위치한 붕대 감은 다리는 신체 구조상 누구의 것으로도 연결될 수 없는 불가능한 위치에 허공을 가로지르듯 배치되어 있다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "지정된 테이블 프레이밍을 시도했으나, 프롬프트에 없는 배경 인물이 등장하고 재판관의 목에 의자가 융합되는 치명적인 물리적 오류가 발생했습니다."
       },
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "지정된 카메라 프레이밍 지시를 완전히 무시했으며, 재판관의 몸통에서 기형적인 다리가 뻗어 나오는 해부학적 불가능성으로 인해 이미지를 사용할 수 없습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선이 바닥에 쓰러진 재판관의 얼굴을 명확하게 향하고 있다.",
        "built_space": "법정 바닥. 오른쪽에 화면 가장자리에 잘린 낡은 테이블이 지정된 대로 낮게 배치되어 있다. 뒤쪽에 쓰러진 의자가 있으나, 화면 왼쪽 배경에 지시되지 않은 책상과 의자가 추가로 존재한다.",
        "entities": "현우(핏발 선 눈, 헝클어진 머리, 양손으로 멱살을 단단히 쥠), 바닥에 쓰러진 재판관. 그러나 왼쪽 배경에 프롬프트에 없는 신원 미상의 인물의 다리와 발이 등장한다.",
        "hard_violations": [
         "프롬프트에 명시되지 않은 인물 추가 (왼쪽 배경에 앉아 있는 사람의 발과 다리)",
         "물리적으로 불가능한 구조 (쓰러진 재판관의 목 부위에 나무 의자 등받이가 관통하듯 융합됨)"
        ],
        "physics": "현우는 쓰러진 재판관 위에 몸을 숙이고 바닥과 상대방의 몸을 통해 체중을 지지하고 있으나, 재판관의 머리를 받치고 있는 의자의 형태가 신체와 기형적으로 결합되어 있어 물리적으로 성립하지 않는다."
       },
       {
        "label": "B",
        "direction": "현우의 시선이 재판관의 얼굴을 향해 내리꽂히고 있다.",
        "built_space": "법정 바닥. 뒤편 배경에 낡은 테이블 전체와 쓰러진 의자가 온전히 보인다. 테이블의 바깥쪽 끝이 화면 가장자리에 잘려야 한다는 카메라 프레이밍 지시를 전혀 따르지 않았다.",
        "entities": "현우(핏발 선 눈, 헝클어진 검은 머리), 쓰러진 재판관. 개에 물려 붕대를 감은 다리가 화면 오른쪽에 나타난다.",
        "hard_violations": [
         "물리적으로 불가능한 해부학 (재판관의 어깨와 몸통 사이에서 붕대를 감은 맨다리가 기형적으로 뻗어 나옴, 현우의 다리라고 가정해도 위치와 각도가 불가능함)"
        ],
        "physics": "현우가 재판관의 옷깃을 틀어쥐고 몸을 지탱하고 있으나, 화면 우측에 위치한 붕대 감은 다리는 신체 구조상 누구의 것으로도 연결될 수 없는 불가능한 위치에 허공을 가로지르듯 배치되어 있다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "지정된 테이블 프레이밍을 시도했으나, 프롬프트에 없는 배경 인물이 등장하고 재판관의 목에 의자가 융합되는 치명적인 물리적 오류가 발생했습니다."
       },
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "지정된 카메라 프레이밍 지시를 완전히 무시했으며, 재판관의 몸통에서 기형적인 다리가 뻗어 나오는 해부학적 불가능성으로 인해 이미지를 사용할 수 없습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선이 바닥에 쓰러진 재판관의 얼굴을 명확하게 향하고 있다.",
        "built_space": "법정 바닥. 오른쪽에 화면 가장자리에 잘린 낡은 테이블이 지정된 대로 낮게 배치되어 있다. 뒤쪽에 쓰러진 의자가 있으나, 화면 왼쪽 배경에 지시되지 않은 책상과 의자가 추가로 존재한다.",
        "entities": "현우(핏발 선 눈, 헝클어진 머리, 양손으로 멱살을 단단히 쥠), 바닥에 쓰러진 재판관. 그러나 왼쪽 배경에 프롬프트에 없는 신원 미상의 인물의 다리와 발이 등장한다.",
        "hard_violations": [
         "프롬프트에 명시되지 않은 인물 추가 (왼쪽 배경에 앉아 있는 사람의 발과 다리)",
         "물리적으로 불가능한 구조 (쓰러진 재판관의 목 부위에 나무 의자 등받이가 관통하듯 융합됨)"
        ],
        "physics": "현우는 쓰러진 재판관 위에 몸을 숙이고 바닥과 상대방의 몸을 통해 체중을 지지하고 있으나, 재판관의 머리를 받치고 있는 의자의 형태가 신체와 기형적으로 결합되어 있어 물리적으로 성립하지 않는다."
       },
       {
        "label": "B",
        "direction": "현우의 시선이 재판관의 얼굴을 향해 내리꽂히고 있다.",
        "built_space": "법정 바닥. 뒤편 배경에 낡은 테이블 전체와 쓰러진 의자가 온전히 보인다. 테이블의 바깥쪽 끝이 화면 가장자리에 잘려야 한다는 카메라 프레이밍 지시를 전혀 따르지 않았다.",
        "entities": "현우(핏발 선 눈, 헝클어진 검은 머리), 쓰러진 재판관. 개에 물려 붕대를 감은 다리가 화면 오른쪽에 나타난다.",
        "hard_violations": [
         "물리적으로 불가능한 해부학 (재판관의 어깨와 몸통 사이에서 붕대를 감은 맨다리가 기형적으로 뻗어 나옴, 현우의 다리라고 가정해도 위치와 각도가 불가능함)"
        ],
        "physics": "현우가 재판관의 옷깃을 틀어쥐고 몸을 지탱하고 있으나, 화면 우측에 위치한 붕대 감은 다리는 신체 구조상 누구의 것으로도 연결될 수 없는 불가능한 위치에 허공을 가로지르듯 배치되어 있다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "바닥에 누운 재판관의 멱살을 양손으로 잡고 얼굴을 노려보는 관계가 정확하지만, 책상과 하체를 지정된 상체 중심 구도보다 많이 보여준다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "양손 멱살잡이와 충혈된 시선은 맞지만, 여분 인물과 중복 의자가 보이고 재판관이 바닥 대신 의자에 기대어 있어 핵심 배치를 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 고개와 눈을 화면 오른쪽 아래의 재판관 얼굴로 향하고 있다. 두 주먹은 재판관 목 양옆의 검은 옷깃을 잡는다. 재판관은 얼굴을 위로 향하고 눈을 거의 감고 있어 뚜렷한 응시 대상은 없다.",
        "built_space": "뒤쪽에 낡은 목제 책상 한 개와 여러 서류 더미가 있고, 오른쪽 바닥에는 넘어진 목제 의자 한 개가 있다. 재판관은 책상 앞 바닥에 누워 있고 현우는 그 위로 몸을 숙였다. 카메라는 책상 상판보다 낮다. 다만 책상의 끝만 배경 가장자리에 걸치는 대신 책상 대부분과 다리까지 보이며, 재판관 주변 바닥도 좁은 띠보다 넓게 노출된다.",
        "entities": "현우는 앳된 동아시아계 남성으로, 헝클어진 검은 머리와 남색 반소매 티셔츠가 인물 참조와 대체로 일치한다. 눈가는 붉고 피부에는 땀이 보인다. 검은 법복과 흰 셔츠, 넥타이를 착용한 중년 남성 재판관이 있으며 추가 인물은 없다. 오른쪽 아래에는 상처와 붕대가 있는 현우의 다리가 보인다. 낡은 목재와 누렇게 닳은 서류는 장소 참조에 부합하며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "재판관의 등과 하체는 바닥에 지지되어 있다. 현우는 낮게 벌린 하체로 체중을 받치며 상체를 숙이고, 양손으로 움켜쥔 옷감에는 당겨지는 주름이 생긴다. 상처 난 다리를 옆으로 뻗은 자세는 다소 불편해 보이지만 불가능한 자세는 아니다. 넘어진 의자는 옆면으로 바닥에 닿고 서류는 책상 위에 놓여 있어 지지 없는 물체는 없다."
       },
       {
        "label": "B",
        "direction": "현우의 충혈된 눈은 화면 왼쪽 아래 재판관의 얼굴을 똑바로 겨냥하며, 재판관도 현우 쪽으로 얼굴을 들고 있다. 두 손은 재판관 목 아래의 자주색 옷깃을 붙잡고 있어 멱살잡이 방향은 맞는다.",
        "built_space": "오른쪽 전경에 서류와 도장이 놓인 큰 책상 한 개, 왼쪽 배경에 별도 책상 한 개가 보인다. 중앙 뒤에는 넘어진 의자 한 개가 있고, 재판관 머리와 등 뒤 전경에는 또 다른 목제 의자의 등받이 기둥이 보인다. 재판관은 바닥에 평평하게 쓰러진 것이 아니라 이 전경 의자에 기대어 상체가 들려 있다. 오른쪽 책상은 배경 끝의 작은 부분이 아니라 화면의 큰 부분을 가리는 전경 구조물이 된다.",
        "entities": "현우의 젊은 동아시아계 외모, 검은 머리, 남색 티셔츠와 땀에 젖은 상태는 참조에 대체로 맞는다. 재판관은 검은 법복에 자주색 깃이 있는 중년 이상 남성이다. 왼쪽 끝에는 두 주인공과 떨어진 검은 바지 차림의 굽힌 다리 일부가 보여 추가 인물이 들어왔다. 서류와 붉은 손잡이 도장은 앞 장면의 소품과 연결되며 읽을 수 있는 문구는 없다. 현우의 부상 다리는 구도 밖이므로 확인할 수 없다.",
        "hard_violations": [
         "화면 왼쪽 끝에 장면이 요구하지 않은 제삼자의 검은 바지 차림 다리 일부가 보인다.",
         "뒤쪽의 넘어진 의자 외에 재판관을 받치는 전경 의자가 추가되어 의자가 중복된다.",
         "바닥에 쓰러져 있어야 할 재판관의 머리와 상체가 전경 의자에 기대어 지지되는 배치로 바뀌었다."
        ],
        "physics": "재판관의 들린 상체는 뒤의 목제 의자와 현우가 잡아당기는 옷깃으로 지지되어 보이므로 공중에 떠 있는 것은 아니다. 그러나 그 지지 방식 자체가 요구된 바닥 대치와 다르다. 현우는 몸을 낮추고 양손으로 옷깃을 당기며, 하체의 바닥 접촉부는 대부분 가려져 있다. 서류와 도장은 상판에 놓여 있고 뒤의 넘어진 의자는 바닥에 닿아 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "바닥에 누운 재판관의 멱살을 양손으로 잡고 얼굴을 노려보는 관계가 정확하지만, 책상과 하체를 지정된 상체 중심 구도보다 많이 보여준다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "양손 멱살잡이와 충혈된 시선은 맞지만, 여분 인물과 중복 의자가 보이고 재판관이 바닥 대신 의자에 기대어 있어 핵심 배치를 위반한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 고개와 눈을 화면 오른쪽 아래의 재판관 얼굴로 향하고 있다. 두 주먹은 재판관 목 양옆의 검은 옷깃을 잡는다. 재판관은 얼굴을 위로 향하고 눈을 거의 감고 있어 뚜렷한 응시 대상은 없다.",
        "built_space": "뒤쪽에 낡은 목제 책상 한 개와 여러 서류 더미가 있고, 오른쪽 바닥에는 넘어진 목제 의자 한 개가 있다. 재판관은 책상 앞 바닥에 누워 있고 현우는 그 위로 몸을 숙였다. 카메라는 책상 상판보다 낮다. 다만 책상의 끝만 배경 가장자리에 걸치는 대신 책상 대부분과 다리까지 보이며, 재판관 주변 바닥도 좁은 띠보다 넓게 노출된다.",
        "entities": "현우는 앳된 동아시아계 남성으로, 헝클어진 검은 머리와 남색 반소매 티셔츠가 인물 참조와 대체로 일치한다. 눈가는 붉고 피부에는 땀이 보인다. 검은 법복과 흰 셔츠, 넥타이를 착용한 중년 남성 재판관이 있으며 추가 인물은 없다. 오른쪽 아래에는 상처와 붕대가 있는 현우의 다리가 보인다. 낡은 목재와 누렇게 닳은 서류는 장소 참조에 부합하며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "재판관의 등과 하체는 바닥에 지지되어 있다. 현우는 낮게 벌린 하체로 체중을 받치며 상체를 숙이고, 양손으로 움켜쥔 옷감에는 당겨지는 주름이 생긴다. 상처 난 다리를 옆으로 뻗은 자세는 다소 불편해 보이지만 불가능한 자세는 아니다. 넘어진 의자는 옆면으로 바닥에 닿고 서류는 책상 위에 놓여 있어 지지 없는 물체는 없다."
       },
       {
        "label": "A",
        "direction": "현우의 충혈된 눈은 화면 왼쪽 아래 재판관의 얼굴을 똑바로 겨냥하며, 재판관도 현우 쪽으로 얼굴을 들고 있다. 두 손은 재판관 목 아래의 자주색 옷깃을 붙잡고 있어 멱살잡이 방향은 맞는다.",
        "built_space": "오른쪽 전경에 서류와 도장이 놓인 큰 책상 한 개, 왼쪽 배경에 별도 책상 한 개가 보인다. 중앙 뒤에는 넘어진 의자 한 개가 있고, 재판관 머리와 등 뒤 전경에는 또 다른 목제 의자의 등받이 기둥이 보인다. 재판관은 바닥에 평평하게 쓰러진 것이 아니라 이 전경 의자에 기대어 상체가 들려 있다. 오른쪽 책상은 배경 끝의 작은 부분이 아니라 화면의 큰 부분을 가리는 전경 구조물이 된다.",
        "entities": "현우의 젊은 동아시아계 외모, 검은 머리, 남색 티셔츠와 땀에 젖은 상태는 참조에 대체로 맞는다. 재판관은 검은 법복에 자주색 깃이 있는 중년 이상 남성이다. 왼쪽 끝에는 두 주인공과 떨어진 검은 바지 차림의 굽힌 다리 일부가 보여 추가 인물이 들어왔다. 서류와 붉은 손잡이 도장은 앞 장면의 소품과 연결되며 읽을 수 있는 문구는 없다. 현우의 부상 다리는 구도 밖이므로 확인할 수 없다.",
        "hard_violations": [
         "화면 왼쪽 끝에 장면이 요구하지 않은 제삼자의 검은 바지 차림 다리 일부가 보인다.",
         "뒤쪽의 넘어진 의자 외에 재판관을 받치는 전경 의자가 추가되어 의자가 중복된다.",
         "바닥에 쓰러져 있어야 할 재판관의 머리와 상체가 전경 의자에 기대어 지지되는 배치로 바뀌었다."
        ],
        "physics": "재판관의 들린 상체는 뒤의 목제 의자와 현우가 잡아당기는 옷깃으로 지지되어 보이므로 공중에 떠 있는 것은 아니다. 그러나 그 지지 방식 자체가 요구된 바닥 대치와 다르다. 현우는 몸을 낮추고 양손으로 옷깃을 당기며, 하체의 바닥 접촉부는 대부분 가려져 있다. 서류와 도장은 상판에 놓여 있고 뒤의 넘어진 의자는 바닥에 닿아 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.25,
    "B": 1.5
   },
   "adjusted": {
    "A": 1.0,
    "B": 1.25
   },
   "violations": {
    "A": [
     "[gemini-pro] 프롬프트에 명시되지 않은 인물 추가 (왼쪽 배경에 앉아 있는 사람의 발과 다리)",
     "[gemini-pro] 물리적으로 불가능한 구조 (쓰러진 재판관의 목 부위에 나무 의자 등받이가 관통하듯 융합됨)",
     "[gpt-high] 화면 왼쪽 끝에 장면이 요구하지 않은 제삼자의 검은 바지 차림 다리 일부가 보인다.",
     "[gpt-high] 뒤쪽의 넘어진 의자 외에 재판관을 받치는 전경 의자가 추가되어 의자가 중복된다.",
     "[gpt-high] 바닥에 쓰러져 있어야 할 재판관의 머리와 상체가 전경 의자에 기대어 지지되는 배치로 바뀌었다."
    ],
    "B": [
     "[gemini-pro] 물리적으로 불가능한 해부학 (재판관의 어깨와 몸통 사이에서 붕대를 감은 맨다리가 기형적으로 뻗어 나옴, 현우의 다리라고 가정해도 위치와 각도가 불가능함)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1000,
   "B": 1250
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1000,
    "verdict_ko": "지정된 테이블 프레이밍을 시도했으나, 프롬프트에 없는 배경 인물이 등장하고 재판관의 목에 의자가 융합되는 치명적인 물리적 오류가 발생했습니다.  ★위반: [gemini-pro] 프롬프트에 명시되지 않은 인물 추가 (왼쪽 배경에 앉아 있는 사람의 발과 다리) / [gemini-pro] 물리적으로 불가능한 구조 (쓰러진 재판관의 목 부위에 나무 의자 등받이가 관통하듯 융합됨) / [gpt-high] 화면 왼쪽 끝에 장면이 요구하지 않은 제삼자의 검은 바지 차림 다리 일부가 보인다. / [gpt-high] 뒤쪽의 넘어진 의자 외에 재판관을 받치는 전경 의자가 추가되어 의자가 중복된다. / [gpt-high] 바닥에 쓰러져 있어야 할 재판관의 머리와 상체가 전경 의자에 기대어 지지되는 배치로 바뀌었다."
   },
   {
    "label": "B",
    "score": 1250,
    "verdict_ko": "지정된 카메라 프레이밍 지시를 완전히 무시했으며, 재판관의 몸통에서 기형적인 다리가 뻗어 나오는 해부학적 불가능성으로 인해 이미지를 사용할 수 없습니다.  ★위반: [gemini-pro] 물리적으로 불가능한 해부학 (재판관의 어깨와 몸통 사이에서 붕대를 감은 맨다리가 기형적으로 뻗어 나옴, 현우의 다리라고 가정해도 위치와 각도가 불가능함)"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features, lighting mood and each person's clothing are LOCKED to this photo; never copy its camera framing. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S8sh5_sel.png",
    "asset_id": "f229b1cd-f3f6-4c3b-b29c-a990d8901eb1",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-b686-7c87-9f29-f9107c84c715",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S8sh5"
  }
 },
 "S8sh8::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:08:54.279125+00:00",
  "fingerprint": "06f41fae607e74056381509a08490efac243d2279643b0af6515cd1e714333c5",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S8sh8_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S8sh8_sel.png",
  "source_sha256": "46be7f205708965d497116cb59153184ce29be1c9a1b7dde02464689550f3df0",
  "file": "S8sh8_cine.png",
  "staged_sha256": "324ca75509a850f946f10856836339e9e01cb213611f33bef6c04d723e88226e",
  "latency_ms": 11320
 },
 "S8sh13::signage": {
  "fp": "b17dd516f0137658",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S8sh13::bgfirst_bg": {
  "input_fingerprint": "3b44ab324d5bcb68",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 극도의 공포에 질린 표정으로 어린 현우와 앰버를 양팔로 빈틈없이 끌어안고 웅크린 미연의 상체.\n\nLOCATION (lock): On a city street at night during a violent riot, among civilians forced to kneel under intermittent gunfire flashes.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: American city street (The family is kneeling during the armed attack) — Only a limited area around and behind the huddled family remains visible; used as Establishes the separate nightmare location without diluting the protective embrace.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Render the nighttime nightmare in black and white with controlled facial readability and abrupt frame-to-frame cadence, without adding an unsupported practical light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 극도의 공포에 질린 표정으로 어린 현우와 앰버를 양팔로 빈틈없이 끌어안고 웅크린 미연의 상체.\n\nLOCATION (lock): On a city street at night during a violent riot, among civilians forced to kneel under intermittent gunfire flashes.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: American city street (The family is kneeling during the armed attack) — Only a limited area around and behind the huddled family remains visible; used as Establishes the separate nightmare location without diluting the protective embrace.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Render the nighttime nightmare in black and white with controlled facial readability and abrupt frame-to-frame cadence, without adding an unsupported practical light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S8sh13__bgfirst_bg.png",
  "asset_id": "8bcd8334-9652-44d4-9e5a-7ec29199889a",
  "input_asset_ids": [
   "863e681c-c732-49e9-9c06-7c759918fd73",
   "dcf94519-9d42-4399-a0db-17ea44888173"
  ]
 },
 "S8sh13": {
  "input_fingerprint": "d2dd963a7cecfe60",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 극도의 공포에 질린 표정으로 어린 현우와 앰버를 양팔로 빈틈없이 끌어안고 웅크린 미연의 상체.\n\nLOCATION (lock): On a city street at night during a violent riot, among civilians forced to kneel under intermittent gunfire flashes. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: American city street (The family is kneeling during the armed attack) — Only a limited area around and behind the huddled family remains visible; used as Establishes the separate nightmare location without diluting the protective embrace.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Render the nighttime nightmare in black and white with controlled facial readability and abrupt frame-to-frame cadence, without adding an unsupported practical light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): This is a black-and-white nighttime memory of a U.S. street, with intermittent gunfire flashes and a large American flag. 미연: Miyeon is kneeling in terror, hunched with both arms wrapped protectively inward. 현우: Hyunwoo appears as a younger child, huddled low within a protective embrace. 앰버: Amber appears as a younger child, huddled low within a protective embrace.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 극도의 공포에 질린 표정으로 어린 현우와 앰버를 양팔로 빈틈없이 끌어안고 웅크린 미연의 상체.\n\nLOCATION (lock): On a city street at night during a violent riot, among civilians forced to kneel under intermittent gunfire flashes. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: American city street (The family is kneeling during the armed attack) — Only a limited area around and behind the huddled family remains visible; used as Establishes the separate nightmare location without diluting the protective embrace.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Render the nighttime nightmare in black and white with controlled facial readability and abrupt frame-to-frame cadence, without adding an unsupported practical light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): This is a black-and-white nighttime memory of a U.S. street, with intermittent gunfire flashes and a large American flag. 미연: Miyeon is kneeling in terror, hunched with both arms wrapped protectively inward. 현우: Hyunwoo appears as a younger child, huddled low within a protective embrace. 앰버: Amber appears as a younger child, huddled low within a protective embrace.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 극도의 공포에 질린 표정으로 어린 현우와 앰버를 양팔로 빈틈없이 끌어안고 웅크린 미연의 상체.\n\nLOCATION (lock): On a city street at night during a violent riot, among civilians forced to kneel under intermittent gunfire flashes. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: American city street (The family is kneeling during the armed attack) — Only a limited area around and behind the huddled family remains visible; used as Establishes the separate nightmare location without diluting the protective embrace.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Render the nighttime nightmare in black and white with controlled facial readability and abrupt frame-to-frame cadence, without adding an unsupported practical light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): This is a black-and-white nighttime memory of a U.S. street, with intermittent gunfire flashes and a large American flag. 미연: Miyeon is kneeling in terror, hunched with both arms wrapped protectively inward. 현우: Hyunwoo appears as a younger child, huddled low within a protective embrace. 앰버: Amber appears as a younger child, huddled low within a protective embrace.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S8sh13__bgfirst_bg.png",
     "asset_id": "8bcd8334-9652-44d4-9e5a-7ec29199889a",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S8sh13.png",
     "asset_id": "863e681c-c732-49e9-9c06-7c759918fd73",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1113064>",
     "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L159B02.png",
     "asset_id": "dcf94519-9d42-4399-a0db-17ea44888173",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1113064>",
     "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "미연은 우측 상단을 향해 공포에 질린 시선을 던지며, 아이들은 미연의 품으로 시선을 거두고 웅크림.",
    "built_space": "레퍼런스와 일치하는 밤거리. 전경의 하수구 덮개, 연석, 우측의 차양과 문이 있는 벽면 구조가 정확히 배치됨. 배경에 무릎 꿇은 시민들이 있음.",
    "entities": "미연(40대 여성), 앰버(금발 10세 여아), 현우(어린아이의 모습으로 묘사됨). 지시된 대형 미국 국기는 없음.",
    "hard_violations": [
     "[gpt-high] 숏 텍스트에 없는 주변 인물 네 명을 거리 양쪽에 추가했다."
    ],
    "physics": "세 인물 모두 보도블록에 무릎을 꿇고 체중을 온전히 싣고 있으며, 미연의 양팔이 아이들의 등과 어깨를 물리적으로 단단히 감싸 지탱함."
   },
   {
    "label": "B",
    "direction": "미연은 정면 위쪽을 향해 소리치듯 응시하고, 아이들은 품에 안겨 시선을 아래로 향함.",
    "built_space": "상점가가 있는 밤거리 구조이나, 로케이션 레퍼런스의 우측 차양 건물과 특정 파이프 구조물이 일치하지 않음.",
    "entities": "미연, 앰버, 현우(청소년에 가깝게 묘사됨). 배경에 무릎 꿇은 인물들과 대형 미국 국기, 총구 화염 빛이 보임.",
    "hard_violations": [
     "[gpt-high] 숏 텍스트가 허용한 미연·현우·앰버 외에 다수의 주변 인물을 추가했다."
    ],
    "physics": "도로 바닥에 무릎을 꿇고 몸을 지탱함. 미연의 팔이 아이들을 감싸고 있으나, 배경 우측 인물의 하반신과 지면의 접촉 묘사가 다소 불분명함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 로케이션(하수구, 우측 차양 건물 등)을 완벽히 재현하고 현우를 어린아이로 정확히 묘사했으나, 대형 미국 국기가 누락되었고 상체 프레이밍보다 다소 넓게 촬영되었습니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "미국 국기는 포함되었으나 로케이션 레퍼런스의 우측 건물 구조가 누락되었고, 현우가 지시보다 성숙하게 묘사되어 로케이션 및 인물 묘사 우선순위에서 밀립니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "미연은 우측 상단을 향해 공포에 질린 시선을 던지며, 아이들은 미연의 품으로 시선을 거두고 웅크림.",
        "built_space": "레퍼런스와 일치하는 밤거리. 전경의 하수구 덮개, 연석, 우측의 차양과 문이 있는 벽면 구조가 정확히 배치됨. 배경에 무릎 꿇은 시민들이 있음.",
        "entities": "미연(40대 여성), 앰버(금발 10세 여아), 현우(어린아이의 모습으로 묘사됨). 지시된 대형 미국 국기는 없음.",
        "hard_violations": [],
        "physics": "세 인물 모두 보도블록에 무릎을 꿇고 체중을 온전히 싣고 있으며, 미연의 양팔이 아이들의 등과 어깨를 물리적으로 단단히 감싸 지탱함."
       },
       {
        "label": "B",
        "direction": "미연은 정면 위쪽을 향해 소리치듯 응시하고, 아이들은 품에 안겨 시선을 아래로 향함.",
        "built_space": "상점가가 있는 밤거리 구조이나, 로케이션 레퍼런스의 우측 차양 건물과 특정 파이프 구조물이 일치하지 않음.",
        "entities": "미연, 앰버, 현우(청소년에 가깝게 묘사됨). 배경에 무릎 꿇은 인물들과 대형 미국 국기, 총구 화염 빛이 보임.",
        "hard_violations": [],
        "physics": "도로 바닥에 무릎을 꿇고 몸을 지탱함. 미연의 팔이 아이들을 감싸고 있으나, 배경 우측 인물의 하반신과 지면의 접촉 묘사가 다소 불분명함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 로케이션(하수구, 우측 차양 건물 등)을 완벽히 재현하고 현우를 어린아이로 정확히 묘사했으나, 대형 미국 국기가 누락되었고 상체 프레이밍보다 다소 넓게 촬영되었습니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "미국 국기는 포함되었으나 로케이션 레퍼런스의 우측 건물 구조가 누락되었고, 현우가 지시보다 성숙하게 묘사되어 로케이션 및 인물 묘사 우선순위에서 밀립니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "미연은 우측 상단을 향해 공포에 질린 시선을 던지며, 아이들은 미연의 품으로 시선을 거두고 웅크림.",
        "built_space": "레퍼런스와 일치하는 밤거리. 전경의 하수구 덮개, 연석, 우측의 차양과 문이 있는 벽면 구조가 정확히 배치됨. 배경에 무릎 꿇은 시민들이 있음.",
        "entities": "미연(40대 여성), 앰버(금발 10세 여아), 현우(어린아이의 모습으로 묘사됨). 지시된 대형 미국 국기는 없음.",
        "hard_violations": [],
        "physics": "세 인물 모두 보도블록에 무릎을 꿇고 체중을 온전히 싣고 있으며, 미연의 양팔이 아이들의 등과 어깨를 물리적으로 단단히 감싸 지탱함."
       },
       {
        "label": "B",
        "direction": "미연은 정면 위쪽을 향해 소리치듯 응시하고, 아이들은 품에 안겨 시선을 아래로 향함.",
        "built_space": "상점가가 있는 밤거리 구조이나, 로케이션 레퍼런스의 우측 차양 건물과 특정 파이프 구조물이 일치하지 않음.",
        "entities": "미연, 앰버, 현우(청소년에 가깝게 묘사됨). 배경에 무릎 꿇은 인물들과 대형 미국 국기, 총구 화염 빛이 보임.",
        "hard_violations": [],
        "physics": "도로 바닥에 무릎을 꿇고 몸을 지탱함. 미연의 팔이 아이들을 감싸고 있으나, 배경 우측 인물의 하반신과 지면의 접촉 묘사가 다소 불분명함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "극도의 공포와 양팔 포옹은 분명하지만, 금지된 주변 인물이 다수 등장하고 구도가 넓으며 현우가 어린아이보다 청소년으로 보인다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "장소의 고정 시설과 어린 두 아이를 감싼 상체 중심 구도가 더 충실하지만, 추가 인물 네 명과 무릎 꿇기보다 주저앉은 자세 때문에 완전한 충족은 아니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "미연은 입을 벌리고 화면 왼쪽 위의 프레임 밖을 바라보며, 정확한 위협 대상은 보이지 않는다. 앰버와 현우는 고개와 시선을 아래로 내려 미연의 품 안을 향한다. 미연의 양팔은 각 아이의 바깥 어깨를 감싸 안쪽으로 끌어당긴다. 총기나 총구는 보이지 않아 배경 섬광의 사격 방향은 확인할 수 없다.",
        "built_space": "가족 뒤 왼쪽에 벽돌 기둥과 판자로 막힌 상점 창들이 있고, 오른쪽 위에 가까운 가로등 하나와 그 뒤 작은 가로등 하나가 보인다. 중앙 뒤에는 성조기 한 장이 있다. 참조의 낡은 거리 재료는 일부 따르지만, 참조에서 두드러지는 오른쪽 출입문과 차양의 관계는 확인되지 않는다. 가족의 다리와 주변 군중까지 보여 상체 중심 미디엄 숏보다 넓다. 흑백 야경은 반복된 악몽 지시에는 맞지만 별도의 낮 시간 잠금에는 맞지 않는다.",
        "entities": "중앙에는 검은 머리의 중년 여성, 금발 여자아이, 검은 머리 남자아이가 있다. 미연의 얼굴과 머리는 참조에 대체로 가깝고 공포 표정도 강하다. 앰버는 어린 금발 여자아이로 맞지만 현우는 상당히 큰 청소년으로 보여 어린 시절이라는 지시가 약하다. 세 사람의 겉옷은 참조의 단순한 둥근목 상의와 다르다. 성조기는 보인다. 가족 외에 화면 양쪽과 뒤로 적어도 여섯 명의 사람이 추가되어 있다. 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [
         "숏 텍스트가 허용한 미연·현우·앰버 외에 다수의 주변 인물을 추가했다."
        ],
        "physics": "앰버의 접힌 다리와 신발은 노면에 닿고, 현우도 접힌 하체를 바닥 가까이 둔다. 미연의 하체는 두 아이 사이에 일부 가려져 있지만 웅크린 무릎 자세와 양립한다. 두 손이 각 아이의 어깨와 상완에 실제로 닿아 포옹을 지탱한다. 지지 없이 떠 있는 신체는 없다. 배경의 작은 불꽃은 충돌이나 총격 섬광으로 해석할 수 있으나 발사원은 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "미연은 눈을 크게 뜨고 화면 오른쪽 위를 바라보며, 그 방향에 밝은 섬광이 보인다. 현우는 얼굴을 안쪽 아래로 돌려 미연과 앰버 쪽을 보고, 앰버는 오른쪽 아래로 고개를 숙인다. 양팔과 손이 두 아이의 몸을 중앙으로 모아 보호한다. 총구나 사수는 보이지 않으므로 섬광이 누구를 향한 사격인지는 판별할 수 없다.",
        "built_space": "왼쪽 전경에 배수구 하나와 연석이 있고, 오른쪽에는 낡은 외벽의 출입문 하나와 그 위 금속 차양 하나가 보인다. 그 뒤 벽돌 건물과 막힌 상점 창, 왼쪽 보도의 나무와 멀어지는 가로등 배열도 참조 위치와 잘 맞는다. 가족이 중앙 전경을 크게 차지해 A보다 보호 포옹 중심의 미디엄 숏에 가깝지만, 접힌 다리와 거리의 긴 원근도 상당히 포함한다. 왼쪽 보도에 두 명, 오른쪽 건물 앞에 두 명의 추가 인물이 있다. 흑백 야경은 악몽 지시와 일치하지만 낮 시간 잠금과는 충돌한다.",
        "entities": "미연은 참조와 유사한 검은 단발의 중년 여성이고, 눈과 입의 긴장으로 공포를 표현한다. 현우는 검은 머리의 어린 남자아이로 표현되어 어린 시절 지시에 A보다 가깝다. 앰버는 금발의 어린 여자아이지만 얼굴이 옆으로 돌아 참조 얼굴 전체의 일치 여부는 제한적으로만 확인된다. 미연은 밝은 긴소매 상의를 입어 참조의 어두운 상의와 다르다. 가족 외 네 명이 보이며, 성조기는 보이지 않는다. 읽을 수 있는 글씨는 없다.",
        "hard_violations": [
         "숏 텍스트에 없는 주변 인물 네 명을 거리 양쪽에 추가했다."
        ],
        "physics": "아이들은 무릎을 가슴 쪽으로 세우고 몸을 낮춘 자세이며, 하체는 화면 아래 노면 쪽으로 이어져 지지 없이 떠 있지는 않다. 미연의 팔은 두 아이의 상체 앞을 가로질러 손으로 옷과 몸을 붙잡고 있다. 포옹의 접촉은 자연스럽지만 세 사람의 자세는 무릎 꿇기보다 바닥에 주저앉아 다리를 모은 상태에 가깝다. 오른쪽 외벽 부근 섬광은 충격 불꽃으로 해석 가능하며, 별도의 떠 있는 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "극도의 공포와 양팔 포옹은 분명하지만, 금지된 주변 인물이 다수 등장하고 구도가 넓으며 현우가 어린아이보다 청소년으로 보인다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "장소의 고정 시설과 어린 두 아이를 감싼 상체 중심 구도가 더 충실하지만, 추가 인물 네 명과 무릎 꿇기보다 주저앉은 자세 때문에 완전한 충족은 아니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "미연은 입을 벌리고 화면 왼쪽 위의 프레임 밖을 바라보며, 정확한 위협 대상은 보이지 않는다. 앰버와 현우는 고개와 시선을 아래로 내려 미연의 품 안을 향한다. 미연의 양팔은 각 아이의 바깥 어깨를 감싸 안쪽으로 끌어당긴다. 총기나 총구는 보이지 않아 배경 섬광의 사격 방향은 확인할 수 없다.",
        "built_space": "가족 뒤 왼쪽에 벽돌 기둥과 판자로 막힌 상점 창들이 있고, 오른쪽 위에 가까운 가로등 하나와 그 뒤 작은 가로등 하나가 보인다. 중앙 뒤에는 성조기 한 장이 있다. 참조의 낡은 거리 재료는 일부 따르지만, 참조에서 두드러지는 오른쪽 출입문과 차양의 관계는 확인되지 않는다. 가족의 다리와 주변 군중까지 보여 상체 중심 미디엄 숏보다 넓다. 흑백 야경은 반복된 악몽 지시에는 맞지만 별도의 낮 시간 잠금에는 맞지 않는다.",
        "entities": "중앙에는 검은 머리의 중년 여성, 금발 여자아이, 검은 머리 남자아이가 있다. 미연의 얼굴과 머리는 참조에 대체로 가깝고 공포 표정도 강하다. 앰버는 어린 금발 여자아이로 맞지만 현우는 상당히 큰 청소년으로 보여 어린 시절이라는 지시가 약하다. 세 사람의 겉옷은 참조의 단순한 둥근목 상의와 다르다. 성조기는 보인다. 가족 외에 화면 양쪽과 뒤로 적어도 여섯 명의 사람이 추가되어 있다. 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [
         "숏 텍스트가 허용한 미연·현우·앰버 외에 다수의 주변 인물을 추가했다."
        ],
        "physics": "앰버의 접힌 다리와 신발은 노면에 닿고, 현우도 접힌 하체를 바닥 가까이 둔다. 미연의 하체는 두 아이 사이에 일부 가려져 있지만 웅크린 무릎 자세와 양립한다. 두 손이 각 아이의 어깨와 상완에 실제로 닿아 포옹을 지탱한다. 지지 없이 떠 있는 신체는 없다. 배경의 작은 불꽃은 충돌이나 총격 섬광으로 해석할 수 있으나 발사원은 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "미연은 눈을 크게 뜨고 화면 오른쪽 위를 바라보며, 그 방향에 밝은 섬광이 보인다. 현우는 얼굴을 안쪽 아래로 돌려 미연과 앰버 쪽을 보고, 앰버는 오른쪽 아래로 고개를 숙인다. 양팔과 손이 두 아이의 몸을 중앙으로 모아 보호한다. 총구나 사수는 보이지 않으므로 섬광이 누구를 향한 사격인지는 판별할 수 없다.",
        "built_space": "왼쪽 전경에 배수구 하나와 연석이 있고, 오른쪽에는 낡은 외벽의 출입문 하나와 그 위 금속 차양 하나가 보인다. 그 뒤 벽돌 건물과 막힌 상점 창, 왼쪽 보도의 나무와 멀어지는 가로등 배열도 참조 위치와 잘 맞는다. 가족이 중앙 전경을 크게 차지해 A보다 보호 포옹 중심의 미디엄 숏에 가깝지만, 접힌 다리와 거리의 긴 원근도 상당히 포함한다. 왼쪽 보도에 두 명, 오른쪽 건물 앞에 두 명의 추가 인물이 있다. 흑백 야경은 악몽 지시와 일치하지만 낮 시간 잠금과는 충돌한다.",
        "entities": "미연은 참조와 유사한 검은 단발의 중년 여성이고, 눈과 입의 긴장으로 공포를 표현한다. 현우는 검은 머리의 어린 남자아이로 표현되어 어린 시절 지시에 A보다 가깝다. 앰버는 금발의 어린 여자아이지만 얼굴이 옆으로 돌아 참조 얼굴 전체의 일치 여부는 제한적으로만 확인된다. 미연은 밝은 긴소매 상의를 입어 참조의 어두운 상의와 다르다. 가족 외 네 명이 보이며, 성조기는 보이지 않는다. 읽을 수 있는 글씨는 없다.",
        "hard_violations": [
         "숏 텍스트에 없는 주변 인물 네 명을 거리 양쪽에 추가했다."
        ],
        "physics": "아이들은 무릎을 가슴 쪽으로 세우고 몸을 낮춘 자세이며, 하체는 화면 아래 노면 쪽으로 이어져 지지 없이 떠 있지는 않다. 미연의 팔은 두 아이의 상체 앞을 가로질러 손으로 옷과 몸을 붙잡고 있다. 포옹의 접촉은 자연스럽지만 세 사람의 자세는 무릎 꿇기보다 바닥에 주저앉아 다리를 모은 상태에 가깝다. 오른쪽 외벽 부근 섬광은 충격 불꽃으로 해석 가능하며, 별도의 떠 있는 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.381
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.131
   },
   "violations": {
    "B": [
     "[gpt-high] 숏 텍스트가 허용한 미연·현우·앰버 외에 다수의 주변 인물을 추가했다."
    ],
    "A": [
     "[gpt-high] 숏 텍스트에 없는 주변 인물 네 명을 거리 양쪽에 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 1131
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "지정된 로케이션(하수구, 우측 차양 건물 등)을 완벽히 재현하고 현우를 어린아이로 정확히 묘사했으나, 대형 미국 국기가 누락되었고 상체 프레이밍보다 다소 넓게 촬영되었습니다.  ★위반: [gpt-high] 숏 텍스트에 없는 주변 인물 네 명을 거리 양쪽에 추가했다."
   },
   {
    "label": "B",
    "score": 1131,
    "verdict_ko": "미국 국기는 포함되었으나 로케이션 레퍼런스의 우측 건물 구조가 누락되었고, 현우가 지시보다 성숙하게 묘사되어 로케이션 및 인물 묘사 우선순위에서 밀립니다.  ★위반: [gpt-high] 숏 텍스트가 허용한 미연·현우·앰버 외에 다수의 주변 인물을 추가했다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L159B02.png",
    "asset_id": "dcf94519-9d42-4399-a0db-17ea44888173",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1113064>",
    "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-b83f-76e3-9e30-d13892012a90",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S8sh13__bgfirst_bg.png",
   "bg_asset_id": "8bcd8334-9652-44d4-9e5a-7ec29199889a",
   "bg_record_key": "S8sh13::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S8sh13::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:10:47.160955+00:00",
  "fingerprint": "a818df0825771e80cacddd670736578093a25bce9f4c3695360a96253afac3c8",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S8sh13_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S8sh13_sel.png",
  "source_sha256": "f9a5da644ac570fa391122d99174fa877961e9172b9703c670d1c8ee3ee4c253",
  "file": "S8sh13_cine.png",
  "staged_sha256": "6ede30ac40f101eb9c384382ef5a9ec83b7e1c5ad3a6fd47192b3ece52c7a64a",
  "latency_ms": 11430
 },
 "S9sh2::signage": {
  "fp": "315f6cf4fa6208e8",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::6fab8fc81ab4f62a": {
  "subjects": [],
  "subject_text": "난민 재판소 밖 쓰레기 더미\n허름한 건물 밖으로 각종 폐기물이 불규칙하게 쌓인 황폐한 공간. 쓰레기 더미가 주변 바닥을 뒤덮고 있다.",
  "identity": "canonical",
  "scope_id": "L161",
  "scope_role": "location_exterior",
  "scope_sha": "d71d8f809e2bbef2"
 },
 "S9sh2::bgfirst_bg": {
  "input_fingerprint": "941c58ae665c2cb8",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 고통으로 일그러진 표정을 한 채 한 손으로 가슴을 강하게 움켜쥐고 크게 숨을 들이마시며 입을 벌린 현우의 상체.\n\nLOCATION (lock): On rubbish-strewn ground outside the refugee tribunal at night, amid desolate heaps of waste.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Rubbish heaps outside the refugee court (Surround the place where 현우 awakens) — Broken-up portions remain behind his shoulders and along the frame edges; used as Locates his awakening in the desolate exterior without inventing individual discarded objects; Ground outside the court (현우 has just raised his body from it) — Visible beneath the lower edge of his seated torso; used as Preserves the physical transition from lying down to sitting up.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use setting-appropriate nighttime ambient light with restrained contrast and natural skin detail, clearly returning from the nightmare to direct observation.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 고통으로 일그러진 표정을 한 채 한 손으로 가슴을 강하게 움켜쥐고 크게 숨을 들이마시며 입을 벌린 현우의 상체.\n\nLOCATION (lock): On rubbish-strewn ground outside the refugee tribunal at night, amid desolate heaps of waste.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Rubbish heaps outside the refugee court (Surround the place where 현우 awakens) — Broken-up portions remain behind his shoulders and along the frame edges; used as Locates his awakening in the desolate exterior without inventing individual discarded objects; Ground outside the court (현우 has just raised his body from it) — Visible beneath the lower edge of his seated torso; used as Preserves the physical transition from lying down to sitting up.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use setting-appropriate nighttime ambient light with restrained contrast and natural skin detail, clearly returning from the nightmare to direct observation.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S9sh2__bgfirst_bg.png",
  "asset_id": "74193d07-10b1-46b2-a9cb-c0e151291eb3",
  "input_asset_ids": [
   "e2420bad-c4d6-4722-a5e2-886e38397c1a",
   "595d4bdd-0afb-4da7-95e6-a460d0259575"
  ]
 },
 "S9sh2": {
  "input_fingerprint": "426ee7ecf112ee2b",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 고통으로 일그러진 표정을 한 채 한 손으로 가슴을 강하게 움켜쥐고 크게 숨을 들이마시며 입을 벌린 현우의 상체.\n\nLOCATION (lock): On rubbish-strewn ground outside the refugee tribunal at night, amid desolate heaps of waste. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Rubbish heaps outside the refugee court (Surround the place where 현우 awakens) — Broken-up portions remain behind his shoulders and along the frame edges; used as Locates his awakening in the desolate exterior without inventing individual discarded objects; Ground outside the court (현우 has just raised his body from it) — Visible beneath the lower edge of his seated torso; used as Preserves the physical transition from lying down to sitting up.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use setting-appropriate nighttime ambient light with restrained contrast and natural skin detail, clearly returning from the nightmare to direct observation.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): It is night outside the refugee court, surrounded by desolate rubbish heaps. 현우: Hyunwoo is waking outside with his dog-bitten leg and injuries from the baton blow still untreated, though his awareness is returning. The stolen wallet containing the judge's identification and cash is already in his pocket.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 고통으로 일그러진 표정을 한 채 한 손으로 가슴을 강하게 움켜쥐고 크게 숨을 들이마시며 입을 벌린 현우의 상체.\n\nLOCATION (lock): On rubbish-strewn ground outside the refugee tribunal at night, amid desolate heaps of waste. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Rubbish heaps outside the refugee court (Surround the place where 현우 awakens) — Broken-up portions remain behind his shoulders and along the frame edges; used as Locates his awakening in the desolate exterior without inventing individual discarded objects; Ground outside the court (현우 has just raised his body from it) — Visible beneath the lower edge of his seated torso; used as Preserves the physical transition from lying down to sitting up.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use setting-appropriate nighttime ambient light with restrained contrast and natural skin detail, clearly returning from the nightmare to direct observation.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): It is night outside the refugee court, surrounded by desolate rubbish heaps. 현우: Hyunwoo is waking outside with his dog-bitten leg and injuries from the baton blow still untreated, though his awareness is returning. The stolen wallet containing the judge's identification and cash is already in his pocket.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 고통으로 일그러진 표정을 한 채 한 손으로 가슴을 강하게 움켜쥐고 크게 숨을 들이마시며 입을 벌린 현우의 상체.\n\nLOCATION (lock): On rubbish-strewn ground outside the refugee tribunal at night, amid desolate heaps of waste. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Rubbish heaps outside the refugee court (Surround the place where 현우 awakens) — Broken-up portions remain behind his shoulders and along the frame edges; used as Locates his awakening in the desolate exterior without inventing individual discarded objects; Ground outside the court (현우 has just raised his body from it) — Visible beneath the lower edge of his seated torso; used as Preserves the physical transition from lying down to sitting up.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use setting-appropriate nighttime ambient light with restrained contrast and natural skin detail, clearly returning from the nightmare to direct observation.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): It is night outside the refugee court, surrounded by desolate rubbish heaps. 현우: Hyunwoo is waking outside with his dog-bitten leg and injuries from the baton blow still untreated, though his awareness is returning. The stolen wallet containing the judge's identification and cash is already in his pocket.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S9sh2__bgfirst_bg.png",
     "asset_id": "74193d07-10b1-46b2-a9cb-c0e151291eb3",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S9sh2.png",
     "asset_id": "e2420bad-c4d6-4722-a5e2-886e38397c1a",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L161B01.png",
     "asset_id": "595d4bdd-0afb-4da7-95e6-a460d0259575",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "시선은 앞쪽 아래 빈 공간을 향함.",
    "built_space": "레퍼런스와 동일한 우측 건물 외벽, 출입문, 좌측의 거대한 쓰레기 더미가 정확히 구현됨.",
    "entities": "현우의 이목구비, 헤어스타일, 그리고 레퍼런스의 남색 반팔 티셔츠가 정확히 일치함.",
    "hard_violations": [],
    "physics": "바닥에 앉아 하체로 체중을 지탱하고 있으며, 한 손으로 가슴을 강하게 움켜쥔 상태임."
   },
   {
    "label": "B",
    "direction": "시선은 앞쪽 아래를 향함.",
    "built_space": "레퍼런스 사진에는 없는 형태의 창문과 다른 구조의 벽면을 가진 건물이 배경에 배치됨.",
    "entities": "현우의 얼굴 형태는 유사하나, 지정된 남색이 아닌 밝은 회색/베이지색 티셔츠를 입고 있음.",
    "hard_violations": [],
    "physics": "바닥에 앉아 자세를 유지하며 한 손을 가슴에 올리고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 로케이션의 건축물과 캐릭터의 남색 의상을 완벽히 유지하며 고통스러워하는 표정과 동작을 정확히 묘사함."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "표정과 동작은 준수하나, 로케이션 배경 건물이 레퍼런스와 전혀 다르고 캐릭터의 의상 색상도 불일치함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 앞쪽 아래 빈 공간을 향함.",
        "built_space": "레퍼런스와 동일한 우측 건물 외벽, 출입문, 좌측의 거대한 쓰레기 더미가 정확히 구현됨.",
        "entities": "현우의 이목구비, 헤어스타일, 그리고 레퍼런스의 남색 반팔 티셔츠가 정확히 일치함.",
        "hard_violations": [],
        "physics": "바닥에 앉아 하체로 체중을 지탱하고 있으며, 한 손으로 가슴을 강하게 움켜쥔 상태임."
       },
       {
        "label": "B",
        "direction": "시선은 앞쪽 아래를 향함.",
        "built_space": "레퍼런스 사진에는 없는 형태의 창문과 다른 구조의 벽면을 가진 건물이 배경에 배치됨.",
        "entities": "현우의 얼굴 형태는 유사하나, 지정된 남색이 아닌 밝은 회색/베이지색 티셔츠를 입고 있음.",
        "hard_violations": [],
        "physics": "바닥에 앉아 자세를 유지하며 한 손을 가슴에 올리고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 로케이션의 건축물과 캐릭터의 남색 의상을 완벽히 유지하며 고통스러워하는 표정과 동작을 정확히 묘사함."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "표정과 동작은 준수하나, 로케이션 배경 건물이 레퍼런스와 전혀 다르고 캐릭터의 의상 색상도 불일치함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 앞쪽 아래 빈 공간을 향함.",
        "built_space": "레퍼런스와 동일한 우측 건물 외벽, 출입문, 좌측의 거대한 쓰레기 더미가 정확히 구현됨.",
        "entities": "현우의 이목구비, 헤어스타일, 그리고 레퍼런스의 남색 반팔 티셔츠가 정확히 일치함.",
        "hard_violations": [],
        "physics": "바닥에 앉아 하체로 체중을 지탱하고 있으며, 한 손으로 가슴을 강하게 움켜쥔 상태임."
       },
       {
        "label": "B",
        "direction": "시선은 앞쪽 아래를 향함.",
        "built_space": "레퍼런스 사진에는 없는 형태의 창문과 다른 구조의 벽면을 가진 건물이 배경에 배치됨.",
        "entities": "현우의 얼굴 형태는 유사하나, 지정된 남색이 아닌 밝은 회색/베이지색 티셔츠를 입고 있음.",
        "hard_violations": [],
        "physics": "바닥에 앉아 자세를 유지하며 한 손을 가슴에 올리고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "상체 중심의 미디엄 숏과 고통스러운 들숨 동작은 충실하지만, 회갈색 티셔츠가 인물 참조와 다르고 건물 및 쓰레기 배치의 장소 일치도가 낮습니다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "가슴을 움켜쥔 채 입을 벌린 착석 상체를 보여주며, 남색 티셔츠와 인물 외형, 야간 법원 외벽·출입문·쓰레기 공터가 참조에 더 충실합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴과 시선은 화면 왼쪽 앞을 향하며 카메라를 똑바로 보지 않습니다. 오른손은 자신의 왼쪽 가슴에 닿아 옷을 움켜쥐고 있습니다. 미간을 찌푸리고 입을 크게 벌려 통증 속에서 숨을 들이마시는 순간으로 읽힙니다. 별도로 조준하거나 바라봐야 하는 외부 대상은 없습니다.",
        "built_space": "인물은 쓰레기가 쌓인 건물 앞 바닥에 앉아 있고, 허리 아래와 양옆에 지면이 보입니다. 뒤쪽 오른편에는 회색 출입문 하나와 그 위 차양 하나, 오른쪽 가장자리에는 수직 배관이 보입니다. 왼쪽에는 셔터형 개구부 하나와 상부 창들이 드러납니다. 낡은 콘크리트 재질은 참조와 유사하지만, 외벽이 배경 대부분을 차지하고 쓰레기가 양옆 가까이 밀집해 참조 공터의 공간적 특징은 덜 유지됩니다.",
        "entities": "인물은 한 명이며, 앳된 동아시아계 남성의 외형과 헝클어진 검은 머리는 현우 설정에 부합합니다. 국적은 외관만으로 확인할 수 없습니다. 얼굴은 참조와 유사하지만 티셔츠가 참조의 남색 대신 회갈색입니다. 얼굴과 옷에 때가 보이며 치료 장비는 없습니다. 개에게 물린 다리 부위와 주머니 속 지갑·신분증·현금은 보이지 않아 확인할 수 없습니다. 쓰레기 더미와 거친 지면은 보이고 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "하단의 굽힌 다리와 낮게 놓인 골반이 지면에 앉은 자세를 이룹니다. 왼팔은 아래로 내려가지만 손의 접촉점은 프레임 밖입니다. 오른손 손가락이 가슴의 천을 잡고 주름을 만들어 힘을 주는 동작이 자연스럽습니다. 몸이나 물체가 근거 없이 떠 있는 모습은 없습니다."
       },
       {
        "label": "B",
        "direction": "고개를 조금 숙여 화면 오른쪽 아래 지면 쪽을 바라봅니다. 오른손은 자신의 왼쪽 가슴에서 티셔츠를 단단히 움켜쥐고 있습니다. 찌푸린 미간과 벌어진 입이 통증을 느끼며 크게 숨을 들이마시는 행동과 맞습니다. 지정된 외부 시선 대상이나 조준 대상은 없습니다.",
        "built_space": "인물은 공터와 건물 옆 포장면의 경계 가까이에 앉아 있습니다. 오른쪽 외벽에는 회색 출입문 하나, 직사각형 차양 하나, 문 오른편의 수직 배관 묶음이 보입니다. 왼쪽 전경과 뒤쪽 울타리 앞에는 쓰레기 더미가 있고, 멀리 나무와 낮은 건물이 배치되어 참조 장소의 구성을 잘 유지합니다. 상체 아래 지면도 보입니다. 중앙 뒤의 녹색 상자는 참조와 배치가 다르지만 고정 시설의 중복이나 불가능한 공간 구성은 보이지 않습니다.",
        "entities": "등장인물은 현우로 읽히는 젊은 동아시아계 남성 한 명입니다. 앳된 얼굴, 헝클어진 검은 머리, 마른 체격과 남색 라운드넥 티셔츠가 인물 참조에 가깝습니다. 한국계 미국인이라는 국적·배경 자체는 영상만으로 확인할 수 없습니다. 얼굴과 옷에는 때와 상처처럼 보이는 흔적이 있고 치료된 모습은 없습니다. 다리의 물린 상처와 주머니 속 지갑 및 내용물은 노출되지 않습니다. 야간 공터, 쓰레기, 법원 외벽이 보이며 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "골반과 굽힌 다리가 지면에 놓이고, 왼손은 화면 아래 오른쪽의 바닥을 짚어 상체를 지탱합니다. 오른손은 가슴의 천을 실제로 집어 당기며, 몸을 막 일으킨 뒤 한 손으로 버티는 자세가 가능합니다. 쓰레기와 상자도 지면이나 더미 위에 놓여 있으며 떠 있는 물체는 없습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "상체 중심의 미디엄 숏과 고통스러운 들숨 동작은 충실하지만, 회갈색 티셔츠가 인물 참조와 다르고 건물 및 쓰레기 배치의 장소 일치도가 낮습니다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "가슴을 움켜쥔 채 입을 벌린 착석 상체를 보여주며, 남색 티셔츠와 인물 외형, 야간 법원 외벽·출입문·쓰레기 공터가 참조에 더 충실합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴과 시선은 화면 왼쪽 앞을 향하며 카메라를 똑바로 보지 않습니다. 오른손은 자신의 왼쪽 가슴에 닿아 옷을 움켜쥐고 있습니다. 미간을 찌푸리고 입을 크게 벌려 통증 속에서 숨을 들이마시는 순간으로 읽힙니다. 별도로 조준하거나 바라봐야 하는 외부 대상은 없습니다.",
        "built_space": "인물은 쓰레기가 쌓인 건물 앞 바닥에 앉아 있고, 허리 아래와 양옆에 지면이 보입니다. 뒤쪽 오른편에는 회색 출입문 하나와 그 위 차양 하나, 오른쪽 가장자리에는 수직 배관이 보입니다. 왼쪽에는 셔터형 개구부 하나와 상부 창들이 드러납니다. 낡은 콘크리트 재질은 참조와 유사하지만, 외벽이 배경 대부분을 차지하고 쓰레기가 양옆 가까이 밀집해 참조 공터의 공간적 특징은 덜 유지됩니다.",
        "entities": "인물은 한 명이며, 앳된 동아시아계 남성의 외형과 헝클어진 검은 머리는 현우 설정에 부합합니다. 국적은 외관만으로 확인할 수 없습니다. 얼굴은 참조와 유사하지만 티셔츠가 참조의 남색 대신 회갈색입니다. 얼굴과 옷에 때가 보이며 치료 장비는 없습니다. 개에게 물린 다리 부위와 주머니 속 지갑·신분증·현금은 보이지 않아 확인할 수 없습니다. 쓰레기 더미와 거친 지면은 보이고 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "하단의 굽힌 다리와 낮게 놓인 골반이 지면에 앉은 자세를 이룹니다. 왼팔은 아래로 내려가지만 손의 접촉점은 프레임 밖입니다. 오른손 손가락이 가슴의 천을 잡고 주름을 만들어 힘을 주는 동작이 자연스럽습니다. 몸이나 물체가 근거 없이 떠 있는 모습은 없습니다."
       },
       {
        "label": "A",
        "direction": "고개를 조금 숙여 화면 오른쪽 아래 지면 쪽을 바라봅니다. 오른손은 자신의 왼쪽 가슴에서 티셔츠를 단단히 움켜쥐고 있습니다. 찌푸린 미간과 벌어진 입이 통증을 느끼며 크게 숨을 들이마시는 행동과 맞습니다. 지정된 외부 시선 대상이나 조준 대상은 없습니다.",
        "built_space": "인물은 공터와 건물 옆 포장면의 경계 가까이에 앉아 있습니다. 오른쪽 외벽에는 회색 출입문 하나, 직사각형 차양 하나, 문 오른편의 수직 배관 묶음이 보입니다. 왼쪽 전경과 뒤쪽 울타리 앞에는 쓰레기 더미가 있고, 멀리 나무와 낮은 건물이 배치되어 참조 장소의 구성을 잘 유지합니다. 상체 아래 지면도 보입니다. 중앙 뒤의 녹색 상자는 참조와 배치가 다르지만 고정 시설의 중복이나 불가능한 공간 구성은 보이지 않습니다.",
        "entities": "등장인물은 현우로 읽히는 젊은 동아시아계 남성 한 명입니다. 앳된 얼굴, 헝클어진 검은 머리, 마른 체격과 남색 라운드넥 티셔츠가 인물 참조에 가깝습니다. 한국계 미국인이라는 국적·배경 자체는 영상만으로 확인할 수 없습니다. 얼굴과 옷에는 때와 상처처럼 보이는 흔적이 있고 치료된 모습은 없습니다. 다리의 물린 상처와 주머니 속 지갑 및 내용물은 노출되지 않습니다. 야간 공터, 쓰레기, 법원 외벽이 보이며 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "골반과 굽힌 다리가 지면에 놓이고, 왼손은 화면 아래 오른쪽의 바닥을 짚어 상체를 지탱합니다. 오른손은 가슴의 천을 실제로 집어 당기며, 몸을 막 일으킨 뒤 한 손으로 버티는 자세가 가능합니다. 쓰레기와 상자도 지면이나 더미 위에 놓여 있으며 떠 있는 물체는 없습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.349
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.349
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1349
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지정된 로케이션의 건축물과 캐릭터의 남색 의상을 완벽히 유지하며 고통스러워하는 표정과 동작을 정확히 묘사함."
   },
   {
    "label": "B",
    "score": 1349,
    "verdict_ko": "표정과 동작은 준수하나, 로케이션 배경 건물이 레퍼런스와 전혀 다르고 캐릭터의 의상 색상도 불일치함."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L161B01.png",
    "asset_id": "595d4bdd-0afb-4da7-95e6-a460d0259575",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-bbcf-753c-8e57-97e4c40ccd94",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S9sh2__bgfirst_bg.png",
   "bg_asset_id": "74193d07-10b1-46b2-a9cb-c0e151291eb3",
   "bg_record_key": "S9sh2::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S9sh2::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:12:20.936175+00:00",
  "fingerprint": "5212cd80285a21c67d39e3c234a07700e57663cd9ffd11f15e69523b91c2d310",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S9sh2_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S9sh2_sel.png",
  "source_sha256": "847a56eef37fa647df29882026412d65023361c08ac275e13c5bd0cb25625d20",
  "file": "S9sh2_cine.png",
  "staged_sha256": "0ec9814716ae0c2d547d3ff778bd345dd57991514688daf54ec0545a00ddd712",
  "latency_ms": 11137
 },
 "S9sh5::signage": {
  "fp": "cc021fcecb804d26",
  "inscriptions": [],
  "cues": [
   {
    "text_native": "",
    "source": "scene_text_implied",
    "source_quote": "신분증"
   }
  ],
  "dropped": []
 },
 "S9sh5": {
  "input_fingerprint": "85e67aeddfb95d2e",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 지폐가 빠져나가 신분증만 덩그러니 남은 빈 가죽 지갑이 허공을 향해 뻗은 현우의 손끝에서 막 떨어져 날아가는 mid-action 순간.\n\nLOCATION (lock): Beside the desolate rubbish heaps outside the refugee tribunal at night, where the emptied wallet is discarded. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Leather wallet (Released from the fingertips with the money removed and identification remaining) — Its open inner side is angled toward the camera during the first instant of separation; used as Marks the release while remaining small enough for hand and body scale to read naturally; Judge's identification (Still inside the discarded wallet) — The identification-bearing face is partly visible within the open wallet; no specific text is supplied; used as Distinguishes the retained identification from the removed money; Rubbish heaps (Remain around the court exterior) — Visible beyond the hand and the wallet's open travel space; used as Maintains exterior continuity and prevents the action from becoming an isolated product image.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding nighttime ambient treatment, using controlled tonal separation to distinguish the fingers, leather wallet, and remaining identification.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The discarded wallet still contains the judge's identification but no longer contains its cash. Nighttime rubbish heaps surround the exterior of the refugee court. 현우: Hyunwoo has retained the cash removed from the wallet and is alert, but still limps on his dog-bitten leg and retains the earlier head injury.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 지폐가 빠져나가 신분증만 덩그러니 남은 빈 가죽 지갑이 허공을 향해 뻗은 현우의 손끝에서 막 떨어져 날아가는 mid-action 순간.\n\nLOCATION (lock): Beside the desolate rubbish heaps outside the refugee tribunal at night, where the emptied wallet is discarded. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Leather wallet (Released from the fingertips with the money removed and identification remaining) — Its open inner side is angled toward the camera during the first instant of separation; used as Marks the release while remaining small enough for hand and body scale to read naturally; Judge's identification (Still inside the discarded wallet) — The identification-bearing face is partly visible within the open wallet; no specific text is supplied; used as Distinguishes the retained identification from the removed money; Rubbish heaps (Remain around the court exterior) — Visible beyond the hand and the wallet's open travel space; used as Maintains exterior continuity and prevents the action from becoming an isolated product image.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding nighttime ambient treatment, using controlled tonal separation to distinguish the fingers, leather wallet, and remaining identification.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The discarded wallet still contains the judge's identification but no longer contains its cash. Nighttime rubbish heaps surround the exterior of the refugee court. 현우: Hyunwoo has retained the cash removed from the wallet and is alert, but still limps on his dog-bitten leg and retains the earlier head injury.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 지폐가 빠져나가 신분증만 덩그러니 남은 빈 가죽 지갑이 허공을 향해 뻗은 현우의 손끝에서 막 떨어져 날아가는 mid-action 순간.\n\nLOCATION (lock): Beside the desolate rubbish heaps outside the refugee tribunal at night, where the emptied wallet is discarded. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Leather wallet (Released from the fingertips with the money removed and identification remaining) — Its open inner side is angled toward the camera during the first instant of separation; used as Marks the release while remaining small enough for hand and body scale to read naturally; Judge's identification (Still inside the discarded wallet) — The identification-bearing face is partly visible within the open wallet; no specific text is supplied; used as Distinguishes the retained identification from the removed money; Rubbish heaps (Remain around the court exterior) — Visible beyond the hand and the wallet's open travel space; used as Maintains exterior continuity and prevents the action from becoming an isolated product image.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding nighttime ambient treatment, using controlled tonal separation to distinguish the fingers, leather wallet, and remaining identification.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The discarded wallet still contains the judge's identification but no longer contains its cash. Nighttime rubbish heaps surround the exterior of the refugee court. 현우: Hyunwoo has retained the cash removed from the wallet and is alert, but still limps on his dog-bitten leg and retains the earlier head injury.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "손은 아래로 향하고 있으며, 지갑은 내부를 카메라 쪽으로 향한 채 떨어지고 있습니다.",
    "built_space": "오른쪽의 콘크리트 건물과 왼쪽의 쓰레기 더미 배치가 이전 샷의 환경과 완벽히 일치합니다.",
    "entities": "현우의 손, 펼쳐진 가죽 지갑, 신분증(현우의 얼굴과 읽을 수 있는 한글 텍스트 포함)이 명확히 보이며 지폐는 없습니다.",
    "hard_violations": [
     "[gpt-high] 신분증의 한글 항목 글자가 판독 가능한 상태로 노출되어, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 명시적 조건을 위반한다."
    ],
    "physics": "지갑은 손끝에서 방금 떨어져 물리적인 지지 없이 공중에 떠 있습니다."
   },
   {
    "label": "B",
    "direction": "손이 아래로 뻗어 있고, 지갑은 공중에서 회전하며 떨어집니다.",
    "built_space": "건물 외벽과 쓰레기장이 배경에 위치하나, 이전 샷의 구조물 디테일과 다소 차이가 있습니다.",
    "entities": "현우의 손, 가죽 지갑, 신분증이 보이며, 우측 가장자리에 파란 셔츠를 입은 몸통 일부가 나타납니다.",
    "hard_violations": [
     "[gemini-pro] physically impossible anatomy or staging (오른쪽 가장자리의 몸통과 지갑을 떨어뜨리는 위쪽 팔이 해부학적으로 연결되지 않는 별개의 신체로 묘사됨)"
    ],
    "physics": "지갑은 공중에 떠 있으며, 손 역시 허공에 위치해 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "신분증의 글씨가 선명하게 렌더링되어 텍스트 불가 지시를 위반했으나, 이전 샷과의 배경 연속성 및 클로즈업 액션 구도를 훌륭하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "화면 우측의 몸통과 지갑을 떨어뜨리는 팔이 해부학적으로 연결되지 않는 심각한 구조적 오류가 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "손은 아래로 향하고 있으며, 지갑은 내부를 카메라 쪽으로 향한 채 떨어지고 있습니다.",
        "built_space": "오른쪽의 콘크리트 건물과 왼쪽의 쓰레기 더미 배치가 이전 샷의 환경과 완벽히 일치합니다.",
        "entities": "현우의 손, 펼쳐진 가죽 지갑, 신분증(현우의 얼굴과 읽을 수 있는 한글 텍스트 포함)이 명확히 보이며 지폐는 없습니다.",
        "hard_violations": [],
        "physics": "지갑은 손끝에서 방금 떨어져 물리적인 지지 없이 공중에 떠 있습니다."
       },
       {
        "label": "B",
        "direction": "손이 아래로 뻗어 있고, 지갑은 공중에서 회전하며 떨어집니다.",
        "built_space": "건물 외벽과 쓰레기장이 배경에 위치하나, 이전 샷의 구조물 디테일과 다소 차이가 있습니다.",
        "entities": "현우의 손, 가죽 지갑, 신분증이 보이며, 우측 가장자리에 파란 셔츠를 입은 몸통 일부가 나타납니다.",
        "hard_violations": [
         "physically impossible anatomy or staging (오른쪽 가장자리의 몸통과 지갑을 떨어뜨리는 위쪽 팔이 해부학적으로 연결되지 않는 별개의 신체로 묘사됨)"
        ],
        "physics": "지갑은 공중에 떠 있으며, 손 역시 허공에 위치해 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "신분증의 글씨가 선명하게 렌더링되어 텍스트 불가 지시를 위반했으나, 이전 샷과의 배경 연속성 및 클로즈업 액션 구도를 훌륭하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "화면 우측의 몸통과 지갑을 떨어뜨리는 팔이 해부학적으로 연결되지 않는 심각한 구조적 오류가 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "손은 아래로 향하고 있으며, 지갑은 내부를 카메라 쪽으로 향한 채 떨어지고 있습니다.",
        "built_space": "오른쪽의 콘크리트 건물과 왼쪽의 쓰레기 더미 배치가 이전 샷의 환경과 완벽히 일치합니다.",
        "entities": "현우의 손, 펼쳐진 가죽 지갑, 신분증(현우의 얼굴과 읽을 수 있는 한글 텍스트 포함)이 명확히 보이며 지폐는 없습니다.",
        "hard_violations": [],
        "physics": "지갑은 손끝에서 방금 떨어져 물리적인 지지 없이 공중에 떠 있습니다."
       },
       {
        "label": "B",
        "direction": "손이 아래로 뻗어 있고, 지갑은 공중에서 회전하며 떨어집니다.",
        "built_space": "건물 외벽과 쓰레기장이 배경에 위치하나, 이전 샷의 구조물 디테일과 다소 차이가 있습니다.",
        "entities": "현우의 손, 가죽 지갑, 신분증이 보이며, 우측 가장자리에 파란 셔츠를 입은 몸통 일부가 나타납니다.",
        "hard_violations": [
         "physically impossible anatomy or staging (오른쪽 가장자리의 몸통과 지갑을 떨어뜨리는 위쪽 팔이 해부학적으로 연결되지 않는 별개의 신체로 묘사됨)"
        ],
        "physics": "지갑은 공중에 떠 있으며, 손 역시 허공에 위치해 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "손끝에서 분리된 지갑과 카메라를 향한 빈 내부·신분증을 근접 촬영으로 구현하고 글자도 흐렸지만, 지갑이 다소 크게 강조되고 배경 건물의 연속성은 약하다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "쓰레기장 배경과 지갑을 놓는 동작은 잘 맞지만, 신분증의 한글 항목이 읽혀 판독 가능한 글자를 전면 금지한 조건을 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "오른쪽 위에서 나온 손의 손가락들이 왼쪽 아래 지갑 쪽으로 펼쳐져 있다. 지갑은 손끝 아래로 분리되어 쓰레기 더미 앞의 빈 공간으로 떨어지는 모습이며, 열린 안쪽과 신분증 면이 카메라를 향한다. 얼굴과 시선은 프레임 밖이다.",
        "built_space": "뒤쪽에 낡은 콘크리트 외벽과 큰 문 하나, 상단 창 일부가 보이고, 아래에는 목재 팔레트·검은 봉투·판재·상자가 쌓여 있다. 인물의 몸통 일부는 오른쪽 가장자리에 있다. 이전 장면의 재료와 야간 외부 환경은 유지하지만, 외벽이 배경 대부분을 차지하고 문의 형태와 배치도 달라 같은 장소라는 연결은 다소 약하다. 중복된 고정 시설이나 반사는 보이지 않는다.",
        "entities": "젊은 남성의 맨손과 팔, 오른쪽 가장자리의 남색 티셔츠와 어두운 바지가 보인다. 얼굴이 없어 현우의 정확한 얼굴·머리·한국계 미국인 정체성은 확인할 수 없지만 보이는 의상과 피부는 참조와 모순되지 않는다. 낡은 갈색 가죽 반지갑 안에 인물 사진이 있는 신분증 한 장이 남아 있고 지폐는 보이지 않는다. 신분증 글자는 흐려 판독하기 어렵다. 보유한 현금과 다리·머리 부상은 이 근접 구도 밖이다.",
        "hard_violations": [],
        "physics": "팔은 화면 오른쪽의 몸으로 이어지고 손은 물건을 놓은 뒤처럼 펼쳐져 있다. 지갑은 손과 떨어져 있지만 바로 위의 펼쳐진 손이 방출의 출발점을 제공하고, 아래 쓰레기 더미와 지면이 낙하 경로에 있어 근거 없는 부유가 아니다. 신분증은 지갑의 투명 수납칸에 지지되어 있다. 지갑은 손에 비해 다소 크게 강조되지만 물리적으로 불가능한 크기나 자세는 아니다."
       },
       {
        "label": "B",
        "direction": "오른쪽 위에서 뻗은 손이 왼쪽 아래의 열린 지갑을 향해 손가락을 펼치고 있다. 지갑은 손끝 바로 아래에서 떨어져 쓰레기장 쪽으로 향하는 순간으로 읽힌다. 신분증이 든 내부는 카메라를 향하며, 얼굴과 시선은 보이지 않는다.",
        "built_space": "왼쪽 쓰레기 더미에 녹색 상자 두 개와 목재 팔레트가 보이고, 오른쪽 중경에는 녹색 상자 하나가 놓여 있다. 뒤에는 울타리와 앙상한 나무, 낮은 건물들이 있으며 오른쪽 끝에 가까운 외벽 일부가 보인다. 이전 장면의 쓰레기장 배치와 밤 분위기를 비교적 잘 이어 간다. 손과 지갑 앞에는 낙하할 공간이 있고, 불가능한 반사나 중복된 고정 시설은 보이지 않는다.",
        "entities": "젊은 남성으로 보이는 맨손과 팔만 등장하며 추가 인물은 없다. 얼굴·머리·의상은 구도 밖이므로 정확한 인물 일치 여부를 확인할 수 없다. 낡은 짙은 갈색 가죽 반지갑에는 사진 신분증 한 장이 남아 있고 현금은 보이지 않는다. 다만 신분증의 제목과 항목에 판독 가능한 한글이 남아 있다. 주변 봉투·상자·팔레트는 실제 쓰레기 더미로 보인다.",
        "hard_violations": [
         "신분증의 한글 항목 글자가 판독 가능한 상태로 노출되어, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 명시적 조건을 위반한다."
        ],
        "physics": "손목과 팔이 화면 밖 인물 쪽으로 자연스럽게 이어지고 다섯 손가락이 지갑을 놓은 형태로 벌어져 있다. 지갑과 손끝 사이의 짧은 간격은 방출 직후를 뒷받침하며, 아래 지면으로 떨어질 수 있어 근거 없는 부유가 아니다. 신분증은 수납칸 안에 고정되어 있고 배경 쓰레기들은 지면이나 다른 쓰레기에 받쳐져 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "손끝에서 분리된 지갑과 카메라를 향한 빈 내부·신분증을 근접 촬영으로 구현하고 글자도 흐렸지만, 지갑이 다소 크게 강조되고 배경 건물의 연속성은 약하다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "쓰레기장 배경과 지갑을 놓는 동작은 잘 맞지만, 신분증의 한글 항목이 읽혀 판독 가능한 글자를 전면 금지한 조건을 위반한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "오른쪽 위에서 나온 손의 손가락들이 왼쪽 아래 지갑 쪽으로 펼쳐져 있다. 지갑은 손끝 아래로 분리되어 쓰레기 더미 앞의 빈 공간으로 떨어지는 모습이며, 열린 안쪽과 신분증 면이 카메라를 향한다. 얼굴과 시선은 프레임 밖이다.",
        "built_space": "뒤쪽에 낡은 콘크리트 외벽과 큰 문 하나, 상단 창 일부가 보이고, 아래에는 목재 팔레트·검은 봉투·판재·상자가 쌓여 있다. 인물의 몸통 일부는 오른쪽 가장자리에 있다. 이전 장면의 재료와 야간 외부 환경은 유지하지만, 외벽이 배경 대부분을 차지하고 문의 형태와 배치도 달라 같은 장소라는 연결은 다소 약하다. 중복된 고정 시설이나 반사는 보이지 않는다.",
        "entities": "젊은 남성의 맨손과 팔, 오른쪽 가장자리의 남색 티셔츠와 어두운 바지가 보인다. 얼굴이 없어 현우의 정확한 얼굴·머리·한국계 미국인 정체성은 확인할 수 없지만 보이는 의상과 피부는 참조와 모순되지 않는다. 낡은 갈색 가죽 반지갑 안에 인물 사진이 있는 신분증 한 장이 남아 있고 지폐는 보이지 않는다. 신분증 글자는 흐려 판독하기 어렵다. 보유한 현금과 다리·머리 부상은 이 근접 구도 밖이다.",
        "hard_violations": [],
        "physics": "팔은 화면 오른쪽의 몸으로 이어지고 손은 물건을 놓은 뒤처럼 펼쳐져 있다. 지갑은 손과 떨어져 있지만 바로 위의 펼쳐진 손이 방출의 출발점을 제공하고, 아래 쓰레기 더미와 지면이 낙하 경로에 있어 근거 없는 부유가 아니다. 신분증은 지갑의 투명 수납칸에 지지되어 있다. 지갑은 손에 비해 다소 크게 강조되지만 물리적으로 불가능한 크기나 자세는 아니다."
       },
       {
        "label": "A",
        "direction": "오른쪽 위에서 뻗은 손이 왼쪽 아래의 열린 지갑을 향해 손가락을 펼치고 있다. 지갑은 손끝 바로 아래에서 떨어져 쓰레기장 쪽으로 향하는 순간으로 읽힌다. 신분증이 든 내부는 카메라를 향하며, 얼굴과 시선은 보이지 않는다.",
        "built_space": "왼쪽 쓰레기 더미에 녹색 상자 두 개와 목재 팔레트가 보이고, 오른쪽 중경에는 녹색 상자 하나가 놓여 있다. 뒤에는 울타리와 앙상한 나무, 낮은 건물들이 있으며 오른쪽 끝에 가까운 외벽 일부가 보인다. 이전 장면의 쓰레기장 배치와 밤 분위기를 비교적 잘 이어 간다. 손과 지갑 앞에는 낙하할 공간이 있고, 불가능한 반사나 중복된 고정 시설은 보이지 않는다.",
        "entities": "젊은 남성으로 보이는 맨손과 팔만 등장하며 추가 인물은 없다. 얼굴·머리·의상은 구도 밖이므로 정확한 인물 일치 여부를 확인할 수 없다. 낡은 짙은 갈색 가죽 반지갑에는 사진 신분증 한 장이 남아 있고 현금은 보이지 않는다. 다만 신분증의 제목과 항목에 판독 가능한 한글이 남아 있다. 주변 봉투·상자·팔레트는 실제 쓰레기 더미로 보인다.",
        "hard_violations": [
         "신분증의 한글 항목 글자가 판독 가능한 상태로 노출되어, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 명시적 조건을 위반한다."
        ],
        "physics": "손목과 팔이 화면 밖 인물 쪽으로 자연스럽게 이어지고 다섯 손가락이 지갑을 놓은 형태로 벌어져 있다. 지갑과 손끝 사이의 짧은 간격은 방출 직후를 뒷받침하며, 아래 지면으로 떨어질 수 있어 근거 없는 부유가 아니다. 신분증은 수납칸 안에 고정되어 있고 배경 쓰레기들은 지면이나 다른 쓰레기에 받쳐져 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.5,
    "B": 1.429
   },
   "adjusted": {
    "A": 1.25,
    "B": 1.179
   },
   "violations": {
    "B": [
     "[gemini-pro] physically impossible anatomy or staging (오른쪽 가장자리의 몸통과 지갑을 떨어뜨리는 위쪽 팔이 해부학적으로 연결되지 않는 별개의 신체로 묘사됨)"
    ],
    "A": [
     "[gpt-high] 신분증의 한글 항목 글자가 판독 가능한 상태로 노출되어, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 명시적 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1250,
   "B": 1179
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1250,
    "verdict_ko": "신분증의 글씨가 선명하게 렌더링되어 텍스트 불가 지시를 위반했으나, 이전 샷과의 배경 연속성 및 클로즈업 액션 구도를 훌륭하게 구현했습니다.  ★위반: [gpt-high] 신분증의 한글 항목 글자가 판독 가능한 상태로 노출되어, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 명시적 조건을 위반한다."
   },
   {
    "label": "B",
    "score": 1179,
    "verdict_ko": "화면 우측의 몸통과 지갑을 떨어뜨리는 팔이 해부학적으로 연결되지 않는 심각한 구조적 오류가 발생했습니다.  ★위반: [gemini-pro] physically impossible anatomy or staging (오른쪽 가장자리의 몸통과 지갑을 떨어뜨리는 위쪽 팔이 해부학적으로 연결되지 않는 별개의 신체로 묘사됨)"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S9sh2_sel.png",
    "asset_id": "269f778c-9601-4c47-8fbf-630c40589a69",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-bf3c-7265-b94e-69c18d756844",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S9sh2"
  }
 },
 "S9sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:13:37.832683+00:00",
  "fingerprint": "91c540930bf97713404cf3cc5a03014e325754e7d4df80ed86614280c212d805",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S9sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S9sh5_sel.png",
  "source_sha256": "0576cc05bc0ecc3c9ba19e15b6cad9297375150e5ab5100549e4b4076f50f0f4",
  "file": "S9sh5_cine.png",
  "staged_sha256": "c7d89d6a3793deaca606d4ff2240a4d92a2431a53a1fae0574ebee85e76005ac",
  "latency_ms": 13335
 },
 "S10sh3::signage": {
  "fp": "c5a6300222d7094b",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::f65f2d6badb3b97d": {
  "subjects": [],
  "subject_text": "허름한 골목 식당 내부\n낡은 탁자와 의자가 놓인 소규모 식당. 출입구에서 떨어진 안쪽 구석에 외진 좌석이 있다.",
  "identity": "canonical",
  "scope_id": "L162",
  "scope_role": "location_interior",
  "scope_sha": "e4027065f7112ec7"
 },
 "S10sh3::bgfirst_bg": {
  "input_fingerprint": "8929de6d3f312bbf",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 돈 봉투를 집어 들려는 남자의 손등 위를 다급한 힘으로 꽉 내리눌러 제압하고 있는 현우의 굳은 손 클로즈업.\n\nLOCATION (lock): At a secluded table inside a shabby alley restaurant, under dim restaurant lighting at night.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Money envelope (Within reach but prevented from being freely taken) — Lies on the tabletop beside the overlapping hands; used as Keeps the stakes of the physical interruption visible without competing with the hands; Secluded restaurant table (Occupied by the two men during the transaction) — The top and near edge are visible from the inherited oblique side; used as Connects the opposing forearms and anchors the continuous conversation axis.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate restaurant ambient light renders the hand pressure and envelope clearly with restrained, continuous contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 돈 봉투를 집어 들려는 남자의 손등 위를 다급한 힘으로 꽉 내리눌러 제압하고 있는 현우의 굳은 손 클로즈업.\n\nLOCATION (lock): At a secluded table inside a shabby alley restaurant, under dim restaurant lighting at night.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Money envelope (Within reach but prevented from being freely taken) — Lies on the tabletop beside the overlapping hands; used as Keeps the stakes of the physical interruption visible without competing with the hands; Secluded restaurant table (Occupied by the two men during the transaction) — The top and near edge are visible from the inherited oblique side; used as Connects the opposing forearms and anchors the continuous conversation axis.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate restaurant ambient light renders the hand pressure and envelope clearly with restrained, continuous contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S10sh3__bgfirst_bg.png",
  "asset_id": "33a03a51-7f51-4fab-9373-3f1601df1a45",
  "input_asset_ids": [
   "2c3ae081-487d-467f-b6d8-4c65de90d776",
   "7a16f2f4-24b6-4a89-a184-9204ad75cbe4"
  ]
 },
 "S10sh3": {
  "input_fingerprint": "1c233810f5e7a87a",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 돈 봉투를 집어 들려는 남자의 손등 위를 다급한 힘으로 꽉 내리눌러 제압하고 있는 현우의 굳은 손 클로즈업.\n\nLOCATION (lock): At a secluded table inside a shabby alley restaurant, under dim restaurant lighting at night. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Money envelope (Within reach but prevented from being freely taken) — Lies on the tabletop beside the overlapping hands; used as Keeps the stakes of the physical interruption visible without competing with the hands; Secluded restaurant table (Occupied by the two men during the transaction) — The top and near edge are visible from the inherited oblique side; used as Connects the opposing forearms and anchors the continuous conversation axis.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate restaurant ambient light renders the hand pressure and envelope clearly with restrained, continuous contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A cash-filled envelope is on the table in an out-of-the-way part of the shabby restaurant. 현우: Hyunwoo is seated at the table and has taken hold of the money envelope again; his dog-bitten leg and earlier head injury remain untreated.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 돈 봉투를 집어 들려는 남자의 손등 위를 다급한 힘으로 꽉 내리눌러 제압하고 있는 현우의 굳은 손 클로즈업.\n\nLOCATION (lock): At a secluded table inside a shabby alley restaurant, under dim restaurant lighting at night. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Money envelope (Within reach but prevented from being freely taken) — Lies on the tabletop beside the overlapping hands; used as Keeps the stakes of the physical interruption visible without competing with the hands; Secluded restaurant table (Occupied by the two men during the transaction) — The top and near edge are visible from the inherited oblique side; used as Connects the opposing forearms and anchors the continuous conversation axis.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate restaurant ambient light renders the hand pressure and envelope clearly with restrained, continuous contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A cash-filled envelope is on the table in an out-of-the-way part of the shabby restaurant. 현우: Hyunwoo is seated at the table and has taken hold of the money envelope again; his dog-bitten leg and earlier head injury remain untreated.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 돈 봉투를 집어 들려는 남자의 손등 위를 다급한 힘으로 꽉 내리눌러 제압하고 있는 현우의 굳은 손 클로즈업.\n\nLOCATION (lock): At a secluded table inside a shabby alley restaurant, under dim restaurant lighting at night. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Money envelope (Within reach but prevented from being freely taken) — Lies on the tabletop beside the overlapping hands; used as Keeps the stakes of the physical interruption visible without competing with the hands; Secluded restaurant table (Occupied by the two men during the transaction) — The top and near edge are visible from the inherited oblique side; used as Connects the opposing forearms and anchors the continuous conversation axis.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate restaurant ambient light renders the hand pressure and envelope clearly with restrained, continuous contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A cash-filled envelope is on the table in an out-of-the-way part of the shabby restaurant. 현우: Hyunwoo is seated at the table and has taken hold of the money envelope again; his dog-bitten leg and earlier head injury remain untreated.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S10sh3__bgfirst_bg.png",
     "asset_id": "33a03a51-7f51-4fab-9373-3f1601df1a45",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S10sh3.png",
     "asset_id": "2c3ae081-487d-467f-b6d8-4c65de90d776",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L162B01.png",
     "asset_id": "7a16f2f4-24b6-4a89-a184-9204ad75cbe4",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "왼쪽 남성(현우)의 시선과 손이 테이블 위의 상대방 손등을 향해 누르고 있음.",
    "built_space": "식당 내부의 나무 테이블, 배경의 의자와 집기들이 레퍼런스 사진과 유사하게 배치됨.",
    "entities": "왼쪽 인물의 얼굴은 현우 레퍼런스와 일치하나, 돈 봉투는 손이 닿지 않은 곳에 방치되어 있음.",
    "hard_violations": [],
    "physics": "두 사람의 팔이 테이블 위에 지지되어 있으며, 누르는 힘이 물리적으로 표현됨."
   },
   {
    "label": "B",
    "direction": "오른쪽에서 뻗어 나온 팔이 돈 봉투 위에 손을 얹은 왼쪽 인물의 팔목을 향해 제압하고 있음.",
    "built_space": "식당 테이블과 배경의 벽면, 메뉴판 등이 레퍼런스 공간 안에서 적절한 각도로 보임.",
    "entities": "오른쪽 팔은 현우의 레퍼런스 의상(네이비 티셔츠)과 일치하며, 왼쪽 인물의 손 아래에 돈 봉투가 놓여 있음.",
    "hard_violations": [
     "[gpt-high] 배경 메뉴판의 ‘공기밥’과 ‘1,000’이 판독 가능하여, 읽을 수 있는 글자를 어디에도 두지 말라는 조건을 위반한다."
    ],
    "physics": "테이블 면을 지지대로 삼아 봉투를 짚은 손과 그 위를 강하게 쥐는 손의 접촉이 자연스러움."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "요구된 손 클로즈업 샷 크기에 더 부합하며, 돈 봉투를 잡으려는 찰나의 동작과 이를 제압하는 상황을 긴장감 있게 잘 구현했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "프레이밍이 손 클로즈업이 아닌 상반신이 포함된 넓은 샷이며, 상대방의 손이 돈 봉투와 떨어져 있어 지시된 상황과 어긋납니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 남성(현우)의 시선과 손이 테이블 위의 상대방 손등을 향해 누르고 있음.",
        "built_space": "식당 내부의 나무 테이블, 배경의 의자와 집기들이 레퍼런스 사진과 유사하게 배치됨.",
        "entities": "왼쪽 인물의 얼굴은 현우 레퍼런스와 일치하나, 돈 봉투는 손이 닿지 않은 곳에 방치되어 있음.",
        "hard_violations": [],
        "physics": "두 사람의 팔이 테이블 위에 지지되어 있으며, 누르는 힘이 물리적으로 표현됨."
       },
       {
        "label": "B",
        "direction": "오른쪽에서 뻗어 나온 팔이 돈 봉투 위에 손을 얹은 왼쪽 인물의 팔목을 향해 제압하고 있음.",
        "built_space": "식당 테이블과 배경의 벽면, 메뉴판 등이 레퍼런스 공간 안에서 적절한 각도로 보임.",
        "entities": "오른쪽 팔은 현우의 레퍼런스 의상(네이비 티셔츠)과 일치하며, 왼쪽 인물의 손 아래에 돈 봉투가 놓여 있음.",
        "hard_violations": [],
        "physics": "테이블 면을 지지대로 삼아 봉투를 짚은 손과 그 위를 강하게 쥐는 손의 접촉이 자연스러움."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "요구된 손 클로즈업 샷 크기에 더 부합하며, 돈 봉투를 잡으려는 찰나의 동작과 이를 제압하는 상황을 긴장감 있게 잘 구현했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "프레이밍이 손 클로즈업이 아닌 상반신이 포함된 넓은 샷이며, 상대방의 손이 돈 봉투와 떨어져 있어 지시된 상황과 어긋납니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 남성(현우)의 시선과 손이 테이블 위의 상대방 손등을 향해 누르고 있음.",
        "built_space": "식당 내부의 나무 테이블, 배경의 의자와 집기들이 레퍼런스 사진과 유사하게 배치됨.",
        "entities": "왼쪽 인물의 얼굴은 현우 레퍼런스와 일치하나, 돈 봉투는 손이 닿지 않은 곳에 방치되어 있음.",
        "hard_violations": [],
        "physics": "두 사람의 팔이 테이블 위에 지지되어 있으며, 누르는 힘이 물리적으로 표현됨."
       },
       {
        "label": "B",
        "direction": "오른쪽에서 뻗어 나온 팔이 돈 봉투 위에 손을 얹은 왼쪽 인물의 팔목을 향해 제압하고 있음.",
        "built_space": "식당 테이블과 배경의 벽면, 메뉴판 등이 레퍼런스 공간 안에서 적절한 각도로 보임.",
        "entities": "오른쪽 팔은 현우의 레퍼런스 의상(네이비 티셔츠)과 일치하며, 왼쪽 인물의 손 아래에 돈 봉투가 놓여 있음.",
        "hard_violations": [],
        "physics": "테이블 면을 지지대로 삼아 봉투를 짚은 손과 그 위를 강하게 쥐는 손의 접촉이 자연스러움."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "손 중심 클로즈업과 봉투를 향한 동작은 더 정확하지만, 배경 메뉴판의 읽을 수 있는 글자와 가격이 명시적 금지 조건을 위반해 탈락한다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "치명적 위반 없이 손을 제압하는 접촉을 구현했지만, 상반신까지 넓어진 구도와 봉투에서 빗나간 상대 손의 방향, 현우의 다른 의상이 충실도를 낮춘다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 남자의 손가락은 오른쪽 아래 돈 봉투를 향하고 실제로 봉투 위에 닿아 있다. 오른쪽에서 뻗은 현우의 손은 상대 손등 중앙보다는 손목과 손등의 경계를 아래로 누르며 감싼다. 눈은 프레임 밖이므로 시선은 확인할 수 없다.",
        "built_space": "전경에 낡은 목재 탁자 하나와 가까운 모서리가 보이고, 뒤에는 갈색 금속틀 의자 등받이 하나가 보인다. 오른쪽 뒤에 금속 수저통 하나와 목재 휴지함 하나, 벽에는 메뉴판과 사진 일부가 있다. 회색 하단과 때 탄 밝은 상단으로 나뉜 벽, 탁자 재질은 장소 참조와 부합한다. 두 사람의 팔은 탁자 양쪽에서 들어오며 좌석 자체는 대부분 화면 밖이다.",
        "entities": "남색 반팔을 입은 현우의 팔과 손, 상대 남자의 팔과 손 및 얼굴 일부, 갈색 종이봉투 하나가 보인다. 상대의 다른 손도 뒤에 일부 보이지만 추가 인물로 읽히지는 않는다. 현우의 옷은 참조와 맞으며 얼굴·머리·상처는 이 구도에서 확인할 수 없다. 봉투 속 현금의 국가와 액면은 보이지 않는다. 배경 메뉴판에는 ‘공기밥’과 ‘1,000’이 읽힌다.",
        "hard_violations": [
         "배경 메뉴판의 ‘공기밥’과 ‘1,000’이 판독 가능하여, 읽을 수 있는 글자를 어디에도 두지 말라는 조건을 위반한다."
        ],
        "physics": "상대의 팔과 손은 탁자 및 그 위 봉투에 지지되고, 현우의 손바닥과 엄지가 상대 손목에 접촉한다. 팔의 긴장과 아래로 누르는 접촉은 실제 제압 동작으로 가능하다. 봉투도 탁자 위에 놓여 있어 지지 없는 물체는 없다. 다만 손등을 꽉 내리누르기보다는 손목을 붙잡는 동작에 가깝다."
       },
       {
        "label": "B",
        "direction": "왼쪽 현우는 겹친 손 쪽을 내려다보며 오른쪽 남자의 손목과 손등 윗부분을 누른다. 상대 손가락은 화면 왼쪽 아래로 뻗어 있고 봉투는 오른쪽 아래에 있어, 봉투를 집으려던 손의 진행 방향이 봉투와 맞지 않는다. 오른쪽 남자는 눈이 잘려 시선을 확인할 수 없다.",
        "built_space": "전경 거래용 탁자 하나의 상판과 가까운 모서리, 뒤쪽 별도 탁자 하나와 갈색 의자 등받이 하나가 보인다. 뒤에는 주방 입구, 병 상자, 달력, 걸린 검은 천, 벽 사진이 보인다. 낡은 투톤 벽과 식당 재료감은 장소 참조를 따른다. 두 사람은 전경 탁자를 사이에 두고 팔을 올려놓았으며, 보이는 범위에서 좌석이나 설비와 충돌하지 않는다. 다만 두 사람의 얼굴 일부와 상반신까지 포함해 요구한 손 클로즈업보다 넓다.",
        "entities": "왼쪽에 젊은 동아시아계 남성 현우, 오른쪽에 거래 상대 남성의 하관과 상반신, 중앙에 겹친 손, 오른쪽 아래에 종이봉투 하나가 보인다. 현우의 검은 머리와 앳된 인상은 대체로 부합하지만 참조의 남색 반팔 대신 짙은 회색 겉옷과 안쪽 티셔츠를 입었다. 봉투 안 현금은 보이지 않는다. 배경 인쇄물은 흐려 뚜렷하게 읽히는 문구가 없다.",
        "hard_violations": [],
        "physics": "상대의 손과 팔은 탁자 위에 지지되고, 현우의 손가락과 손바닥은 그 손목 및 손등에 밀착한다. 상대의 다른 팔꿈치도 상판에 놓여 있다. 봉투는 자체 무게로 탁자 위에 놓여 있으며 떠 있는 신체나 물체는 없다. 누르고 붙잡는 동작은 가능하지만, 봉투를 향한 집기 동작이 중단된 순간이라는 인과는 약하다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "손 중심 클로즈업과 봉투를 향한 동작은 더 정확하지만, 배경 메뉴판의 읽을 수 있는 글자와 가격이 명시적 금지 조건을 위반해 탈락한다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "치명적 위반 없이 손을 제압하는 접촉을 구현했지만, 상반신까지 넓어진 구도와 봉투에서 빗나간 상대 손의 방향, 현우의 다른 의상이 충실도를 낮춘다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽 남자의 손가락은 오른쪽 아래 돈 봉투를 향하고 실제로 봉투 위에 닿아 있다. 오른쪽에서 뻗은 현우의 손은 상대 손등 중앙보다는 손목과 손등의 경계를 아래로 누르며 감싼다. 눈은 프레임 밖이므로 시선은 확인할 수 없다.",
        "built_space": "전경에 낡은 목재 탁자 하나와 가까운 모서리가 보이고, 뒤에는 갈색 금속틀 의자 등받이 하나가 보인다. 오른쪽 뒤에 금속 수저통 하나와 목재 휴지함 하나, 벽에는 메뉴판과 사진 일부가 있다. 회색 하단과 때 탄 밝은 상단으로 나뉜 벽, 탁자 재질은 장소 참조와 부합한다. 두 사람의 팔은 탁자 양쪽에서 들어오며 좌석 자체는 대부분 화면 밖이다.",
        "entities": "남색 반팔을 입은 현우의 팔과 손, 상대 남자의 팔과 손 및 얼굴 일부, 갈색 종이봉투 하나가 보인다. 상대의 다른 손도 뒤에 일부 보이지만 추가 인물로 읽히지는 않는다. 현우의 옷은 참조와 맞으며 얼굴·머리·상처는 이 구도에서 확인할 수 없다. 봉투 속 현금의 국가와 액면은 보이지 않는다. 배경 메뉴판에는 ‘공기밥’과 ‘1,000’이 읽힌다.",
        "hard_violations": [
         "배경 메뉴판의 ‘공기밥’과 ‘1,000’이 판독 가능하여, 읽을 수 있는 글자를 어디에도 두지 말라는 조건을 위반한다."
        ],
        "physics": "상대의 팔과 손은 탁자 및 그 위 봉투에 지지되고, 현우의 손바닥과 엄지가 상대 손목에 접촉한다. 팔의 긴장과 아래로 누르는 접촉은 실제 제압 동작으로 가능하다. 봉투도 탁자 위에 놓여 있어 지지 없는 물체는 없다. 다만 손등을 꽉 내리누르기보다는 손목을 붙잡는 동작에 가깝다."
       },
       {
        "label": "A",
        "direction": "왼쪽 현우는 겹친 손 쪽을 내려다보며 오른쪽 남자의 손목과 손등 윗부분을 누른다. 상대 손가락은 화면 왼쪽 아래로 뻗어 있고 봉투는 오른쪽 아래에 있어, 봉투를 집으려던 손의 진행 방향이 봉투와 맞지 않는다. 오른쪽 남자는 눈이 잘려 시선을 확인할 수 없다.",
        "built_space": "전경 거래용 탁자 하나의 상판과 가까운 모서리, 뒤쪽 별도 탁자 하나와 갈색 의자 등받이 하나가 보인다. 뒤에는 주방 입구, 병 상자, 달력, 걸린 검은 천, 벽 사진이 보인다. 낡은 투톤 벽과 식당 재료감은 장소 참조를 따른다. 두 사람은 전경 탁자를 사이에 두고 팔을 올려놓았으며, 보이는 범위에서 좌석이나 설비와 충돌하지 않는다. 다만 두 사람의 얼굴 일부와 상반신까지 포함해 요구한 손 클로즈업보다 넓다.",
        "entities": "왼쪽에 젊은 동아시아계 남성 현우, 오른쪽에 거래 상대 남성의 하관과 상반신, 중앙에 겹친 손, 오른쪽 아래에 종이봉투 하나가 보인다. 현우의 검은 머리와 앳된 인상은 대체로 부합하지만 참조의 남색 반팔 대신 짙은 회색 겉옷과 안쪽 티셔츠를 입었다. 봉투 안 현금은 보이지 않는다. 배경 인쇄물은 흐려 뚜렷하게 읽히는 문구가 없다.",
        "hard_violations": [],
        "physics": "상대의 손과 팔은 탁자 위에 지지되고, 현우의 손가락과 손바닥은 그 손목 및 손등에 밀착한다. 상대의 다른 팔꿈치도 상판에 놓여 있다. 봉투는 자체 무게로 탁자 위에 놓여 있으며 떠 있는 신체나 물체는 없다. 누르고 붙잡는 동작은 가능하지만, 봉투를 향한 집기 동작이 중단된 순간이라는 인과는 약하다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.667,
    "B": 1.5
   },
   "adjusted": {
    "A": 1.667,
    "B": 1.25
   },
   "violations": {
    "B": [
     "[gpt-high] 배경 메뉴판의 ‘공기밥’과 ‘1,000’이 판독 가능하여, 읽을 수 있는 글자를 어디에도 두지 말라는 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1250,
   "A": 1667
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1250,
    "verdict_ko": "요구된 손 클로즈업 샷 크기에 더 부합하며, 돈 봉투를 잡으려는 찰나의 동작과 이를 제압하는 상황을 긴장감 있게 잘 구현했습니다.  ★위반: [gpt-high] 배경 메뉴판의 ‘공기밥’과 ‘1,000’이 판독 가능하여, 읽을 수 있는 글자를 어디에도 두지 말라는 조건을 위반한다."
   },
   {
    "label": "A",
    "score": 1667,
    "verdict_ko": "프레이밍이 손 클로즈업이 아닌 상반신이 포함된 넓은 샷이며, 상대방의 손이 돈 봉투와 떨어져 있어 지시된 상황과 어긋납니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L162B01.png",
    "asset_id": "7a16f2f4-24b6-4a89-a184-9204ad75cbe4",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-c0f5-7b1c-9c26-02ca315f704c",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S10sh3__bgfirst_bg.png",
   "bg_asset_id": "33a03a51-7f51-4fab-9373-3f1601df1a45",
   "bg_record_key": "S10sh3::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S10sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:15:22.139584+00:00",
  "fingerprint": "04916b94fa7ddf964a1631052c7dad6134065822e7909398227324c22cfe4aa0",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S10sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S10sh3_sel.png",
  "source_sha256": "60b5556697cc08017e9ca2ed253dc385464b8861a8fe554a97df9c71408c55d3",
  "file": "S10sh3_cine.png",
  "staged_sha256": "3368ee62345d015675120792732ca22dc64df35af3b36f3337bd07fb39ab540a",
  "latency_ms": 10301
 },
 "S10sh6::signage": {
  "fp": "7794cd3d370fa59e",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S10sh6": {
  "input_fingerprint": "554129c56eaacbe6",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 사이렌 소리가 울리는 가운데 현우를 향해 조롱하듯 여유롭게 웃고 있는 남자의 얼굴.\n\nLOCATION (lock): At the secluded dining table inside the shabby alley restaurant, with dim lighting during the nighttime conversation. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Secluded restaurant table (Remains between the two seated men) — Only a narrow edge is retained below the face and foreground shoulder; used as Maintains continuity with the interrupted transaction while allowing the mocking expression to lead.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the restaurant's ambient illumination and restrained contrast, with no flashing or colored light inferred from the siren.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The nighttime meeting remains at the table in the shabby restaurant's secluded area.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 사이렌 소리가 울리는 가운데 현우를 향해 조롱하듯 여유롭게 웃고 있는 남자의 얼굴.\n\nLOCATION (lock): At the secluded dining table inside the shabby alley restaurant, with dim lighting during the nighttime conversation. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Secluded restaurant table (Remains between the two seated men) — Only a narrow edge is retained below the face and foreground shoulder; used as Maintains continuity with the interrupted transaction while allowing the mocking expression to lead.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the restaurant's ambient illumination and restrained contrast, with no flashing or colored light inferred from the siren.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The nighttime meeting remains at the table in the shabby restaurant's secluded area.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 사이렌 소리가 울리는 가운데 현우를 향해 조롱하듯 여유롭게 웃고 있는 남자의 얼굴.\n\nLOCATION (lock): At the secluded dining table inside the shabby alley restaurant, with dim lighting during the nighttime conversation. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Secluded restaurant table (Remains between the two seated men) — Only a narrow edge is retained below the face and foreground shoulder; used as Maintains continuity with the interrupted transaction while allowing the mocking expression to lead.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the restaurant's ambient illumination and restrained contrast, with no flashing or colored light inferred from the siren.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The nighttime meeting remains at the table in the shabby restaurant's secluded area.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "흰 셔츠의 남자가 우측 전경의 인물을 바라봄.",
    "built_space": "인물의 좌우 위치가 뒤바뀜. 남자가 좌측에 있고, 이전 샷에서 현우 뒤에 있던 달력과 앞치마가 남자 뒤에 잘못 배치됨.",
    "entities": "흰 셔츠의 남자와 어두운 셔츠를 입은 현우의 어깨가 보임.",
    "hard_violations": [
     "[gemini-pro] 인물 좌우 위치 및 배경 배치가 뒤바뀐 공간 연속성 위반",
     "[gemini-pro] 지지대 없이 비정상적으로 세워져 있는 전경의 서류 봉투",
     "[gpt-high] 왼쪽 아래 파란 병 상자에 읽을 수 있는 'Cass' 상표가 보여, 읽을 수 있는 글자와 로고를 모두 금지한 지시를 위반한다."
    ],
    "physics": "전경 중앙의 서류 봉투가 잡는 손이나 지지대 없이 세로로 서 있음."
   },
   {
    "label": "B",
    "direction": "흰 셔츠의 남자가 좌측 전경의 현우를 향해 시선을 던짐.",
    "built_space": "현우가 좌측, 남자가 우측에 배치되어 이전 샷의 구도와 배경(좌측 달력, 우측 액자)을 정확히 유지함.",
    "entities": "조롱하는 표정의 흰 셔츠 남자와 어두운 셔츠, 헝클어진 머리의 현우가 정확히 구현됨.",
    "hard_violations": [],
    "physics": "인물들이 테이블 구조에 맞게 자연스럽게 앉아 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "이전 샷의 인물 배치와 배경 구조를 완벽히 유지하며 조롱하는 미소를 정확히 포착했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "인물의 좌우 위치가 바뀌는 심각한 공간 연속성 위반이 있으며, 전경의 봉투가 부자연스럽게 서 있습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "흰 셔츠의 남자가 우측 전경의 인물을 바라봄.",
        "built_space": "인물의 좌우 위치가 뒤바뀜. 남자가 좌측에 있고, 이전 샷에서 현우 뒤에 있던 달력과 앞치마가 남자 뒤에 잘못 배치됨.",
        "entities": "흰 셔츠의 남자와 어두운 셔츠를 입은 현우의 어깨가 보임.",
        "hard_violations": [
         "인물 좌우 위치 및 배경 배치가 뒤바뀐 공간 연속성 위반",
         "지지대 없이 비정상적으로 세워져 있는 전경의 서류 봉투"
        ],
        "physics": "전경 중앙의 서류 봉투가 잡는 손이나 지지대 없이 세로로 서 있음."
       },
       {
        "label": "B",
        "direction": "흰 셔츠의 남자가 좌측 전경의 현우를 향해 시선을 던짐.",
        "built_space": "현우가 좌측, 남자가 우측에 배치되어 이전 샷의 구도와 배경(좌측 달력, 우측 액자)을 정확히 유지함.",
        "entities": "조롱하는 표정의 흰 셔츠 남자와 어두운 셔츠, 헝클어진 머리의 현우가 정확히 구현됨.",
        "hard_violations": [],
        "physics": "인물들이 테이블 구조에 맞게 자연스럽게 앉아 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "이전 샷의 인물 배치와 배경 구조를 완벽히 유지하며 조롱하는 미소를 정확히 포착했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "인물의 좌우 위치가 바뀌는 심각한 공간 연속성 위반이 있으며, 전경의 봉투가 부자연스럽게 서 있습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "흰 셔츠의 남자가 우측 전경의 인물을 바라봄.",
        "built_space": "인물의 좌우 위치가 뒤바뀜. 남자가 좌측에 있고, 이전 샷에서 현우 뒤에 있던 달력과 앞치마가 남자 뒤에 잘못 배치됨.",
        "entities": "흰 셔츠의 남자와 어두운 셔츠를 입은 현우의 어깨가 보임.",
        "hard_violations": [
         "인물 좌우 위치 및 배경 배치가 뒤바뀐 공간 연속성 위반",
         "지지대 없이 비정상적으로 세워져 있는 전경의 서류 봉투"
        ],
        "physics": "전경 중앙의 서류 봉투가 잡는 손이나 지지대 없이 세로로 서 있음."
       },
       {
        "label": "B",
        "direction": "흰 셔츠의 남자가 좌측 전경의 현우를 향해 시선을 던짐.",
        "built_space": "현우가 좌측, 남자가 우측에 배치되어 이전 샷의 구도와 배경(좌측 달력, 우측 액자)을 정확히 유지함.",
        "entities": "조롱하는 표정의 흰 셔츠 남자와 어두운 셔츠, 헝클어진 머리의 현우가 정확히 구현됨.",
        "hard_violations": [],
        "physics": "인물들이 테이블 구조에 맞게 자연스럽게 앉아 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "현우에게 향한 시선과 여유로운 비웃음, 기존 식당의 조명을 잘 유지하며 B보다 얼굴 중심 구도에 가깝지만, 상반신 비중이 크고 두 사람 사이의 좁은 탁자 가장자리가 보이지 않는다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "현우를 보며 웃는 행동은 맞지만 얼굴보다 상반신과 전경 어깨의 비중이 크고, 병 상자의 읽을 수 있는 상표가 명시적인 문자·로고 금지를 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "남자의 얼굴과 두 눈은 화면 왼쪽 전경의 현우를 향한다. 현우의 뒤통수와 귀도 상대 남자를 마주하는 방향이며, 렌즈를 바라보는 구도는 아니다. 좁힌 눈과 살짝 드러낸 치아가 여유로운 조롱에 부합한다.",
        "built_space": "왼쪽 뒤에 냉장고 한 대, 중앙 왼쪽에 달력 하나와 걸린 검은 천 하나, 아래에 파란 병 상자들이 보인다. 오른쪽에는 벽 액자 일부와 배경 탁자 하나가 있다. 낡은 투톤 벽과 기존 설비의 재질은 이전 장면과 이어진다. 두 사람은 서로 마주하지만, 요구된 두 사람 사이 탁자의 좁은 가장자리는 전경과 화면 하단 밖으로 가려졌다.",
        "entities": "주인공 남자는 검은 머리의 한국인으로 보이는 성인 남성이며 이전 장면의 밝은 회베이지색 셔츠를 유지한다. 현우는 검은 머리와 어두운 겉옷을 입은 뒤통수·어깨만 보여 얼굴과 정확한 나이는 확인할 수 없다. 추가 인물은 없고, 봉투와 손은 프레임 밖이다. 배경 표기는 흐려 읽기 어렵다. 조명은 따뜻하고 어두우며 사이렌을 이유로 한 색광은 없다.",
        "hard_violations": [],
        "physics": "남자의 머리는 목과 상체에 자연스럽게 연결되고, 앞으로 기울인 상체와 아래로 내려간 팔은 탁자 맞은편 착석 자세와 양립한다. 좌면과 손의 접점은 크롭 밖이므로 직접 확인되지 않지만, 떠 있는 몸이나 무지지 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "남자는 머리를 기울여 오른쪽 전경의 현우를 바라보며 이를 드러내 웃는다. 현우도 남자를 향해 앉아 있어 시선의 대상은 맞는다. 표정은 조롱으로 해석할 수 있지만 A보다 환한 웃음에 가깝다.",
        "built_space": "왼쪽에 냉장고 한 대, 달력 하나, 걸린 검은 천 하나와 파란 병 상자들이 있고 오른쪽 벽에는 풍경 액자 하나와 위쪽 액자 가장자리가 보인다. 낡은 투톤 벽은 이전 식당과 유사하다. 맞은편 착석 관계는 유지되지만, 얼굴뿐 아니라 몸통과 현우의 큰 어깨까지 포함해 요구된 얼굴 클로즈업보다 넓다. 두 사람 사이 탁자 가장자리도 분명하게 드러나지 않는다.",
        "entities": "밝은 회베이지색 셔츠의 한국인으로 보이는 성인 남성과 검은 머리·어두운 겉옷의 현우 일부가 보인다. 현우의 얼굴은 가려져 신원 세부를 검증할 수 없다. 하단 중앙에는 이전 장면의 갈색 봉투로 보이는 물체가 있으나, 평평하게 놓였던 상태와 달리 세워진 모습이다. 왼쪽 아래 병 상자에는 읽을 수 있는 영문 상표가 있다. 추가 인물이나 사이렌 색광은 없다.",
        "hard_violations": [
         "왼쪽 아래 파란 병 상자에 읽을 수 있는 'Cass' 상표가 보여, 읽을 수 있는 글자와 로고를 모두 금지한 지시를 위반한다."
        ],
        "physics": "남자의 기울어진 머리와 앞으로 숙인 상체는 자연스럽고 현우의 전경 어깨도 정상적인 착석 자세로 보인다. 봉투 하단과 손은 프레임 밖이라 봉투를 받치는 정확한 접점은 확인되지 않는다. 화면 밖 탁자에 놓이거나 손에 들릴 수 있는 위치이므로 무지지 부유라고 단정할 근거는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "현우에게 향한 시선과 여유로운 비웃음, 기존 식당의 조명을 잘 유지하며 B보다 얼굴 중심 구도에 가깝지만, 상반신 비중이 크고 두 사람 사이의 좁은 탁자 가장자리가 보이지 않는다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "현우를 보며 웃는 행동은 맞지만 얼굴보다 상반신과 전경 어깨의 비중이 크고, 병 상자의 읽을 수 있는 상표가 명시적인 문자·로고 금지를 위반한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "남자의 얼굴과 두 눈은 화면 왼쪽 전경의 현우를 향한다. 현우의 뒤통수와 귀도 상대 남자를 마주하는 방향이며, 렌즈를 바라보는 구도는 아니다. 좁힌 눈과 살짝 드러낸 치아가 여유로운 조롱에 부합한다.",
        "built_space": "왼쪽 뒤에 냉장고 한 대, 중앙 왼쪽에 달력 하나와 걸린 검은 천 하나, 아래에 파란 병 상자들이 보인다. 오른쪽에는 벽 액자 일부와 배경 탁자 하나가 있다. 낡은 투톤 벽과 기존 설비의 재질은 이전 장면과 이어진다. 두 사람은 서로 마주하지만, 요구된 두 사람 사이 탁자의 좁은 가장자리는 전경과 화면 하단 밖으로 가려졌다.",
        "entities": "주인공 남자는 검은 머리의 한국인으로 보이는 성인 남성이며 이전 장면의 밝은 회베이지색 셔츠를 유지한다. 현우는 검은 머리와 어두운 겉옷을 입은 뒤통수·어깨만 보여 얼굴과 정확한 나이는 확인할 수 없다. 추가 인물은 없고, 봉투와 손은 프레임 밖이다. 배경 표기는 흐려 읽기 어렵다. 조명은 따뜻하고 어두우며 사이렌을 이유로 한 색광은 없다.",
        "hard_violations": [],
        "physics": "남자의 머리는 목과 상체에 자연스럽게 연결되고, 앞으로 기울인 상체와 아래로 내려간 팔은 탁자 맞은편 착석 자세와 양립한다. 좌면과 손의 접점은 크롭 밖이므로 직접 확인되지 않지만, 떠 있는 몸이나 무지지 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "남자는 머리를 기울여 오른쪽 전경의 현우를 바라보며 이를 드러내 웃는다. 현우도 남자를 향해 앉아 있어 시선의 대상은 맞는다. 표정은 조롱으로 해석할 수 있지만 A보다 환한 웃음에 가깝다.",
        "built_space": "왼쪽에 냉장고 한 대, 달력 하나, 걸린 검은 천 하나와 파란 병 상자들이 있고 오른쪽 벽에는 풍경 액자 하나와 위쪽 액자 가장자리가 보인다. 낡은 투톤 벽은 이전 식당과 유사하다. 맞은편 착석 관계는 유지되지만, 얼굴뿐 아니라 몸통과 현우의 큰 어깨까지 포함해 요구된 얼굴 클로즈업보다 넓다. 두 사람 사이 탁자 가장자리도 분명하게 드러나지 않는다.",
        "entities": "밝은 회베이지색 셔츠의 한국인으로 보이는 성인 남성과 검은 머리·어두운 겉옷의 현우 일부가 보인다. 현우의 얼굴은 가려져 신원 세부를 검증할 수 없다. 하단 중앙에는 이전 장면의 갈색 봉투로 보이는 물체가 있으나, 평평하게 놓였던 상태와 달리 세워진 모습이다. 왼쪽 아래 병 상자에는 읽을 수 있는 영문 상표가 있다. 추가 인물이나 사이렌 색광은 없다.",
        "hard_violations": [
         "왼쪽 아래 파란 병 상자에 읽을 수 있는 'Cass' 상표가 보여, 읽을 수 있는 글자와 로고를 모두 금지한 지시를 위반한다."
        ],
        "physics": "남자의 기울어진 머리와 앞으로 숙인 상체는 자연스럽고 현우의 전경 어깨도 정상적인 착석 자세로 보인다. 봉투 하단과 손은 프레임 밖이라 봉투를 받치는 정확한 접점은 확인되지 않는다. 화면 밖 탁자에 놓이거나 손에 들릴 수 있는 위치이므로 무지지 부유라고 단정할 근거는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.0,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.75,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 인물 좌우 위치 및 배경 배치가 뒤바뀐 공간 연속성 위반",
     "[gemini-pro] 지지대 없이 비정상적으로 세워져 있는 전경의 서류 봉투",
     "[gpt-high] 왼쪽 아래 파란 병 상자에 읽을 수 있는 'Cass' 상표가 보여, 읽을 수 있는 글자와 로고를 모두 금지한 지시를 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 750
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "이전 샷의 인물 배치와 배경 구조를 완벽히 유지하며 조롱하는 미소를 정확히 포착했습니다."
   },
   {
    "label": "A",
    "score": 750,
    "verdict_ko": "인물의 좌우 위치가 바뀌는 심각한 공간 연속성 위반이 있으며, 전경의 봉투가 부자연스럽게 서 있습니다.  ★위반: [gemini-pro] 인물 좌우 위치 및 배경 배치가 뒤바뀐 공간 연속성 위반 / [gemini-pro] 지지대 없이 비정상적으로 세워져 있는 전경의 서류 봉투 / [gpt-high] 왼쪽 아래 파란 병 상자에 읽을 수 있는 'Cass' 상표가 보여, 읽을 수 있는 글자와 로고를 모두 금지한 지시를 위반한다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features, lighting mood and each person's clothing are LOCKED to this photo; never copy its camera framing. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S10sh3_sel.png",
    "asset_id": "9a40ce7b-0049-4eca-9da3-2dba92dc6ccb",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-c464-7c94-9e48-50b45bc16236",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S10sh3"
  }
 },
 "S10sh6::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:16:31.763603+00:00",
  "fingerprint": "251bc24b0409be0e6d3cd7207b66d47f02ed2e94f0182af5d4ecdd4f5474b403",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S10sh6_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S10sh6_sel.png",
  "source_sha256": "2f4da26f43d3cf218aa0d8ab0ccc77ca5a4a728c15c7190c814fe55c1a73ab06",
  "file": "S10sh6_cine.png",
  "staged_sha256": "e143826d268fe01e482bf00ca54ada468e6c63ad3ca2f4273f4f2a29f4f4bcde",
  "latency_ms": 10366
 },
 "S10sh7::signage": {
  "fp": "b5111a583a2f90ae",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S10sh7": {
  "input_fingerprint": "e157ad74e82fc352",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 희미한 식당 조명 아래, 두 주먹을 꽉 쥔 채 무력하게 고개를 푹 숙이고 얼어붙은 현우의 전신.\n\nLOCATION (lock): In the secluded seating area of a shabby alley restaurant, under its faint interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Restaurant table (Beside 현우 at the secluded seating position) — Its side and part of its top are visible to the left of his body; used as Provides spatial context without concealing his fists or feet; 현우's seat (Occupied) — Seen obliquely from its open side; used as Supports the full-body silhouette and makes his collapsed seated posture legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Faint restaurant illumination preserves subdued facial and hand detail within restrained tonal contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The meeting remains in the secluded part of the shabby restaurant at night. 현우: Hyunwoo remains at the table, visibly frustrated and powerless, with his leg wound and earlier head injury unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 희미한 식당 조명 아래, 두 주먹을 꽉 쥔 채 무력하게 고개를 푹 숙이고 얼어붙은 현우의 전신.\n\nLOCATION (lock): In the secluded seating area of a shabby alley restaurant, under its faint interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Restaurant table (Beside 현우 at the secluded seating position) — Its side and part of its top are visible to the left of his body; used as Provides spatial context without concealing his fists or feet; 현우's seat (Occupied) — Seen obliquely from its open side; used as Supports the full-body silhouette and makes his collapsed seated posture legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Faint restaurant illumination preserves subdued facial and hand detail within restrained tonal contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The meeting remains in the secluded part of the shabby restaurant at night. 현우: Hyunwoo remains at the table, visibly frustrated and powerless, with his leg wound and earlier head injury unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 희미한 식당 조명 아래, 두 주먹을 꽉 쥔 채 무력하게 고개를 푹 숙이고 얼어붙은 현우의 전신.\n\nLOCATION (lock): In the secluded seating area of a shabby alley restaurant, under its faint interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Restaurant table (Beside 현우 at the secluded seating position) — Its side and part of its top are visible to the left of his body; used as Provides spatial context without concealing his fists or feet; 현우's seat (Occupied) — Seen obliquely from its open side; used as Supports the full-body silhouette and makes his collapsed seated posture legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Faint restaurant illumination preserves subdued facial and hand detail within restrained tonal contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The meeting remains in the secluded part of the shabby restaurant at night. 현우: Hyunwoo remains at the table, visibly frustrated and powerless, with his leg wound and earlier head injury unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "고개를 무력하게 바닥 쪽으로 푹 숙이고 있으며, 두 주먹을 허벅지 위에 꽉 쥐고 있음.",
    "built_space": "식당 내부. 화면 좌측(인물의 오른쪽)에 테이블의 측면과 상단이 보이며, 배경 우측 벽에는 레퍼런스와 동일한 달력, 검은 앞치마, 파란색 플라스틱 상자가 위치함. 인물은 등받이가 없는 의자 정면에 앉아 있음.",
    "entities": "현우의 외모(얼굴, 머리스타일)와 일치함. 레퍼런스의 밝은 셔츠를 입고 있으며, 바지와 얼굴에 명시된 다리 부상 및 머리 부상(핏자국)이 묘사됨.",
    "hard_violations": [],
    "physics": "의자에 엉덩이를 대고 체중을 싣고 있으며, 두 발은 바닥을 단단히 딛고 있고 두 손은 허벅지 위에 안정적으로 올려져 있음."
   },
   {
    "label": "B",
    "direction": "고개를 아래로 푹 숙이고 허벅지 근처 허공이나 무릎 쪽을 향하고 있으며, 두 주먹을 쥐고 있음.",
    "built_space": "식당 내부. 화면 좌측에 테이블이 위치하고, 인물 뒤쪽 벽면에 달력과 앞치마가 보임. 인물은 측면 방향으로 스툴에 앉아 있음.",
    "entities": "현우의 외모와 의상은 일치하나, 요구된 다리 부상의 흔적이 보이지 않음.",
    "hard_violations": [
     "[gpt-high] 파란 병 상자에 'CASS' 상표가 읽히므로, 읽을 수 있는 글자와 로고를 금지한 조건을 위반한다."
    ],
    "physics": "스툴에 앉아 상체를 숙이고 있으며 체중은 의자에 지탱됨. 발은 프레임 밖으로 잘려 지지면을 확인할 수 없음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "전신 샷 및 발 노출 요구사항을 정확히 충족하며, 지시된 다리와 머리의 부상을 바지의 핏자국과 얼굴의 상처로 잘 유지하여 프롬프트 충실도가 가장 높습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "화면 하단이 무릎 아래에서 잘려 '전신' 및 '발을 가리지 않을 것'이라는 주요 프레이밍 지시를 어겼으며, 명시된 다리 부상도 누락되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "고개를 무력하게 바닥 쪽으로 푹 숙이고 있으며, 두 주먹을 허벅지 위에 꽉 쥐고 있음.",
        "built_space": "식당 내부. 화면 좌측(인물의 오른쪽)에 테이블의 측면과 상단이 보이며, 배경 우측 벽에는 레퍼런스와 동일한 달력, 검은 앞치마, 파란색 플라스틱 상자가 위치함. 인물은 등받이가 없는 의자 정면에 앉아 있음.",
        "entities": "현우의 외모(얼굴, 머리스타일)와 일치함. 레퍼런스의 밝은 셔츠를 입고 있으며, 바지와 얼굴에 명시된 다리 부상 및 머리 부상(핏자국)이 묘사됨.",
        "hard_violations": [],
        "physics": "의자에 엉덩이를 대고 체중을 싣고 있으며, 두 발은 바닥을 단단히 딛고 있고 두 손은 허벅지 위에 안정적으로 올려져 있음."
       },
       {
        "label": "B",
        "direction": "고개를 아래로 푹 숙이고 허벅지 근처 허공이나 무릎 쪽을 향하고 있으며, 두 주먹을 쥐고 있음.",
        "built_space": "식당 내부. 화면 좌측에 테이블이 위치하고, 인물 뒤쪽 벽면에 달력과 앞치마가 보임. 인물은 측면 방향으로 스툴에 앉아 있음.",
        "entities": "현우의 외모와 의상은 일치하나, 요구된 다리 부상의 흔적이 보이지 않음.",
        "hard_violations": [],
        "physics": "스툴에 앉아 상체를 숙이고 있으며 체중은 의자에 지탱됨. 발은 프레임 밖으로 잘려 지지면을 확인할 수 없음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "전신 샷 및 발 노출 요구사항을 정확히 충족하며, 지시된 다리와 머리의 부상을 바지의 핏자국과 얼굴의 상처로 잘 유지하여 프롬프트 충실도가 가장 높습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "화면 하단이 무릎 아래에서 잘려 '전신' 및 '발을 가리지 않을 것'이라는 주요 프레이밍 지시를 어겼으며, 명시된 다리 부상도 누락되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "고개를 무력하게 바닥 쪽으로 푹 숙이고 있으며, 두 주먹을 허벅지 위에 꽉 쥐고 있음.",
        "built_space": "식당 내부. 화면 좌측(인물의 오른쪽)에 테이블의 측면과 상단이 보이며, 배경 우측 벽에는 레퍼런스와 동일한 달력, 검은 앞치마, 파란색 플라스틱 상자가 위치함. 인물은 등받이가 없는 의자 정면에 앉아 있음.",
        "entities": "현우의 외모(얼굴, 머리스타일)와 일치함. 레퍼런스의 밝은 셔츠를 입고 있으며, 바지와 얼굴에 명시된 다리 부상 및 머리 부상(핏자국)이 묘사됨.",
        "hard_violations": [],
        "physics": "의자에 엉덩이를 대고 체중을 싣고 있으며, 두 발은 바닥을 단단히 딛고 있고 두 손은 허벅지 위에 안정적으로 올려져 있음."
       },
       {
        "label": "B",
        "direction": "고개를 아래로 푹 숙이고 허벅지 근처 허공이나 무릎 쪽을 향하고 있으며, 두 주먹을 쥐고 있음.",
        "built_space": "식당 내부. 화면 좌측에 테이블이 위치하고, 인물 뒤쪽 벽면에 달력과 앞치마가 보임. 인물은 측면 방향으로 스툴에 앉아 있음.",
        "entities": "현우의 외모와 의상은 일치하나, 요구된 다리 부상의 흔적이 보이지 않음.",
        "hard_violations": [],
        "physics": "스툴에 앉아 상체를 숙이고 있으며 체중은 의자에 지탱됨. 발은 프레임 밖으로 잘려 지지면을 확인할 수 없음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "고개를 숙이고 두 주먹을 쥔 동작은 맞지만, 하퇴와 발을 잘라 전신 와이드숏을 어겼고 상자의 읽히는 상표도 금지 조건을 위반한다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "두 주먹과 양발이 모두 드러나는 전신 구도와 무력하게 숙인 자세를 충족하지만, 좌석의 사선 시점과 이전 장면의 의상·공간 연속성은 덜 정확하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 머리를 깊이 숙여 얼굴과 시선이 무릎 앞의 바닥 쪽을 향한다. 두 주먹은 무릎 위쪽에서 서로 가까이 쥐고 있으며 카메라나 다른 인물을 향하지 않는다. 무기나 이동하는 물체는 없다.",
        "built_space": "인물 왼쪽에 나무 탁자 한 개의 상판과 측면이 보이고, 오른쪽 뒤에도 탁자 한 개가 있다. 점유한 좌석 한 개, 왼쪽 탁자 아래 좌석, 왼쪽 아래 잘린 좌석, 오른쪽 뒤의 빈 좌석들이 보인다. 현우의 좌석은 옆에서 비스듬히 보이며 탁자가 주먹을 가리지 않는다. 뒤에는 냉장고 한 대, 달력 한 장, 파란 병 상자 더미와 걸린 검은 천이 있어 참고 장소의 요소를 상당 부분 따른다. 다만 화면이 하퇴와 발을 잘라 요구된 전신과 바닥 접촉을 보여주지 못한다.",
        "entities": "사람은 젊은 동아시아계 남성 한 명뿐이며 검은 머리와 마른 체격은 현우의 설정에 대체로 맞는다. 숙인 얼굴만으로 정확한 동일인 여부나 국적은 판별하기 어렵다. 밝은 베이지색 셔츠는 이전 장면 전경 현우의 어두운 회녹색 옷 및 인물 참고의 남색 티셔츠와 다르다. 짙은 바지의 보이는 부분에는 다리 부상이 뚜렷하지 않고 머리 부상도 확인되지 않는다. 뒤쪽 파란 상자에 읽을 수 있는 영문 상표가 있다.",
        "hard_violations": [
         "파란 병 상자에 'CASS' 상표가 읽히므로, 읽을 수 있는 글자와 로고를 금지한 조건을 위반한다."
        ],
        "physics": "엉덩이는 갈색 좌석에 놓여 있고 좌석 다리는 바닥으로 이어진다. 앞으로 굽힌 상체와 허벅지 가까이 내린 팔, 손목에 연결된 두 주먹은 가능한 자세다. 발은 프레임 밖이므로 접촉 상태를 확인할 수 없지만, 몸이 지지 없이 떠 있지는 않다."
       },
       {
        "label": "B",
        "direction": "현우의 머리는 두 주먹 사이와 무릎 쪽으로 깊이 떨어져 있고 시선도 아래를 향한다. 양 주먹을 무릎 사이 위쪽에 꽉 모은 동작이 선명하다. 카메라를 바라보거나 다른 대상을 겨누는 행동은 없다.",
        "built_space": "현우의 머리부터 양 신발까지 모두 들어오는 전신 구도다. 왼쪽 탁자 한 개의 측면과 상판 일부가 몸 옆에 보이고 오른쪽 가장자리에도 탁자 한 개가 있다. 점유한 의자 한 개와 왼쪽의 빈 원형 좌석 한 개가 식별되며 뒤쪽에는 다른 좌석 일부가 겹친다. 탁자가 주먹이나 발을 가리지 않는다. 다만 의자는 요구한 사선보다 정면에 가깝게 보인다. 오른쪽 뒤의 냉장고 한 대, 달력 한 장, 검은 앞치마와 파란 상자 더미, 낡은 투톤 벽은 참고 장소와 연결되지만 넓게 드러난 주방과 전경 칸막이 때문에 정확한 공간 연속성은 확신하기 어렵다.",
        "entities": "젊은 동아시아계 남성 한 명만 있으며 헝클어진 검은 머리와 체격은 현우 설정에 대체로 부합한다. 얼굴은 머리카락과 숙인 각도에 가려 정확한 얼굴 일치는 제한적으로만 확인된다. 베이지색 셔츠는 이전 장면 현우의 옷 및 인물 참고의 남색 티셔츠와 다르다. 갈색 바지 양쪽 무릎과 정강이에 혈흔이 보여 다리 부상 상태를 표현한다. 머리 부상은 가려져 확인할 수 없다. 벽의 인쇄물은 보이지만 명확히 읽히는 문구는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "엉덩이가 의자 좌판에 놓이고 의자 다리와 양 신발 밑창이 바닥에 닿아 체중을 지지한다. 벌린 무릎 위로 상체를 숙이고 팔꿈치를 허벅지 가까이에 둔 채 두 주먹을 쥔 자세는 물리적으로 가능하다. 지지 없이 떠 있는 사람이나 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "고개를 숙이고 두 주먹을 쥔 동작은 맞지만, 하퇴와 발을 잘라 전신 와이드숏을 어겼고 상자의 읽히는 상표도 금지 조건을 위반한다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "두 주먹과 양발이 모두 드러나는 전신 구도와 무력하게 숙인 자세를 충족하지만, 좌석의 사선 시점과 이전 장면의 의상·공간 연속성은 덜 정확하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 머리를 깊이 숙여 얼굴과 시선이 무릎 앞의 바닥 쪽을 향한다. 두 주먹은 무릎 위쪽에서 서로 가까이 쥐고 있으며 카메라나 다른 인물을 향하지 않는다. 무기나 이동하는 물체는 없다.",
        "built_space": "인물 왼쪽에 나무 탁자 한 개의 상판과 측면이 보이고, 오른쪽 뒤에도 탁자 한 개가 있다. 점유한 좌석 한 개, 왼쪽 탁자 아래 좌석, 왼쪽 아래 잘린 좌석, 오른쪽 뒤의 빈 좌석들이 보인다. 현우의 좌석은 옆에서 비스듬히 보이며 탁자가 주먹을 가리지 않는다. 뒤에는 냉장고 한 대, 달력 한 장, 파란 병 상자 더미와 걸린 검은 천이 있어 참고 장소의 요소를 상당 부분 따른다. 다만 화면이 하퇴와 발을 잘라 요구된 전신과 바닥 접촉을 보여주지 못한다.",
        "entities": "사람은 젊은 동아시아계 남성 한 명뿐이며 검은 머리와 마른 체격은 현우의 설정에 대체로 맞는다. 숙인 얼굴만으로 정확한 동일인 여부나 국적은 판별하기 어렵다. 밝은 베이지색 셔츠는 이전 장면 전경 현우의 어두운 회녹색 옷 및 인물 참고의 남색 티셔츠와 다르다. 짙은 바지의 보이는 부분에는 다리 부상이 뚜렷하지 않고 머리 부상도 확인되지 않는다. 뒤쪽 파란 상자에 읽을 수 있는 영문 상표가 있다.",
        "hard_violations": [
         "파란 병 상자에 'CASS' 상표가 읽히므로, 읽을 수 있는 글자와 로고를 금지한 조건을 위반한다."
        ],
        "physics": "엉덩이는 갈색 좌석에 놓여 있고 좌석 다리는 바닥으로 이어진다. 앞으로 굽힌 상체와 허벅지 가까이 내린 팔, 손목에 연결된 두 주먹은 가능한 자세다. 발은 프레임 밖이므로 접촉 상태를 확인할 수 없지만, 몸이 지지 없이 떠 있지는 않다."
       },
       {
        "label": "A",
        "direction": "현우의 머리는 두 주먹 사이와 무릎 쪽으로 깊이 떨어져 있고 시선도 아래를 향한다. 양 주먹을 무릎 사이 위쪽에 꽉 모은 동작이 선명하다. 카메라를 바라보거나 다른 대상을 겨누는 행동은 없다.",
        "built_space": "현우의 머리부터 양 신발까지 모두 들어오는 전신 구도다. 왼쪽 탁자 한 개의 측면과 상판 일부가 몸 옆에 보이고 오른쪽 가장자리에도 탁자 한 개가 있다. 점유한 의자 한 개와 왼쪽의 빈 원형 좌석 한 개가 식별되며 뒤쪽에는 다른 좌석 일부가 겹친다. 탁자가 주먹이나 발을 가리지 않는다. 다만 의자는 요구한 사선보다 정면에 가깝게 보인다. 오른쪽 뒤의 냉장고 한 대, 달력 한 장, 검은 앞치마와 파란 상자 더미, 낡은 투톤 벽은 참고 장소와 연결되지만 넓게 드러난 주방과 전경 칸막이 때문에 정확한 공간 연속성은 확신하기 어렵다.",
        "entities": "젊은 동아시아계 남성 한 명만 있으며 헝클어진 검은 머리와 체격은 현우 설정에 대체로 부합한다. 얼굴은 머리카락과 숙인 각도에 가려 정확한 얼굴 일치는 제한적으로만 확인된다. 베이지색 셔츠는 이전 장면 현우의 옷 및 인물 참고의 남색 티셔츠와 다르다. 갈색 바지 양쪽 무릎과 정강이에 혈흔이 보여 다리 부상 상태를 표현한다. 머리 부상은 가려져 확인할 수 없다. 벽의 인쇄물은 보이지만 명확히 읽히는 문구는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "엉덩이가 의자 좌판에 놓이고 의자 다리와 양 신발 밑창이 바닥에 닿아 체중을 지지한다. 벌린 무릎 위로 상체를 숙이고 팔꿈치를 허벅지 가까이에 둔 채 두 주먹을 쥔 자세는 물리적으로 가능하다. 지지 없이 떠 있는 사람이나 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.946
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.696
   },
   "violations": {
    "B": [
     "[gpt-high] 파란 병 상자에 'CASS' 상표가 읽히므로, 읽을 수 있는 글자와 로고를 금지한 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 696
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "전신 샷 및 발 노출 요구사항을 정확히 충족하며, 지시된 다리와 머리의 부상을 바지의 핏자국과 얼굴의 상처로 잘 유지하여 프롬프트 충실도가 가장 높습니다."
   },
   {
    "label": "B",
    "score": 696,
    "verdict_ko": "화면 하단이 무릎 아래에서 잘려 '전신' 및 '발을 가리지 않을 것'이라는 주요 프레이밍 지시를 어겼으며, 명시된 다리 부상도 누락되었습니다.  ★위반: [gpt-high] 파란 병 상자에 'CASS' 상표가 읽히므로, 읽을 수 있는 글자와 로고를 금지한 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features, lighting mood and each person's clothing are LOCKED to this photo; never copy its camera framing. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S10sh6_sel.png",
    "asset_id": "ecb50f26-c610-4a5f-98f2-015a55d9b1f6",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-c61b-750f-899a-a33614fc7f3f",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S10sh6"
  }
 },
 "S10sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:17:46.393270+00:00",
  "fingerprint": "aeb5279cdd783f62b7a171823644b86eeed8351224581708bb535262e00832d4",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S10sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S10sh7_sel.png",
  "source_sha256": "7655598cd6063d1b1ae888b84549397022e426d6873e693ea8cd13787acae019",
  "file": "S10sh7_cine.png",
  "staged_sha256": "cebe7b8a0d4cbab501434e8e92138a29d1cca66cd1f1cdc491212344f2fc33d2",
  "latency_ms": 11395
 },
 "S11sh1::signage": {
  "fp": "3880848d190c97db",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S11sh1::ab_conti": {
  "input_fingerprint": "72106a82e86350ac",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 어둠이 내린 좁은 골목길을 향해 닫히는 도중인 육중한 철창문 틈새로 몸을 비틀어 구겨 넣은 자세의 현우의 다급한 전신.\n\nLOCATION (lock): At the narrowing opening of the refugee settlement's heavy barred entrance gate, opening onto a dark alley. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Closing entrance gate in the middle-left of the frame, midground; Alley beyond the entrance in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Heavy barred entrance gate (Closing, leaving a narrowing passage) — Seen obliquely from inside the entrance, with its closing edge beside 현우; used as Creates the constricting left boundary without obscuring his complete body; Entrance threshold and alley (현우 is crossing into the settlement) — The threshold crosses the lower field, and the alley continues toward screen right; used as Establishes an unambiguous inward route for the subsequent pan.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The nighttime entrance remains dim after the settlement's lights go out, with only enough ambient visibility to distinguish his body from the gate.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camp entrance is closing for curfew; shop doors are shutting and the container homes and streetlights are being extinguished, leaving the streets dark. 현우: Hyunwoo is hurrying through the entrance before it closes, still carrying the untreated dog-bite injury and earlier head injury.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "shot_run_spend_attempt_count": 3,
  "readings": [
   {
    "label": "A",
    "direction": "시선은 골목길 안쪽(화면 우측)을 향하고 있으며, 몸통은 좁아지는 문틈을 통과하는 방향으로 움직이고 있음.",
    "built_space": "마름모 장식의 철제 대문, 바닥 레일, 좌측 컨테이너 구조물과 우측 원형 조명 기둥이 레퍼런스와 일치하며 지정된 야간 환경에 맞게 배치됨.",
    "entities": "현우(젊은 아시아인 남성, 헝클어진 검은 머리)의 외형이 일치하며, 얼굴의 상처와 바지 다리 부분의 핏자국(개에게 물린 상처)이 확인됨.",
    "hard_violations": [],
    "physics": "양발로 지면을 딛고 오른손으로 닫히는 철문을 짚으며 틈새를 비집고 들어가는 체중 이동과 지지가 물리적으로 자연스러움."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "지정된 야간 조명 속에서 닫히는 문틈을 빠져나오는 다급한 동작과 프레이밍, 그리고 머리와 다리의 상처 세부 사항까지 샷 텍스트를 충실히 구현함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 골목길 안쪽(화면 우측)을 향하고 있으며, 몸통은 좁아지는 문틈을 통과하는 방향으로 움직이고 있음.",
        "built_space": "마름모 장식의 철제 대문, 바닥 레일, 좌측 컨테이너 구조물과 우측 원형 조명 기둥이 레퍼런스와 일치하며 지정된 야간 환경에 맞게 배치됨.",
        "entities": "현우(젊은 아시아인 남성, 헝클어진 검은 머리)의 외형이 일치하며, 얼굴의 상처와 바지 다리 부분의 핏자국(개에게 물린 상처)이 확인됨.",
        "hard_violations": [],
        "physics": "양발로 지면을 딛고 오른손으로 닫히는 철문을 짚으며 틈새를 비집고 들어가는 체중 이동과 지지가 물리적으로 자연스러움."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "지정된 야간 조명 속에서 닫히는 문틈을 빠져나오는 다급한 동작과 프레이밍, 그리고 머리와 다리의 상처 세부 사항까지 샷 텍스트를 충실히 구현함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 골목길 안쪽(화면 우측)을 향하고 있으며, 몸통은 좁아지는 문틈을 통과하는 방향으로 움직이고 있음.",
        "built_space": "마름모 장식의 철제 대문, 바닥 레일, 좌측 컨테이너 구조물과 우측 원형 조명 기둥이 레퍼런스와 일치하며 지정된 야간 환경에 맞게 배치됨.",
        "entities": "현우(젊은 아시아인 남성, 헝클어진 검은 머리)의 외형이 일치하며, 얼굴의 상처와 바지 다리 부분의 핏자국(개에게 물린 상처)이 확인됨.",
        "hard_violations": [],
        "physics": "양발로 지면을 딛고 오른손으로 닫히는 철문을 짚으며 틈새를 비집고 들어가는 체중 이동과 지지가 물리적으로 자연스러움."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "전신 와이드 구도와 오른쪽 골목, 문을 붙잡고 비집고 들어오는 동작은 부합하지만 철문이 장소 사진보다 지나치게 높으며, B가 제공되지 않아 비교 승자는 잠정적입니다."
       },
       {
        "label": "B",
        "score": 0,
        "verdict_ko": "이미지가 제공되지 않아 평가할 수 없으며, 0점과 최하위 배치는 실제 충실도 판정이 아닌 미제공 표시입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 얼굴과 시선은 화면 오른쪽 골목 방향을 향한다. 상체와 앞다리도 오른쪽으로 나가고 뒤로 뻗은 손은 철문 가장자리를 잡는다. 문을 통과하는 동작은 보이지만, 배경에 정착촌 컨테이너들이 보여 정착촌 안으로 들어오는 방향인지는 명확하지 않다.",
        "built_space": "왼쪽의 큰 철문 면과 중앙의 비스듬한 철문 면 사이에 현우가 있다. 하단에는 문턱과 문 바퀴가 있고, 오른쪽 문기둥에는 구형 등 하나와 확성기 하나가 보인다. 왼쪽 가장자리에는 볼록거울 하나와 일부 잘린 구형 등이 있다. 오른쪽 뒤로 컨테이너 사이 통로가 이어진다. 마름모 장식과 재료는 참고 장소와 유사하지만, 철문 높이와 장식 비례는 사진보다 크게 확대되었다.",
        "entities": "인물은 한 명으로, 앳된 동아시아계 남성의 외모와 헝클어진 검은 머리, 짙은 반팔 티셔츠가 현우 참고 이미지에 대체로 맞는다. 국적은 외관만으로 확인할 수 없다. 이마의 상처와 바지 정강이 부근의 혈흔은 보이지만 혈흔만으로 개에 물린 상처인지 확정할 수 없다. 철창문과 어두운 골목이 있으며 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "뒤로 뻗은 손이 문 가장자리를 잡고 있어 몸을 비틀 때의 버팀점이 있다. 앞발은 지면 가까이 내딛는 중이고 뒷발의 접지는 문 가장자리에 가려 확인하기 어렵다. 허리와 무릎의 굽힘은 틈을 통과하는 보행 동작으로 가능하며, 명백한 무지지 부유로 단정할 근거는 없다. 철문은 하단 바퀴와 문기둥 쪽 구조로 지지된다."
       },
       {
        "label": "B",
        "direction": "이미지가 제공되지 않아 시선과 이동 방향을 관찰할 수 없다.",
        "built_space": "이미지가 제공되지 않아 문과 통로의 구조 및 인물 배치를 관찰할 수 없다.",
        "entities": "이미지가 제공되지 않아 인물과 사물을 확인할 수 없다.",
        "hard_violations": [],
        "physics": "이미지가 제공되지 않아 접지, 지지점 및 동작의 물리적 타당성을 확인할 수 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": null,
     "ok": false
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gpt-high"
   ],
   "route": "single_forward"
  },
  "totals": {
   "A": 8
  },
  "selected": "A",
  "ranking": [
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 8,
    "verdict_ko": "지정된 야간 조명 속에서 닫히는 문틈을 빠져나오는 다급한 동작과 프레이밍, 그리고 머리와 다리의 상처 세부 사항까지 샷 텍스트를 충실히 구현함."
   }
  ],
  "refs": [
   {
    "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_refugee_gate_sel.png",
    "asset_id": "331ee13b-1a46-4cba-9a37-07f52e1d1493",
    "role": "location_seed_bg"
   },
   {
    "label": "LAYOUT SKETCH — a bare thin-line layout guide, a REFERENCE ONLY: take from it ONLY the camera framing, figure placement, pose and size/depth order. It carries ZERO visual style — every texture, material, light and all realism come from the text and the photographic reference. Never let any line-drawing quality leak into the output.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S11sh1.png",
    "asset_id": "37dd0d43-2fdf-429b-82e4-045c520df701",
    "role": "conti_light"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": true,
  "shot_run_uid": "06aae5a9-44fb-76cb-b780-94b77fadbb95"
 },
 "S11sh4::signage": {
  "fp": "15ac91a737b02265",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::0e33b57ac265c530": {
  "subjects": [],
  "subject_text": "인천 난민촌 입구와 철창문 앞\n철창 출입문과 ‘난민 거주지역’ 표지판이 있는 진입로. 안쪽으로 낡은 철제 컨테이너들이 조밀하게 늘어서고 방송용 스피커가 설치돼 있다.",
  "identity": "canonical",
  "scope_id": "L163",
  "scope_role": "location_exterior",
  "scope_sha": "823ab6b5b4be054b"
 },
 "S11sh4::bgfirst_bg": {
  "input_fingerprint": "03c257130aebaf1b",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 열린 컨테이너 문 안으로 한 발을 내딛은 현우의 뒷모습.\n\nLOCATION (lock): At the open front threshold of a family's container home in the dark refugee settlement, viewed from the alley outside.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Container entrance door (Open as 현우 enters) — The open leaf is seen obliquely beside the entrance, leaving his back and leading foot visible; used as Frames the crossing without concealing the body's movement; Container number 7-31 (Identifies 현우's home) — The exterior-facing marking bearing 7-31 is visible beside the entrance; used as Provides a restrained location identifier outside the central action; Container doorway threshold (Being crossed) — Viewed from behind and above 현우's trailing foot; used as Separates the exterior passage from the interior destination.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the settlement's lights-out nighttime darkness without introducing an illuminated interior beyond the doorway.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 열린 컨테이너 문 안으로 한 발을 내딛은 현우의 뒷모습.\n\nLOCATION (lock): At the open front threshold of a family's container home in the dark refugee settlement, viewed from the alley outside.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Container entrance door (Open as 현우 enters) — The open leaf is seen obliquely beside the entrance, leaving his back and leading foot visible; used as Frames the crossing without concealing the body's movement; Container number 7-31 (Identifies 현우's home) — The exterior-facing marking bearing 7-31 is visible beside the entrance; used as Provides a restrained location identifier outside the central action; Container doorway threshold (Being crossed) — Viewed from behind and above 현우's trailing foot; used as Separates the exterior passage from the interior destination.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the settlement's lights-out nighttime darkness without introducing an illuminated interior beyond the doorway.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S11sh4__bgfirst_bg.png",
  "asset_id": "ec7f53f2-f10f-4cb4-9c93-d5c6f78750b1",
  "input_asset_ids": [
   "478420ea-6c5c-40ae-a832-b5c0b14b89aa",
   "331ee13b-1a46-4cba-9a37-07f52e1d1493"
  ]
 },
 "S11sh4": {
  "input_fingerprint": "22b4a9386dba13a9",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 열린 컨테이너 문 안으로 한 발을 내딛은 현우의 뒷모습.\n\nLOCATION (lock): At the open front threshold of a family's container home in the dark refugee settlement, viewed from the alley outside. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Container entrance door (Open as 현우 enters) — The open leaf is seen obliquely beside the entrance, leaving his back and leading foot visible; used as Frames the crossing without concealing the body's movement; Container number 7-31 (Identifies 현우's home) — The exterior-facing marking bearing 7-31 is visible beside the entrance; used as Provides a restrained location identifier outside the central action; Container doorway threshold (Being crossed) — Viewed from behind and above 현우's trailing foot; used as Separates the exterior passage from the interior destination.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the settlement's lights-out nighttime darkness without introducing an illuminated interior beyond the doorway.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The home container is marked 7-31, and its door is open for entry. The surrounding camp streets remain dark after curfew lights-out. 현우: Hyunwoo is entering his home with his upper garment still on; his face and leg remain injured, and the leg wound has not yet been treated.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 열린 컨테이너 문 안으로 한 발을 내딛은 현우의 뒷모습.\n\nLOCATION (lock): At the open front threshold of a family's container home in the dark refugee settlement, viewed from the alley outside. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Container entrance door (Open as 현우 enters) — The open leaf is seen obliquely beside the entrance, leaving his back and leading foot visible; used as Frames the crossing without concealing the body's movement; Container number 7-31 (Identifies 현우's home) — The exterior-facing marking bearing 7-31 is visible beside the entrance; used as Provides a restrained location identifier outside the central action; Container doorway threshold (Being crossed) — Viewed from behind and above 현우's trailing foot; used as Separates the exterior passage from the interior destination.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the settlement's lights-out nighttime darkness without introducing an illuminated interior beyond the doorway.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The home container is marked 7-31, and its door is open for entry. The surrounding camp streets remain dark after curfew lights-out. 현우: Hyunwoo is entering his home with his upper garment still on; his face and leg remain injured, and the leg wound has not yet been treated.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 열린 컨테이너 문 안으로 한 발을 내딛은 현우의 뒷모습.\n\nLOCATION (lock): At the open front threshold of a family's container home in the dark refugee settlement, viewed from the alley outside. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Container entrance door (Open as 현우 enters) — The open leaf is seen obliquely beside the entrance, leaving his back and leading foot visible; used as Frames the crossing without concealing the body's movement; Container number 7-31 (Identifies 현우's home) — The exterior-facing marking bearing 7-31 is visible beside the entrance; used as Provides a restrained location identifier outside the central action; Container doorway threshold (Being crossed) — Viewed from behind and above 현우's trailing foot; used as Separates the exterior passage from the interior destination.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the settlement's lights-out nighttime darkness without introducing an illuminated interior beyond the doorway.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The home container is marked 7-31, and its door is open for entry. The surrounding camp streets remain dark after curfew lights-out. 현우: Hyunwoo is entering his home with his upper garment still on; his face and leg remain injured, and the leg wound has not yet been treated.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S11sh4__bgfirst_bg.png",
     "asset_id": "ec7f53f2-f10f-4cb4-9c93-d5c6f78750b1",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S11sh4.png",
     "asset_id": "478420ea-6c5c-40ae-a832-b5c0b14b89aa",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_refugee_gate_sel.png",
     "asset_id": "331ee13b-1a46-4cba-9a37-07f52e1d1493",
     "role": "location_seed_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선과 몸의 방향이 컨테이너 내부를 향하고 있음.",
    "built_space": "레퍼런스와 일치하는 포장된 넓은 길 양옆으로 컨테이너들이 늘어서 있음. 우측 컨테이너 문이 열려 있고 그 옆에 '7-31' 번호가 보임. 가로등이 켜져 있음.",
    "entities": "뒷모습의 현우(검은 머리의 젊은 남성)와 외투가 보임. 지정된 7-31 컨테이너와 열린 문이 확인됨.",
    "hard_violations": [
     "[gpt-high] 문 옆의 ‘7-31’이 명확하게 읽혀 최종 지시의 판독 가능한 문자 금지 조건을 위반한다."
    ],
    "physics": "오른발은 바닥을 딛고 왼발은 문턱을 넘어가는 자연스러운 보행 자세로 지지됨."
   },
   {
    "label": "B",
    "direction": "현우가 컨테이너 안쪽을 향해 똑바로 진입하고 있음.",
    "built_space": "레퍼런스와 구조가 다른 매우 좁고 어두운 자갈길 골목. 열린 컨테이너 문과 희미하게 '7-31'이 보임.",
    "entities": "어둠 속에 서 있는 뒷모습의 현우. 컨테이너 번호 일부 확인 가능.",
    "hard_violations": [
     "[gpt-high] 문 옆의 ‘7-31’이 명확하게 읽혀 최종 지시의 판독 가능한 문자 금지 조건을 위반한다."
    ],
    "physics": "두 발이 바닥과 문턱에 닿아 있어 체중을 지탱함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 앵글과 로케이션의 구조(넓은 포장도로, 컨테이너 배치)를 정확히 구현했으나, 소등(lights-out) 지시와 달리 가로등이 밝게 켜져 있어 아쉽습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "어두운 야간 조명 지시는 따랐으나, 레퍼런스의 넓은 포장도로 대신 좁은 자갈길을 렌더링하여 로케이션 구조를 완전히 잘못 묘사했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선과 몸의 방향이 컨테이너 내부를 향하고 있음.",
        "built_space": "레퍼런스와 일치하는 포장된 넓은 길 양옆으로 컨테이너들이 늘어서 있음. 우측 컨테이너 문이 열려 있고 그 옆에 '7-31' 번호가 보임. 가로등이 켜져 있음.",
        "entities": "뒷모습의 현우(검은 머리의 젊은 남성)와 외투가 보임. 지정된 7-31 컨테이너와 열린 문이 확인됨.",
        "hard_violations": [],
        "physics": "오른발은 바닥을 딛고 왼발은 문턱을 넘어가는 자연스러운 보행 자세로 지지됨."
       },
       {
        "label": "B",
        "direction": "현우가 컨테이너 안쪽을 향해 똑바로 진입하고 있음.",
        "built_space": "레퍼런스와 구조가 다른 매우 좁고 어두운 자갈길 골목. 열린 컨테이너 문과 희미하게 '7-31'이 보임.",
        "entities": "어둠 속에 서 있는 뒷모습의 현우. 컨테이너 번호 일부 확인 가능.",
        "hard_violations": [],
        "physics": "두 발이 바닥과 문턱에 닿아 있어 체중을 지탱함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 앵글과 로케이션의 구조(넓은 포장도로, 컨테이너 배치)를 정확히 구현했으나, 소등(lights-out) 지시와 달리 가로등이 밝게 켜져 있어 아쉽습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "어두운 야간 조명 지시는 따랐으나, 레퍼런스의 넓은 포장도로 대신 좁은 자갈길을 렌더링하여 로케이션 구조를 완전히 잘못 묘사했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선과 몸의 방향이 컨테이너 내부를 향하고 있음.",
        "built_space": "레퍼런스와 일치하는 포장된 넓은 길 양옆으로 컨테이너들이 늘어서 있음. 우측 컨테이너 문이 열려 있고 그 옆에 '7-31' 번호가 보임. 가로등이 켜져 있음.",
        "entities": "뒷모습의 현우(검은 머리의 젊은 남성)와 외투가 보임. 지정된 7-31 컨테이너와 열린 문이 확인됨.",
        "hard_violations": [],
        "physics": "오른발은 바닥을 딛고 왼발은 문턱을 넘어가는 자연스러운 보행 자세로 지지됨."
       },
       {
        "label": "B",
        "direction": "현우가 컨테이너 안쪽을 향해 똑바로 진입하고 있음.",
        "built_space": "레퍼런스와 구조가 다른 매우 좁고 어두운 자갈길 골목. 열린 컨테이너 문과 희미하게 '7-31'이 보임.",
        "entities": "어둠 속에 서 있는 뒷모습의 현우. 컨테이너 번호 일부 확인 가능.",
        "hard_violations": [],
        "physics": "두 발이 바닥과 문턱에 닿아 있어 체중을 지탱함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "등을 보인 채 문턱을 넘는 동작과 어두운 실내는 맞지만, 자갈 골목과 철창 중심의 공간이 장소 참조에서 벗어나며 읽히는 번호도 금지 조건을 위반합니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "외부 와이드숏의 진입 동작과 포장도로·컨테이너 배치가 더 충실하지만, 읽히는 번호와 켜진 가로등, 참조에 없는 겉옷 때문에 재촬영이 필요합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 등은 카메라를 향하고 머리와 몸의 진행 방향은 열린 출입구 안쪽이다. 눈은 보이지 않아 실제 시선은 확인할 수 없다. 앞발은 실내 문턱 쪽으로 들어가고 뒷발은 골목에 남아 있어 입장 방향이 맞는다.",
        "built_space": "중앙에 출입구 하나, 오른쪽으로 열린 문짝 하나, 그 아래 높은 문턱 하나가 보인다. 문짝은 몸과 앞발을 가리지 않는다. 왼쪽 전경에는 다른 컨테이너의 닫힌 문이, 양옆에는 큰 철창 창문들이 보인다. 카메라는 골목에서 뒷발보다 높은 위치에 있다. 다만 참조의 포장된 통로 대신 좁은 자갈 통로를 만들었고, 철창이 강조된 벽면과 주택 간격도 참조 장소와 상당히 다르다.",
        "entities": "사람은 한 명뿐이며 헝클어진 검은 머리와 마른 젊은 남성의 뒷모습은 현우 설정과 대체로 맞는다. 얼굴이 가려져 정확한 나이·한국계 정체성·얼굴 상처는 확인할 수 없다. 짙은 남색 상의는 참조와 색이 가깝지만 긴소매로 보인다. 바지가 다리를 덮어 미처치 상처 여부는 확인되지 않는다. 집 번호는 문 오른쪽에서 선명하게 읽힌다. 실내는 불이 꺼져 있지만 골목에는 켜진 작은 등이 보인다.",
        "hard_violations": [
         "문 옆의 ‘7-31’이 명확하게 읽혀 최종 지시의 판독 가능한 문자 금지 조건을 위반한다."
        ],
        "physics": "앞쪽 신발은 문턱 또는 바로 안쪽 바닥에 놓이고, 뒷발은 자갈 바닥에서 밀어 들어가는 자세다. 다리와 몸통의 연결 및 체중 이동은 입장 동작으로 가능하다. 문은 문틀 쪽 경첩으로 지지되는 구조이며, 지지 없이 떠 있는 사람이나 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "현우는 카메라에 등을 보이고 약간 오른쪽의 실내를 향해 몸과 머리를 돌린다. 얼굴과 눈은 보이지 않는다. 앞발은 출입구 안쪽으로 향하고 뒷발은 외부 보도에 있어 문 안으로 한 발 들어서는 방향이 명확하다.",
        "built_space": "오른쪽 주택에 출입구 하나와 바깥으로 열린 문짝 하나, 낮은 문턱 하나가 있다. 입구 왼쪽에는 작은 창 하나와 벽 부착 상자 하나, 지면 가까운 설비함 하나가 보이고 오른쪽 가장자리에는 창 하나가 일부 보인다. 열린 문짝이 현우의 등과 발을 가리지 않으며 카메라는 골목에서 뒷발을 내려다본다. 포장도로, 낮은 컨테이너 열, 창문과 외부 설비는 참조와 더 가깝다. 다만 참조에서 번호는 주택 전면 모서리에 있는데 여기서는 긴 측벽의 출입구 옆에 놓여 정확한 입면 배치는 달라졌다.",
        "entities": "추가 인물 없이 검은 헝클어진 머리의 젊은 남성 한 명이 있다. 뒷모습만으로 얼굴 동일성이나 정확한 나이·한국계 정체성은 검증할 수 없다. 남색 상의 위에 회색 재킷을 입어 참조 의상과 다르다. 드러난 뒷발목에는 상처처럼 보이는 어두운 붉은 흔적이 있고 붕대는 보이지 않는다. 문 왼쪽 번호는 읽을 수 있다. 밤이지만 거리의 여러 가로등이 켜져 노면을 밝히므로 소등 후의 어두운 거리 조건과 맞지 않는다. 실내 자체의 켜진 조명은 보이지 않는다.",
        "hard_violations": [
         "문 옆의 ‘7-31’이 명확하게 읽혀 최종 지시의 판독 가능한 문자 금지 조건을 위반한다."
        ],
        "physics": "뒷발 신발은 외부 보도에 닿고 굽힌 앞다리의 신발은 문턱 안쪽으로 올라가 있어 지면의 지지를 받는 자연스러운 진입 동작이다. 양팔은 몸 옆에서 자연스럽게 내려와 있다. 문짝은 오른쪽 문틀의 경첩으로 지지되며, 공중에 지지 없이 떠 있는 몸이나 물체는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "등을 보인 채 문턱을 넘는 동작과 어두운 실내는 맞지만, 자갈 골목과 철창 중심의 공간이 장소 참조에서 벗어나며 읽히는 번호도 금지 조건을 위반합니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "외부 와이드숏의 진입 동작과 포장도로·컨테이너 배치가 더 충실하지만, 읽히는 번호와 켜진 가로등, 참조에 없는 겉옷 때문에 재촬영이 필요합니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 등은 카메라를 향하고 머리와 몸의 진행 방향은 열린 출입구 안쪽이다. 눈은 보이지 않아 실제 시선은 확인할 수 없다. 앞발은 실내 문턱 쪽으로 들어가고 뒷발은 골목에 남아 있어 입장 방향이 맞는다.",
        "built_space": "중앙에 출입구 하나, 오른쪽으로 열린 문짝 하나, 그 아래 높은 문턱 하나가 보인다. 문짝은 몸과 앞발을 가리지 않는다. 왼쪽 전경에는 다른 컨테이너의 닫힌 문이, 양옆에는 큰 철창 창문들이 보인다. 카메라는 골목에서 뒷발보다 높은 위치에 있다. 다만 참조의 포장된 통로 대신 좁은 자갈 통로를 만들었고, 철창이 강조된 벽면과 주택 간격도 참조 장소와 상당히 다르다.",
        "entities": "사람은 한 명뿐이며 헝클어진 검은 머리와 마른 젊은 남성의 뒷모습은 현우 설정과 대체로 맞는다. 얼굴이 가려져 정확한 나이·한국계 정체성·얼굴 상처는 확인할 수 없다. 짙은 남색 상의는 참조와 색이 가깝지만 긴소매로 보인다. 바지가 다리를 덮어 미처치 상처 여부는 확인되지 않는다. 집 번호는 문 오른쪽에서 선명하게 읽힌다. 실내는 불이 꺼져 있지만 골목에는 켜진 작은 등이 보인다.",
        "hard_violations": [
         "문 옆의 ‘7-31’이 명확하게 읽혀 최종 지시의 판독 가능한 문자 금지 조건을 위반한다."
        ],
        "physics": "앞쪽 신발은 문턱 또는 바로 안쪽 바닥에 놓이고, 뒷발은 자갈 바닥에서 밀어 들어가는 자세다. 다리와 몸통의 연결 및 체중 이동은 입장 동작으로 가능하다. 문은 문틀 쪽 경첩으로 지지되는 구조이며, 지지 없이 떠 있는 사람이나 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "현우는 카메라에 등을 보이고 약간 오른쪽의 실내를 향해 몸과 머리를 돌린다. 얼굴과 눈은 보이지 않는다. 앞발은 출입구 안쪽으로 향하고 뒷발은 외부 보도에 있어 문 안으로 한 발 들어서는 방향이 명확하다.",
        "built_space": "오른쪽 주택에 출입구 하나와 바깥으로 열린 문짝 하나, 낮은 문턱 하나가 있다. 입구 왼쪽에는 작은 창 하나와 벽 부착 상자 하나, 지면 가까운 설비함 하나가 보이고 오른쪽 가장자리에는 창 하나가 일부 보인다. 열린 문짝이 현우의 등과 발을 가리지 않으며 카메라는 골목에서 뒷발을 내려다본다. 포장도로, 낮은 컨테이너 열, 창문과 외부 설비는 참조와 더 가깝다. 다만 참조에서 번호는 주택 전면 모서리에 있는데 여기서는 긴 측벽의 출입구 옆에 놓여 정확한 입면 배치는 달라졌다.",
        "entities": "추가 인물 없이 검은 헝클어진 머리의 젊은 남성 한 명이 있다. 뒷모습만으로 얼굴 동일성이나 정확한 나이·한국계 정체성은 검증할 수 없다. 남색 상의 위에 회색 재킷을 입어 참조 의상과 다르다. 드러난 뒷발목에는 상처처럼 보이는 어두운 붉은 흔적이 있고 붕대는 보이지 않는다. 문 왼쪽 번호는 읽을 수 있다. 밤이지만 거리의 여러 가로등이 켜져 노면을 밝히므로 소등 후의 어두운 거리 조건과 맞지 않는다. 실내 자체의 켜진 조명은 보이지 않는다.",
        "hard_violations": [
         "문 옆의 ‘7-31’이 명확하게 읽혀 최종 지시의 판독 가능한 문자 금지 조건을 위반한다."
        ],
        "physics": "뒷발 신발은 외부 보도에 닿고 굽힌 앞다리의 신발은 문턱 안쪽으로 올라가 있어 지면의 지지를 받는 자연스러운 진입 동작이다. 양팔은 몸 옆에서 자연스럽게 내려와 있다. 문짝은 오른쪽 문틀의 경첩으로 지지되며, 공중에 지지 없이 떠 있는 몸이나 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.321
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.071
   },
   "violations": {
    "B": [
     "[gpt-high] 문 옆의 ‘7-31’이 명확하게 읽혀 최종 지시의 판독 가능한 문자 금지 조건을 위반한다."
    ],
    "A": [
     "[gpt-high] 문 옆의 ‘7-31’이 명확하게 읽혀 최종 지시의 판독 가능한 문자 금지 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 1071
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "지정된 앵글과 로케이션의 구조(넓은 포장도로, 컨테이너 배치)를 정확히 구현했으나, 소등(lights-out) 지시와 달리 가로등이 밝게 켜져 있어 아쉽습니다.  ★위반: [gpt-high] 문 옆의 ‘7-31’이 명확하게 읽혀 최종 지시의 판독 가능한 문자 금지 조건을 위반한다."
   },
   {
    "label": "B",
    "score": 1071,
    "verdict_ko": "어두운 야간 조명 지시는 따랐으나, 레퍼런스의 넓은 포장도로 대신 좁은 자갈길을 렌더링하여 로케이션 구조를 완전히 잘못 묘사했습니다.  ★위반: [gpt-high] 문 옆의 ‘7-31’이 명확하게 읽혀 최종 지시의 판독 가능한 문자 금지 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_refugee_gate_sel.png",
    "asset_id": "331ee13b-1a46-4cba-9a37-07f52e1d1493",
    "role": "location_seed_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-cb26-7ca0-8dc7-ce6f6218d7ac",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S11sh4__bgfirst_bg.png",
   "bg_asset_id": "ec7f53f2-f10f-4cb4-9c93-d5c6f78750b1",
   "bg_record_key": "S11sh4::bgfirst_bg",
   "chain_winner": true,
   "authority": "seed_bg"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  },
  "lane_policy": "ab_select_ready"
 },
 "S11sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:20:26.471461+00:00",
  "fingerprint": "167891726e9ff2a53bce44859c02afc6fd642adce69bf63f02944e1b365f590d",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S11sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S11sh4_sel.png",
  "source_sha256": "e741ecae56373379630dc7c1dd5653f65a305257306d7e43b377bd86c18cebab",
  "file": "S11sh4_cine.png",
  "staged_sha256": "805a3f69d22932688fd013f75a081071ddc33ba2ca6a44d47f29fd93a705de17",
  "latency_ms": 10261
 },
 "S12sh7::signage": {
  "fp": "b0d7431472acf0b6",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::f424efd76bc0d46e": {
  "subjects": [],
  "subject_text": "현우의 컨테이너 내부\n식탁과 생활용품을 둔 비좁은 컨테이너 주거 공간. 안쪽에 작은 방으로 연결되는 문이 있고, 밤에는 식탁의 촛불이 실내를 밝힌다.",
  "identity": "canonical",
  "scope_id": "L166",
  "scope_role": "location_interior",
  "scope_sha": "ad03e7fae5aaa48f"
 },
 "S12sh7::bgfirst_bg": {
  "input_fingerprint": "bb68b70ba912526a",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 상체를 뒤로 물린 자세로 미연이 내민 약병을 차가운 눈빛으로 노려보는 현우.\n\nLOCATION (lock): In the dining area inside a refugee family's container home, lit by a candle on the table.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Medicine bottle (Held out by 미연) — Seen from the side in her extended hand, below 현우's chest; used as Provides the focus of refusal while remaining small enough for the intervening empty space to read; Dining table (현우's meal remains set out) — An oblique portion of the tabletop crosses the lower frame; used as Anchors the domestic context beneath the separated hand and torso.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Tabletop candlelight gives the hands and faces restrained warmth while the surrounding interior remains subdued.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 상체를 뒤로 물린 자세로 미연이 내민 약병을 차가운 눈빛으로 노려보는 현우.\n\nLOCATION (lock): In the dining area inside a refugee family's container home, lit by a candle on the table.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Medicine bottle (Held out by 미연) — Seen from the side in her extended hand, below 현우's chest; used as Provides the focus of refusal while remaining small enough for the intervening empty space to read; Dining table (현우's meal remains set out) — An oblique portion of the tabletop crosses the lower frame; used as Anchors the domestic context beneath the separated hand and torso.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Tabletop candlelight gives the hands and faces restrained warmth while the surrounding interior remains subdued.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S12sh7__bgfirst_bg.png",
  "asset_id": "520d1ce5-4724-4ae7-a50a-649a26087ab2",
  "input_asset_ids": [
   "70be5339-3395-4c59-bd14-9f3f92278784",
   "a4110c6d-4f09-489a-91bf-d6c8e7837b22"
  ]
 },
 "S12sh7": {
  "input_fingerprint": "97bc164f56e4f9e4",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 상체를 뒤로 물린 자세로 미연이 내민 약병을 차가운 눈빛으로 노려보는 현우.\n\nLOCATION (lock): In the dining area inside a refugee family's container home, lit by a candle on the table. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Medicine bottle (Held out by 미연) — Seen from the side in her extended hand, below 현우's chest; used as Provides the focus of refusal while remaining small enough for the intervening empty space to read; Dining table (현우's meal remains set out) — An oblique portion of the tabletop crosses the lower frame; used as Anchors the domestic context beneath the separated hand and torso.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Tabletop candlelight gives the hands and faces restrained warmth while the surrounding interior remains subdued.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A candle continues to light the dining table, where the meal was served; the mending clothes remain in the room and the medicine container is now open. 현우: Hyunwoo has removed his upper garment and is seated at the table with visible facial injuries and a bleeding dog-bite wound on his leg. 미연: Miyeon is at the dining table with the opened medicine container, preparing to examine the leg injury.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 미연 right now, so 미연's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 미연: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 상체를 뒤로 물린 자세로 미연이 내민 약병을 차가운 눈빛으로 노려보는 현우.\n\nLOCATION (lock): In the dining area inside a refugee family's container home, lit by a candle on the table. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Medicine bottle (Held out by 미연) — Seen from the side in her extended hand, below 현우's chest; used as Provides the focus of refusal while remaining small enough for the intervening empty space to read; Dining table (현우's meal remains set out) — An oblique portion of the tabletop crosses the lower frame; used as Anchors the domestic context beneath the separated hand and torso.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Tabletop candlelight gives the hands and faces restrained warmth while the surrounding interior remains subdued.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A candle continues to light the dining table, where the meal was served; the mending clothes remain in the room and the medicine container is now open. 현우: Hyunwoo has removed his upper garment and is seated at the table with visible facial injuries and a bleeding dog-bite wound on his leg. 미연: Miyeon is at the dining table with the opened medicine container, preparing to examine the leg injury.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 미연 right now, so 미연's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 미연: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 상체를 뒤로 물린 자세로 미연이 내민 약병을 차가운 눈빛으로 노려보는 현우.\n\nLOCATION (lock): In the dining area inside a refugee family's container home, lit by a candle on the table. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Medicine bottle (Held out by 미연) — Seen from the side in her extended hand, below 현우's chest; used as Provides the focus of refusal while remaining small enough for the intervening empty space to read; Dining table (현우's meal remains set out) — An oblique portion of the tabletop crosses the lower frame; used as Anchors the domestic context beneath the separated hand and torso.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Tabletop candlelight gives the hands and faces restrained warmth while the surrounding interior remains subdued.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A candle continues to light the dining table, where the meal was served; the mending clothes remain in the room and the medicine container is now open. 현우: Hyunwoo has removed his upper garment and is seated at the table with visible facial injuries and a bleeding dog-bite wound on his leg. 미연: Miyeon is at the dining table with the opened medicine container, preparing to examine the leg injury.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 미연 right now, so 미연's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 미연: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S12sh7__bgfirst_bg.png",
     "asset_id": "520d1ce5-4724-4ae7-a50a-649a26087ab2",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S12sh7.png",
     "asset_id": "70be5339-3395-4c59-bd14-9f3f92278784",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1113064>",
     "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L166B01.png",
     "asset_id": "a4110c6d-4f09-489a-91bf-d6c8e7837b22",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1113064>",
     "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "미연은 현우에게 약병을 내밀지만, 현우의 시선은 약병이 아닌 아래쪽 무릎이나 식탁을 향함.",
    "built_space": "컨테이너 내부 구조와 사선(oblique)으로 하단을 가로지르는 식탁 배치가 지시와 정확히 일치함.",
    "entities": "현우(상의 탈의, 얼굴과 다리 상처), 미연, 약병이 지시된 인물 외형과 일치하게 존재하나 약병 뚜껑은 닫혀 있음.",
    "hard_violations": [],
    "physics": "현우가 뒤로 물러나 안착한 자세와 미연이 약병을 든 손의 무게 중심이 모두 자연스러움."
   },
   {
    "label": "B",
    "direction": "미연이 내민 약병을 현우가 정확히 바라보며 응시함.",
    "built_space": "컨테이너 내부이나 식탁이 사선이 아닌 수평으로 배치되어 프레이밍 지시를 따르지 않음.",
    "entities": "현우와 미연의 외형은 일치하나 다리 상처가 보이지 않으며, 약병이 손과 식탁에 각각 하나씩 중복 생성됨.",
    "hard_violations": [
     "[gemini-pro] 중복된 사물(약병이 미연의 손과 식탁 위에 두 개 생성됨)",
     "[gpt-high] 미연이 내민 약병 외에 식탁 위에 두 번째 약병을 추가하여, 하나의 중심 약병을 중복 소품으로 표현했다."
    ],
    "physics": "미연이 뻗은 팔의 비율과 각도가 어색하며, 현우가 무언가를 짚으려는 팔의 위치가 다소 부자연스러움."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "프레이밍 지시(사선 식탁)와 인물의 자세, 다리 상처 등을 충실히 구현했으나, 현우의 시선이 약병을 향하지 않고 약병이 닫혀 있는 점이 아쉽습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "현우의 시선 방향은 정확하나, 식탁의 수평 배치로 프레이밍 지시를 위반했으며 약병이 중복 생성되는 결정적인 오류가 있습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "미연은 현우에게 약병을 내밀지만, 현우의 시선은 약병이 아닌 아래쪽 무릎이나 식탁을 향함.",
        "built_space": "컨테이너 내부 구조와 사선(oblique)으로 하단을 가로지르는 식탁 배치가 지시와 정확히 일치함.",
        "entities": "현우(상의 탈의, 얼굴과 다리 상처), 미연, 약병이 지시된 인물 외형과 일치하게 존재하나 약병 뚜껑은 닫혀 있음.",
        "hard_violations": [],
        "physics": "현우가 뒤로 물러나 안착한 자세와 미연이 약병을 든 손의 무게 중심이 모두 자연스러움."
       },
       {
        "label": "B",
        "direction": "미연이 내민 약병을 현우가 정확히 바라보며 응시함.",
        "built_space": "컨테이너 내부이나 식탁이 사선이 아닌 수평으로 배치되어 프레이밍 지시를 따르지 않음.",
        "entities": "현우와 미연의 외형은 일치하나 다리 상처가 보이지 않으며, 약병이 손과 식탁에 각각 하나씩 중복 생성됨.",
        "hard_violations": [
         "중복된 사물(약병이 미연의 손과 식탁 위에 두 개 생성됨)"
        ],
        "physics": "미연이 뻗은 팔의 비율과 각도가 어색하며, 현우가 무언가를 짚으려는 팔의 위치가 다소 부자연스러움."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "프레이밍 지시(사선 식탁)와 인물의 자세, 다리 상처 등을 충실히 구현했으나, 현우의 시선이 약병을 향하지 않고 약병이 닫혀 있는 점이 아쉽습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "현우의 시선 방향은 정확하나, 식탁의 수평 배치로 프레이밍 지시를 위반했으며 약병이 중복 생성되는 결정적인 오류가 있습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "미연은 현우에게 약병을 내밀지만, 현우의 시선은 약병이 아닌 아래쪽 무릎이나 식탁을 향함.",
        "built_space": "컨테이너 내부 구조와 사선(oblique)으로 하단을 가로지르는 식탁 배치가 지시와 정확히 일치함.",
        "entities": "현우(상의 탈의, 얼굴과 다리 상처), 미연, 약병이 지시된 인물 외형과 일치하게 존재하나 약병 뚜껑은 닫혀 있음.",
        "hard_violations": [],
        "physics": "현우가 뒤로 물러나 안착한 자세와 미연이 약병을 든 손의 무게 중심이 모두 자연스러움."
       },
       {
        "label": "B",
        "direction": "미연이 내민 약병을 현우가 정확히 바라보며 응시함.",
        "built_space": "컨테이너 내부이나 식탁이 사선이 아닌 수평으로 배치되어 프레이밍 지시를 따르지 않음.",
        "entities": "현우와 미연의 외형은 일치하나 다리 상처가 보이지 않으며, 약병이 손과 식탁에 각각 하나씩 중복 생성됨.",
        "hard_violations": [
         "중복된 사물(약병이 미연의 손과 식탁 위에 두 개 생성됨)"
        ],
        "physics": "미연이 뻗은 팔의 비율과 각도가 어색하며, 현우가 무언가를 짚으려는 팔의 위치가 다소 부자연스러움."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "현우가 약병보다 미연의 얼굴 쪽을 바라보고 상체도 충분히 물리지 않았으며, 별도의 약병까지 추가되어 핵심 거절 동작과 소품 연속성이 어긋난다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "뒤로 물린 상체, 약병을 노려보는 시선, 손과 몸통 사이의 빈 공간 및 사선 식탁이 지시를 잘 구현하지만, 약병이 닫혀 있고 미연의 의상도 참조와 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 눈은 오른쪽 위 미연의 얼굴 방향을 향하며, 그보다 아래에 있는 약병을 직접 노려보지 않는다. 미연은 현우 쪽을 보고 오른팔을 뻗어 병을 내민다. 현우도 병 가까이 손을 들어, 상체를 뒤로 빼며 거리를 유지하는 순간보다 병을 받거나 막는 순간에 가깝다.",
        "built_space": "컨테이너 패널 벽과 천장, 뒤쪽 출입구 하나, 냉장고 하나, 전자레인지가 놓인 선반 하나가 보인다. 왼쪽에는 문 하나와 창들이 있고, 현우 뒤 의자 등받이와 왼쪽 전경의 빈 의자가 보인다. 식탁 하나가 화면 아래를 크게 차지하며 촛불 하나 외에 천장등도 켜져 있어 요구된 절제된 촛불 중심 조명보다 밝다. 참조의 생활공간 재료와 주요 설비는 대체로 유지된다.",
        "entities": "상의가 없는 젊은 동아시아계 남성의 검은 머리와 얼굴 상처는 현우의 설정에 부합하며 얼굴도 참조에 비교적 가깝다. 오른쪽의 중년 동아시아계 여성은 검은 머리와 남색 소매를 보여 미연 설정과 맞는다. 미연이 든 작은 약병에는 흰 뚜껑이 붙어 있고, 식탁에는 별도의 열린 약병과 분리된 뚜껑이 있다. 식사와 촛불, 오른쪽의 접힌 옷감은 보인다. 다리 상처는 프레임 밖이므로 평가하지 않는다. 판독 가능한 글자는 확인되지 않는다.",
        "hard_violations": [
         "미연이 내민 약병 외에 식탁 위에 두 번째 약병을 추가하여, 하나의 중심 약병을 중복 소품으로 표현했다."
        ],
        "physics": "현우는 등받이가 보이는 의자에 앉아 있으며 몸이 공중에 떠 있지는 않다. 미연의 손가락이 병을 감싸 지지하고 팔도 몸에 자연스럽게 연결된다. 그릇과 약병, 촛대 및 옷감은 식탁이나 좌석 위에 놓여 있다. 다만 현우의 몸통은 대체로 곧게 서 있어 요구된 뒤로 물러나는 자세가 약하다."
       },
       {
        "label": "B",
        "direction": "현우는 고개와 눈을 왼쪽 아래로 향해 미연의 손에 든 갈색 약병을 노려본다. 미연의 얼굴도 현우와 약병 사이를 향하며, 뻗은 손은 병을 현우의 가슴 아래 높이에 제시한다. 병과 뒤로 물린 몸통 사이에 분명한 빈 공간이 남는다.",
        "built_space": "오른쪽 패널 벽을 따라 놓인 등받이 벤치 하나에 현우가 앉아 있다. 왼쪽 뒤에는 냉장고 하나와 전자레인지 하나를 포함한 수납 선반 하나가 있고, 벽의 사진들과 노출 배관 및 스위치가 참조 장소와 대응한다. 식탁 하나의 사선 가장자리가 하단을 가로지르며 촛불 하나가 놓여 있다. 미연은 왼쪽 전경에서 식탁 너머로 팔을 뻗는다. 중간 거리의 인물 중심 구도이며 실내는 촛불의 따뜻한 빛 외에는 어둡게 유지된다.",
        "entities": "현우는 검은 헝클어진 머리, 벗은 상체, 얼굴 상처를 가진 젊은 동아시아계 남성이다. 다만 참조보다 얼굴이 더 갸름하고 성숙해 보인다. 미연은 검은 단발머리의 여성으로 보이지만 얼굴 대부분이 가려져 정확한 얼굴 일치는 확인하기 어렵고, 갈색 소매는 참조의 남색 상의와 다르다. 손에 든 갈색 약병은 하나이며 뚜껑이 닫혀 있다. 식탁의 식사와 촛불, 현우 다리의 피 묻은 붕대는 보이나 개에게 물린 상처 자체는 붕대에 가려져 있다. 수선 중인 옷은 별도로 식별되지 않으며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우의 골반과 허벅지는 벤치 좌판에 지지되고, 뒤로 기운 상체는 등받이 앞에 자연스럽게 놓인다. 한 손은 허벅지 부근에, 다른 손은 좌석 가장자리 부근에 있어 거절하며 몸을 물린 자세가 물리적으로 가능하다. 미연은 손가락으로 약병을 확실히 잡고 있다. 식기와 촛대는 식탁 위에, 옷감과 붕대는 신체 위에 놓여 있어 지지 없이 떠 있는 물체는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "현우가 약병보다 미연의 얼굴 쪽을 바라보고 상체도 충분히 물리지 않았으며, 별도의 약병까지 추가되어 핵심 거절 동작과 소품 연속성이 어긋난다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "뒤로 물린 상체, 약병을 노려보는 시선, 손과 몸통 사이의 빈 공간 및 사선 식탁이 지시를 잘 구현하지만, 약병이 닫혀 있고 미연의 의상도 참조와 다르다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 눈은 오른쪽 위 미연의 얼굴 방향을 향하며, 그보다 아래에 있는 약병을 직접 노려보지 않는다. 미연은 현우 쪽을 보고 오른팔을 뻗어 병을 내민다. 현우도 병 가까이 손을 들어, 상체를 뒤로 빼며 거리를 유지하는 순간보다 병을 받거나 막는 순간에 가깝다.",
        "built_space": "컨테이너 패널 벽과 천장, 뒤쪽 출입구 하나, 냉장고 하나, 전자레인지가 놓인 선반 하나가 보인다. 왼쪽에는 문 하나와 창들이 있고, 현우 뒤 의자 등받이와 왼쪽 전경의 빈 의자가 보인다. 식탁 하나가 화면 아래를 크게 차지하며 촛불 하나 외에 천장등도 켜져 있어 요구된 절제된 촛불 중심 조명보다 밝다. 참조의 생활공간 재료와 주요 설비는 대체로 유지된다.",
        "entities": "상의가 없는 젊은 동아시아계 남성의 검은 머리와 얼굴 상처는 현우의 설정에 부합하며 얼굴도 참조에 비교적 가깝다. 오른쪽의 중년 동아시아계 여성은 검은 머리와 남색 소매를 보여 미연 설정과 맞는다. 미연이 든 작은 약병에는 흰 뚜껑이 붙어 있고, 식탁에는 별도의 열린 약병과 분리된 뚜껑이 있다. 식사와 촛불, 오른쪽의 접힌 옷감은 보인다. 다리 상처는 프레임 밖이므로 평가하지 않는다. 판독 가능한 글자는 확인되지 않는다.",
        "hard_violations": [
         "미연이 내민 약병 외에 식탁 위에 두 번째 약병을 추가하여, 하나의 중심 약병을 중복 소품으로 표현했다."
        ],
        "physics": "현우는 등받이가 보이는 의자에 앉아 있으며 몸이 공중에 떠 있지는 않다. 미연의 손가락이 병을 감싸 지지하고 팔도 몸에 자연스럽게 연결된다. 그릇과 약병, 촛대 및 옷감은 식탁이나 좌석 위에 놓여 있다. 다만 현우의 몸통은 대체로 곧게 서 있어 요구된 뒤로 물러나는 자세가 약하다."
       },
       {
        "label": "A",
        "direction": "현우는 고개와 눈을 왼쪽 아래로 향해 미연의 손에 든 갈색 약병을 노려본다. 미연의 얼굴도 현우와 약병 사이를 향하며, 뻗은 손은 병을 현우의 가슴 아래 높이에 제시한다. 병과 뒤로 물린 몸통 사이에 분명한 빈 공간이 남는다.",
        "built_space": "오른쪽 패널 벽을 따라 놓인 등받이 벤치 하나에 현우가 앉아 있다. 왼쪽 뒤에는 냉장고 하나와 전자레인지 하나를 포함한 수납 선반 하나가 있고, 벽의 사진들과 노출 배관 및 스위치가 참조 장소와 대응한다. 식탁 하나의 사선 가장자리가 하단을 가로지르며 촛불 하나가 놓여 있다. 미연은 왼쪽 전경에서 식탁 너머로 팔을 뻗는다. 중간 거리의 인물 중심 구도이며 실내는 촛불의 따뜻한 빛 외에는 어둡게 유지된다.",
        "entities": "현우는 검은 헝클어진 머리, 벗은 상체, 얼굴 상처를 가진 젊은 동아시아계 남성이다. 다만 참조보다 얼굴이 더 갸름하고 성숙해 보인다. 미연은 검은 단발머리의 여성으로 보이지만 얼굴 대부분이 가려져 정확한 얼굴 일치는 확인하기 어렵고, 갈색 소매는 참조의 남색 상의와 다르다. 손에 든 갈색 약병은 하나이며 뚜껑이 닫혀 있다. 식탁의 식사와 촛불, 현우 다리의 피 묻은 붕대는 보이나 개에게 물린 상처 자체는 붕대에 가려져 있다. 수선 중인 옷은 별도로 식별되지 않으며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우의 골반과 허벅지는 벤치 좌판에 지지되고, 뒤로 기운 상체는 등받이 앞에 자연스럽게 놓인다. 한 손은 허벅지 부근에, 다른 손은 좌석 가장자리 부근에 있어 거절하며 몸을 물린 자세가 물리적으로 가능하다. 미연은 손가락으로 약병을 확실히 잡고 있다. 식기와 촛대는 식탁 위에, 옷감과 붕대는 신체 위에 놓여 있어 지지 없이 떠 있는 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.042
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.792
   },
   "violations": {
    "B": [
     "[gemini-pro] 중복된 사물(약병이 미연의 손과 식탁 위에 두 개 생성됨)",
     "[gpt-high] 미연이 내민 약병 외에 식탁 위에 두 번째 약병을 추가하여, 하나의 중심 약병을 중복 소품으로 표현했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 792
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "프레이밍 지시(사선 식탁)와 인물의 자세, 다리 상처 등을 충실히 구현했으나, 현우의 시선이 약병을 향하지 않고 약병이 닫혀 있는 점이 아쉽습니다."
   },
   {
    "label": "B",
    "score": 792,
    "verdict_ko": "현우의 시선 방향은 정확하나, 식탁의 수평 배치로 프레이밍 지시를 위반했으며 약병이 중복 생성되는 결정적인 오류가 있습니다.  ★위반: [gemini-pro] 중복된 사물(약병이 미연의 손과 식탁 위에 두 개 생성됨) / [gpt-high] 미연이 내민 약병 외에 식탁 위에 두 번째 약병을 추가하여, 하나의 중심 약병을 중복 소품으로 표현했다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L166B01.png",
    "asset_id": "a4110c6d-4f09-489a-91bf-d6c8e7837b22",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1113064>",
    "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-ceaf-7a72-9aa9-d27e834fb596",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S12sh7__bgfirst_bg.png",
   "bg_asset_id": "520d1ce5-4724-4ae7-a50a-649a26087ab2",
   "bg_record_key": "S12sh7::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S12sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:22:35.320065+00:00",
  "fingerprint": "ef02db34bddd679c8ef6b8dc5756ad3a3e7d34985f08f5de0360304bad3e6577",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S12sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S12sh7_sel.png",
  "source_sha256": "a6320b2342589af84c237f8af301b96473c5d4a059173cac78d705f0f0e9bfe1",
  "file": "S12sh7_cine.png",
  "staged_sha256": "8a0f4079c38d60ae59de9c7aa8bf81a21b0e28b3694f8ad0f72c62d0fb9c6847",
  "latency_ms": 9235
 },
 "S12sh8::signage": {
  "fp": "be121b0c2b1fc2f0",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S12sh8": {
  "input_fingerprint": "622b0f200a2aeaab",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 양손이 식탁에 강하게 맞닿은 충격으로 촛불의 불꽃이 크게 꺾인 찰나, 양팔을 뻗은 채 굳어 있는 현우의 상체.\n\nLOCATION (lock): At the dining table inside the container home, illuminated by the candle whose flame bends with the impact. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Dining table (Both of 현우's hands have struck its top) — The tabletop is seen at a shallow angle beneath his extended arms; used as Connects the two hand contacts and gives the impact a stable spatial base; Tabletop candle (Flame sharply bent at the instant of impact) — Seen side-on in the lower-left field, separate from the hand silhouettes; used as Provides a small, readable physical consequence of the impact.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The established candlelight retains its restrained warmth, with a momentary unevenness as the flame bends.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the container dining area's table, interior surfaces, and candlelit nighttime atmosphere. Exclude the medicine bottle as a hand-held foreground element; do not transfer it from the earlier interaction.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The table candle remains the room's established light source, with the meal setting and opened medicine container still present. There is no established electrical lighting after curfew. 현우: Hyunwoo still has his upper garment off, with injured facial skin and a bleeding leg wound; no dressing has been applied.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 양손이 식탁에 강하게 맞닿은 충격으로 촛불의 불꽃이 크게 꺾인 찰나, 양팔을 뻗은 채 굳어 있는 현우의 상체.\n\nLOCATION (lock): At the dining table inside the container home, illuminated by the candle whose flame bends with the impact. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Dining table (Both of 현우's hands have struck its top) — The tabletop is seen at a shallow angle beneath his extended arms; used as Connects the two hand contacts and gives the impact a stable spatial base; Tabletop candle (Flame sharply bent at the instant of impact) — Seen side-on in the lower-left field, separate from the hand silhouettes; used as Provides a small, readable physical consequence of the impact.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The established candlelight retains its restrained warmth, with a momentary unevenness as the flame bends.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the container dining area's table, interior surfaces, and candlelit nighttime atmosphere. Exclude the medicine bottle as a hand-held foreground element; do not transfer it from the earlier interaction.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The table candle remains the room's established light source, with the meal setting and opened medicine container still present. There is no established electrical lighting after curfew. 현우: Hyunwoo still has his upper garment off, with injured facial skin and a bleeding leg wound; no dressing has been applied.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 양손이 식탁에 강하게 맞닿은 충격으로 촛불의 불꽃이 크게 꺾인 찰나, 양팔을 뻗은 채 굳어 있는 현우의 상체.\n\nLOCATION (lock): At the dining table inside the container home, illuminated by the candle whose flame bends with the impact. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Dining table (Both of 현우's hands have struck its top) — The tabletop is seen at a shallow angle beneath his extended arms; used as Connects the two hand contacts and gives the impact a stable spatial base; Tabletop candle (Flame sharply bent at the instant of impact) — Seen side-on in the lower-left field, separate from the hand silhouettes; used as Provides a small, readable physical consequence of the impact.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The established candlelight retains its restrained warmth, with a momentary unevenness as the flame bends.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the container dining area's table, interior surfaces, and candlelit nighttime atmosphere. Exclude the medicine bottle as a hand-held foreground element; do not transfer it from the earlier interaction.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The table candle remains the room's established light source, with the meal setting and opened medicine container still present. There is no established electrical lighting after curfew. 현우: Hyunwoo still has his upper garment off, with injured facial skin and a bleeding leg wound; no dressing has been applied.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선이 식탁 위를 향하고 있으며, 두 팔은 식탁 앞쪽으로 뻗어 있음.",
    "built_space": "컨테이너 내부의 구조와 가구 배치는 전반적으로 유사하나, 현우 뒤쪽 벽면에 있던 사진들이 사라졌으며, 배경 좌측에 프롬프트에 없는 두 번째 촛불이 놓여 있음.",
    "entities": "현우의 인물 특징과 상의 탈의 상태, 다리의 상처 및 붕대(참조 이미지 요소)가 나타남. 식탁 위에 식기가 있으나 구급약 통이 누락됨. 촛불의 불꽃은 오른쪽으로 크게 꺾여 있음.",
    "hard_violations": [
     "[gemini-pro] 지시된 소품(구급약 통) 누락",
     "[gemini-pro] 프롬프트 설정에 어긋나는 임의의 추가 조명(배경 좌측의 두 번째 촛불) 생성",
     "[gpt-high] 이전 장면에 없던 별도 식탁과 좌석을 왼쪽 통로에 추가하고, 그 위에 두 번째 켜진 양초를 만들어 고정된 공간과 기존 단일 촛불 광원 설정을 바꾸었다."
    ],
    "physics": "현우의 두 손이 식탁 위에 놓여 몸을 지탱하고 있으며, 촛불은 물리적 충격이나 바람에 의해 오른쪽으로 휘어진 상태로 묘사됨."
   },
   {
    "label": "B",
    "direction": "현우의 시선이 아래쪽 식탁을 향하고 있으며, 두 팔이 촛불 양옆을 향해 뻗어 있음.",
    "built_space": "컨테이너 내부 구조와 조명 분위기가 참조 이미지와 완벽히 일치하며, 현우 등 뒤 벽면의 사진들과 배관 스위치 등 고정 요소들이 제자리에 정확히 구현됨.",
    "entities": "현우의 얼굴과 체형이 참조 이미지와 일치하며, 지시된 식기류와 구급약 통이 식탁 위에 올바르게 배치되어 있음. 촛불의 불꽃이 왼쪽으로 강하게 꺾인 찰나를 잘 포착함.",
    "hard_violations": [],
    "physics": "현우의 양손이 식탁 표면에 닿아 상체를 안정적으로 지지하고 있으며, 양팔을 뻗은 충격으로 인해 촛불의 불꽃이 밀려나듯 왼쪽으로 꺾이는 물리적 현상이 자연스럽게 표현됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "참조 이미지의 공간적 디테일(벽면의 사진 등)을 정확히 유지하고, 촛불의 꺾임과 식탁 위의 구급약 통 등 프롬프트가 지시한 세부 소품 및 상황을 매우 충실하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "요구된 소품인 구급약 통을 누락하고 배경 벽면의 디테일을 지우는 등 공간 일치도가 떨어지며, 배경 좌측에 지시되지 않은 추가 촛불을 생성하여 조명 설정을 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선이 식탁 위를 향하고 있으며, 두 팔은 식탁 앞쪽으로 뻗어 있음.",
        "built_space": "컨테이너 내부의 구조와 가구 배치는 전반적으로 유사하나, 현우 뒤쪽 벽면에 있던 사진들이 사라졌으며, 배경 좌측에 프롬프트에 없는 두 번째 촛불이 놓여 있음.",
        "entities": "현우의 인물 특징과 상의 탈의 상태, 다리의 상처 및 붕대(참조 이미지 요소)가 나타남. 식탁 위에 식기가 있으나 구급약 통이 누락됨. 촛불의 불꽃은 오른쪽으로 크게 꺾여 있음.",
        "hard_violations": [
         "지시된 소품(구급약 통) 누락",
         "프롬프트 설정에 어긋나는 임의의 추가 조명(배경 좌측의 두 번째 촛불) 생성"
        ],
        "physics": "현우의 두 손이 식탁 위에 놓여 몸을 지탱하고 있으며, 촛불은 물리적 충격이나 바람에 의해 오른쪽으로 휘어진 상태로 묘사됨."
       },
       {
        "label": "B",
        "direction": "현우의 시선이 아래쪽 식탁을 향하고 있으며, 두 팔이 촛불 양옆을 향해 뻗어 있음.",
        "built_space": "컨테이너 내부 구조와 조명 분위기가 참조 이미지와 완벽히 일치하며, 현우 등 뒤 벽면의 사진들과 배관 스위치 등 고정 요소들이 제자리에 정확히 구현됨.",
        "entities": "현우의 얼굴과 체형이 참조 이미지와 일치하며, 지시된 식기류와 구급약 통이 식탁 위에 올바르게 배치되어 있음. 촛불의 불꽃이 왼쪽으로 강하게 꺾인 찰나를 잘 포착함.",
        "hard_violations": [],
        "physics": "현우의 양손이 식탁 표면에 닿아 상체를 안정적으로 지지하고 있으며, 양팔을 뻗은 충격으로 인해 촛불의 불꽃이 밀려나듯 왼쪽으로 꺾이는 물리적 현상이 자연스럽게 표현됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "참조 이미지의 공간적 디테일(벽면의 사진 등)을 정확히 유지하고, 촛불의 꺾임과 식탁 위의 구급약 통 등 프롬프트가 지시한 세부 소품 및 상황을 매우 충실하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "요구된 소품인 구급약 통을 누락하고 배경 벽면의 디테일을 지우는 등 공간 일치도가 떨어지며, 배경 좌측에 지시되지 않은 추가 촛불을 생성하여 조명 설정을 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선이 식탁 위를 향하고 있으며, 두 팔은 식탁 앞쪽으로 뻗어 있음.",
        "built_space": "컨테이너 내부의 구조와 가구 배치는 전반적으로 유사하나, 현우 뒤쪽 벽면에 있던 사진들이 사라졌으며, 배경 좌측에 프롬프트에 없는 두 번째 촛불이 놓여 있음.",
        "entities": "현우의 인물 특징과 상의 탈의 상태, 다리의 상처 및 붕대(참조 이미지 요소)가 나타남. 식탁 위에 식기가 있으나 구급약 통이 누락됨. 촛불의 불꽃은 오른쪽으로 크게 꺾여 있음.",
        "hard_violations": [
         "지시된 소품(구급약 통) 누락",
         "프롬프트 설정에 어긋나는 임의의 추가 조명(배경 좌측의 두 번째 촛불) 생성"
        ],
        "physics": "현우의 두 손이 식탁 위에 놓여 몸을 지탱하고 있으며, 촛불은 물리적 충격이나 바람에 의해 오른쪽으로 휘어진 상태로 묘사됨."
       },
       {
        "label": "B",
        "direction": "현우의 시선이 아래쪽 식탁을 향하고 있으며, 두 팔이 촛불 양옆을 향해 뻗어 있음.",
        "built_space": "컨테이너 내부 구조와 조명 분위기가 참조 이미지와 완벽히 일치하며, 현우 등 뒤 벽면의 사진들과 배관 스위치 등 고정 요소들이 제자리에 정확히 구현됨.",
        "entities": "현우의 얼굴과 체형이 참조 이미지와 일치하며, 지시된 식기류와 구급약 통이 식탁 위에 올바르게 배치되어 있음. 촛불의 불꽃이 왼쪽으로 강하게 꺾인 찰나를 잘 포착함.",
        "hard_violations": [],
        "physics": "현우의 양손이 식탁 표면에 닿아 상체를 안정적으로 지지하고 있으며, 양팔을 뻗은 충격으로 인해 촛불의 불꽃이 밀려나듯 왼쪽으로 꺾이는 물리적 현상이 자연스럽게 표현됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "양손의 식탁 접촉과 좌하단에서 크게 꺾인 촛불, 기존 공간을 충실히 담았지만 팔꿈치가 굽어 있고 약병이 닫혀 있다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "뻗은 양팔의 동작은 더 정확하지만 배경에 별도 식탁과 두 번째 촛불을 추가했으며, 드레싱이 없어야 하는 다리에 붕대가 보인다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선은 화면 왼쪽 아래 식탁 쪽을 향하며 카메라를 보지 않는다. 양손은 손바닥을 아래로 향해 같은 상판에 닿아 있다. 좌하단 촛불은 심지에서 화면 왼쪽으로 크게 꺾이며 손 윤곽과 분리되어 보인다. 불꽃의 특정 좌우 방향은 요구되지 않았다.",
        "built_space": "전경 식탁 하나와 오른쪽 벽을 따라 놓인 긴 등받이 좌석 하나가 보인다. 뒤에는 냉장고 하나, 금속 수납 선반 하나와 그 안의 전자레인지 하나, 오른쪽 벽의 전선관과 전기함 하나가 있어 이전 공간의 주요 구조를 유지한다. 현우는 좌석 앞쪽에서 식탁으로 몸을 기울인다. 낮은 각도의 상판이 두 손 접촉점을 연결하고, 상체 중심의 미디엄 구도를 이룬다. 불가능한 반사는 없다.",
        "entities": "인물은 현우 한 명뿐이다. 앳된 동아시아계 남성 외모, 헝클어진 검은 머리, 마른 체격, 벗은 상의와 얼굴 찰과상이 참조와 대체로 맞는다. 국적은 외형만으로 확인할 수 없다. 나무 식탁, 받침 위 양초 하나, 식기와 음식, 컵, 갈색 약병이 보이며 약병은 손에 들려 있지 않지만 뚜껑이 닫혀 있다. 다리 상처는 옷과 화면 경계로 확인하기 어렵다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "양 손바닥과 손가락이 상판에 붙어 있고 상체의 전방 하중을 받는다. 하체는 뒤의 좌석 위치와 자연스럽게 연결되며 떠 있는 몸이나 물체는 없다. 양초와 식기는 식탁에 놓여 있다. 불꽃은 심지에 붙은 채 옆으로 휘어 충격 직후의 공기 흐름으로 해석할 수 있다. 다만 팔꿈치가 굽고 전완이 상판 가까이 내려와 있어, 양팔을 뻗은 채 굳은 순간보다 기대고 있는 자세에 가깝다."
       },
       {
        "label": "B",
        "direction": "현우는 두 손 사이의 식탁 쪽으로 시선을 낮춘다. 두 팔은 앞으로 거의 곧게 뻗고 손바닥은 아래를 향해 상판에 닿는다. 전경 좌하단 불꽃은 화면 오른쪽 위로 뚜렷하게 기울며 손과 겹치지 않는다. 배경 왼쪽의 또 다른 촛불은 수직으로 타고 있다.",
        "built_space": "주 식탁 하나와 오른쪽의 등받이 좌석, 뒤쪽 냉장고 하나, 수납 선반 하나 및 전자레인지 하나, 벽의 전기함과 전선관은 기존 공간과 대응한다. 그러나 왼쪽 통로에 별도 작은 식탁과 좌석을 만들고 그 위에 두 번째 켜진 양초까지 배치했다. 현우의 좌석과 주 식탁의 관계 및 낮은 상판 시점은 자연스럽다. 무릎까지 보이는 구도이며 불가능한 반사는 없다.",
        "entities": "현우로 보이는 젊은 동아시아계 남성 한 명이 있으며 검은 머리, 얼굴 상처, 벗은 상의와 체격은 참조에 대체로 부합한다. 주 식탁에는 양초와 받침, 밥그릇, 다른 그릇과 젓가락이 있다. 열린 약통은 보이지 않으나 상판 일부는 화면 밖이다. 드러난 무릎에는 피 묻은 붕대가 있어 명시된 무드레싱 상태와 다르다. 별도 배경 양초가 추가되었고 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "이전 장면에 없던 별도 식탁과 좌석을 왼쪽 통로에 추가하고, 그 위에 두 번째 켜진 양초를 만들어 고정된 공간과 기존 단일 촛불 광원 설정을 바꾸었다."
        ],
        "physics": "두 손바닥이 상판에 확실히 닿고 거의 펴진 팔이 기울어진 상체를 지지한다. 골반은 뒤의 좌석에 놓인 것으로 자연스럽게 읽히며, 충격 직후 몸을 멈춘 자세가 가능하다. 양초와 식기는 각각 식탁에 받쳐져 있고 불꽃은 심지에 이어져 있어 지지 없는 물체는 없다. 전경 불꽃의 기울어짐도 물리적으로 가능한 범위다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "양손의 식탁 접촉과 좌하단에서 크게 꺾인 촛불, 기존 공간을 충실히 담았지만 팔꿈치가 굽어 있고 약병이 닫혀 있다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "뻗은 양팔의 동작은 더 정확하지만 배경에 별도 식탁과 두 번째 촛불을 추가했으며, 드레싱이 없어야 하는 다리에 붕대가 보인다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 시선은 화면 왼쪽 아래 식탁 쪽을 향하며 카메라를 보지 않는다. 양손은 손바닥을 아래로 향해 같은 상판에 닿아 있다. 좌하단 촛불은 심지에서 화면 왼쪽으로 크게 꺾이며 손 윤곽과 분리되어 보인다. 불꽃의 특정 좌우 방향은 요구되지 않았다.",
        "built_space": "전경 식탁 하나와 오른쪽 벽을 따라 놓인 긴 등받이 좌석 하나가 보인다. 뒤에는 냉장고 하나, 금속 수납 선반 하나와 그 안의 전자레인지 하나, 오른쪽 벽의 전선관과 전기함 하나가 있어 이전 공간의 주요 구조를 유지한다. 현우는 좌석 앞쪽에서 식탁으로 몸을 기울인다. 낮은 각도의 상판이 두 손 접촉점을 연결하고, 상체 중심의 미디엄 구도를 이룬다. 불가능한 반사는 없다.",
        "entities": "인물은 현우 한 명뿐이다. 앳된 동아시아계 남성 외모, 헝클어진 검은 머리, 마른 체격, 벗은 상의와 얼굴 찰과상이 참조와 대체로 맞는다. 국적은 외형만으로 확인할 수 없다. 나무 식탁, 받침 위 양초 하나, 식기와 음식, 컵, 갈색 약병이 보이며 약병은 손에 들려 있지 않지만 뚜껑이 닫혀 있다. 다리 상처는 옷과 화면 경계로 확인하기 어렵다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "양 손바닥과 손가락이 상판에 붙어 있고 상체의 전방 하중을 받는다. 하체는 뒤의 좌석 위치와 자연스럽게 연결되며 떠 있는 몸이나 물체는 없다. 양초와 식기는 식탁에 놓여 있다. 불꽃은 심지에 붙은 채 옆으로 휘어 충격 직후의 공기 흐름으로 해석할 수 있다. 다만 팔꿈치가 굽고 전완이 상판 가까이 내려와 있어, 양팔을 뻗은 채 굳은 순간보다 기대고 있는 자세에 가깝다."
       },
       {
        "label": "A",
        "direction": "현우는 두 손 사이의 식탁 쪽으로 시선을 낮춘다. 두 팔은 앞으로 거의 곧게 뻗고 손바닥은 아래를 향해 상판에 닿는다. 전경 좌하단 불꽃은 화면 오른쪽 위로 뚜렷하게 기울며 손과 겹치지 않는다. 배경 왼쪽의 또 다른 촛불은 수직으로 타고 있다.",
        "built_space": "주 식탁 하나와 오른쪽의 등받이 좌석, 뒤쪽 냉장고 하나, 수납 선반 하나 및 전자레인지 하나, 벽의 전기함과 전선관은 기존 공간과 대응한다. 그러나 왼쪽 통로에 별도 작은 식탁과 좌석을 만들고 그 위에 두 번째 켜진 양초까지 배치했다. 현우의 좌석과 주 식탁의 관계 및 낮은 상판 시점은 자연스럽다. 무릎까지 보이는 구도이며 불가능한 반사는 없다.",
        "entities": "현우로 보이는 젊은 동아시아계 남성 한 명이 있으며 검은 머리, 얼굴 상처, 벗은 상의와 체격은 참조에 대체로 부합한다. 주 식탁에는 양초와 받침, 밥그릇, 다른 그릇과 젓가락이 있다. 열린 약통은 보이지 않으나 상판 일부는 화면 밖이다. 드러난 무릎에는 피 묻은 붕대가 있어 명시된 무드레싱 상태와 다르다. 별도 배경 양초가 추가되었고 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "이전 장면에 없던 별도 식탁과 좌석을 왼쪽 통로에 추가하고, 그 위에 두 번째 켜진 양초를 만들어 고정된 공간과 기존 단일 촛불 광원 설정을 바꾸었다."
        ],
        "physics": "두 손바닥이 상판에 확실히 닿고 거의 펴진 팔이 기울어진 상체를 지지한다. 골반은 뒤의 좌석에 놓인 것으로 자연스럽게 읽히며, 충격 직후 몸을 멈춘 자세가 가능하다. 양초와 식기는 각각 식탁에 받쳐져 있고 불꽃은 심지에 이어져 있어 지지 없는 물체는 없다. 전경 불꽃의 기울어짐도 물리적으로 가능한 범위다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.042,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.792,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 지시된 소품(구급약 통) 누락",
     "[gemini-pro] 프롬프트 설정에 어긋나는 임의의 추가 조명(배경 좌측의 두 번째 촛불) 생성",
     "[gpt-high] 이전 장면에 없던 별도 식탁과 좌석을 왼쪽 통로에 추가하고, 그 위에 두 번째 켜진 양초를 만들어 고정된 공간과 기존 단일 촛불 광원 설정을 바꾸었다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 792
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "참조 이미지의 공간적 디테일(벽면의 사진 등)을 정확히 유지하고, 촛불의 꺾임과 식탁 위의 구급약 통 등 프롬프트가 지시한 세부 소품 및 상황을 매우 충실하게 구현했습니다."
   },
   {
    "label": "A",
    "score": 792,
    "verdict_ko": "요구된 소품인 구급약 통을 누락하고 배경 벽면의 디테일을 지우는 등 공간 일치도가 떨어지며, 배경 좌측에 지시되지 않은 추가 촛불을 생성하여 조명 설정을 위반했습니다.  ★위반: [gemini-pro] 지시된 소품(구급약 통) 누락 / [gemini-pro] 프롬프트 설정에 어긋나는 임의의 추가 조명(배경 좌측의 두 번째 촛불) 생성 / [gpt-high] 이전 장면에 없던 별도 식탁과 좌석을 왼쪽 통로에 추가하고, 그 위에 두 번째 켜진 양초를 만들어 고정된 공간과 기존 단일 촛불 광원 설정을 바꾸었다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S12sh7_sel.png",
    "asset_id": "49c37d3b-0355-4215-af1b-a656ae6df7a1",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-d232-7f33-bbfd-d265f4a98a9c",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S12sh7"
  }
 },
 "S12sh8::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:23:54.217469+00:00",
  "fingerprint": "c42080b1311accbc37f6415fb94f376c8f298943eb3be6818375463dc9d8116d",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S12sh8_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S12sh8_sel.png",
  "source_sha256": "38dbc435e005648cc1d48eb2fccf028e72a2a0f6f3f26b729a3c44afc87222d7",
  "file": "S12sh8_cine.png",
  "staged_sha256": "0c28ed121088c245f9b83df89b04962d8b8edf98b9b48ce4cd21394b2afa79de",
  "latency_ms": 11776
 },
 "S12sh12::signage": {
  "fp": "2695a1071ee832be",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S12sh12": {
  "input_fingerprint": "91f8cee5af05abd8",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 텅 빈 식탁 앞에 홀로 앉아 굳은 표정으로 닫힌 문 쪽을 응시하는 미연의 상체.\n\nLOCATION (lock): At the candlelit dining table inside the container home, facing the closed front door. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Unoccupied dining-table space (No other person remains at the table with 미연) — The tabletop extends from beside her toward the lower-right field; used as Makes the absent relationship visible through empty space rather than an additional figure.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the established intimate candlelight and subdued surrounding contrast without adding a new source after 현우's departure.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the same dining table, candle, and surrounding container interior under nighttime candlelight. Exclude the young man, the girl, and the momentarily bent candle flame from the reference.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The container's exterior door is now shut after being slammed. The table candle remains lit, with the meal setting and opened medicine container not yet cleared. 미연: Miyeon remains at the dining table after the argument.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 텅 빈 식탁 앞에 홀로 앉아 굳은 표정으로 닫힌 문 쪽을 응시하는 미연의 상체.\n\nLOCATION (lock): At the candlelit dining table inside the container home, facing the closed front door. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Unoccupied dining-table space (No other person remains at the table with 미연) — The tabletop extends from beside her toward the lower-right field; used as Makes the absent relationship visible through empty space rather than an additional figure.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the established intimate candlelight and subdued surrounding contrast without adding a new source after 현우's departure.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the same dining table, candle, and surrounding container interior under nighttime candlelight. Exclude the young man, the girl, and the momentarily bent candle flame from the reference.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The container's exterior door is now shut after being slammed. The table candle remains lit, with the meal setting and opened medicine container not yet cleared. 미연: Miyeon remains at the dining table after the argument.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 텅 빈 식탁 앞에 홀로 앉아 굳은 표정으로 닫힌 문 쪽을 응시하는 미연의 상체.\n\nLOCATION (lock): At the candlelit dining table inside the container home, facing the closed front door. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Unoccupied dining-table space (No other person remains at the table with 미연) — The tabletop extends from beside her toward the lower-right field; used as Makes the absent relationship visible through empty space rather than an additional figure.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the established intimate candlelight and subdued surrounding contrast without adding a new source after 현우's departure.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the same dining table, candle, and surrounding container interior under nighttime candlelight. Exclude the young man, the girl, and the momentarily bent candle flame from the reference.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The container's exterior door is now shut after being slammed. The table candle remains lit, with the meal setting and opened medicine container not yet cleared. 미연: Miyeon remains at the dining table after the argument.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "미연은 화면 우측 밖을 응시하고 있으며, 프롬프트가 지시한 타겟인 닫힌 문(배경 좌측 위치)을 향하지 않음.",
    "built_space": "컨테이너 내부. 레퍼런스 이미지에서 우측 벽에 있던 배관, 스위치, 사진들이 좌측 벽에 잘못 옮겨져 배치됨.",
    "entities": "미연의 얼굴과 의상이 캐릭터 레퍼런스와 일치함. 촛불, 약병, 식기류가 레퍼런스와 동일하게 식탁 위에 존재함.",
    "hard_violations": [],
    "physics": "벤치에 앉아 있는 신체 하중과 자세가 안정적이며, 지지대 없이 떠 있는 객체가 없음."
   },
   {
    "label": "B",
    "direction": "미연은 화면 우측을 향해 시선을 두고 있어, 배경 좌측에 있는 닫힌 문을 보지 않음.",
    "built_space": "좌측 벽의 창문 등 레퍼런스의 본래 공간 구조를 비교적 정확히 유지하고 있음.",
    "entities": "미연의 인물 형태와 촛불, 식기류가 지정된 대로 존재함.",
    "hard_violations": [
     "[gemini-pro] 해부학적으로 불가능한 신체 구조 (왼쪽 팔에 오른쪽 손이 연결됨)"
    ],
    "physics": "식탁 위에 올려둔 팔의 구조가 인체공학적으로 불가능한 형태임."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "시선이 배경의 문을 향하지 않고 벽면 고정물 위치가 뒤바뀌는 공간 오류가 있으나, 해부학적 훼손이 없어 B보다 우위입니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "왼쪽 팔에 오른쪽 손이 달려 있는 치명적인 해부학적 오류가 발생하여 규정 위반으로 탈락입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "미연은 화면 우측 밖을 응시하고 있으며, 프롬프트가 지시한 타겟인 닫힌 문(배경 좌측 위치)을 향하지 않음.",
        "built_space": "컨테이너 내부. 레퍼런스 이미지에서 우측 벽에 있던 배관, 스위치, 사진들이 좌측 벽에 잘못 옮겨져 배치됨.",
        "entities": "미연의 얼굴과 의상이 캐릭터 레퍼런스와 일치함. 촛불, 약병, 식기류가 레퍼런스와 동일하게 식탁 위에 존재함.",
        "hard_violations": [],
        "physics": "벤치에 앉아 있는 신체 하중과 자세가 안정적이며, 지지대 없이 떠 있는 객체가 없음."
       },
       {
        "label": "B",
        "direction": "미연은 화면 우측을 향해 시선을 두고 있어, 배경 좌측에 있는 닫힌 문을 보지 않음.",
        "built_space": "좌측 벽의 창문 등 레퍼런스의 본래 공간 구조를 비교적 정확히 유지하고 있음.",
        "entities": "미연의 인물 형태와 촛불, 식기류가 지정된 대로 존재함.",
        "hard_violations": [
         "해부학적으로 불가능한 신체 구조 (왼쪽 팔에 오른쪽 손이 연결됨)"
        ],
        "physics": "식탁 위에 올려둔 팔의 구조가 인체공학적으로 불가능한 형태임."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "시선이 배경의 문을 향하지 않고 벽면 고정물 위치가 뒤바뀌는 공간 오류가 있으나, 해부학적 훼손이 없어 B보다 우위입니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "왼쪽 팔에 오른쪽 손이 달려 있는 치명적인 해부학적 오류가 발생하여 규정 위반으로 탈락입니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "미연은 화면 우측 밖을 응시하고 있으며, 프롬프트가 지시한 타겟인 닫힌 문(배경 좌측 위치)을 향하지 않음.",
        "built_space": "컨테이너 내부. 레퍼런스 이미지에서 우측 벽에 있던 배관, 스위치, 사진들이 좌측 벽에 잘못 옮겨져 배치됨.",
        "entities": "미연의 얼굴과 의상이 캐릭터 레퍼런스와 일치함. 촛불, 약병, 식기류가 레퍼런스와 동일하게 식탁 위에 존재함.",
        "hard_violations": [],
        "physics": "벤치에 앉아 있는 신체 하중과 자세가 안정적이며, 지지대 없이 떠 있는 객체가 없음."
       },
       {
        "label": "B",
        "direction": "미연은 화면 우측을 향해 시선을 두고 있어, 배경 좌측에 있는 닫힌 문을 보지 않음.",
        "built_space": "좌측 벽의 창문 등 레퍼런스의 본래 공간 구조를 비교적 정확히 유지하고 있음.",
        "entities": "미연의 인물 형태와 촛불, 식기류가 지정된 대로 존재함.",
        "hard_violations": [
         "해부학적으로 불가능한 신체 구조 (왼쪽 팔에 오른쪽 손이 연결됨)"
        ],
        "physics": "식탁 위에 올려둔 팔의 구조가 인체공학적으로 불가능한 형태임."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "상체 중심의 미디엄 숏과 오른쪽 아래로 이어지는 빈 식탁은 더 충실하지만, 미연의 시선이 뒤쪽의 닫힌 문을 향하지 않고 약병도 닫혀 있다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "홀로 남은 미연과 촛불은 맞지만, 문을 빗나간 시선에 더해 인물이 작고 좌석까지 넓게 보여 지정된 상체 미디엄 숏에서 더 멀어진다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "미연의 얼굴과 눈은 화면 오른쪽, 식탁 건너편 빈 좌석과 벽 쪽을 향한다. 닫힌 문은 인물 뒤쪽의 화면 중앙에 있어, 보이는 시선이 그 문에 도달하지 않는다. 손은 식기 옆에 머물며 특정 대상을 가리키지 않는다.",
        "built_space": "뒤쪽 닫힌 문 한 개, 그 오른쪽 냉장고 한 대와 개방형 수납 선반 한 조, 선반의 전자레인지 한 대가 보인다. 왼쪽에는 창들이, 오른쪽 벽에는 사진 세 장이 있어 참고 장소의 주요 배치를 대체로 유지한다. 미연은 식탁 왼쪽 긴 좌석에 앉고 건너편 좌석은 비어 있다. 식탁이 인물 옆에서 오른쪽 아래로 넓게 이어지며, 상체가 비교적 크게 잡힌다. 불가능한 거울 반사는 없다.",
        "entities": "등장인물은 검은 단발머리와 남색 상의를 입은 중년 여성 한 명으로, 미연 참고의 얼굴 윤곽·머리·복장과 대체로 부합한다. 청년이나 소녀는 없다. 나무 식탁, 받침 위 촛불 한 개, 컵 한 개, 그릇 두 개와 음식 접시, 젓가락, 갈색 약병이 보인다. 촛불은 곧게 타지만 약병에는 뚜껑이 있어 열린 약통 조건과 다르다. 표정은 굳어 있으며 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "미연의 몸은 뒤에 보이는 나무 좌석에 지지되고 팔과 손은 식탁 위에 자연스럽게 놓인다. 식기와 약병은 상판에, 초는 받침에 놓여 있다. 떠 있는 신체나 지지 없는 물체는 없고 촛불의 상판 반사도 가능하다."
       },
       {
        "label": "B",
        "direction": "미연은 옆얼굴로 화면 오른쪽의 식탁 건너편을 바라본다. 닫힌 문은 얼굴 뒤쪽 방향인 화면 중앙 깊숙한 곳에 있어 문을 응시하는 동작으로 읽히지 않는다. 이동하거나 무언가를 겨누는 대상은 없다.",
        "built_space": "뒤쪽 문 한 개, 오른쪽 냉장고 한 대, 수납 선반 한 조와 전자레인지 한 대, 왼쪽 창들이 보인다. 다만 참고에서 오른쪽 벽에 있던 사진군과 전기 배관·스위치가 여기서는 왼쪽 벽에 나타나 장소 연속성이 약해진다. 미연은 왼쪽 긴 좌석에 홀로 앉아 있고 식탁은 오른쪽 아래로 이어진다. 좌석과 넓은 실내까지 보여 인물 상체의 화면 비중이 A보다 작다. 불가능한 반사는 보이지 않는다.",
        "entities": "검은 단발머리의 중년 여성 한 명만 있으며 남색 상의와 얼굴은 미연 참고에 대체로 부합한다. 나무 식탁 위에는 촛불 한 개, 음식 접시 한 개, 컵 한 개, 크기가 다른 그릇 세 개, 젓가락과 갈색 약병이 보인다. 약병은 뚜껑이 닫혀 있어 열린 상태의 연속성 조건을 충족하지 않는다. 굳은 표정과 야간의 따뜻한 조명은 유지되며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "엉덩이는 보이는 긴 좌석에 놓여 있고 등은 등받이 쪽을 향한다. 팔은 몸 아래로 내려가며 손은 식탁에 가려져 있어 접촉 상태를 단정할 수 없다. 식기·약병·초는 상판이나 받침에 지지되고, 부유하거나 해부학적으로 불가능한 자세는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "상체 중심의 미디엄 숏과 오른쪽 아래로 이어지는 빈 식탁은 더 충실하지만, 미연의 시선이 뒤쪽의 닫힌 문을 향하지 않고 약병도 닫혀 있다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "홀로 남은 미연과 촛불은 맞지만, 문을 빗나간 시선에 더해 인물이 작고 좌석까지 넓게 보여 지정된 상체 미디엄 숏에서 더 멀어진다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "미연의 얼굴과 눈은 화면 오른쪽, 식탁 건너편 빈 좌석과 벽 쪽을 향한다. 닫힌 문은 인물 뒤쪽의 화면 중앙에 있어, 보이는 시선이 그 문에 도달하지 않는다. 손은 식기 옆에 머물며 특정 대상을 가리키지 않는다.",
        "built_space": "뒤쪽 닫힌 문 한 개, 그 오른쪽 냉장고 한 대와 개방형 수납 선반 한 조, 선반의 전자레인지 한 대가 보인다. 왼쪽에는 창들이, 오른쪽 벽에는 사진 세 장이 있어 참고 장소의 주요 배치를 대체로 유지한다. 미연은 식탁 왼쪽 긴 좌석에 앉고 건너편 좌석은 비어 있다. 식탁이 인물 옆에서 오른쪽 아래로 넓게 이어지며, 상체가 비교적 크게 잡힌다. 불가능한 거울 반사는 없다.",
        "entities": "등장인물은 검은 단발머리와 남색 상의를 입은 중년 여성 한 명으로, 미연 참고의 얼굴 윤곽·머리·복장과 대체로 부합한다. 청년이나 소녀는 없다. 나무 식탁, 받침 위 촛불 한 개, 컵 한 개, 그릇 두 개와 음식 접시, 젓가락, 갈색 약병이 보인다. 촛불은 곧게 타지만 약병에는 뚜껑이 있어 열린 약통 조건과 다르다. 표정은 굳어 있으며 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "미연의 몸은 뒤에 보이는 나무 좌석에 지지되고 팔과 손은 식탁 위에 자연스럽게 놓인다. 식기와 약병은 상판에, 초는 받침에 놓여 있다. 떠 있는 신체나 지지 없는 물체는 없고 촛불의 상판 반사도 가능하다."
       },
       {
        "label": "A",
        "direction": "미연은 옆얼굴로 화면 오른쪽의 식탁 건너편을 바라본다. 닫힌 문은 얼굴 뒤쪽 방향인 화면 중앙 깊숙한 곳에 있어 문을 응시하는 동작으로 읽히지 않는다. 이동하거나 무언가를 겨누는 대상은 없다.",
        "built_space": "뒤쪽 문 한 개, 오른쪽 냉장고 한 대, 수납 선반 한 조와 전자레인지 한 대, 왼쪽 창들이 보인다. 다만 참고에서 오른쪽 벽에 있던 사진군과 전기 배관·스위치가 여기서는 왼쪽 벽에 나타나 장소 연속성이 약해진다. 미연은 왼쪽 긴 좌석에 홀로 앉아 있고 식탁은 오른쪽 아래로 이어진다. 좌석과 넓은 실내까지 보여 인물 상체의 화면 비중이 A보다 작다. 불가능한 반사는 보이지 않는다.",
        "entities": "검은 단발머리의 중년 여성 한 명만 있으며 남색 상의와 얼굴은 미연 참고에 대체로 부합한다. 나무 식탁 위에는 촛불 한 개, 음식 접시 한 개, 컵 한 개, 크기가 다른 그릇 세 개, 젓가락과 갈색 약병이 보인다. 약병은 뚜껑이 닫혀 있어 열린 상태의 연속성 조건을 충족하지 않는다. 굳은 표정과 야간의 따뜻한 조명은 유지되며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "엉덩이는 보이는 긴 좌석에 놓여 있고 등은 등받이 쪽을 향한다. 팔은 몸 아래로 내려가며 손은 식탁에 가려져 있어 접촉 상태를 단정할 수 없다. 식기·약병·초는 상판이나 받침에 지지되고, 부유하거나 해부학적으로 불가능한 자세는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.667,
    "B": 1.6
   },
   "adjusted": {
    "A": 1.667,
    "B": 1.35
   },
   "violations": {
    "B": [
     "[gemini-pro] 해부학적으로 불가능한 신체 구조 (왼쪽 팔에 오른쪽 손이 연결됨)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1667,
   "B": 1350
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1667,
    "verdict_ko": "시선이 배경의 문을 향하지 않고 벽면 고정물 위치가 뒤바뀌는 공간 오류가 있으나, 해부학적 훼손이 없어 B보다 우위입니다."
   },
   {
    "label": "B",
    "score": 1350,
    "verdict_ko": "왼쪽 팔에 오른쪽 손이 달려 있는 치명적인 해부학적 오류가 발생하여 규정 위반으로 탈락입니다.  ★위반: [gemini-pro] 해부학적으로 불가능한 신체 구조 (왼쪽 팔에 오른쪽 손이 연결됨)"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S12sh8_sel.png",
    "asset_id": "f2b624d2-8fff-4095-a127-d178777ec46c",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1113064>",
    "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-d3f6-7ee7-b8c2-4b51e3d1031a",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S12sh8"
  }
 },
 "S12sh12::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:25:35.172515+00:00",
  "fingerprint": "ad8f2e586e6a8934593aa390427fee48ed30a330c0a52614fa13b8ec732da746",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S12sh12_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S12sh12_sel.png",
  "source_sha256": "41510b6a73b128cd946d1bee6b62465805aa1b28473aa2e4e80451dead1cd359",
  "file": "S12sh12_cine.png",
  "staged_sha256": "c66bea9ca103e726b7c5ba6e76b5f75e2fef77955158df380350f98556b28859",
  "latency_ms": 10538
 },
 "S13sh10::signage": {
  "fp": "d23f70edea146926",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "groupbg::roadside_hideout": {
  "input_fingerprint": "84a69df1b2a9c646",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "roadside_hideout",
    "tags": [
     "S13sh10"
    ]
   },
   "context_sig": "6e556fbde0803929"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: In an outdoor gang gathering spot beside the refugee settlement's seawall, near burning metal barrels and a barbecue grill.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n인천 난민촌 인공제방과 공사장: 바다를 막고 있으나 곳곳에 심한 균열이 간 거대한 콘크리트 장벽과 그 앞의 작업 구역. (특징: 표면에 물이 스며들고 굵은 금이 간 거대한 제방 벽; 자재들이 쌓여 있는 공사 현장; 물웅덩이가 파인 질척이는 흙바닥; 지게차와 주차된 군용 트럭)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 인공제방 바로 옆 노상에 쿠마의 아지트가 있다.\n- 쿠마, 현우를 아랑곳하지 않고 인공제방 쪽으로 성큼성큼 앞서 걷는다.\n\nTIME OF DAY (lock): night.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: In an outdoor gang gathering spot beside the refugee settlement's seawall, near burning metal barrels and a barbecue grill.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n인천 난민촌 인공제방과 공사장: 바다를 막고 있으나 곳곳에 심한 균열이 간 거대한 콘크리트 장벽과 그 앞의 작업 구역. (특징: 표면에 물이 스며들고 굵은 금이 간 거대한 제방 벽; 자재들이 쌓여 있는 공사 현장; 물웅덩이가 파인 질척이는 흙바닥; 지게차와 주차된 군용 트럭)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 인공제방 바로 옆 노상에 쿠마의 아지트가 있다.\n- 쿠마, 현우를 아랑곳하지 않고 인공제방 쪽으로 성큼성큼 앞서 걷는다.\n\nTIME OF DAY (lock): night.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_roadside_hideout_f5126d.png",
  "asset_id": "3da48160-9267-4344-8758-06815d5af5e3",
  "input_asset_ids": [
   "a2b7d59b-6427-4f49-8773-4ed140b22cee"
  ],
  "origin_tag": "S13sh10",
  "place_text": "In an outdoor gang gathering spot beside the refugee settlement's seawall, near burning metal barrels and a barbecue grill.",
  "origin_inputs": {
   "place_text": "In an outdoor gang gathering spot beside the refugee settlement's seawall, near burning metal barrels and a barbecue grill.",
   "time_of_day_en": "night",
   "conti_asset_id": "a2b7d59b-6427-4f49-8773-4ed140b22cee"
  }
 },
 "S13sh10::bgfirst_bg": {
  "input_fingerprint": "cc716612c2959c56",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 빼앗은 권총의 차가운 총구를 쿠마의 이마 정중앙에 맞닿게 댄 현우의 단호한 옆얼굴.\n\nLOCATION (lock): In an outdoor gang gathering spot beside the refugee settlement's seawall, near burning metal barrels and a barbecue grill.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Seized handgun (Held with its muzzle against 쿠마's forehead) — Seen laterally, with the barrel axis running from 현우's hand toward 쿠마; used as Connects the two profiles while leaving the contact point readable and the weapon proportionate to the hands.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The hideout's established barrel fire provides selective warmth and controlled facial contrast within the nighttime darkness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 빼앗은 권총의 차가운 총구를 쿠마의 이마 정중앙에 맞닿게 댄 현우의 단호한 옆얼굴.\n\nLOCATION (lock): In an outdoor gang gathering spot beside the refugee settlement's seawall, near burning metal barrels and a barbecue grill.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Seized handgun (Held with its muzzle against 쿠마's forehead) — Seen laterally, with the barrel axis running from 현우's hand toward 쿠마; used as Connects the two profiles while leaving the contact point readable and the weapon proportionate to the hands.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The hideout's established barrel fire provides selective warmth and controlled facial contrast within the nighttime darkness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S13sh10__bgfirst_bg.png",
  "asset_id": "51a6f4c4-1510-44cd-986d-c3a6e8933313",
  "input_asset_ids": [
   "a2b7d59b-6427-4f49-8773-4ed140b22cee",
   "3da48160-9267-4344-8758-06815d5af5e3"
  ]
 },
 "S13sh10": {
  "input_fingerprint": "0c2e435c617fdd5e",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 빼앗은 권총의 차가운 총구를 쿠마의 이마 정중앙에 맞닿게 댄 현우의 단호한 옆얼굴.\n\nLOCATION (lock): In an outdoor gang gathering spot beside the refugee settlement's seawall, near burning metal barrels and a barbecue grill. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Seized handgun (Held with its muzzle against 쿠마's forehead) — Seen laterally, with the barrel axis running from 현우's hand toward 쿠마; used as Connects the two profiles while leaving the contact point readable and the weapon proportionate to the hands.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The hideout's established barrel fire provides selective warmth and controlled facial contrast within the nighttime darkness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A fire burns in a metal drum at the gang's roadside hideout, and chicken skewers are on the barbecue grill. The adjoining seawall remains cracked and leaking. 현우: Hyunwoo wears his upper garment and holds a seized gun in an aiming posture; his facial injuries and dog-bitten leg remain untreated. 쿠마: Kuma remains at the barbecue gathering with a chicken skewer in hand.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 빼앗은 권총의 차가운 총구를 쿠마의 이마 정중앙에 맞닿게 댄 현우의 단호한 옆얼굴.\n\nLOCATION (lock): In an outdoor gang gathering spot beside the refugee settlement's seawall, near burning metal barrels and a barbecue grill. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Seized handgun (Held with its muzzle against 쿠마's forehead) — Seen laterally, with the barrel axis running from 현우's hand toward 쿠마; used as Connects the two profiles while leaving the contact point readable and the weapon proportionate to the hands.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The hideout's established barrel fire provides selective warmth and controlled facial contrast within the nighttime darkness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A fire burns in a metal drum at the gang's roadside hideout, and chicken skewers are on the barbecue grill. The adjoining seawall remains cracked and leaking. 현우: Hyunwoo wears his upper garment and holds a seized gun in an aiming posture; his facial injuries and dog-bitten leg remain untreated. 쿠마: Kuma remains at the barbecue gathering with a chicken skewer in hand.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 빼앗은 권총의 차가운 총구를 쿠마의 이마 정중앙에 맞닿게 댄 현우의 단호한 옆얼굴.\n\nLOCATION (lock): In an outdoor gang gathering spot beside the refugee settlement's seawall, near burning metal barrels and a barbecue grill. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Seized handgun (Held with its muzzle against 쿠마's forehead) — Seen laterally, with the barrel axis running from 현우's hand toward 쿠마; used as Connects the two profiles while leaving the contact point readable and the weapon proportionate to the hands.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The hideout's established barrel fire provides selective warmth and controlled facial contrast within the nighttime darkness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A fire burns in a metal drum at the gang's roadside hideout, and chicken skewers are on the barbecue grill. The adjoining seawall remains cracked and leaking. 현우: Hyunwoo wears his upper garment and holds a seized gun in an aiming posture; his facial injuries and dog-bitten leg remain untreated. 쿠마: Kuma remains at the barbecue gathering with a chicken skewer in hand.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S13sh10__bgfirst_bg.png",
     "asset_id": "51a6f4c4-1510-44cd-986d-c3a6e8933313",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S13sh10.png",
     "asset_id": "a2b7d59b-6427-4f49-8773-4ed140b22cee",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 쿠마: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1218434>",
     "asset_id": "aed54006-63aa-42f8-8d50-5ff0f4385074",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:842741>",
     "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
     "role": "prop_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_roadside_hideout_f5126d.png",
     "asset_id": "3da48160-9267-4344-8758-06815d5af5e3",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 쿠마: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1218434>",
     "asset_id": "aed54006-63aa-42f8-8d50-5ff0f4385074",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:842741>",
     "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
     "role": "prop_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우의 권총 총구가 쿠마의 이마를 겨누고 맞닿아 있음.",
    "built_space": "방파제, 바비큐 그릴, 불타는 드럼통 등이 위치한 야외 배경이나, 배경 공간에 다수의 엑스트라가 배치됨.",
    "entities": "현우와 쿠마가 등장하지만, 배경에 지문에 없는 4명 이상의 추가 인물(앉아있는 사람 1명, 서있는 사람 3명 등)이 존재함.",
    "hard_violations": [
     "[gemini-pro] invented people (지문에 없는 추가 인물 다수 등장)",
     "[gpt-high] 샷에 허용된 현우와 쿠마 외에 배경 인물 네 명을 추가했다."
    ],
    "physics": "현우의 손이 권총을 쥐고 있고, 쿠마의 손이 꼬치를 들고 있어 두 사물 모두 적절히 지탱됨."
   },
   {
    "label": "B",
    "direction": "현우가 든 권총의 총구가 쿠마의 이마 정중앙을 정확히 겨누고 맞닿아 있음.",
    "built_space": "야간의 방파제 배경이며, 불이 피워진 드럼통 화로가 적절한 위치에 배치됨.",
    "entities": "현우(왼쪽)와 쿠마(오른쪽)만 등장하여 레퍼런스와 일치함. 현우의 얼굴 상처, 권총, 쿠마가 든 닭꼬치 모두 지시사항에 부합함.",
    "hard_violations": [],
    "physics": "현우의 손이 권총을 제대로 쥐어 지탱하고 있으며, 쿠마의 손 역시 꼬치 막대를 안정적으로 들고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "요구된 클로즈업 샷 크기를 완벽히 준수했으며, 명시된 인물만 프레임에 담아 총구를 겨눈 팽팽한 대치 상황을 매우 훌륭하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "프레이밍이 요구된 클로즈업보다 훨씬 넓게 잡혔으며, 지문에 없는 다수의 엑스트라 인물을 배경에 추가하는 치명적인 위반이 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "현우가 든 권총의 총구가 쿠마의 이마 정중앙을 정확히 겨누고 맞닿아 있음.",
        "built_space": "야간의 방파제 배경이며, 불이 피워진 드럼통 화로가 적절한 위치에 배치됨.",
        "entities": "현우(왼쪽)와 쿠마(오른쪽)만 등장하여 레퍼런스와 일치함. 현우의 얼굴 상처, 권총, 쿠마가 든 닭꼬치 모두 지시사항에 부합함.",
        "hard_violations": [],
        "physics": "현우의 손이 권총을 제대로 쥐어 지탱하고 있으며, 쿠마의 손 역시 꼬치 막대를 안정적으로 들고 있음."
       },
       {
        "label": "A",
        "direction": "현우의 권총 총구가 쿠마의 이마를 겨누고 맞닿아 있음.",
        "built_space": "방파제, 바비큐 그릴, 불타는 드럼통 등이 위치한 야외 배경이나, 배경 공간에 다수의 엑스트라가 배치됨.",
        "entities": "현우와 쿠마가 등장하지만, 배경에 지문에 없는 4명 이상의 추가 인물(앉아있는 사람 1명, 서있는 사람 3명 등)이 존재함.",
        "hard_violations": [
         "invented people (지문에 없는 추가 인물 다수 등장)"
        ],
        "physics": "현우의 손이 권총을 쥐고 있고, 쿠마의 손이 꼬치를 들고 있어 두 사물 모두 적절히 지탱됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "요구된 클로즈업 샷 크기를 완벽히 준수했으며, 명시된 인물만 프레임에 담아 총구를 겨눈 팽팽한 대치 상황을 매우 훌륭하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "프레이밍이 요구된 클로즈업보다 훨씬 넓게 잡혔으며, 지문에 없는 다수의 엑스트라 인물을 배경에 추가하는 치명적인 위반이 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우가 든 권총의 총구가 쿠마의 이마 정중앙을 정확히 겨누고 맞닿아 있음.",
        "built_space": "야간의 방파제 배경이며, 불이 피워진 드럼통 화로가 적절한 위치에 배치됨.",
        "entities": "현우(왼쪽)와 쿠마(오른쪽)만 등장하여 레퍼런스와 일치함. 현우의 얼굴 상처, 권총, 쿠마가 든 닭꼬치 모두 지시사항에 부합함.",
        "hard_violations": [],
        "physics": "현우의 손이 권총을 제대로 쥐어 지탱하고 있으며, 쿠마의 손 역시 꼬치 막대를 안정적으로 들고 있음."
       },
       {
        "label": "A",
        "direction": "현우의 권총 총구가 쿠마의 이마를 겨누고 맞닿아 있음.",
        "built_space": "방파제, 바비큐 그릴, 불타는 드럼통 등이 위치한 야외 배경이나, 배경 공간에 다수의 엑스트라가 배치됨.",
        "entities": "현우와 쿠마가 등장하지만, 배경에 지문에 없는 4명 이상의 추가 인물(앉아있는 사람 1명, 서있는 사람 3명 등)이 존재함.",
        "hard_violations": [
         "invented people (지문에 없는 추가 인물 다수 등장)"
        ],
        "physics": "현우의 손이 권총을 쥐고 있고, 쿠마의 손이 꼬치를 들고 있어 두 사물 모두 적절히 지탱됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "현우의 단호한 옆얼굴과 쿠마의 이마에 닿은 총구를 근접 구도로 정확히 연결하지만, 현우의 겉옷과 개머리판 없는 권총은 참조와 다르다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "허용되지 않은 배경 인물 네 명을 추가했고, 클로즈업 대신 장소와 여러 사람을 넓게 보여주며 쿠마의 외형도 참조에서 크게 벗어난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 현우는 오른쪽 쿠마를 응시하고 쿠마도 현우 쪽을 바라본다. 현우의 손에서 오른쪽으로 이어지는 권총의 총열이 쿠마의 이마 중앙 부위에 닿으며, 측면에서 접촉점이 읽힌다. 쿠마가 든 꼬치는 위를 향한다.",
        "built_space": "두 사람의 얼굴과 현우의 손·팔이 전경을 크게 차지한다. 뒤에는 불타는 금속 드럼통 한 개와 그 앞의 사각 그릴 한 개가 보이고, 오른쪽으로 콘크리트 방파제, 왼쪽으로 천막과 적재물이 이어진다. 젖은 통로와 야간 항만 조명이 장소 참조에 부합하며 시설 중복이나 불가능한 반사는 보이지 않는다.",
        "entities": "인물은 현우와 쿠마 두 명뿐이다. 현우는 앳된 동아시아계 남성의 얼굴, 헝클어진 검은 머리, 치료하지 않은 얼굴 상처를 보인다. 다만 참조의 남색 티셔츠와 달리 회녹색 칼라 겉옷을 입었다. 쿠마는 짧은 검은 머리와 옅은 수염이 있는 젊은 성인 남성으로 참조에 비교적 가깝고 남색 상의를 입었다. 국적과 혼혈 여부 자체는 외모만으로 확인할 수 없다. 쿠마의 손에는 닭고기 꼬치가 있으나 그릴 위 음식은 선명하게 식별되지 않는다. 권총은 금속 재질과 측면 형태가 유사하지만 참조의 긴 개머리판이 없다. 다리 부상은 프레임 밖이며 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "현우의 손이 권총 손잡이를 감싸고 손목과 팔뚝이 연결되어 총을 지지한다. 팔을 들어 상대의 이마에 총구를 대는 자세가 물리적으로 가능하며 총과 손의 크기도 자연스럽다. 쿠마의 꼬치는 화면 아래쪽 손에 잡혀 있다. 두 사람의 하체는 잘렸지만 상체가 공중에 떠 있다는 증거는 없고, 드럼통과 그릴은 바닥에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "왼쪽 현우는 낮은 위치의 쿠마를 내려다보고 쿠마는 현우를 올려다본다. 권총은 현우의 손에서 오른쪽 아래로 향해 쿠마의 이마에 닿으며 접촉점은 보인다. 쿠마의 꼬치는 얼굴 앞에서 오른쪽 위로 기울어 있다. 배경 인물들은 대체로 두 사람의 대치 쪽을 향한다.",
        "built_space": "현우의 상체와 길게 뻗은 팔, 낮은 위치의 쿠마, 넓은 항만 통로를 함께 보여주는 구도로 요구된 클로즈업보다 넓다. 중앙 아래에 불타는 드럼통 한 개, 그 앞에 사각 그릴 한 개가 있고 오른쪽 방파제와 바다, 왼쪽 천막·지게차·트럭이 보인다. 장소의 주요 재료와 배치는 참조와 가깝지만 배경에 서 있는 사람 세 명과 앉아 있는 사람 한 명을 추가했다.",
        "entities": "현우는 검은 머리의 어린 동아시아계 남성이고 얼굴에 상처가 있으나 회녹색 겉옷은 참조 의상과 다르다. 쿠마는 남색 상의를 입고 꼬치를 들었지만, 참조보다 앳되고 머리가 길며 수염이 거의 없어 지정된 얼굴과 헤어스타일의 일치도가 낮다. 쿠마의 뺨에도 상처가 보인다. 두 주인공 외에 배경 인물 네 명이 있다. 그릴에는 여러 고기 꼬치가 놓여 있다. 권총에는 참조의 개머리판이 없고 총열·슬라이드 형태도 다르다. 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [
         "샷에 허용된 현우와 쿠마 외에 배경 인물 네 명을 추가했다."
        ],
        "physics": "현우의 뻗은 팔과 손이 권총을 지지하며 총구를 상대의 이마에 대는 동작은 가능하다. 쿠마는 손으로 꼬치를 잡고 있다. 쿠마의 하체와 좌석 여부는 화면 밖이어서 낮은 자세의 지지 방식을 확정할 수 없지만, 부유한다고 볼 근거도 없다. 배경의 앉은 사람은 상자형 좌석에 지지되고 서 있는 사람들은 통로에 위치한다. 그릴은 다리로, 드럼통은 밑면으로 바닥에 지지된다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "현우의 단호한 옆얼굴과 쿠마의 이마에 닿은 총구를 근접 구도로 정확히 연결하지만, 현우의 겉옷과 개머리판 없는 권총은 참조와 다르다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "허용되지 않은 배경 인물 네 명을 추가했고, 클로즈업 대신 장소와 여러 사람을 넓게 보여주며 쿠마의 외형도 참조에서 크게 벗어난다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽 현우는 오른쪽 쿠마를 응시하고 쿠마도 현우 쪽을 바라본다. 현우의 손에서 오른쪽으로 이어지는 권총의 총열이 쿠마의 이마 중앙 부위에 닿으며, 측면에서 접촉점이 읽힌다. 쿠마가 든 꼬치는 위를 향한다.",
        "built_space": "두 사람의 얼굴과 현우의 손·팔이 전경을 크게 차지한다. 뒤에는 불타는 금속 드럼통 한 개와 그 앞의 사각 그릴 한 개가 보이고, 오른쪽으로 콘크리트 방파제, 왼쪽으로 천막과 적재물이 이어진다. 젖은 통로와 야간 항만 조명이 장소 참조에 부합하며 시설 중복이나 불가능한 반사는 보이지 않는다.",
        "entities": "인물은 현우와 쿠마 두 명뿐이다. 현우는 앳된 동아시아계 남성의 얼굴, 헝클어진 검은 머리, 치료하지 않은 얼굴 상처를 보인다. 다만 참조의 남색 티셔츠와 달리 회녹색 칼라 겉옷을 입었다. 쿠마는 짧은 검은 머리와 옅은 수염이 있는 젊은 성인 남성으로 참조에 비교적 가깝고 남색 상의를 입었다. 국적과 혼혈 여부 자체는 외모만으로 확인할 수 없다. 쿠마의 손에는 닭고기 꼬치가 있으나 그릴 위 음식은 선명하게 식별되지 않는다. 권총은 금속 재질과 측면 형태가 유사하지만 참조의 긴 개머리판이 없다. 다리 부상은 프레임 밖이며 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "현우의 손이 권총 손잡이를 감싸고 손목과 팔뚝이 연결되어 총을 지지한다. 팔을 들어 상대의 이마에 총구를 대는 자세가 물리적으로 가능하며 총과 손의 크기도 자연스럽다. 쿠마의 꼬치는 화면 아래쪽 손에 잡혀 있다. 두 사람의 하체는 잘렸지만 상체가 공중에 떠 있다는 증거는 없고, 드럼통과 그릴은 바닥에 놓여 있다."
       },
       {
        "label": "A",
        "direction": "왼쪽 현우는 낮은 위치의 쿠마를 내려다보고 쿠마는 현우를 올려다본다. 권총은 현우의 손에서 오른쪽 아래로 향해 쿠마의 이마에 닿으며 접촉점은 보인다. 쿠마의 꼬치는 얼굴 앞에서 오른쪽 위로 기울어 있다. 배경 인물들은 대체로 두 사람의 대치 쪽을 향한다.",
        "built_space": "현우의 상체와 길게 뻗은 팔, 낮은 위치의 쿠마, 넓은 항만 통로를 함께 보여주는 구도로 요구된 클로즈업보다 넓다. 중앙 아래에 불타는 드럼통 한 개, 그 앞에 사각 그릴 한 개가 있고 오른쪽 방파제와 바다, 왼쪽 천막·지게차·트럭이 보인다. 장소의 주요 재료와 배치는 참조와 가깝지만 배경에 서 있는 사람 세 명과 앉아 있는 사람 한 명을 추가했다.",
        "entities": "현우는 검은 머리의 어린 동아시아계 남성이고 얼굴에 상처가 있으나 회녹색 겉옷은 참조 의상과 다르다. 쿠마는 남색 상의를 입고 꼬치를 들었지만, 참조보다 앳되고 머리가 길며 수염이 거의 없어 지정된 얼굴과 헤어스타일의 일치도가 낮다. 쿠마의 뺨에도 상처가 보인다. 두 주인공 외에 배경 인물 네 명이 있다. 그릴에는 여러 고기 꼬치가 놓여 있다. 권총에는 참조의 개머리판이 없고 총열·슬라이드 형태도 다르다. 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [
         "샷에 허용된 현우와 쿠마 외에 배경 인물 네 명을 추가했다."
        ],
        "physics": "현우의 뻗은 팔과 손이 권총을 지지하며 총구를 상대의 이마에 대는 동작은 가능하다. 쿠마는 손으로 꼬치를 잡고 있다. 쿠마의 하체와 좌석 여부는 화면 밖이어서 낮은 자세의 지지 방식을 확정할 수 없지만, 부유한다고 볼 근거도 없다. 배경의 앉은 사람은 상자형 좌석에 지지되고 서 있는 사람들은 통로에 위치한다. 그릴은 다리로, 드럼통은 밑면으로 바닥에 지지된다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.625,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.375,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] invented people (지문에 없는 추가 인물 다수 등장)",
     "[gpt-high] 샷에 허용된 현우와 쿠마 외에 배경 인물 네 명을 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 375
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "요구된 클로즈업 샷 크기를 완벽히 준수했으며, 명시된 인물만 프레임에 담아 총구를 겨눈 팽팽한 대치 상황을 매우 훌륭하게 구현했습니다."
   },
   {
    "label": "A",
    "score": 375,
    "verdict_ko": "프레이밍이 요구된 클로즈업보다 훨씬 넓게 잡혔으며, 지문에 없는 다수의 엑스트라 인물을 배경에 추가하는 치명적인 위반이 발생했습니다.  ★위반: [gemini-pro] invented people (지문에 없는 추가 인물 다수 등장) / [gpt-high] 샷에 허용된 현우와 쿠마 외에 배경 인물 네 명을 추가했다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_roadside_hideout_f5126d.png",
    "asset_id": "3da48160-9267-4344-8758-06815d5af5e3",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 쿠마: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1218434>",
    "asset_id": "aed54006-63aa-42f8-8d50-5ff0f4385074",
    "role": "character_ref"
   },
   {
    "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
    "path": "<bytes:842741>",
    "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
    "role": "prop_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-d5bc-7c5b-97fa-2001aea503eb",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S13sh10__bgfirst_bg.png",
   "bg_asset_id": "51a6f4c4-1510-44cd-986d-c3a6e8933313",
   "bg_record_key": "S13sh10::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "roadside_hideout",
   "groupbg_asset_id": "3da48160-9267-4344-8758-06815d5af5e3"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S13sh10::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:27:44.958926+00:00",
  "fingerprint": "fd2b4d89f23b2b14a1f6231e8ad9fb5811ffbabd260ea1547b904e09e04718ad",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S13sh10_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S13sh10_sel.png",
  "source_sha256": "831a992006206886d706d1f5fb77674d6a3094acf9631a45fa22b5106b77a2da",
  "file": "S13sh10_cine.png",
  "staged_sha256": "8e86bbe76cacf2f69b58491c340d97479bbb4b3c7e74c29360bb84e705328ec5",
  "latency_ms": 17494
 },
 "S13sh13::signage": {
  "fp": "1bd17acfabc10faf",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S13sh13": {
  "input_fingerprint": "9450b5b667e185fd",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 갈라진 콘크리트 틈새로 탁한 물이 줄줄 흘러내리는 거대한 인공제방의 젖은 벽면 클로즈업.\n\nLOCATION (lock): Directly against the exterior face of the refugee settlement's concrete seawall, where water seeps through deep cracks at night. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Cracked embankment wall (Wet, with murky water running through the concrete cracks) — Its exposed face is viewed obliquely, revealing the crack and the wall's recession; used as Provides the environmental field around the narrow band of active seepage; Water flowing from the cracks (Running downward over the wall); used as Forms the primary moving detail without relying on a reflection or an added atmospheric effect; Embankment supports (Water also runs over the supports) — Only a peripheral portion is visible along the receding wall edge; used as Maintains structural context at the boundary of the close view.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the wall within the scene's established nighttime darkness, using restrained tonal separation to reveal the flowing water without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall has visible cracks with seawater seeping through, and water runs along its supports in the nighttime darkness. The nearby hideout's drum fire and barbecue remain established.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 갈라진 콘크리트 틈새로 탁한 물이 줄줄 흘러내리는 거대한 인공제방의 젖은 벽면 클로즈업.\n\nLOCATION (lock): Directly against the exterior face of the refugee settlement's concrete seawall, where water seeps through deep cracks at night. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Cracked embankment wall (Wet, with murky water running through the concrete cracks) — Its exposed face is viewed obliquely, revealing the crack and the wall's recession; used as Provides the environmental field around the narrow band of active seepage; Water flowing from the cracks (Running downward over the wall); used as Forms the primary moving detail without relying on a reflection or an added atmospheric effect; Embankment supports (Water also runs over the supports) — Only a peripheral portion is visible along the receding wall edge; used as Maintains structural context at the boundary of the close view.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the wall within the scene's established nighttime darkness, using restrained tonal separation to reveal the flowing water without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall has visible cracks with seawater seeping through, and water runs along its supports in the nighttime darkness. The nearby hideout's drum fire and barbecue remain established.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 갈라진 콘크리트 틈새로 탁한 물이 줄줄 흘러내리는 거대한 인공제방의 젖은 벽면 클로즈업.\n\nLOCATION (lock): Directly against the exterior face of the refugee settlement's concrete seawall, where water seeps through deep cracks at night. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Cracked embankment wall (Wet, with murky water running through the concrete cracks) — Its exposed face is viewed obliquely, revealing the crack and the wall's recession; used as Provides the environmental field around the narrow band of active seepage; Water flowing from the cracks (Running downward over the wall); used as Forms the primary moving detail without relying on a reflection or an added atmospheric effect; Embankment supports (Water also runs over the supports) — Only a peripheral portion is visible along the receding wall edge; used as Maintains structural context at the boundary of the close view.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the wall within the scene's established nighttime darkness, using restrained tonal separation to reveal the flowing water without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall has visible cracks with seawater seeping through, and water runs along its supports in the nighttime darkness. The nearby hideout's drum fire and barbecue remain established.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "탁한 물줄기가 콘크리트 벽의 깊게 갈라진 틈에서 빠져나와 수직으로 떨어지고 있음.",
    "built_space": "수직에 가까운 거대한 콘크리트 벽면이 사선 구도로 배치되어 있으며, 우측 후경을 따라 벽면을 지탱하는 수직 돌출 구조물(지지대)들이 규칙적으로 배열되어 있음. 멀리 가로등이 보임.",
    "entities": "인물은 등장하지 않음. 젖고 갈라진 콘크리트 제방 벽, 탁한 물줄기가 프롬프트의 지시대로 정확히 존재함.",
    "hard_violations": [
     "[gpt-high] 참조에 확립되지 않은 가로등 두 개를 발광하는 상태로 추가하여, 새로운 광원을 도입하지 말라는 명시적 제한을 위반합니다."
    ],
    "physics": "물이 중력의 영향을 받아 틈새에서 아래로 자연스럽게 쏟아져 내리며, 벽면의 질감과 물의 반사가 물리적으로 타당함."
   },
   {
    "label": "B",
    "direction": "물줄기가 갈라진 틈에서 시작되어 경사진 벽면을 타고 아래로 흐름.",
    "built_space": "상당히 기울어진 사면 형태의 콘크리트 제방이 보이며, 하단에는 물이 고이거나 흐를 수 있는 턱이 있음. 프롬프트가 요구한 벽 가장자리의 수직 지지대(supports)는 보이지 않음.",
    "entities": "인물은 등장하지 않음. 갈라진 콘크리트 벽과 흐르는 탁한 물이 존재함.",
    "hard_violations": [
     "[gpt-high] 참조에 확립되지 않은 인공 점광원들을 배경에 추가하여, 새로운 광원을 도입하지 말라는 명시적 제한을 위반합니다."
    ],
    "physics": "물이 경사면을 따라 중력 방향으로 자연스럽게 흘러내리고 있으며, 젖은 표면의 반사 등 물리적 표현이 무난함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "프롬프트가 요구한 야간의 조명 상태와 갈라진 틈에서 쏟아지는 물줄기, 그리고 후경으로 물러나며 보이는 제방 지지대(supports) 구조를 레퍼런스에 맞춰 충실히 구현했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "갈라진 콘크리트 틈과 흐르는 물은 표현되었으나, 프롬프트에 명시된 가장자리의 제방 지지대(supports) 구조가 누락되었고 벽면의 각도가 레퍼런스와 다소 차이가 있습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "탁한 물줄기가 콘크리트 벽의 깊게 갈라진 틈에서 빠져나와 수직으로 떨어지고 있음.",
        "built_space": "수직에 가까운 거대한 콘크리트 벽면이 사선 구도로 배치되어 있으며, 우측 후경을 따라 벽면을 지탱하는 수직 돌출 구조물(지지대)들이 규칙적으로 배열되어 있음. 멀리 가로등이 보임.",
        "entities": "인물은 등장하지 않음. 젖고 갈라진 콘크리트 제방 벽, 탁한 물줄기가 프롬프트의 지시대로 정확히 존재함.",
        "hard_violations": [],
        "physics": "물이 중력의 영향을 받아 틈새에서 아래로 자연스럽게 쏟아져 내리며, 벽면의 질감과 물의 반사가 물리적으로 타당함."
       },
       {
        "label": "B",
        "direction": "물줄기가 갈라진 틈에서 시작되어 경사진 벽면을 타고 아래로 흐름.",
        "built_space": "상당히 기울어진 사면 형태의 콘크리트 제방이 보이며, 하단에는 물이 고이거나 흐를 수 있는 턱이 있음. 프롬프트가 요구한 벽 가장자리의 수직 지지대(supports)는 보이지 않음.",
        "entities": "인물은 등장하지 않음. 갈라진 콘크리트 벽과 흐르는 탁한 물이 존재함.",
        "hard_violations": [],
        "physics": "물이 경사면을 따라 중력 방향으로 자연스럽게 흘러내리고 있으며, 젖은 표면의 반사 등 물리적 표현이 무난함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "프롬프트가 요구한 야간의 조명 상태와 갈라진 틈에서 쏟아지는 물줄기, 그리고 후경으로 물러나며 보이는 제방 지지대(supports) 구조를 레퍼런스에 맞춰 충실히 구현했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "갈라진 콘크리트 틈과 흐르는 물은 표현되었으나, 프롬프트에 명시된 가장자리의 제방 지지대(supports) 구조가 누락되었고 벽면의 각도가 레퍼런스와 다소 차이가 있습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "탁한 물줄기가 콘크리트 벽의 깊게 갈라진 틈에서 빠져나와 수직으로 떨어지고 있음.",
        "built_space": "수직에 가까운 거대한 콘크리트 벽면이 사선 구도로 배치되어 있으며, 우측 후경을 따라 벽면을 지탱하는 수직 돌출 구조물(지지대)들이 규칙적으로 배열되어 있음. 멀리 가로등이 보임.",
        "entities": "인물은 등장하지 않음. 젖고 갈라진 콘크리트 제방 벽, 탁한 물줄기가 프롬프트의 지시대로 정확히 존재함.",
        "hard_violations": [],
        "physics": "물이 중력의 영향을 받아 틈새에서 아래로 자연스럽게 쏟아져 내리며, 벽면의 질감과 물의 반사가 물리적으로 타당함."
       },
       {
        "label": "B",
        "direction": "물줄기가 갈라진 틈에서 시작되어 경사진 벽면을 타고 아래로 흐름.",
        "built_space": "상당히 기울어진 사면 형태의 콘크리트 제방이 보이며, 하단에는 물이 고이거나 흐를 수 있는 턱이 있음. 프롬프트가 요구한 벽 가장자리의 수직 지지대(supports)는 보이지 않음.",
        "entities": "인물은 등장하지 않음. 갈라진 콘크리트 벽과 흐르는 탁한 물이 존재함.",
        "hard_violations": [],
        "physics": "물이 경사면을 따라 중력 방향으로 자연스럽게 흘러내리고 있으며, 젖은 표면의 반사 등 물리적 표현이 무난함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "탁한 물의 하향 흐름과 야간 근접 촬영은 맞지만, 크게 드러난 배수로와 경사진 벽체가 참조 장소에서 벗어나고 새 조명까지 보입니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "참조의 심하게 갈라진 콘크리트와 균열에서 쏟아지는 탁수를 더 충실하게 근접 포착하지만, 가로등을 추가하고 지지대 구간을 지나치게 넓게 보여줍니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "화면 중앙 왼쪽의 깊은 균열에서 갈색 물이 나와 벽면을 따라 오른쪽 아래로 내려갑니다. 하단의 수평 턱에 도달한 물은 그 가장자리로 넘칩니다. 사람의 시선이나 겨냥하는 물체는 없습니다.",
        "built_space": "비스듬한 콘크리트 벽 한 면이 오른쪽 먼 곳으로 후퇴합니다. 벽 아래에는 연속된 넓은 턱 하나와 그 바깥의 깊은 수로 하나가 보이며, 오른쪽 끝에는 수로를 가로지르는 부재 일부가 있습니다. 참조의 크게 파손된 벽면과 달리 넓고 비교적 연속적인 경사면이며, 주변부에 그쳐야 할 구조물이 하단을 크게 차지합니다. 배경에는 흰 점광원이 적어도 네 개 보입니다.",
        "entities": "젖고 갈라진 콘크리트, 탁한 갈색 물, 제방 구조는 식별됩니다. 사람·얼굴·읽을 수 있는 글자는 없습니다. 콘크리트의 파손 양상은 참조보다 덜 박락되어 있고 벽의 형태도 다릅니다. 드럼통 불과 바비큐는 근접 프레임 밖이므로 미노출 자체는 문제가 아닙니다.",
        "hard_violations": [
         "참조에 확립되지 않은 인공 점광원들을 배경에 추가하여, 새로운 광원을 도입하지 말라는 명시적 제한을 위반합니다."
        ],
        "physics": "물은 균열이라는 출구에서 나와 중력 방향으로 벽에 붙어 흐르고, 아래 턱에 모인 뒤 가장자리로 넘칩니다. 벽과 턱이 물의 접촉면을 제공하며 흐름은 물리적으로 가능합니다. 지지 없이 떠 있는 물체나 신체는 없습니다."
       },
       {
        "label": "B",
        "direction": "중앙 왼쪽의 벌어진 콘크리트 틈에서 탁수가 나와 거의 수직으로 화면 아래를 향해 떨어집니다. 오른쪽 지지대 주변의 물도 벽을 따라 바닥으로 내려갑니다. 요구한 균열에서 아래로 흐르는 방향이 명확합니다.",
        "built_space": "크게 갈라지고 박락된 벽 한 면이 화면 대부분을 채우며 오른쪽으로 후퇴합니다. 가까운 돌출 지지대 하나와 원근으로 작아지는 지지대들이 최소 네 개 더 보입니다. 오른쪽에는 젖은 지면, 상단에는 난간 한 줄, 배경에는 가로등 두 개가 보입니다. 주된 벽면의 재질과 손상은 참조에 더 가깝지만, 지지대와 지면이 단순한 주변부보다 넓게 노출됩니다.",
        "entities": "거친 회색 콘크리트, 깊은 균열, 갈색을 띠는 탁수, 물에 젖은 지지대가 모두 식별됩니다. 사람·얼굴·읽을 수 있는 글자는 없습니다. 참조의 덩어리째 갈라진 콘크리트 외관을 A보다 잘 유지합니다. 드럼통 불과 바비큐는 이 근접 구도에서 보이지 않아도 됩니다.",
        "hard_violations": [
         "참조에 확립되지 않은 가로등 두 개를 발광하는 상태로 추가하여, 새로운 광원을 도입하지 말라는 명시적 제한을 위반합니다."
        ],
        "physics": "물줄기는 벽의 균열에서 시작하여 아래로 떨어지고, 일부는 콘크리트 표면에 붙어 흐릅니다. 오른쪽의 물은 지지대 표면을 타고 내려가 바닥의 고인 물에 합류합니다. 균열 출구와 접촉면, 낙하 방향이 연결되어 있으며 지지 없이 떠 있는 물체는 없습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "탁한 물의 하향 흐름과 야간 근접 촬영은 맞지만, 크게 드러난 배수로와 경사진 벽체가 참조 장소에서 벗어나고 새 조명까지 보입니다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "참조의 심하게 갈라진 콘크리트와 균열에서 쏟아지는 탁수를 더 충실하게 근접 포착하지만, 가로등을 추가하고 지지대 구간을 지나치게 넓게 보여줍니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "화면 중앙 왼쪽의 깊은 균열에서 갈색 물이 나와 벽면을 따라 오른쪽 아래로 내려갑니다. 하단의 수평 턱에 도달한 물은 그 가장자리로 넘칩니다. 사람의 시선이나 겨냥하는 물체는 없습니다.",
        "built_space": "비스듬한 콘크리트 벽 한 면이 오른쪽 먼 곳으로 후퇴합니다. 벽 아래에는 연속된 넓은 턱 하나와 그 바깥의 깊은 수로 하나가 보이며, 오른쪽 끝에는 수로를 가로지르는 부재 일부가 있습니다. 참조의 크게 파손된 벽면과 달리 넓고 비교적 연속적인 경사면이며, 주변부에 그쳐야 할 구조물이 하단을 크게 차지합니다. 배경에는 흰 점광원이 적어도 네 개 보입니다.",
        "entities": "젖고 갈라진 콘크리트, 탁한 갈색 물, 제방 구조는 식별됩니다. 사람·얼굴·읽을 수 있는 글자는 없습니다. 콘크리트의 파손 양상은 참조보다 덜 박락되어 있고 벽의 형태도 다릅니다. 드럼통 불과 바비큐는 근접 프레임 밖이므로 미노출 자체는 문제가 아닙니다.",
        "hard_violations": [
         "참조에 확립되지 않은 인공 점광원들을 배경에 추가하여, 새로운 광원을 도입하지 말라는 명시적 제한을 위반합니다."
        ],
        "physics": "물은 균열이라는 출구에서 나와 중력 방향으로 벽에 붙어 흐르고, 아래 턱에 모인 뒤 가장자리로 넘칩니다. 벽과 턱이 물의 접촉면을 제공하며 흐름은 물리적으로 가능합니다. 지지 없이 떠 있는 물체나 신체는 없습니다."
       },
       {
        "label": "A",
        "direction": "중앙 왼쪽의 벌어진 콘크리트 틈에서 탁수가 나와 거의 수직으로 화면 아래를 향해 떨어집니다. 오른쪽 지지대 주변의 물도 벽을 따라 바닥으로 내려갑니다. 요구한 균열에서 아래로 흐르는 방향이 명확합니다.",
        "built_space": "크게 갈라지고 박락된 벽 한 면이 화면 대부분을 채우며 오른쪽으로 후퇴합니다. 가까운 돌출 지지대 하나와 원근으로 작아지는 지지대들이 최소 네 개 더 보입니다. 오른쪽에는 젖은 지면, 상단에는 난간 한 줄, 배경에는 가로등 두 개가 보입니다. 주된 벽면의 재질과 손상은 참조에 더 가깝지만, 지지대와 지면이 단순한 주변부보다 넓게 노출됩니다.",
        "entities": "거친 회색 콘크리트, 깊은 균열, 갈색을 띠는 탁수, 물에 젖은 지지대가 모두 식별됩니다. 사람·얼굴·읽을 수 있는 글자는 없습니다. 참조의 덩어리째 갈라진 콘크리트 외관을 A보다 잘 유지합니다. 드럼통 불과 바비큐는 이 근접 구도에서 보이지 않아도 됩니다.",
        "hard_violations": [
         "참조에 확립되지 않은 가로등 두 개를 발광하는 상태로 추가하여, 새로운 광원을 도입하지 말라는 명시적 제한을 위반합니다."
        ],
        "physics": "물줄기는 벽의 균열에서 시작하여 아래로 떨어지고, 일부는 콘크리트 표면에 붙어 흐릅니다. 오른쪽의 물은 지지대 표면을 타고 내려가 바닥의 고인 물에 합류합니다. 균열 출구와 접촉면, 낙하 방향이 연결되어 있으며 지지 없이 떠 있는 물체는 없습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.238
   },
   "adjusted": {
    "A": 1.75,
    "B": 0.988
   },
   "violations": {
    "B": [
     "[gpt-high] 참조에 확립되지 않은 인공 점광원들을 배경에 추가하여, 새로운 광원을 도입하지 말라는 명시적 제한을 위반합니다."
    ],
    "A": [
     "[gpt-high] 참조에 확립되지 않은 가로등 두 개를 발광하는 상태로 추가하여, 새로운 광원을 도입하지 말라는 명시적 제한을 위반합니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 988
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "프롬프트가 요구한 야간의 조명 상태와 갈라진 틈에서 쏟아지는 물줄기, 그리고 후경으로 물러나며 보이는 제방 지지대(supports) 구조를 레퍼런스에 맞춰 충실히 구현했습니다.  ★위반: [gpt-high] 참조에 확립되지 않은 가로등 두 개를 발광하는 상태로 추가하여, 새로운 광원을 도입하지 말라는 명시적 제한을 위반합니다."
   },
   {
    "label": "B",
    "score": 988,
    "verdict_ko": "갈라진 콘크리트 틈과 흐르는 물은 표현되었으나, 프롬프트에 명시된 가장자리의 제방 지지대(supports) 구조가 누락되었고 벽면의 각도가 레퍼런스와 다소 차이가 있습니다.  ★위반: [gpt-high] 참조에 확립되지 않은 인공 점광원들을 배경에 추가하여, 새로운 광원을 도입하지 말라는 명시적 제한을 위반합니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S7sh11_sel.png",
    "asset_id": "8463927d-a43d-4e9f-a68f-6bce56177ade",
    "role": "prev_still"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-dac4-7db9-96b2-79b3557248dd",
  "ref_mode": "prev만 (배경 전용·공유 계획)",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S7sh11"
  },
  "lane_policy": "ab_select_bypass:bg_only:share_plan_prev_bgonly"
 },
 "S13sh13::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:28:46.993276+00:00",
  "fingerprint": "bc13d0942d70c08541d98e50bc71c48af48b49b4f3e5fb8ee382a9aade2c1fb7",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S13sh13_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S13sh13_sel.png",
  "source_sha256": "cf87c8cabfa5158f14d6e3c6f40cc2ebfc49dcf2eb2e15ea66d64aa6d912bb38",
  "file": "S13sh13_cine.png",
  "staged_sha256": "e4a7571ddcc68170ba6f1217da4920cb9ee7236d134ffafdcae7b1d20ce75cd7",
  "latency_ms": 10543
 },
 "S13sh20::signage": {
  "fp": "e96c4612775fbad4",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S13sh20": {
  "input_fingerprint": "6370cf93b115caf5",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 쿠마의 손에 들린 지폐 뭉치를 재빠르게 움켜쥔 찰나의 현우 손 클로즈업.\n\nLOCATION (lock): At the foot of the leaking seawall beside the refugee settlement, where the nighttime negotiation concludes. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Banknotes (Still supported by 쿠마 as 현우 grips them) — Overlapping note faces are seen obliquely between the hands; no denomination needs to be legible; used as Marks the exact transfer of possession while remaining subordinate in size to the hands; Embankment wall (Cracked and wet in the surrounding location) — A soft, partial section remains beyond the cropped torsos; used as Retains the negotiation's location without competing with the hand action.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain subdued nighttime ambient visibility at the embankment without carrying the hideout's firelight into this location.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the cracked concrete wall, persistent seawater seepage, and dark nighttime coloration as the environmental reference. Exclude the barbecue grill, beer, and burning barrel from the separate hideout.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall remains wet and cracked, with water running along the supports under nighttime darkness. The nearby drum fire and barbecue have not been extinguished. 현우: Hyunwoo is taking possession of the cash payment and still retains the seized gun; his upper garment is on and his facial and leg injuries persist. 쿠마: Kuma stands beside the seawall with one palm still wet, completing the cash handover from inside his clothing.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 쿠마, 현우 right now, so 쿠마, 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 쿠마, 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 쿠마의 손에 들린 지폐 뭉치를 재빠르게 움켜쥔 찰나의 현우 손 클로즈업.\n\nLOCATION (lock): At the foot of the leaking seawall beside the refugee settlement, where the nighttime negotiation concludes. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Banknotes (Still supported by 쿠마 as 현우 grips them) — Overlapping note faces are seen obliquely between the hands; no denomination needs to be legible; used as Marks the exact transfer of possession while remaining subordinate in size to the hands; Embankment wall (Cracked and wet in the surrounding location) — A soft, partial section remains beyond the cropped torsos; used as Retains the negotiation's location without competing with the hand action.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain subdued nighttime ambient visibility at the embankment without carrying the hideout's firelight into this location.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the cracked concrete wall, persistent seawater seepage, and dark nighttime coloration as the environmental reference. Exclude the barbecue grill, beer, and burning barrel from the separate hideout.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall remains wet and cracked, with water running along the supports under nighttime darkness. The nearby drum fire and barbecue have not been extinguished. 현우: Hyunwoo is taking possession of the cash payment and still retains the seized gun; his upper garment is on and his facial and leg injuries persist. 쿠마: Kuma stands beside the seawall with one palm still wet, completing the cash handover from inside his clothing.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 쿠마, 현우 right now, so 쿠마, 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 쿠마, 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 쿠마의 손에 들린 지폐 뭉치를 재빠르게 움켜쥔 찰나의 현우 손 클로즈업.\n\nLOCATION (lock): At the foot of the leaking seawall beside the refugee settlement, where the nighttime negotiation concludes. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Banknotes (Still supported by 쿠마 as 현우 grips them) — Overlapping note faces are seen obliquely between the hands; no denomination needs to be legible; used as Marks the exact transfer of possession while remaining subordinate in size to the hands; Embankment wall (Cracked and wet in the surrounding location) — A soft, partial section remains beyond the cropped torsos; used as Retains the negotiation's location without competing with the hand action.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain subdued nighttime ambient visibility at the embankment without carrying the hideout's firelight into this location.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the cracked concrete wall, persistent seawater seepage, and dark nighttime coloration as the environmental reference. Exclude the barbecue grill, beer, and burning barrel from the separate hideout.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall remains wet and cracked, with water running along the supports under nighttime darkness. The nearby drum fire and barbecue have not been extinguished. 현우: Hyunwoo is taking possession of the cash payment and still retains the seized gun; his upper garment is on and his facial and leg injuries persist. 쿠마: Kuma stands beside the seawall with one palm still wet, completing the cash handover from inside his clothing.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 쿠마, 현우 right now, so 쿠마, 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 쿠마, 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우와 쿠마의 손이 중앙의 지폐 뭉치를 향해 정확히 뻗어 교환하고 있음.",
    "built_space": "물이 흐르고 균열이 간 방파제 벽이 배경에 올바르게 배치됨.",
    "entities": "현우의 얼굴(상처 포함)과 쿠마의 젖은 손바닥이 지침과 일치하며, 지폐 문자는 비식별화됨.",
    "hard_violations": [],
    "physics": "손이 지폐를 단단히 쥐고 있으며 물리적 지지 관계가 자연스러움."
   },
   {
    "label": "B",
    "direction": "두 손이 프레임 중앙의 지폐 뭉치를 향해 마주보고 있음.",
    "built_space": "갈라진 방파제 벽과 물줄기가 배경에 적절히 묘사됨.",
    "entities": "손과 지폐가 묘사되었으나, 지폐 위에 명확한 텍스트가 식별됨.",
    "hard_violations": [
     "[gemini-pro] No readable writing 지침 위반: 지폐에 'UH GDOSCC', '오안' 등 읽을 수 있는 텍스트가 선명하게 렌더링됨.",
     "[gpt-high] 지폐 표면에 '한국은행' 등 읽을 수 있는 문자가 노출되어, 이미지 어디에도 읽을 수 있는 글자가 없어야 한다는 조건을 위반한다."
    ],
    "physics": "지폐를 잡고 있는 양손의 파지가 물리적으로 타당함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "지폐의 텍스트를 적절히 비식별화하고 현우의 상처 및 쿠마의 젖은 손 등 주요 디테일을 성공적으로 구현함."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "지폐에 명확히 읽히는 문자를 생성하여 '읽을 수 있는 텍스트 금지' 지침을 심각하게 위반함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우와 쿠마의 손이 중앙의 지폐 뭉치를 향해 정확히 뻗어 교환하고 있음.",
        "built_space": "물이 흐르고 균열이 간 방파제 벽이 배경에 올바르게 배치됨.",
        "entities": "현우의 얼굴(상처 포함)과 쿠마의 젖은 손바닥이 지침과 일치하며, 지폐 문자는 비식별화됨.",
        "hard_violations": [],
        "physics": "손이 지폐를 단단히 쥐고 있으며 물리적 지지 관계가 자연스러움."
       },
       {
        "label": "B",
        "direction": "두 손이 프레임 중앙의 지폐 뭉치를 향해 마주보고 있음.",
        "built_space": "갈라진 방파제 벽과 물줄기가 배경에 적절히 묘사됨.",
        "entities": "손과 지폐가 묘사되었으나, 지폐 위에 명확한 텍스트가 식별됨.",
        "hard_violations": [
         "No readable writing 지침 위반: 지폐에 'UH GDOSCC', '오안' 등 읽을 수 있는 텍스트가 선명하게 렌더링됨."
        ],
        "physics": "지폐를 잡고 있는 양손의 파지가 물리적으로 타당함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "지폐의 텍스트를 적절히 비식별화하고 현우의 상처 및 쿠마의 젖은 손 등 주요 디테일을 성공적으로 구현함."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "지폐에 명확히 읽히는 문자를 생성하여 '읽을 수 있는 텍스트 금지' 지침을 심각하게 위반함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우와 쿠마의 손이 중앙의 지폐 뭉치를 향해 정확히 뻗어 교환하고 있음.",
        "built_space": "물이 흐르고 균열이 간 방파제 벽이 배경에 올바르게 배치됨.",
        "entities": "현우의 얼굴(상처 포함)과 쿠마의 젖은 손바닥이 지침과 일치하며, 지폐 문자는 비식별화됨.",
        "hard_violations": [],
        "physics": "손이 지폐를 단단히 쥐고 있으며 물리적 지지 관계가 자연스러움."
       },
       {
        "label": "B",
        "direction": "두 손이 프레임 중앙의 지폐 뭉치를 향해 마주보고 있음.",
        "built_space": "갈라진 방파제 벽과 물줄기가 배경에 적절히 묘사됨.",
        "entities": "손과 지폐가 묘사되었으나, 지폐 위에 명확한 텍스트가 식별됨.",
        "hard_violations": [
         "No readable writing 지침 위반: 지폐에 'UH GDOSCC', '오안' 등 읽을 수 있는 텍스트가 선명하게 렌더링됨."
        ],
        "physics": "지폐를 잡고 있는 양손의 파지가 물리적으로 타당함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "양손이 지폐를 함께 지지하는 순간은 맞지만, 지폐의 읽을 수 있는 글자가 명시적 금지 조건을 위반하고 방벽이 손보다 과도하게 강조된다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "쿠마의 젖은 손바닥이 받친 지폐 뭉치를 현우가 움켜쥐는 관계와 손 중심 근접 구도가 더 정확하지만, 얼굴과 의복 노출이 많고 참고 의상과 차이가 있다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽의 젖은 손이 지폐를 오른쪽으로 내밀고, 오른쪽 손이 그 지폐 윗면을 엄지와 손가락으로 집는다. 접촉 대상은 정확히 같은 지폐 묶음이다. 얼굴과 눈, 총은 보이지 않아 시선이나 총구 방향은 평가할 수 없다.",
        "built_space": "뒤에는 갈라진 콘크리트 방벽 한 면, 중앙 균열에서 떨어지는 물줄기 하나, 오른쪽으로 이어지는 버팀벽들과 가로등 불빛 하나가 보인다. 두 사람의 몸통은 좌우 가장자리에 잘려 있다. 장소의 재질과 야간 색조는 참고와 맞지만, 방벽과 젖은 바닥이 넓고 비교적 선명해 요청한 부드러운 부분 배경보다 강하게 드러난다.",
        "entities": "두 사람의 손과 팔, 왼쪽 사람의 아래로 내린 다른 손이 보이며 추가 인물은 없다. 젖은 왼손은 쿠마, 오른쪽에서 잡는 손은 현우로 해석된다. 손의 피부와 형태는 젊은 성인에 부합하지만 얼굴이 없어 정확한 나이·민족적 배경·인물 동일성은 확인할 수 없다. 왼쪽에는 남색 상의와 회녹색 겉옷, 오른쪽에는 검은 긴소매가 보여 참고의 남색 티셔츠 차림과 차이가 있다. 지폐는 한국 화폐 도안을 닮았으며 '한국은행' 등 글자가 읽혀 무문자 조건에 어긋난다. 얼굴·다리 부상과 소지한 총은 프레임 밖이므로 불일치로 판단하지 않는다.",
        "hard_violations": [
         "지폐 표면에 '한국은행' 등 읽을 수 있는 문자가 노출되어, 이미지 어디에도 읽을 수 있는 글자가 없어야 한다는 조건을 위반한다."
        ],
        "physics": "지폐는 왼쪽 손가락이 아래에서 받치고 오른쪽 엄지가 위에서 누르므로 떠 있지 않다. 두 팔은 화면 가장자리의 몸통 방향으로 자연스럽게 이어진다. 아래로 내린 손 역시 팔에 연결되어 있다. 집는 동작은 가능하지만 손 전체로 재빨리 움켜쥐기보다는 끝을 집어 받는 모습에 가깝다. 물줄기는 벽의 균열에서 아래로 떨어진다."
       },
       {
        "label": "B",
        "direction": "왼쪽 현우의 손이 중앙 지폐 뭉치를 감싸 잡고, 오른쪽 쿠마의 펼친 손바닥은 같은 뭉치를 현우 쪽으로 내밀며 받친다. 양손의 행동 대상이 일치한다. 현우의 얼굴은 손 쪽으로 숙여져 있지만 눈은 잘려 정확한 시선은 확인할 수 없다. 총구는 보이지 않는다.",
        "built_space": "배경에는 갈라지고 젖은 콘크리트 방벽 한 면, 중앙 왼쪽의 누수 물줄기 하나, 오른쪽으로 물러나는 버팀벽들과 상단의 가로등 불빛 하나가 보인다. 현우의 잘린 얼굴과 상체가 왼쪽, 쿠마의 잘린 몸통과 팔이 오른쪽에 배치된다. 구조물의 중복이나 불가능한 반사는 없다. 손 뒤의 방벽은 흐려져 장소를 유지하면서 상대적으로 덜 경쟁하지만, 얼굴과 상체가 손 클로즈업에 필요한 범위보다 많이 포함된다.",
        "entities": "현우의 얼굴 하부와 손, 쿠마의 손과 팔이 보이며 추가 인물은 없다. 현우는 앳된 남성으로 보이고 볼의 상처가 유지된다. 얼굴 일부만 보여 참고 인물과의 완전한 동일성 및 한국계 미국인이라는 배경은 확인할 수 없다. 쿠마의 손은 젊은 성인 남성에게 가능한 크기와 피부이며 손바닥에 물기가 보이지만, 얼굴이 없어 중국계 혼혈 정체성은 확인할 수 없다. 현우의 회녹색 겉옷·남색 안쪽 옷·밝은 속옷과 쿠마의 검은 긴소매는 참고의 단순한 남색 티셔츠와 다르다. 여러 장이 겹친 실제 종이 지폐 뭉치가 비스듬히 보이며 국가와 액면은 확정할 만큼 읽히지 않는다. 총과 다리 부상은 프레임 밖이므로 감점 근거로 삼지 않는다.",
        "hard_violations": [],
        "physics": "쿠마의 젖은 손바닥과 굽힌 손가락이 지폐 아래를 계속 받치고, 현우의 엄지와 나머지 손가락이 반대편을 감싸 압박한다. 지폐의 휨과 겹친 가장자리가 양손의 압력에 부합하며 지지 없이 뜬 물체는 없다. 손목과 팔의 연결도 자연스럽고, 양쪽이 동시에 잡은 인계 순간으로 성립한다. 배경의 물은 균열에서 중력 방향으로 흐른다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "양손이 지폐를 함께 지지하는 순간은 맞지만, 지폐의 읽을 수 있는 글자가 명시적 금지 조건을 위반하고 방벽이 손보다 과도하게 강조된다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "쿠마의 젖은 손바닥이 받친 지폐 뭉치를 현우가 움켜쥐는 관계와 손 중심 근접 구도가 더 정확하지만, 얼굴과 의복 노출이 많고 참고 의상과 차이가 있다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽의 젖은 손이 지폐를 오른쪽으로 내밀고, 오른쪽 손이 그 지폐 윗면을 엄지와 손가락으로 집는다. 접촉 대상은 정확히 같은 지폐 묶음이다. 얼굴과 눈, 총은 보이지 않아 시선이나 총구 방향은 평가할 수 없다.",
        "built_space": "뒤에는 갈라진 콘크리트 방벽 한 면, 중앙 균열에서 떨어지는 물줄기 하나, 오른쪽으로 이어지는 버팀벽들과 가로등 불빛 하나가 보인다. 두 사람의 몸통은 좌우 가장자리에 잘려 있다. 장소의 재질과 야간 색조는 참고와 맞지만, 방벽과 젖은 바닥이 넓고 비교적 선명해 요청한 부드러운 부분 배경보다 강하게 드러난다.",
        "entities": "두 사람의 손과 팔, 왼쪽 사람의 아래로 내린 다른 손이 보이며 추가 인물은 없다. 젖은 왼손은 쿠마, 오른쪽에서 잡는 손은 현우로 해석된다. 손의 피부와 형태는 젊은 성인에 부합하지만 얼굴이 없어 정확한 나이·민족적 배경·인물 동일성은 확인할 수 없다. 왼쪽에는 남색 상의와 회녹색 겉옷, 오른쪽에는 검은 긴소매가 보여 참고의 남색 티셔츠 차림과 차이가 있다. 지폐는 한국 화폐 도안을 닮았으며 '한국은행' 등 글자가 읽혀 무문자 조건에 어긋난다. 얼굴·다리 부상과 소지한 총은 프레임 밖이므로 불일치로 판단하지 않는다.",
        "hard_violations": [
         "지폐 표면에 '한국은행' 등 읽을 수 있는 문자가 노출되어, 이미지 어디에도 읽을 수 있는 글자가 없어야 한다는 조건을 위반한다."
        ],
        "physics": "지폐는 왼쪽 손가락이 아래에서 받치고 오른쪽 엄지가 위에서 누르므로 떠 있지 않다. 두 팔은 화면 가장자리의 몸통 방향으로 자연스럽게 이어진다. 아래로 내린 손 역시 팔에 연결되어 있다. 집는 동작은 가능하지만 손 전체로 재빨리 움켜쥐기보다는 끝을 집어 받는 모습에 가깝다. 물줄기는 벽의 균열에서 아래로 떨어진다."
       },
       {
        "label": "A",
        "direction": "왼쪽 현우의 손이 중앙 지폐 뭉치를 감싸 잡고, 오른쪽 쿠마의 펼친 손바닥은 같은 뭉치를 현우 쪽으로 내밀며 받친다. 양손의 행동 대상이 일치한다. 현우의 얼굴은 손 쪽으로 숙여져 있지만 눈은 잘려 정확한 시선은 확인할 수 없다. 총구는 보이지 않는다.",
        "built_space": "배경에는 갈라지고 젖은 콘크리트 방벽 한 면, 중앙 왼쪽의 누수 물줄기 하나, 오른쪽으로 물러나는 버팀벽들과 상단의 가로등 불빛 하나가 보인다. 현우의 잘린 얼굴과 상체가 왼쪽, 쿠마의 잘린 몸통과 팔이 오른쪽에 배치된다. 구조물의 중복이나 불가능한 반사는 없다. 손 뒤의 방벽은 흐려져 장소를 유지하면서 상대적으로 덜 경쟁하지만, 얼굴과 상체가 손 클로즈업에 필요한 범위보다 많이 포함된다.",
        "entities": "현우의 얼굴 하부와 손, 쿠마의 손과 팔이 보이며 추가 인물은 없다. 현우는 앳된 남성으로 보이고 볼의 상처가 유지된다. 얼굴 일부만 보여 참고 인물과의 완전한 동일성 및 한국계 미국인이라는 배경은 확인할 수 없다. 쿠마의 손은 젊은 성인 남성에게 가능한 크기와 피부이며 손바닥에 물기가 보이지만, 얼굴이 없어 중국계 혼혈 정체성은 확인할 수 없다. 현우의 회녹색 겉옷·남색 안쪽 옷·밝은 속옷과 쿠마의 검은 긴소매는 참고의 단순한 남색 티셔츠와 다르다. 여러 장이 겹친 실제 종이 지폐 뭉치가 비스듬히 보이며 국가와 액면은 확정할 만큼 읽히지 않는다. 총과 다리 부상은 프레임 밖이므로 감점 근거로 삼지 않는다.",
        "hard_violations": [],
        "physics": "쿠마의 젖은 손바닥과 굽힌 손가락이 지폐 아래를 계속 받치고, 현우의 엄지와 나머지 손가락이 반대편을 감싸 압박한다. 지폐의 휨과 겹친 가장자리가 양손의 압력에 부합하며 지지 없이 뜬 물체는 없다. 손목과 팔의 연결도 자연스럽고, 양쪽이 동시에 잡은 인계 순간으로 성립한다. 배경의 물은 균열에서 중력 방향으로 흐른다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.625
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.375
   },
   "violations": {
    "B": [
     "[gemini-pro] No readable writing 지침 위반: 지폐에 'UH GDOSCC', '오안' 등 읽을 수 있는 텍스트가 선명하게 렌더링됨.",
     "[gpt-high] 지폐 표면에 '한국은행' 등 읽을 수 있는 문자가 노출되어, 이미지 어디에도 읽을 수 있는 글자가 없어야 한다는 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 375
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지폐의 텍스트를 적절히 비식별화하고 현우의 상처 및 쿠마의 젖은 손 등 주요 디테일을 성공적으로 구현함."
   },
   {
    "label": "B",
    "score": 375,
    "verdict_ko": "지폐에 명확히 읽히는 문자를 생성하여 '읽을 수 있는 텍스트 금지' 지침을 심각하게 위반함.  ★위반: [gemini-pro] No readable writing 지침 위반: 지폐에 'UH GDOSCC', '오안' 등 읽을 수 있는 텍스트가 선명하게 렌더링됨. / [gpt-high] 지폐 표면에 '한국은행' 등 읽을 수 있는 문자가 노출되어, 이미지 어디에도 읽을 수 있는 글자가 없어야 한다는 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S13sh13_sel.png",
    "asset_id": "f01bde92-a444-4c1f-ae8e-069d4483136c",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 쿠마: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1218434>",
    "asset_id": "aed54006-63aa-42f8-8d50-5ff0f4385074",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-dc81-767a-b961-9a76c53383f2",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S13sh13"
  }
 },
 "S13sh20::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:30:28.993046+00:00",
  "fingerprint": "afc778568bb323941f385d0218d45ca366113ec0f81913a7d2b67eb4653429db",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S13sh20_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S13sh20_sel.png",
  "source_sha256": "a8797158c805c741eb2789d1370a9ca4f37529a50db6b68bc58c7c04c42aafdb",
  "file": "S13sh20_cine.png",
  "staged_sha256": "85c0b0a7d01e2a8e7a70c9c8c08248228b0a3acd4e011e6e9f7750ce3bb84ede",
  "latency_ms": 30575
 },
 "S14sh3::signage": {
  "fp": "3862b78c3466387e",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S14sh3": {
  "input_fingerprint": "cea9f573b95c3e8d",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 깨진 화병 조각 위로 발끝을 슬며시 밀고 있는 mid-action 자세의 찰리 발밑 클로즈업.\n\nLOCATION (lock): On the floor of the container home's shared living area, beside broken vase fragments in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Broken vase fragments (Being pushed along the floor by 찰리's foot); used as A small scattered group ahead of the toe makes the concealment attempt readable without exaggerated fragment scale; Container interior floor (Supporting 찰리's feet and the broken fragments) — Viewed at a downward oblique angle from beside his feet; used as Provides uninterrupted spatial context for the direction and extent of the sliding action.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained daytime ambient illumination appropriate to the container interior, preserving readable floor-level detail without specifying an unsupported source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Broken vase fragments are on the container floor, being pushed aside rather than removed. Charlie retains his dirty aged gorilla-shaped casing, blue-lit eyes, worn Ubik chest logo, netting, and previously donned old coat and hat.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 깨진 화병 조각 위로 발끝을 슬며시 밀고 있는 mid-action 자세의 찰리 발밑 클로즈업.\n\nLOCATION (lock): On the floor of the container home's shared living area, beside broken vase fragments in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Broken vase fragments (Being pushed along the floor by 찰리's foot); used as A small scattered group ahead of the toe makes the concealment attempt readable without exaggerated fragment scale; Container interior floor (Supporting 찰리's feet and the broken fragments) — Viewed at a downward oblique angle from beside his feet; used as Provides uninterrupted spatial context for the direction and extent of the sliding action.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained daytime ambient illumination appropriate to the container interior, preserving readable floor-level detail without specifying an unsupported source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Broken vase fragments are on the container floor, being pushed aside rather than removed. Charlie retains his dirty aged gorilla-shaped casing, blue-lit eyes, worn Ubik chest logo, netting, and previously donned old coat and hat.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 깨진 화병 조각 위로 발끝을 슬며시 밀고 있는 mid-action 자세의 찰리 발밑 클로즈업.\n\nLOCATION (lock): On the floor of the container home's shared living area, beside broken vase fragments in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Broken vase fragments (Being pushed along the floor by 찰리's foot); used as A small scattered group ahead of the toe makes the concealment attempt readable without exaggerated fragment scale; Container interior floor (Supporting 찰리's feet and the broken fragments) — Viewed at a downward oblique angle from beside his feet; used as Provides uninterrupted spatial context for the direction and extent of the sliding action.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained daytime ambient illumination appropriate to the container interior, preserving readable floor-level detail without specifying an unsupported source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Broken vase fragments are on the container floor, being pushed aside rather than removed. Charlie retains his dirty aged gorilla-shaped casing, blue-lit eyes, worn Ubik chest logo, netting, and previously donned old coat and hat.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "기계 발이 화면 왼쪽을 향해 파편을 밀고 있으며, 바닥에 뚜렷한 쓸림 자국이 형성됨.",
    "built_space": "컨테이너 내부. 배경에 기준 이미지와 유사한 나무 벤치와 탁자가 배치되었으나 왼쪽에 철제 선반이 추가됨. 밝은 주간 채광.",
    "entities": "찰리의 기계 발(베이지색 장갑 및 피스톤 구조 일치), 낡은 코트 자락, 청화백자 파편.",
    "hard_violations": [],
    "physics": "발이 바닥을 단단히 딛고 파편을 미는 중이며, 쓸린 궤적이 물리적인 움직임을 명확히 뒷받침함."
   },
   {
    "label": "B",
    "direction": "기계 발이 파편들 사이에 정적으로 놓여 특정 방향으로의 움직임이 보이지 않음.",
    "built_space": "컨테이너 내부. 왼쪽에 배관이 있는 패널 벽이 배치되었으나, 오른쪽에 기준 이미지에 없는 거대한 철골 기둥이 세워져 있음. 나무 바닥.",
    "entities": "찰리의 기계 발(구조 일치), 다양한 색상의 화병 파편들. 요구된 코트는 보이지 않음.",
    "hard_violations": [],
    "physics": "발이 나무 바닥에 지지되어 있으나, 밀거나 미끄러지는 동적인 힘의 작용이나 흔적이 없음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "프롬프트가 강조한 발로 화병 파편을 슬며시 미는 동작과 바닥의 쓸림 궤적을 훌륭하게 연출했으며, 코트 자락의 디테일도 잘 반영되었습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "파편을 미는 슬라이딩 액션이 생략된 채 정적으로 서 있으며, 배경에 기준 이미지에 존재하지 않는 굵은 철골 기둥이 추가되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "기계 발이 화면 왼쪽을 향해 파편을 밀고 있으며, 바닥에 뚜렷한 쓸림 자국이 형성됨.",
        "built_space": "컨테이너 내부. 배경에 기준 이미지와 유사한 나무 벤치와 탁자가 배치되었으나 왼쪽에 철제 선반이 추가됨. 밝은 주간 채광.",
        "entities": "찰리의 기계 발(베이지색 장갑 및 피스톤 구조 일치), 낡은 코트 자락, 청화백자 파편.",
        "hard_violations": [],
        "physics": "발이 바닥을 단단히 딛고 파편을 미는 중이며, 쓸린 궤적이 물리적인 움직임을 명확히 뒷받침함."
       },
       {
        "label": "B",
        "direction": "기계 발이 파편들 사이에 정적으로 놓여 특정 방향으로의 움직임이 보이지 않음.",
        "built_space": "컨테이너 내부. 왼쪽에 배관이 있는 패널 벽이 배치되었으나, 오른쪽에 기준 이미지에 없는 거대한 철골 기둥이 세워져 있음. 나무 바닥.",
        "entities": "찰리의 기계 발(구조 일치), 다양한 색상의 화병 파편들. 요구된 코트는 보이지 않음.",
        "hard_violations": [],
        "physics": "발이 나무 바닥에 지지되어 있으나, 밀거나 미끄러지는 동적인 힘의 작용이나 흔적이 없음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "프롬프트가 강조한 발로 화병 파편을 슬며시 미는 동작과 바닥의 쓸림 궤적을 훌륭하게 연출했으며, 코트 자락의 디테일도 잘 반영되었습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "파편을 미는 슬라이딩 액션이 생략된 채 정적으로 서 있으며, 배경에 기준 이미지에 존재하지 않는 굵은 철골 기둥이 추가되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "기계 발이 화면 왼쪽을 향해 파편을 밀고 있으며, 바닥에 뚜렷한 쓸림 자국이 형성됨.",
        "built_space": "컨테이너 내부. 배경에 기준 이미지와 유사한 나무 벤치와 탁자가 배치되었으나 왼쪽에 철제 선반이 추가됨. 밝은 주간 채광.",
        "entities": "찰리의 기계 발(베이지색 장갑 및 피스톤 구조 일치), 낡은 코트 자락, 청화백자 파편.",
        "hard_violations": [],
        "physics": "발이 바닥을 단단히 딛고 파편을 미는 중이며, 쓸린 궤적이 물리적인 움직임을 명확히 뒷받침함."
       },
       {
        "label": "B",
        "direction": "기계 발이 파편들 사이에 정적으로 놓여 특정 방향으로의 움직임이 보이지 않음.",
        "built_space": "컨테이너 내부. 왼쪽에 배관이 있는 패널 벽이 배치되었으나, 오른쪽에 기준 이미지에 없는 거대한 철골 기둥이 세워져 있음. 나무 바닥.",
        "entities": "찰리의 기계 발(구조 일치), 다양한 색상의 화병 파편들. 요구된 코트는 보이지 않음.",
        "hard_violations": [],
        "physics": "발이 나무 바닥에 지지되어 있으나, 밀거나 미끄러지는 동적인 힘의 작용이나 흔적이 없음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "발끝 앞 파편의 미끄러짐으로 밀어내는 중간 동작이 분명하고 목재 벤치·식탁도 장소 연속성을 살리지만, 파편 무리는 지시보다 넓고 많다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "발밑 클로즈업과 장갑 발의 외형은 맞지만, 파편을 슬며시 밀기보다 밟고 있는 모습에 가깝고 배경의 장소 일치도도 낮다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "중앙의 장갑 발끝은 화면 왼쪽 아래의 파편 무리를 향한다. 일부 곡면 파편이 발가락 바로 아래에 있어 목표물과 접촉하지만, 옆으로 밀리는 방향이나 이동 흔적은 뚜렷하지 않다. 얼굴·시선·무기는 프레임에 없다.",
        "built_space": "바닥 가까이에서 비스듬히 내려다본 구도로, 중앙의 한쪽 발과 오른쪽 가장자리의 반대쪽 발 일부가 보인다. 배경에는 세로 패널 벽, 왼쪽 전기함류 두 개와 배관, 가는 수직 지지대 하나, 오른쪽의 굵은 철제 기둥 하나가 보인다. 이전 장면의 목재 벤치와 식탁은 보이지 않아 같은 생활 공간이라는 연결이 약하다. 이전 사진만으로 이 기둥들의 부재까지 확정할 수는 없다. 반사는 없다.",
        "entities": "보이는 신체는 찰리의 하퇴와 분절된 기계식 발이며, 긁히고 더러워진 샌드 베이지 장갑과 검은 관절은 인물 참조에 부합한다. 얼굴과 상체가 없어 눈빛·마스크·로고·모자는 평가 대상이 아니다. 발 앞에는 십여 개의 실제 도자기 같은 곡면 파편이 있으나, 청색·갈색·옅은 녹색 면이 섞여 한 화병의 조각이라는 통일성은 약하다. 다른 사람이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "중앙 발의 뒤쪽과 바깥쪽 발가락은 바닥에 닿고, 앞쪽 발가락은 일부 파편 위에 놓여 있다. 파편들은 바닥이나 발가락과 접촉해 지지되며 부유하지 않는다. 체중을 실어 파편을 누르는 자세는 가능하지만, 가볍게 밀어 숨기는 동작보다는 밟는 순간으로 읽힌다."
       },
       {
        "label": "B",
        "direction": "주동작 발의 발끝은 화면 왼쪽의 깨진 화병 무리를 향한다. 발끝 바로 앞의 큰 청백색 파편에 왼쪽으로 번진 움직임과 바닥 먼지가 보여, 발이 목표 파편을 왼쪽으로 밀어내는 관계가 읽힌다. 얼굴·시선·무기는 보이지 않는다.",
        "built_space": "발 옆의 낮은 사선 시점에서 바닥과 두 발을 잡았다. 배경에는 목재 벤치 하나, 목재 식탁 하나와 보이는 다리 세 개, 왼쪽 금속 선반 한 구획, 세로 패널 벽과 창 일부가 있다. 벤치는 벽 앞, 식탁은 그 앞에 놓여 이전 장면의 생활 공간과 재료·가구 관계가 더 잘 이어진다. 파편까지 이어지는 바닥 면도 끊기지 않는다. 다만 배경과 하퇴가 차지하는 면적이 커 발끝에만 집중한 구도보다는 조금 넓다.",
        "entities": "찰리의 두 장갑 발과 하퇴 일부가 보이며, 낡은 베이지 장갑판·검은 기계 관절·분절 발가락이 참조와 맞는다. 위쪽에는 낡은 외투 자락도 보인다. 프레임 밖의 얼굴·모자·가슴 표식은 평가하지 않는다. 청백색 무늬의 화병 목 부분과 같은 무늬의 파편들이 있어 깨진 화병이라는 정체가 명확하지만, 작은 산재 무리라는 요구에 비해 조각 수와 분포가 크다. 다른 사람과 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "오른쪽 아래의 반대 발이 바닥에 단단히 놓여 지지점이 되고, 주동작 발도 뒤쪽 접지부를 유지한 채 앞쪽으로 파편을 민다. 움직이는 파편은 바닥 높이에서 흐려져 보이며 발의 밀기라는 힘의 원인이 있다. 나머지 파편과 화병 목 부분도 바닥에 놓여 있어 지지 없는 부유나 불가능한 자세는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "발끝 앞 파편의 미끄러짐으로 밀어내는 중간 동작이 분명하고 목재 벤치·식탁도 장소 연속성을 살리지만, 파편 무리는 지시보다 넓고 많다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "발밑 클로즈업과 장갑 발의 외형은 맞지만, 파편을 슬며시 밀기보다 밟고 있는 모습에 가깝고 배경의 장소 일치도도 낮다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "중앙의 장갑 발끝은 화면 왼쪽 아래의 파편 무리를 향한다. 일부 곡면 파편이 발가락 바로 아래에 있어 목표물과 접촉하지만, 옆으로 밀리는 방향이나 이동 흔적은 뚜렷하지 않다. 얼굴·시선·무기는 프레임에 없다.",
        "built_space": "바닥 가까이에서 비스듬히 내려다본 구도로, 중앙의 한쪽 발과 오른쪽 가장자리의 반대쪽 발 일부가 보인다. 배경에는 세로 패널 벽, 왼쪽 전기함류 두 개와 배관, 가는 수직 지지대 하나, 오른쪽의 굵은 철제 기둥 하나가 보인다. 이전 장면의 목재 벤치와 식탁은 보이지 않아 같은 생활 공간이라는 연결이 약하다. 이전 사진만으로 이 기둥들의 부재까지 확정할 수는 없다. 반사는 없다.",
        "entities": "보이는 신체는 찰리의 하퇴와 분절된 기계식 발이며, 긁히고 더러워진 샌드 베이지 장갑과 검은 관절은 인물 참조에 부합한다. 얼굴과 상체가 없어 눈빛·마스크·로고·모자는 평가 대상이 아니다. 발 앞에는 십여 개의 실제 도자기 같은 곡면 파편이 있으나, 청색·갈색·옅은 녹색 면이 섞여 한 화병의 조각이라는 통일성은 약하다. 다른 사람이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "중앙 발의 뒤쪽과 바깥쪽 발가락은 바닥에 닿고, 앞쪽 발가락은 일부 파편 위에 놓여 있다. 파편들은 바닥이나 발가락과 접촉해 지지되며 부유하지 않는다. 체중을 실어 파편을 누르는 자세는 가능하지만, 가볍게 밀어 숨기는 동작보다는 밟는 순간으로 읽힌다."
       },
       {
        "label": "A",
        "direction": "주동작 발의 발끝은 화면 왼쪽의 깨진 화병 무리를 향한다. 발끝 바로 앞의 큰 청백색 파편에 왼쪽으로 번진 움직임과 바닥 먼지가 보여, 발이 목표 파편을 왼쪽으로 밀어내는 관계가 읽힌다. 얼굴·시선·무기는 보이지 않는다.",
        "built_space": "발 옆의 낮은 사선 시점에서 바닥과 두 발을 잡았다. 배경에는 목재 벤치 하나, 목재 식탁 하나와 보이는 다리 세 개, 왼쪽 금속 선반 한 구획, 세로 패널 벽과 창 일부가 있다. 벤치는 벽 앞, 식탁은 그 앞에 놓여 이전 장면의 생활 공간과 재료·가구 관계가 더 잘 이어진다. 파편까지 이어지는 바닥 면도 끊기지 않는다. 다만 배경과 하퇴가 차지하는 면적이 커 발끝에만 집중한 구도보다는 조금 넓다.",
        "entities": "찰리의 두 장갑 발과 하퇴 일부가 보이며, 낡은 베이지 장갑판·검은 기계 관절·분절 발가락이 참조와 맞는다. 위쪽에는 낡은 외투 자락도 보인다. 프레임 밖의 얼굴·모자·가슴 표식은 평가하지 않는다. 청백색 무늬의 화병 목 부분과 같은 무늬의 파편들이 있어 깨진 화병이라는 정체가 명확하지만, 작은 산재 무리라는 요구에 비해 조각 수와 분포가 크다. 다른 사람과 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "오른쪽 아래의 반대 발이 바닥에 단단히 놓여 지지점이 되고, 주동작 발도 뒤쪽 접지부를 유지한 채 앞쪽으로 파편을 민다. 움직이는 파편은 바닥 높이에서 흐려져 보이며 발의 밀기라는 힘의 원인이 있다. 나머지 파편과 화병 목 부분도 바닥에 놓여 있어 지지 없는 부유나 불가능한 자세는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.321
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.321
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1321
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "프롬프트가 강조한 발로 화병 파편을 슬며시 미는 동작과 바닥의 쓸림 궤적을 훌륭하게 연출했으며, 코트 자락의 디테일도 잘 반영되었습니다."
   },
   {
    "label": "B",
    "score": 1321,
    "verdict_ko": "파편을 미는 슬라이딩 액션이 생략된 채 정적으로 서 있으며, 배경에 기준 이미지에 존재하지 않는 굵은 철골 기둥이 추가되었습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S12sh12_sel.png",
    "asset_id": "88c6035a-eee8-485d-a089-be8000003130",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-de51-7eea-8edc-78821f20b30d",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S12sh12"
  }
 },
 "S14sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:00:24.398668+00:00",
  "fingerprint": "d1db0a5bc3f6356ad8cc89ce34167e29a947228009e48c5e823c5f637311de5b",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S14sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S14sh3_sel.png",
  "source_sha256": "7544968496052579b845549670361798fb492245f1e0b5dd56496b8bf39327b0",
  "file": "S14sh3_cine.png",
  "staged_sha256": "ffb50c9529997967e2bab584adad0fbcd30c502649477121884c5626d321dc91",
  "latency_ms": 10624
 },
 "S14sh5::signage": {
  "fp": "dd22c80933257b71",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S14sh5": {
  "input_fingerprint": "cc2390848340c63d",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 현우를 향해 양팔을 벌린 채 앞을 가로막고 서 있는 앰버의 굳은 전신.\n\nLOCATION (lock): In the shared living area inside the refugee family's container home, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Container interior (Interior of 현우's home) — The floor and interior boundaries remain visible around Amber's full figure; used as Establish the space she is blocking without adding foreground obstructions.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral ambient illumination with controlled contrast, preserving the tension in her face without introducing a new lighting cue.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Broken vase fragments remain on the container floor, partly pushed aside rather than removed. Charlie is an old, dirt-covered gorilla-shaped robot with blue-lit eyes and a worn Ubik chest logo, still wearing the old coat and hat used as a disguise. 앰버: She still wears her mask and waist tool pouch and has a recurring cough.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 현우를 향해 양팔을 벌린 채 앞을 가로막고 서 있는 앰버의 굳은 전신.\n\nLOCATION (lock): In the shared living area inside the refugee family's container home, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Container interior (Interior of 현우's home) — The floor and interior boundaries remain visible around Amber's full figure; used as Establish the space she is blocking without adding foreground obstructions.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral ambient illumination with controlled contrast, preserving the tension in her face without introducing a new lighting cue.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Broken vase fragments remain on the container floor, partly pushed aside rather than removed. Charlie is an old, dirt-covered gorilla-shaped robot with blue-lit eyes and a worn Ubik chest logo, still wearing the old coat and hat used as a disguise. 앰버: She still wears her mask and waist tool pouch and has a recurring cough.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 현우를 향해 양팔을 벌린 채 앞을 가로막고 서 있는 앰버의 굳은 전신.\n\nLOCATION (lock): In the shared living area inside the refugee family's container home, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Container interior (Interior of 현우's home) — The floor and interior boundaries remain visible around Amber's full figure; used as Establish the space she is blocking without adding foreground obstructions.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral ambient illumination with controlled contrast, preserving the tension in her face without introducing a new lighting cue.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Broken vase fragments remain on the container floor, partly pushed aside rather than removed. Charlie is an old, dirt-covered gorilla-shaped robot with blue-lit eyes and a worn Ubik chest logo, still wearing the old coat and hat used as a disguise. 앰버: She still wears her mask and waist tool pouch and has a recurring cough.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "B",
    "direction": "앰버는 화면 왼쪽 전경의 성인 쪽을 올려다보며 몸을 그쪽으로 향한다. 양팔은 좌우로, 어깨보다 조금 낮게 벌려 통행을 막는다. 시선의 대상은 분명하지만, 그 대상을 성인 신체로 화면에 넣는 것은 허용되지 않는다.",
    "built_space": "세로 골이 있는 컨테이너 벽, 콘크리트 바닥, 왼쪽 창 두 곳, 나무 벤치 하나와 식탁 하나, 뒤쪽 문 하나, 오른쪽 냉장고 하나와 금속 작업대·선반이 보인다. 앰버는 식탁 앞 빈 바닥에 서 있고 전신 주변의 바닥과 벽 경계가 보인다. 참조의 주요 재질과 가구는 이어지지만, 왼쪽의 큰 성인 신체가 공간을 가린다.",
    "entities": "금발의 어린 여자아이 한 명이 보이며 큰 눈과 머리 형태는 앰버 참조에 대체로 부합한다. 혼혈 배경 자체는 외모만으로 확정할 수 없다. 마스크는 코와 입을 덮고 허리에 도구가 든 주머니가 있다. 남색 반팔 대신 회색 긴팔과 회색 바지를 입었다. 청백색 꽃병 파편은 바닥에 남아 있다. 왼쪽에는 앰버가 아닌 성인의 몸통·손·다리가 추가되었다. 찰리는 보이지 않으며 이번 인물 구성상 추가할 필요가 없다.",
    "hard_violations": [
     "앰버만 허용된 화면에 성인의 몸통·손·다리를 추가했으며, 이 신체가 금지된 전경 가림 요소로 크게 들어왔다."
    ],
    "physics": "앰버는 벌린 두 발을 바닥에 붙여 몸을 지탱하고 있으며 양팔을 유지하는 자세도 가능하다. 공구 주머니는 허리띠에 매달려 있고 꽃병 파편과 가구는 바닥에 놓여 있다. 성인의 발은 프레임 밖이지만 몸은 아래로 자연스럽게 이어져 떠 있는 신체로 보이지는 않는다."
   },
   {
    "label": "A",
    "direction": "앰버는 정면의 카메라 가까운 화면 밖 지점을 응시한다. 현우 자체는 보이지 않으므로 대상을 직접 확인할 수 없지만, 화면 밖 현우를 막는 대면 구도로 성립한다. 양팔은 어깨 높이에서 좌우로 뻗고 손가락도 벌려 길을 막는 동작이 명확하다.",
    "built_space": "컨테이너 벽과 넓은 콘크리트 바닥, 왼쪽 창 두 곳, 나무 벤치 하나와 식탁 하나, 왼쪽 금속 선반 하나, 오른쪽 냉장고 하나와 작업대 하나, 오른쪽 끝 문 하나가 보인다. 앰버는 가구와 겹쳐 서지 않고 식탁 앞 통로를 차지한다. 전신과 그 주위의 바닥·실내 경계가 모두 보이며 전경 인물 가림이 없다. 참조의 벤치·식탁·선반과 낡은 실내 재질이 잘 이어진다.",
    "entities": "금발, 큰 눈, 둥근 얼굴의 어린 여자아이 한 명만 등장하며 참조의 얼굴·체격과 남색 반팔을 잘 따른다. 혼혈 배경은 외모만으로 단정할 수 없다. 허리 공구 주머니와 청백색 꽃병 파편이 있다. 마스크는 존재하지만 턱 아래로 내려가 코와 입이 드러나므로 얼굴 가림 유지 지시와 다르다. 기침 동작은 보이지 않으나 양팔을 벌리고 멈춘 순간과 모순되지는 않는다. 읽을 수 있는 문구는 뚜렷하지 않다.",
    "hard_violations": [],
    "physics": "두 신발이 바닥에 닿아 벌린 다리로 체중을 지탱한다. 수평으로 뻗은 팔과 펼친 손은 실제 사람이 취할 수 있는 차단 자세다. 주머니는 허리띠로, 내려간 마스크는 귀걸이 끈으로 지지된다. 꽃병 파편은 바닥에 놓여 있으며 지지 없이 떠 있는 물체나 신체는 없다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": null,
     "normalized": null,
     "ok": false
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "전신과 마스크·공구 주머니는 구현했지만, 허용되지 않은 성인 신체를 전경에 추가해 인물 제한과 전경 가림 없는 구도 지시를 위반했다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "앰버만의 굳은 전신과 양팔로 길을 막는 와이드 구도를 충실히 구현했으나, 마스크를 턱 아래로 내려 얼굴 가림 유지 조건을 어겼다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버는 화면 왼쪽 전경의 성인 쪽을 올려다보며 몸을 그쪽으로 향한다. 양팔은 좌우로, 어깨보다 조금 낮게 벌려 통행을 막는다. 시선의 대상은 분명하지만, 그 대상을 성인 신체로 화면에 넣는 것은 허용되지 않는다.",
        "built_space": "세로 골이 있는 컨테이너 벽, 콘크리트 바닥, 왼쪽 창 두 곳, 나무 벤치 하나와 식탁 하나, 뒤쪽 문 하나, 오른쪽 냉장고 하나와 금속 작업대·선반이 보인다. 앰버는 식탁 앞 빈 바닥에 서 있고 전신 주변의 바닥과 벽 경계가 보인다. 참조의 주요 재질과 가구는 이어지지만, 왼쪽의 큰 성인 신체가 공간을 가린다.",
        "entities": "금발의 어린 여자아이 한 명이 보이며 큰 눈과 머리 형태는 앰버 참조에 대체로 부합한다. 혼혈 배경 자체는 외모만으로 확정할 수 없다. 마스크는 코와 입을 덮고 허리에 도구가 든 주머니가 있다. 남색 반팔 대신 회색 긴팔과 회색 바지를 입었다. 청백색 꽃병 파편은 바닥에 남아 있다. 왼쪽에는 앰버가 아닌 성인의 몸통·손·다리가 추가되었다. 찰리는 보이지 않으며 이번 인물 구성상 추가할 필요가 없다.",
        "hard_violations": [
         "앰버만 허용된 화면에 성인의 몸통·손·다리를 추가했으며, 이 신체가 금지된 전경 가림 요소로 크게 들어왔다."
        ],
        "physics": "앰버는 벌린 두 발을 바닥에 붙여 몸을 지탱하고 있으며 양팔을 유지하는 자세도 가능하다. 공구 주머니는 허리띠에 매달려 있고 꽃병 파편과 가구는 바닥에 놓여 있다. 성인의 발은 프레임 밖이지만 몸은 아래로 자연스럽게 이어져 떠 있는 신체로 보이지는 않는다."
       },
       {
        "label": "B",
        "direction": "앰버는 정면의 카메라 가까운 화면 밖 지점을 응시한다. 현우 자체는 보이지 않으므로 대상을 직접 확인할 수 없지만, 화면 밖 현우를 막는 대면 구도로 성립한다. 양팔은 어깨 높이에서 좌우로 뻗고 손가락도 벌려 길을 막는 동작이 명확하다.",
        "built_space": "컨테이너 벽과 넓은 콘크리트 바닥, 왼쪽 창 두 곳, 나무 벤치 하나와 식탁 하나, 왼쪽 금속 선반 하나, 오른쪽 냉장고 하나와 작업대 하나, 오른쪽 끝 문 하나가 보인다. 앰버는 가구와 겹쳐 서지 않고 식탁 앞 통로를 차지한다. 전신과 그 주위의 바닥·실내 경계가 모두 보이며 전경 인물 가림이 없다. 참조의 벤치·식탁·선반과 낡은 실내 재질이 잘 이어진다.",
        "entities": "금발, 큰 눈, 둥근 얼굴의 어린 여자아이 한 명만 등장하며 참조의 얼굴·체격과 남색 반팔을 잘 따른다. 혼혈 배경은 외모만으로 단정할 수 없다. 허리 공구 주머니와 청백색 꽃병 파편이 있다. 마스크는 존재하지만 턱 아래로 내려가 코와 입이 드러나므로 얼굴 가림 유지 지시와 다르다. 기침 동작은 보이지 않으나 양팔을 벌리고 멈춘 순간과 모순되지는 않는다. 읽을 수 있는 문구는 뚜렷하지 않다.",
        "hard_violations": [],
        "physics": "두 신발이 바닥에 닿아 벌린 다리로 체중을 지탱한다. 수평으로 뻗은 팔과 펼친 손은 실제 사람이 취할 수 있는 차단 자세다. 주머니는 허리띠로, 내려간 마스크는 귀걸이 끈으로 지지된다. 꽃병 파편은 바닥에 놓여 있으며 지지 없이 떠 있는 물체나 신체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "전신과 마스크·공구 주머니는 구현했지만, 허용되지 않은 성인 신체를 전경에 추가해 인물 제한과 전경 가림 없는 구도 지시를 위반했다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "앰버만의 굳은 전신과 양팔로 길을 막는 와이드 구도를 충실히 구현했으나, 마스크를 턱 아래로 내려 얼굴 가림 유지 조건을 어겼다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "앰버는 화면 왼쪽 전경의 성인 쪽을 올려다보며 몸을 그쪽으로 향한다. 양팔은 좌우로, 어깨보다 조금 낮게 벌려 통행을 막는다. 시선의 대상은 분명하지만, 그 대상을 성인 신체로 화면에 넣는 것은 허용되지 않는다.",
        "built_space": "세로 골이 있는 컨테이너 벽, 콘크리트 바닥, 왼쪽 창 두 곳, 나무 벤치 하나와 식탁 하나, 뒤쪽 문 하나, 오른쪽 냉장고 하나와 금속 작업대·선반이 보인다. 앰버는 식탁 앞 빈 바닥에 서 있고 전신 주변의 바닥과 벽 경계가 보인다. 참조의 주요 재질과 가구는 이어지지만, 왼쪽의 큰 성인 신체가 공간을 가린다.",
        "entities": "금발의 어린 여자아이 한 명이 보이며 큰 눈과 머리 형태는 앰버 참조에 대체로 부합한다. 혼혈 배경 자체는 외모만으로 확정할 수 없다. 마스크는 코와 입을 덮고 허리에 도구가 든 주머니가 있다. 남색 반팔 대신 회색 긴팔과 회색 바지를 입었다. 청백색 꽃병 파편은 바닥에 남아 있다. 왼쪽에는 앰버가 아닌 성인의 몸통·손·다리가 추가되었다. 찰리는 보이지 않으며 이번 인물 구성상 추가할 필요가 없다.",
        "hard_violations": [
         "앰버만 허용된 화면에 성인의 몸통·손·다리를 추가했으며, 이 신체가 금지된 전경 가림 요소로 크게 들어왔다."
        ],
        "physics": "앰버는 벌린 두 발을 바닥에 붙여 몸을 지탱하고 있으며 양팔을 유지하는 자세도 가능하다. 공구 주머니는 허리띠에 매달려 있고 꽃병 파편과 가구는 바닥에 놓여 있다. 성인의 발은 프레임 밖이지만 몸은 아래로 자연스럽게 이어져 떠 있는 신체로 보이지는 않는다."
       },
       {
        "label": "A",
        "direction": "앰버는 정면의 카메라 가까운 화면 밖 지점을 응시한다. 현우 자체는 보이지 않으므로 대상을 직접 확인할 수 없지만, 화면 밖 현우를 막는 대면 구도로 성립한다. 양팔은 어깨 높이에서 좌우로 뻗고 손가락도 벌려 길을 막는 동작이 명확하다.",
        "built_space": "컨테이너 벽과 넓은 콘크리트 바닥, 왼쪽 창 두 곳, 나무 벤치 하나와 식탁 하나, 왼쪽 금속 선반 하나, 오른쪽 냉장고 하나와 작업대 하나, 오른쪽 끝 문 하나가 보인다. 앰버는 가구와 겹쳐 서지 않고 식탁 앞 통로를 차지한다. 전신과 그 주위의 바닥·실내 경계가 모두 보이며 전경 인물 가림이 없다. 참조의 벤치·식탁·선반과 낡은 실내 재질이 잘 이어진다.",
        "entities": "금발, 큰 눈, 둥근 얼굴의 어린 여자아이 한 명만 등장하며 참조의 얼굴·체격과 남색 반팔을 잘 따른다. 혼혈 배경은 외모만으로 단정할 수 없다. 허리 공구 주머니와 청백색 꽃병 파편이 있다. 마스크는 존재하지만 턱 아래로 내려가 코와 입이 드러나므로 얼굴 가림 유지 지시와 다르다. 기침 동작은 보이지 않으나 양팔을 벌리고 멈춘 순간과 모순되지는 않는다. 읽을 수 있는 문구는 뚜렷하지 않다.",
        "hard_violations": [],
        "physics": "두 신발이 바닥에 닿아 벌린 다리로 체중을 지탱한다. 수평으로 뻗은 팔과 펼친 손은 실제 사람이 취할 수 있는 차단 자세다. 주머니는 허리띠로, 내려간 마스크는 귀걸이 끈으로 지지된다. 꽃병 파편은 바닥에 놓여 있으며 지지 없이 떠 있는 물체나 신체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gemini-pro"
   ],
   "route": "single_reverse"
  },
  "totals": {
   "B": 2,
   "A": 8
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2,
    "verdict_ko": "전신과 마스크·공구 주머니는 구현했지만, 허용되지 않은 성인 신체를 전경에 추가해 인물 제한과 전경 가림 없는 구도 지시를 위반했다."
   },
   {
    "label": "A",
    "score": 8,
    "verdict_ko": "앰버만의 굳은 전신과 양팔로 길을 막는 와이드 구도를 충실히 구현했으나, 마스크를 턱 아래로 내려 얼굴 가림 유지 조건을 어겼다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S14sh3_sel.png",
    "asset_id": "47514c3d-b173-4169-a30a-8e2ecb55cedb",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-e01a-7693-9ce9-e5345eabfda0",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S14sh3"
  }
 },
 "S14sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:01:15.854459+00:00",
  "fingerprint": "b8c19b76c63518bc7205c9891e5aa3dc43c54fbdfee072051cce1d290197c387",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S14sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S14sh5_sel.png",
  "source_sha256": "86a38ef9e34cdacdd0e7ca378d802b6f1e2883f4a8aed93e0a6a468b2ddeab0c",
  "file": "S14sh5_cine.png",
  "staged_sha256": "c86dc85f10a2fb103f6d6572a02b83caea5ca7c4b47c601c93c543fda807cafe",
  "latency_ms": 9255
 },
 "S14sh9::signage": {
  "fp": "5602a8de271d82eb",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S14sh9": {
  "input_fingerprint": "441c2f92fd89e699",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 험상궂은 표정으로 찰리의 팔을 거칠게 움켜쥔 현우의 밀착된 상체.\n\nLOCATION (lock): Inside the container home's shared living area, at the spot where the robot has been handling household objects in daytime light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container interior (Same interior as the blocking exchange) — Only a narrow portion of the interior remains behind the two upper bodies; used as Maintain location continuity without competing with the grip.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the container's neutral ambient illumination and restrained contrast, keeping the face and gripping hand equally legible.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the same container interior, daylight, and household fixtures, including the broken vase fragments that remain on the floor. Exclude an intact replacement vase and any shards suspended in the air.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The broken vase fragments remain on the floor where they were pushed aside. Charlie retains his dirty, worn gorilla-shaped body, blue-lit eyes, worn Ubik chest logo, old coat and hat. 현우: He is now standing after rising from his seat, with his outer shirt on. His facial injuries and untreated dog-bite leg injury remain.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 험상궂은 표정으로 찰리의 팔을 거칠게 움켜쥔 현우의 밀착된 상체.\n\nLOCATION (lock): Inside the container home's shared living area, at the spot where the robot has been handling household objects in daytime light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container interior (Same interior as the blocking exchange) — Only a narrow portion of the interior remains behind the two upper bodies; used as Maintain location continuity without competing with the grip.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the container's neutral ambient illumination and restrained contrast, keeping the face and gripping hand equally legible.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the same container interior, daylight, and household fixtures, including the broken vase fragments that remain on the floor. Exclude an intact replacement vase and any shards suspended in the air.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The broken vase fragments remain on the floor where they were pushed aside. Charlie retains his dirty, worn gorilla-shaped body, blue-lit eyes, worn Ubik chest logo, old coat and hat. 현우: He is now standing after rising from his seat, with his outer shirt on. His facial injuries and untreated dog-bite leg injury remain.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 험상궂은 표정으로 찰리의 팔을 거칠게 움켜쥔 현우의 밀착된 상체.\n\nLOCATION (lock): Inside the container home's shared living area, at the spot where the robot has been handling household objects in daytime light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container interior (Same interior as the blocking exchange) — Only a narrow portion of the interior remains behind the two upper bodies; used as Maintain location continuity without competing with the grip.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the container's neutral ambient illumination and restrained contrast, keeping the face and gripping hand equally legible.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the same container interior, daylight, and household fixtures, including the broken vase fragments that remain on the floor. Exclude an intact replacement vase and any shards suspended in the air.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The broken vase fragments remain on the floor where they were pushed aside. Charlie retains his dirty, worn gorilla-shaped body, blue-lit eyes, worn Ubik chest logo, old coat and hat. 현우: He is now standing after rising from his seat, with his outer shirt on. His facial injuries and untreated dog-bite leg injury remain.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우와 찰리가 서로의 얼굴을 정면으로 마주보고 있음.",
    "built_space": "컨테이너 내부. 배경에 창문과 선반이 있으며, 바닥에 흐릿하게 깨진 도자기 파편이 흩어져 있음.",
    "entities": "현우(얼굴 일치, 얼굴 상처 있음, 겉옷 셔츠 착용), 찰리(로봇 외형 일치, 파란 눈, Ubik 로고 확인됨, 지정된 낡은 코트는 입었으나 모자는 없음).",
    "hard_violations": [
     "[gpt-high] 찰리의 가슴에 ‘UBIK’ 문자와 로고가 판독 가능하게 노출되어, 읽을 수 있는 글씨와 로고를 금지한 지시를 위반한다."
    ],
    "physics": "찰리의 기계 팔이 현우의 목과 턱 부근을 누르듯 뻗어 있고, 현우의 손은 찰리의 팔 위에 힘없이 가볍게 얹혀 있음. 지면 위에 서서 안정적인 자세를 유지함."
   },
   {
    "label": "B",
    "direction": "현우의 시선이 강한 적의를 띠고 찰리의 팔과 몸통 부근을 향해 내리꽂히듯 향해 있음.",
    "built_space": "컨테이너 내부. 좌측 뒤로 랙 선반이 보이고, 바닥에는 이전 샷에 등장했던 파란 무늬의 깨진 도자기 파편이 동일한 위치에 놓여 있음.",
    "entities": "현우(레퍼런스와 일치하는 얼굴 및 헤어스타일, 얼굴 상처 있음, 단 겉옷 없이 남색 티셔츠만 착용함), 찰리(기계 팔 및 몸체 일치, 지정된 코트와 모자는 미착용, 얼굴은 프레임 밖으로 크롭됨).",
    "hard_violations": [],
    "physics": "현우가 체중을 실어 찰리의 기계 팔 관절부를 손으로 거칠게 움켜쥐고 있으며, 근육의 긴장감과 맞닿은 힘의 균형이 물리적으로 자연스럽게 지탱됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "제시문의 핵심인 '험상궂은 표정'과 '거칠게 움켜쥔' 동작을 완벽하게 구현하고 이전 샷의 깨진 도자기 파편을 정확히 배치했으나, 두 인물의 지정된 겉옷(코트 등)이 누락되었습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "의상 설정은 비교적 잘 지켰으나, 샷 텍스트가 요구하는 '험상궂은 표정'과 '거칠게 움켜쥔' 강렬한 동작이 전혀 없이 차분하고 정적인 자세로 연출되어 우선순위가 높은 샷 연출 조건을 크게 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "현우의 시선이 강한 적의를 띠고 찰리의 팔과 몸통 부근을 향해 내리꽂히듯 향해 있음.",
        "built_space": "컨테이너 내부. 좌측 뒤로 랙 선반이 보이고, 바닥에는 이전 샷에 등장했던 파란 무늬의 깨진 도자기 파편이 동일한 위치에 놓여 있음.",
        "entities": "현우(레퍼런스와 일치하는 얼굴 및 헤어스타일, 얼굴 상처 있음, 단 겉옷 없이 남색 티셔츠만 착용함), 찰리(기계 팔 및 몸체 일치, 지정된 코트와 모자는 미착용, 얼굴은 프레임 밖으로 크롭됨).",
        "hard_violations": [],
        "physics": "현우가 체중을 실어 찰리의 기계 팔 관절부를 손으로 거칠게 움켜쥐고 있으며, 근육의 긴장감과 맞닿은 힘의 균형이 물리적으로 자연스럽게 지탱됨."
       },
       {
        "label": "A",
        "direction": "현우와 찰리가 서로의 얼굴을 정면으로 마주보고 있음.",
        "built_space": "컨테이너 내부. 배경에 창문과 선반이 있으며, 바닥에 흐릿하게 깨진 도자기 파편이 흩어져 있음.",
        "entities": "현우(얼굴 일치, 얼굴 상처 있음, 겉옷 셔츠 착용), 찰리(로봇 외형 일치, 파란 눈, Ubik 로고 확인됨, 지정된 낡은 코트는 입었으나 모자는 없음).",
        "hard_violations": [],
        "physics": "찰리의 기계 팔이 현우의 목과 턱 부근을 누르듯 뻗어 있고, 현우의 손은 찰리의 팔 위에 힘없이 가볍게 얹혀 있음. 지면 위에 서서 안정적인 자세를 유지함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "제시문의 핵심인 '험상궂은 표정'과 '거칠게 움켜쥔' 동작을 완벽하게 구현하고 이전 샷의 깨진 도자기 파편을 정확히 배치했으나, 두 인물의 지정된 겉옷(코트 등)이 누락되었습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "의상 설정은 비교적 잘 지켰으나, 샷 텍스트가 요구하는 '험상궂은 표정'과 '거칠게 움켜쥔' 강렬한 동작이 전혀 없이 차분하고 정적인 자세로 연출되어 우선순위가 높은 샷 연출 조건을 크게 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 시선이 강한 적의를 띠고 찰리의 팔과 몸통 부근을 향해 내리꽂히듯 향해 있음.",
        "built_space": "컨테이너 내부. 좌측 뒤로 랙 선반이 보이고, 바닥에는 이전 샷에 등장했던 파란 무늬의 깨진 도자기 파편이 동일한 위치에 놓여 있음.",
        "entities": "현우(레퍼런스와 일치하는 얼굴 및 헤어스타일, 얼굴 상처 있음, 단 겉옷 없이 남색 티셔츠만 착용함), 찰리(기계 팔 및 몸체 일치, 지정된 코트와 모자는 미착용, 얼굴은 프레임 밖으로 크롭됨).",
        "hard_violations": [],
        "physics": "현우가 체중을 실어 찰리의 기계 팔 관절부를 손으로 거칠게 움켜쥐고 있으며, 근육의 긴장감과 맞닿은 힘의 균형이 물리적으로 자연스럽게 지탱됨."
       },
       {
        "label": "A",
        "direction": "현우와 찰리가 서로의 얼굴을 정면으로 마주보고 있음.",
        "built_space": "컨테이너 내부. 배경에 창문과 선반이 있으며, 바닥에 흐릿하게 깨진 도자기 파편이 흩어져 있음.",
        "entities": "현우(얼굴 일치, 얼굴 상처 있음, 겉옷 셔츠 착용), 찰리(로봇 외형 일치, 파란 눈, Ubik 로고 확인됨, 지정된 낡은 코트는 입었으나 모자는 없음).",
        "hard_violations": [],
        "physics": "찰리의 기계 팔이 현우의 목과 턱 부근을 누르듯 뻗어 있고, 현우의 손은 찰리의 팔 위에 힘없이 가볍게 얹혀 있음. 지면 위에 서서 안정적인 자세를 유지함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "험상궂은 현우의 얼굴과 찰리의 팔을 거칠게 움켜쥔 손이 밀착 구도로 잘 드러나지만, 현우의 겉셔츠와 찰리의 낡은 코트가 누락되고 배경이 다소 넓다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "겉셔츠·코트와 파란 눈은 반영했지만, 가슴의 읽을 수 있는 문자와 로고가 명시적 금지사항을 위반하고 현우의 험상궂은 표정도 약하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 몸을 찰리 쪽으로 숙이고 오른쪽 위의 화면 밖 찰리 얼굴 방향을 노려본다. 현우의 손은 찰리의 팔꿈치에 가까운 팔 장갑과 관절을 감싸고 있다. 찰리의 팔은 현우 가슴 앞을 가로질러 왼쪽으로 굽혀져 있으며, 얼굴은 잘려 시선은 확인할 수 없다.",
        "built_space": "왼쪽에 창 하나와 검은 금속 선반 하나, 뒤쪽에 목재 가구 일부가 보이며 밝은 패널 벽과 콘크리트 바닥이 이전 장소와 이어진다. 바닥에는 청백색 화병 파편이 놓여 있다. 두 상체는 가까이 붙어 있지만 왼쪽 선반과 바닥까지 비교적 넓게 보여, 배경을 좁게 남기라는 요구에는 덜 밀착된다. 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "현우는 앳된 동아시아계 남성으로 보이고 헝클어진 검은 머리와 얼굴 형태가 참조와 가깝다. 뺨 상처와 찡그린 눈썹, 굳은 입매가 보인다. 남색 상의는 있으나 요구된 겉셔츠는 없다. 찰리의 육중한 팔과 샌드 베이지 장갑판, 관절과 표면 마모는 참조에 부합하지만 드러난 몸통과 어깨에 코트가 없다. 얼굴·눈·모자는 화면 밖이므로 평가하지 않는다. 추가 인물, 온전한 대체 화병, 읽을 수 있는 문자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "현우의 손바닥과 굽힌 손가락이 찰리의 팔에 실제로 접촉하여 움켜쥐는 동작을 만든다. 찰리의 수평 팔은 몸통에 연결된 어깨와 굽힌 팔꿈치 관절로 지지된다. 현우는 상체를 앞으로 숙인 상태이며 하체가 잘려 발의 접지는 확인되지 않지만, 몸이 공중에 떠 있다는 증거는 없다. 화병 파편은 바닥에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "현우는 오른쪽 찰리의 얼굴을 바라보고, 찰리도 얼굴을 왼쪽 아래 현우 쪽으로 돌리고 있다. 현우의 손은 가슴 높이로 가로놓인 찰리의 전완 장갑을 잡는다. 찰리의 주먹은 현우의 어깨와 목 아래 가까이에 있어, 현우가 먼저 거칠게 붙드는 순간보다는 찰리의 뻗은 팔을 붙잡아 버티는 인상도 난다.",
        "built_space": "뒤에는 창 두 개, 왼쪽 금속 선반 하나와 보온용기, 중앙의 흐릿한 목재 가구 일부가 보인다. 밝은 패널 벽과 바닥 재질은 이전 컨테이너 실내와 대체로 일치하며 파편도 바닥에 남아 있다. 두 인물의 상체가 화면을 차지하지만 중앙 창과 바닥까지 노출되어 배경은 요구보다 다소 넓다. 중복 설비나 불가능한 반사는 확인되지 않는다.",
        "entities": "현우는 검은 머리의 젊은 동아시아계 남성으로 참조와 대체로 일치하고 얼굴 상처, 남색 속옷과 낡은 겉셔츠가 보인다. 다만 표정은 험상궂다기보다 긴장한 옆얼굴에 가깝다. 찰리는 흰 마스크형 얼굴, 파란 눈, 마모된 베이지 장갑과 낡은 코트를 갖췄다. 머리 위는 잘려 모자 유무를 판단하지 않는다. 가슴에는 읽을 수 있는 ‘UBIK’ 문자와 선명한 로고가 드러난다. 추가 인물이나 온전한 화병은 없다.",
        "hard_violations": [
         "찰리의 가슴에 ‘UBIK’ 문자와 로고가 판독 가능하게 노출되어, 읽을 수 있는 글씨와 로고를 금지한 지시를 위반한다."
        ],
        "physics": "현우의 손가락과 엄지가 찰리의 전완 장갑을 감싸 실제 접촉을 이룬다. 찰리의 팔은 어깨와 팔꿈치의 기계 관절에 연결되어 수평 자세를 유지하며, 주먹은 현우의 어깨 가까이에 놓여 있다. 두 상체의 자세는 물리적으로 가능하고, 하체가 잘렸다는 이유만으로 부유했다고 볼 근거는 없다. 코트는 어깨에서 자연스럽게 늘어지고 파편은 바닥이 받친다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "험상궂은 현우의 얼굴과 찰리의 팔을 거칠게 움켜쥔 손이 밀착 구도로 잘 드러나지만, 현우의 겉셔츠와 찰리의 낡은 코트가 누락되고 배경이 다소 넓다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "겉셔츠·코트와 파란 눈은 반영했지만, 가슴의 읽을 수 있는 문자와 로고가 명시적 금지사항을 위반하고 현우의 험상궂은 표정도 약하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 몸을 찰리 쪽으로 숙이고 오른쪽 위의 화면 밖 찰리 얼굴 방향을 노려본다. 현우의 손은 찰리의 팔꿈치에 가까운 팔 장갑과 관절을 감싸고 있다. 찰리의 팔은 현우 가슴 앞을 가로질러 왼쪽으로 굽혀져 있으며, 얼굴은 잘려 시선은 확인할 수 없다.",
        "built_space": "왼쪽에 창 하나와 검은 금속 선반 하나, 뒤쪽에 목재 가구 일부가 보이며 밝은 패널 벽과 콘크리트 바닥이 이전 장소와 이어진다. 바닥에는 청백색 화병 파편이 놓여 있다. 두 상체는 가까이 붙어 있지만 왼쪽 선반과 바닥까지 비교적 넓게 보여, 배경을 좁게 남기라는 요구에는 덜 밀착된다. 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "현우는 앳된 동아시아계 남성으로 보이고 헝클어진 검은 머리와 얼굴 형태가 참조와 가깝다. 뺨 상처와 찡그린 눈썹, 굳은 입매가 보인다. 남색 상의는 있으나 요구된 겉셔츠는 없다. 찰리의 육중한 팔과 샌드 베이지 장갑판, 관절과 표면 마모는 참조에 부합하지만 드러난 몸통과 어깨에 코트가 없다. 얼굴·눈·모자는 화면 밖이므로 평가하지 않는다. 추가 인물, 온전한 대체 화병, 읽을 수 있는 문자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "현우의 손바닥과 굽힌 손가락이 찰리의 팔에 실제로 접촉하여 움켜쥐는 동작을 만든다. 찰리의 수평 팔은 몸통에 연결된 어깨와 굽힌 팔꿈치 관절로 지지된다. 현우는 상체를 앞으로 숙인 상태이며 하체가 잘려 발의 접지는 확인되지 않지만, 몸이 공중에 떠 있다는 증거는 없다. 화병 파편은 바닥에 놓여 있다."
       },
       {
        "label": "A",
        "direction": "현우는 오른쪽 찰리의 얼굴을 바라보고, 찰리도 얼굴을 왼쪽 아래 현우 쪽으로 돌리고 있다. 현우의 손은 가슴 높이로 가로놓인 찰리의 전완 장갑을 잡는다. 찰리의 주먹은 현우의 어깨와 목 아래 가까이에 있어, 현우가 먼저 거칠게 붙드는 순간보다는 찰리의 뻗은 팔을 붙잡아 버티는 인상도 난다.",
        "built_space": "뒤에는 창 두 개, 왼쪽 금속 선반 하나와 보온용기, 중앙의 흐릿한 목재 가구 일부가 보인다. 밝은 패널 벽과 바닥 재질은 이전 컨테이너 실내와 대체로 일치하며 파편도 바닥에 남아 있다. 두 인물의 상체가 화면을 차지하지만 중앙 창과 바닥까지 노출되어 배경은 요구보다 다소 넓다. 중복 설비나 불가능한 반사는 확인되지 않는다.",
        "entities": "현우는 검은 머리의 젊은 동아시아계 남성으로 참조와 대체로 일치하고 얼굴 상처, 남색 속옷과 낡은 겉셔츠가 보인다. 다만 표정은 험상궂다기보다 긴장한 옆얼굴에 가깝다. 찰리는 흰 마스크형 얼굴, 파란 눈, 마모된 베이지 장갑과 낡은 코트를 갖췄다. 머리 위는 잘려 모자 유무를 판단하지 않는다. 가슴에는 읽을 수 있는 ‘UBIK’ 문자와 선명한 로고가 드러난다. 추가 인물이나 온전한 화병은 없다.",
        "hard_violations": [
         "찰리의 가슴에 ‘UBIK’ 문자와 로고가 판독 가능하게 노출되어, 읽을 수 있는 글씨와 로고를 금지한 지시를 위반한다."
        ],
        "physics": "현우의 손가락과 엄지가 찰리의 전완 장갑을 감싸 실제 접촉을 이룬다. 찰리의 팔은 어깨와 팔꿈치의 기계 관절에 연결되어 수평 자세를 유지하며, 주먹은 현우의 어깨 가까이에 놓여 있다. 두 상체의 자세는 물리적으로 가능하고, 하체가 잘렸다는 이유만으로 부유했다고 볼 근거는 없다. 코트는 어깨에서 자연스럽게 늘어지고 파편은 바닥이 받친다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.929,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.679,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gpt-high] 찰리의 가슴에 ‘UBIK’ 문자와 로고가 판독 가능하게 노출되어, 읽을 수 있는 글씨와 로고를 금지한 지시를 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 679
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "제시문의 핵심인 '험상궂은 표정'과 '거칠게 움켜쥔' 동작을 완벽하게 구현하고 이전 샷의 깨진 도자기 파편을 정확히 배치했으나, 두 인물의 지정된 겉옷(코트 등)이 누락되었습니다."
   },
   {
    "label": "A",
    "score": 679,
    "verdict_ko": "의상 설정은 비교적 잘 지켰으나, 샷 텍스트가 요구하는 '험상궂은 표정'과 '거칠게 움켜쥔' 강렬한 동작이 전혀 없이 차분하고 정적인 자세로 연출되어 우선순위가 높은 샷 연출 조건을 크게 위반했습니다.  ★위반: [gpt-high] 찰리의 가슴에 ‘UBIK’ 문자와 로고가 판독 가능하게 노출되어, 읽을 수 있는 글씨와 로고를 금지한 지시를 위반한다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S14sh5_sel.png",
    "asset_id": "82b50d5c-6da4-46d0-83ed-f250642db0d6",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-e1dc-7022-a338-5a044be92f0f",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S14sh5"
  }
 },
 "S14sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:02:25.413723+00:00",
  "fingerprint": "3464d17621a4f75c19f9a8848619b8af31be1839e146fc30185b62314d83249d",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S14sh9_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S14sh9_sel.png",
  "source_sha256": "077ec0fe7bfba9eb77de32d6c206f8bf348a05b7cc12514cee3befc4f518c92f",
  "file": "S14sh9_cine.png",
  "staged_sha256": "f2ecd5a4d47800be3e897a52419d7d312b09cc51005e671aac202751dca1b7c9",
  "latency_ms": 11127
 },
 "S15sh4::signage": {
  "fp": "507064929f8094a7",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::c685222b67bf4235": {
  "subjects": [],
  "subject_text": "마츠다의 고물상 내부\n온갖 낡은 가전제품과 고물이 빽빽하게 쌓인 상점. 전기제품 수리용 책상과 텔레비전, 컴퓨터가 비좁은 실내에 놓여 있다.",
  "identity": "canonical",
  "scope_id": "L168",
  "scope_role": "location_interior",
  "scope_sha": "49a46d2906d2c018"
 },
 "S15sh4::bgfirst_bg": {
  "input_fingerprint": "0c214deaf052740c",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리의 가슴 장갑 표면에 남은 로고를 손가락으로 가리키는 마츠다의 팔.\n\nLOCATION (lock): In the customer area inside a refugee settlement junk shop, amid piled appliances and scrap in subdued daytime light.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Worn 유빅 chest logo (Partially worn away but still identifiable) — The marked outward face of Charlie's chest is visible obliquely, with the surviving 유빅 lettering beside Matsuda's fingertip; used as Provide the precise evidence being identified, as a small detail within the chest rather than an enlarged isolated graphic; Piled appliances and scrap (Accumulated throughout the shop) — Partial outlines remain around the inspected upper body; used as Preserve the shop context and realistic scale behind the close inspection.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral shop illumination with gentle tonal separation so the worn lettering and pointing finger remain readable without a fabricated spotlight.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리의 가슴 장갑 표면에 남은 로고를 손가락으로 가리키는 마츠다의 팔.\n\nLOCATION (lock): In the customer area inside a refugee settlement junk shop, amid piled appliances and scrap in subdued daytime light.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Worn 유빅 chest logo (Partially worn away but still identifiable) — The marked outward face of Charlie's chest is visible obliquely, with the surviving 유빅 lettering beside Matsuda's fingertip; used as Provide the precise evidence being identified, as a small detail within the chest rather than an enlarged isolated graphic; Piled appliances and scrap (Accumulated throughout the shop) — Partial outlines remain around the inspected upper body; used as Preserve the shop context and realistic scale behind the close inspection.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral shop illumination with gentle tonal separation so the worn lettering and pointing finger remain readable without a fabricated spotlight.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S15sh4__bgfirst_bg.png",
  "asset_id": "08f633d6-c3f3-4ebc-a4c1-0fb688cbc402",
  "input_asset_ids": [
   "8a256e3b-52f6-4af6-aa77-ef9e992d170a",
   "cb825975-b374-47cb-9da4-9a5bcc3765b2"
  ]
 },
 "S15sh4": {
  "input_fingerprint": "aa83d7e77644bd4d",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 가슴 장갑 표면에 남은 로고를 손가락으로 가리키는 마츠다의 팔.\n\nLOCATION (lock): In the customer area inside a refugee settlement junk shop, amid piled appliances and scrap in subdued daytime light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Worn 유빅 chest logo (Partially worn away but still identifiable) — The marked outward face of Charlie's chest is visible obliquely, with the surviving 유빅 lettering beside Matsuda's fingertip; used as Provide the precise evidence being identified, as a small detail within the chest rather than an enlarged isolated graphic; Piled appliances and scrap (Accumulated throughout the shop) — Partial outlines remain around the inspected upper body; used as Preserve the shop context and realistic scale behind the close inspection.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral shop illumination with gentle tonal separation so the worn lettering and pointing finger remain readable without a fabricated spotlight.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Appliances and scrap are piled throughout the shop. Charlie remains dirty and worn, with blue-lit eyes, an old coat and hat, and a visibly worn Ubik logo on his chest. 마츠다: He has moved away from his repair work and is standing in the shop's inspection area with an arm extended.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 마츠다 (일본인 남성, 60대, 나이 든 얼굴, 눈가 잔주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 가슴 장갑 표면에 남은 로고를 손가락으로 가리키는 마츠다의 팔.\n\nLOCATION (lock): In the customer area inside a refugee settlement junk shop, amid piled appliances and scrap in subdued daytime light. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Worn 유빅 chest logo (Partially worn away but still identifiable) — The marked outward face of Charlie's chest is visible obliquely, with the surviving 유빅 lettering beside Matsuda's fingertip; used as Provide the precise evidence being identified, as a small detail within the chest rather than an enlarged isolated graphic; Piled appliances and scrap (Accumulated throughout the shop) — Partial outlines remain around the inspected upper body; used as Preserve the shop context and realistic scale behind the close inspection.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral shop illumination with gentle tonal separation so the worn lettering and pointing finger remain readable without a fabricated spotlight.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Appliances and scrap are piled throughout the shop. Charlie remains dirty and worn, with blue-lit eyes, an old coat and hat, and a visibly worn Ubik logo on his chest. 마츠다: He has moved away from his repair work and is standing in the shop's inspection area with an arm extended.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 마츠다 (일본인 남성, 60대, 나이 든 얼굴, 눈가 잔주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 가슴 장갑 표면에 남은 로고를 손가락으로 가리키는 마츠다의 팔.\n\nLOCATION (lock): In the customer area inside a refugee settlement junk shop, amid piled appliances and scrap in subdued daytime light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Worn 유빅 chest logo (Partially worn away but still identifiable) — The marked outward face of Charlie's chest is visible obliquely, with the surviving 유빅 lettering beside Matsuda's fingertip; used as Provide the precise evidence being identified, as a small detail within the chest rather than an enlarged isolated graphic; Piled appliances and scrap (Accumulated throughout the shop) — Partial outlines remain around the inspected upper body; used as Preserve the shop context and realistic scale behind the close inspection.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral shop illumination with gentle tonal separation so the worn lettering and pointing finger remain readable without a fabricated spotlight.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Appliances and scrap are piled throughout the shop. Charlie remains dirty and worn, with blue-lit eyes, an old coat and hat, and a visibly worn Ubik logo on his chest. 마츠다: He has moved away from his repair work and is standing in the shop's inspection area with an arm extended.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 마츠다 (일본인 남성, 60대, 나이 든 얼굴, 눈가 잔주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S15sh4__bgfirst_bg.png",
     "asset_id": "08f633d6-c3f3-4ebc-a4c1-0fb688cbc402",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S15sh4.png",
     "asset_id": "8a256e3b-52f6-4af6-aa77-ef9e992d170a",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 마츠다: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1426352>",
     "asset_id": "36f5f867-079a-4d08-baa1-280e22734416",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L168B03.png",
     "asset_id": "cb825975-b374-47cb-9da4-9a5bcc3765b2",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 마츠다: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1426352>",
     "asset_id": "36f5f867-079a-4d08-baa1-280e22734416",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "마츠다의 손가락이 찰리의 가슴에 있는 둥근 원형 심볼을 향해 가리키고 있음.",
    "built_space": "배경의 선반, 장비, 고철 가전제품들이 레퍼런스 사진의 중고품 상점 환경을 적절히 구성하고 있음.",
    "entities": "마츠다는 노년의 아시아계 남성으로 보이나, 찰리는 모자와 코트가 없고 장갑판이 회색조를 띠며 요구된 글자 대신 임의의 심볼이 그려져 있음.",
    "hard_violations": [],
    "physics": "인물들의 자세가 중력에 맞게 자연스러우며, 마츠다의 내민 팔 역시 어깨를 통해 정상적으로 지지됨."
   },
   {
    "label": "B",
    "direction": "마츠다의 손가락이 찰리의 가슴 장갑에 적힌 텍스트('유빅'과 유사한 형태)를 정확히 향해 가리키고 있음.",
    "built_space": "중고품 상점의 선반, 낡은 가전제품, 고철 더미가 레퍼런스 이미지와 일치하는 공간적 스케일로 배경에 배치되어 있음.",
    "entities": "찰리는 샌드 베이지색 장갑, 파란 눈, 낡은 코트와 모자를 착용하고 있으며 가슴에 한글 로고가 있음. 마츠다는 연륜 있는 아시아계 남성의 외형을 갖춤.",
    "hard_violations": [
     "[gpt-high] 가슴의 ‘유빅’ 문자가 명확히 판독되어, 이미지 어디에도 읽을 수 있는 문자를 두지 말라는 최종 지시를 위반한다."
    ],
    "physics": "두 인물 모두 지면을 딛고 안정적으로 서 있으며, 마츠다의 팔과 가리키는 동작이 자연스럽게 지지되고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지시된 '유빅' 글자와 찰리의 의상(모자, 낡은 코트), 장갑 색상 및 파란 눈 등 프롬프트의 세부 요구사항을 매우 충실히 구현했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "찰리의 장갑 색상이 다르고 지정된 모자와 코트가 누락되었으며, 가슴의 '유빅' 글자 대신 임의의 원형 마크가 생성되어 프롬프트 반영도가 낮습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "마츠다의 손가락이 찰리의 가슴 장갑에 적힌 텍스트('유빅'과 유사한 형태)를 정확히 향해 가리키고 있음.",
        "built_space": "중고품 상점의 선반, 낡은 가전제품, 고철 더미가 레퍼런스 이미지와 일치하는 공간적 스케일로 배경에 배치되어 있음.",
        "entities": "찰리는 샌드 베이지색 장갑, 파란 눈, 낡은 코트와 모자를 착용하고 있으며 가슴에 한글 로고가 있음. 마츠다는 연륜 있는 아시아계 남성의 외형을 갖춤.",
        "hard_violations": [],
        "physics": "두 인물 모두 지면을 딛고 안정적으로 서 있으며, 마츠다의 팔과 가리키는 동작이 자연스럽게 지지되고 있음."
       },
       {
        "label": "A",
        "direction": "마츠다의 손가락이 찰리의 가슴에 있는 둥근 원형 심볼을 향해 가리키고 있음.",
        "built_space": "배경의 선반, 장비, 고철 가전제품들이 레퍼런스 사진의 중고품 상점 환경을 적절히 구성하고 있음.",
        "entities": "마츠다는 노년의 아시아계 남성으로 보이나, 찰리는 모자와 코트가 없고 장갑판이 회색조를 띠며 요구된 글자 대신 임의의 심볼이 그려져 있음.",
        "hard_violations": [],
        "physics": "인물들의 자세가 중력에 맞게 자연스러우며, 마츠다의 내민 팔 역시 어깨를 통해 정상적으로 지지됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지시된 '유빅' 글자와 찰리의 의상(모자, 낡은 코트), 장갑 색상 및 파란 눈 등 프롬프트의 세부 요구사항을 매우 충실히 구현했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "찰리의 장갑 색상이 다르고 지정된 모자와 코트가 누락되었으며, 가슴의 '유빅' 글자 대신 임의의 원형 마크가 생성되어 프롬프트 반영도가 낮습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "마츠다의 손가락이 찰리의 가슴 장갑에 적힌 텍스트('유빅'과 유사한 형태)를 정확히 향해 가리키고 있음.",
        "built_space": "중고품 상점의 선반, 낡은 가전제품, 고철 더미가 레퍼런스 이미지와 일치하는 공간적 스케일로 배경에 배치되어 있음.",
        "entities": "찰리는 샌드 베이지색 장갑, 파란 눈, 낡은 코트와 모자를 착용하고 있으며 가슴에 한글 로고가 있음. 마츠다는 연륜 있는 아시아계 남성의 외형을 갖춤.",
        "hard_violations": [],
        "physics": "두 인물 모두 지면을 딛고 안정적으로 서 있으며, 마츠다의 팔과 가리키는 동작이 자연스럽게 지지되고 있음."
       },
       {
        "label": "A",
        "direction": "마츠다의 손가락이 찰리의 가슴에 있는 둥근 원형 심볼을 향해 가리키고 있음.",
        "built_space": "배경의 선반, 장비, 고철 가전제품들이 레퍼런스 사진의 중고품 상점 환경을 적절히 구성하고 있음.",
        "entities": "마츠다는 노년의 아시아계 남성으로 보이나, 찰리는 모자와 코트가 없고 장갑판이 회색조를 띠며 요구된 글자 대신 임의의 심볼이 그려져 있음.",
        "hard_violations": [],
        "physics": "인물들의 자세가 중력에 맞게 자연스러우며, 마츠다의 내민 팔 역시 어깨를 통해 정상적으로 지지됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "손끝의 지시 대상과 장소·의상은 잘 맞지만, 가슴과 팔의 클로즈업보다 넓게 잡았고 ‘유빅’이 명확히 읽혀 최종 문자 금지 조건을 위반한다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "비스듬한 가슴 장갑과 이를 가리키는 팔의 근접 구도가 더 충실하지만, 원형 마크는 유빅으로 식별하기 어렵고 장갑 색상·설계와 마츠다의 의상이 참조와 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "오른쪽에서 뻗은 마츠다의 검지는 왼쪽의 찰리 가슴을 향하며, 손끝이 ‘유빅’ 글자의 오른쪽 끝에 닿는다. 찰리의 얼굴은 마츠다 쪽으로 약간 돌아가 아래를 향한다. 마츠다는 가슴 쪽을 향하지만 눈이 프레임 밖이어서 정확한 시선은 확인할 수 없다.",
        "built_space": "왼쪽에 전면 세탁기 한 대, 그 위의 소형 가전, 뒤쪽의 높은 냉장고 한 대와 소형 브라운관 모니터 한 대가 보인다. 오른쪽에는 원형 팬 장치 한 대와 그 아래 브라운관 모니터 한 대, 양옆 금속 선반 및 뒤쪽 공구 벽이 있다. 참조의 가전 배치와 재질을 밀접하게 따른다. 찰리는 중앙 통로에, 마츠다는 오른쪽 가장자리에 있어 집기와 충돌하지 않는다. 다만 머리부터 골반까지 보여 요구한 근접 구도보다 넓다.",
        "entities": "찰리의 흰 각진 마스크, 모래색 장갑, 파란 눈, 낡은 모자와 때 묻은 코트가 보인다. 노출된 장갑은 참조의 형태와 대체로 연결된다. 마츠다는 회색 머리의 노년 남성 일부와 주름진 손으로 나타나며, 남색 정장과 흰 소매도 참조에 부합한다. 가슴의 ‘유빅’은 작은 표면 인쇄로 표현되었으나 글자가 거의 온전히 읽힌다. 별도의 인물은 없다.",
        "hard_violations": [
         "가슴의 ‘유빅’ 문자가 명확히 판독되어, 이미지 어디에도 읽을 수 있는 문자를 두지 말라는 최종 지시를 위반한다."
        ],
        "physics": "마츠다의 손은 손목과 소매로 이어지고 팔이 화면 오른쪽 몸통에 연결되어 있다. 검지로 가슴 표면을 짚는 동작이 가능하며 손끝 접촉도 자연스럽다. 코트는 찰리의 어깨에 걸리고 장갑은 몸통 구조에 연결된다. 발은 잘렸지만 공중에 뜬 자세나 지지 없는 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "왼쪽 마츠다의 팔이 오른쪽 위로 뻗고, 검지 끝은 찰리 가슴의 마모된 원형 마크 왼쪽 위에 닿는다. 지시 대상은 분명히 가슴 표면의 흔적이다. 마츠다의 고개는 손끝과 마크 쪽으로 숙여져 있고 찰리의 얼굴도 아래쪽을 향하지만, 두 인물의 정확한 눈동자 방향은 확인하기 어렵다.",
        "built_space": "왼쪽 배경에 높은 냉장고 한 대와 브라운관 모니터 한 대, 그 아래 쌓인 사각 가전들이 보인다. 뒤에는 작업대 하나와 공구 벽, 상단 선반이 있으며 오른쪽 끝에도 선반 일부가 남는다. 참조 장소의 재질과 주요 배치가 유지된다. 마츠다는 왼쪽 전경, 찰리는 오른쪽 검사 공간을 차지한다. 비스듬한 가슴과 팔을 크게 잡고 주변 가전은 부분 윤곽만 남겨 클로즈업 요구에 더 가깝다.",
        "entities": "찰리는 육중한 기계 몸통, 파란 눈 일부, 흰 마스크와 낡은 천 겉옷으로 표현된다. 장갑은 참조의 모래색보다 회백색이며 가슴 판과 어깨 구조도 다르다. 모자와 하체는 프레임 밖이라 평가할 수 없다. 마츠다는 회색 머리, 나이 든 옆얼굴과 주름진 손을 가진 남성이지만, 검은 작업복은 참조의 정장과 다르다. 가슴의 마모된 원형 도안에는 읽기 어려운 흔적만 있어 ‘유빅’이라는 식별성은 부족하다. 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "마츠다의 손과 팔은 왼쪽 어깨 및 몸통에 연속적으로 연결되고, 검지는 가슴 장갑에 자연스럽게 닿는다. 찰리의 장갑과 관절은 몸통에 결합되어 있으며 겉옷은 어깨에 걸쳐 아래로 처진다. 마크는 긁히고 벗겨진 도료처럼 표면에 붙어 있다. 하체는 프레임 밖이며, 보이는 범위에는 지지 없이 떠 있는 신체나 물체가 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "손끝의 지시 대상과 장소·의상은 잘 맞지만, 가슴과 팔의 클로즈업보다 넓게 잡았고 ‘유빅’이 명확히 읽혀 최종 문자 금지 조건을 위반한다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "비스듬한 가슴 장갑과 이를 가리키는 팔의 근접 구도가 더 충실하지만, 원형 마크는 유빅으로 식별하기 어렵고 장갑 색상·설계와 마츠다의 의상이 참조와 다르다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "오른쪽에서 뻗은 마츠다의 검지는 왼쪽의 찰리 가슴을 향하며, 손끝이 ‘유빅’ 글자의 오른쪽 끝에 닿는다. 찰리의 얼굴은 마츠다 쪽으로 약간 돌아가 아래를 향한다. 마츠다는 가슴 쪽을 향하지만 눈이 프레임 밖이어서 정확한 시선은 확인할 수 없다.",
        "built_space": "왼쪽에 전면 세탁기 한 대, 그 위의 소형 가전, 뒤쪽의 높은 냉장고 한 대와 소형 브라운관 모니터 한 대가 보인다. 오른쪽에는 원형 팬 장치 한 대와 그 아래 브라운관 모니터 한 대, 양옆 금속 선반 및 뒤쪽 공구 벽이 있다. 참조의 가전 배치와 재질을 밀접하게 따른다. 찰리는 중앙 통로에, 마츠다는 오른쪽 가장자리에 있어 집기와 충돌하지 않는다. 다만 머리부터 골반까지 보여 요구한 근접 구도보다 넓다.",
        "entities": "찰리의 흰 각진 마스크, 모래색 장갑, 파란 눈, 낡은 모자와 때 묻은 코트가 보인다. 노출된 장갑은 참조의 형태와 대체로 연결된다. 마츠다는 회색 머리의 노년 남성 일부와 주름진 손으로 나타나며, 남색 정장과 흰 소매도 참조에 부합한다. 가슴의 ‘유빅’은 작은 표면 인쇄로 표현되었으나 글자가 거의 온전히 읽힌다. 별도의 인물은 없다.",
        "hard_violations": [
         "가슴의 ‘유빅’ 문자가 명확히 판독되어, 이미지 어디에도 읽을 수 있는 문자를 두지 말라는 최종 지시를 위반한다."
        ],
        "physics": "마츠다의 손은 손목과 소매로 이어지고 팔이 화면 오른쪽 몸통에 연결되어 있다. 검지로 가슴 표면을 짚는 동작이 가능하며 손끝 접촉도 자연스럽다. 코트는 찰리의 어깨에 걸리고 장갑은 몸통 구조에 연결된다. 발은 잘렸지만 공중에 뜬 자세나 지지 없는 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "왼쪽 마츠다의 팔이 오른쪽 위로 뻗고, 검지 끝은 찰리 가슴의 마모된 원형 마크 왼쪽 위에 닿는다. 지시 대상은 분명히 가슴 표면의 흔적이다. 마츠다의 고개는 손끝과 마크 쪽으로 숙여져 있고 찰리의 얼굴도 아래쪽을 향하지만, 두 인물의 정확한 눈동자 방향은 확인하기 어렵다.",
        "built_space": "왼쪽 배경에 높은 냉장고 한 대와 브라운관 모니터 한 대, 그 아래 쌓인 사각 가전들이 보인다. 뒤에는 작업대 하나와 공구 벽, 상단 선반이 있으며 오른쪽 끝에도 선반 일부가 남는다. 참조 장소의 재질과 주요 배치가 유지된다. 마츠다는 왼쪽 전경, 찰리는 오른쪽 검사 공간을 차지한다. 비스듬한 가슴과 팔을 크게 잡고 주변 가전은 부분 윤곽만 남겨 클로즈업 요구에 더 가깝다.",
        "entities": "찰리는 육중한 기계 몸통, 파란 눈 일부, 흰 마스크와 낡은 천 겉옷으로 표현된다. 장갑은 참조의 모래색보다 회백색이며 가슴 판과 어깨 구조도 다르다. 모자와 하체는 프레임 밖이라 평가할 수 없다. 마츠다는 회색 머리, 나이 든 옆얼굴과 주름진 손을 가진 남성이지만, 검은 작업복은 참조의 정장과 다르다. 가슴의 마모된 원형 도안에는 읽기 어려운 흔적만 있어 ‘유빅’이라는 식별성은 부족하다. 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "마츠다의 손과 팔은 왼쪽 어깨 및 몸통에 연속적으로 연결되고, 검지는 가슴 장갑에 자연스럽게 닿는다. 찰리의 장갑과 관절은 몸통에 결합되어 있으며 겉옷은 어깨에 걸쳐 아래로 처진다. 마크는 긁히고 벗겨진 도료처럼 표면에 붙어 있다. 하체는 프레임 밖이며, 보이는 범위에는 지지 없이 떠 있는 신체나 물체가 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.571,
    "B": 1.571
   },
   "adjusted": {
    "A": 1.571,
    "B": 1.321
   },
   "violations": {
    "B": [
     "[gpt-high] 가슴의 ‘유빅’ 문자가 명확히 판독되어, 이미지 어디에도 읽을 수 있는 문자를 두지 말라는 최종 지시를 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1321,
   "A": 1571
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1321,
    "verdict_ko": "지시된 '유빅' 글자와 찰리의 의상(모자, 낡은 코트), 장갑 색상 및 파란 눈 등 프롬프트의 세부 요구사항을 매우 충실히 구현했습니다.  ★위반: [gpt-high] 가슴의 ‘유빅’ 문자가 명확히 판독되어, 이미지 어디에도 읽을 수 있는 문자를 두지 말라는 최종 지시를 위반한다."
   },
   {
    "label": "A",
    "score": 1571,
    "verdict_ko": "찰리의 장갑 색상이 다르고 지정된 모자와 코트가 누락되었으며, 가슴의 '유빅' 글자 대신 임의의 원형 마크가 생성되어 프롬프트 반영도가 낮습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L168B03.png",
    "asset_id": "cb825975-b374-47cb-9da4-9a5bcc3765b2",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 마츠다: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1426352>",
    "asset_id": "36f5f867-079a-4d08-baa1-280e22734416",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-e3b8-736a-b8b5-9742d78128cd",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S15sh4__bgfirst_bg.png",
   "bg_asset_id": "08f633d6-c3f3-4ebc-a4c1-0fb688cbc402",
   "bg_record_key": "S15sh4::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S15sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:36:26.234731+00:00",
  "fingerprint": "78d997736286f1f66ed023eea402ba4509efae21f69261153e307a4ad36c7520",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S15sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S15sh4_sel.png",
  "source_sha256": "c0388acda1039922c9c24cb75fc7e575bb90614a22d41cb852012eaa919c71a2",
  "file": "S15sh4_cine.png",
  "staged_sha256": "4de89dbba4e5e99b27e0cda62fdd1650815edfd6bc89e4795889edff264a1276",
  "latency_ms": 11089
 },
 "S15sh10::signage": {
  "fp": "44823a96f0df2315",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S15sh10": {
  "input_fingerprint": "398bad604bdb8fbf",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 어두운 상점 안, 낡은 전화기 수화기를 귀에 바짝 댄 채 입꼬리를 올린 마츠다의 얼굴.\n\nLOCATION (lock): At the telephone area inside the cluttered junk shop, in its dim daytime interior. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old telephone receiver (Held tightly against Matsuda's ear during the call) — Seen along his near cheek, with the mouthpiece beside rather than across his smile; used as Connect the private expression to the act of reporting Charlie while remaining small relative to the face; Piled appliances and scrap (Still accumulated inside the shop) — Indistinct portions remain behind Matsuda; used as Anchor the private call in the same shop without drawing attention away from his expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the described dim shop ambience with controlled facial detail and no newly introduced practical light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 마츠다 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The shop remains piled with appliances and scrap, and the computer browser has the retrieved news about the unique missing Ubik robot and its 500-million-won reward. Charlie is outside the shop, still dirty and worn, with the old coat and hat and worn chest logo. 마츠다: He remains inside the shop, using the telephone.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 마츠다 right now, so 마츠다's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 마츠다: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 마츠다 (일본인 남성, 60대, 나이 든 얼굴, 눈가 잔주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 어두운 상점 안, 낡은 전화기 수화기를 귀에 바짝 댄 채 입꼬리를 올린 마츠다의 얼굴.\n\nLOCATION (lock): At the telephone area inside the cluttered junk shop, in its dim daytime interior. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old telephone receiver (Held tightly against Matsuda's ear during the call) — Seen along his near cheek, with the mouthpiece beside rather than across his smile; used as Connect the private expression to the act of reporting Charlie while remaining small relative to the face; Piled appliances and scrap (Still accumulated inside the shop) — Indistinct portions remain behind Matsuda; used as Anchor the private call in the same shop without drawing attention away from his expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the described dim shop ambience with controlled facial detail and no newly introduced practical light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 마츠다 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The shop remains piled with appliances and scrap, and the computer browser has the retrieved news about the unique missing Ubik robot and its 500-million-won reward. Charlie is outside the shop, still dirty and worn, with the old coat and hat and worn chest logo. 마츠다: He remains inside the shop, using the telephone.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 마츠다 right now, so 마츠다's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 마츠다: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 마츠다 (일본인 남성, 60대, 나이 든 얼굴, 눈가 잔주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 어두운 상점 안, 낡은 전화기 수화기를 귀에 바짝 댄 채 입꼬리를 올린 마츠다의 얼굴.\n\nLOCATION (lock): At the telephone area inside the cluttered junk shop, in its dim daytime interior. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old telephone receiver (Held tightly against Matsuda's ear during the call) — Seen along his near cheek, with the mouthpiece beside rather than across his smile; used as Connect the private expression to the act of reporting Charlie while remaining small relative to the face; Piled appliances and scrap (Still accumulated inside the shop) — Indistinct portions remain behind Matsuda; used as Anchor the private call in the same shop without drawing attention away from his expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the described dim shop ambience with controlled facial detail and no newly introduced practical light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 마츠다 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The shop remains piled with appliances and scrap, and the computer browser has the retrieved news about the unique missing Ubik robot and its 500-million-won reward. Charlie is outside the shop, still dirty and worn, with the old coat and hat and worn chest logo. 마츠다: He remains inside the shop, using the telephone.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 마츠다 right now, so 마츠다's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 마츠다: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 마츠다 (일본인 남성, 60대, 나이 든 얼굴, 눈가 잔주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "시선은 화면 밖 왼쪽 아래를 자연스럽게 향하고 있으며 통화 중인 상황에 부합함. 전화기 송화기는 입을 가리지 않고 뺨 옆에 올바르게 위치함.",
    "built_space": "어두운 상점 내부로, 배경에 낡은 모니터와 냉장고 등 가전제품이 쌓여 있음. 프레이밍은 요구된 클로즈업 샷을 정확히 따름.",
    "entities": "마츠다(레퍼런스와 일치하는 노령의 남성)가 이전 샷과 동일한 스타일의 검은색 재킷을 입고 있음. 요구된 낡은 전화기 수화기가 등장함.",
    "hard_violations": [],
    "physics": "왼손이 수화기를 귀에 바짝 밀착시켜 단단히 쥐고 있으며, 손가락의 구조와 파지법이 물리적으로 자연스러움. 전화선 역시 아래로 자연스럽게 늘어져 있음."
   },
   {
    "label": "B",
    "direction": "시선이 카메라 렌즈를 정면으로 향하고 있어 지시문이 금지한 부자연스러운 정면 응시를 보여줌. 송화기는 턱 옆에 위치함.",
    "built_space": "어두운 상점 내부를 배경으로 낡은 전자제품 폐기물들이 배치되어 있음. 인물에 맞춘 클로즈업 샷임.",
    "entities": "마츠다(레퍼런스 인물과 일치)가 등장하며 검은색 재킷을 착용하고 있음. 낡은 전화기를 들고 있음.",
    "hard_violations": [],
    "physics": "오른손으로 전화기를 쥐고 귀에 대고 있으나, 수화기를 쥔 손가락 관절 형태가 약간 뭉개져 있어 다소 부자연스러움."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "지시된 '입꼬리를 올린' 표정과 자연스러운 시선 처리를 잘 구현했으며, 이전 샷의 조명 분위기와 배경의 디테일을 훌륭하게 유지했습니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "표정은 잘 묘사되었으나 카메라 렌즈를 정면으로 응시하여 '렌즈를 응시하지 말라'는 자연스러운 연기 지침을 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 화면 밖 왼쪽 아래를 자연스럽게 향하고 있으며 통화 중인 상황에 부합함. 전화기 송화기는 입을 가리지 않고 뺨 옆에 올바르게 위치함.",
        "built_space": "어두운 상점 내부로, 배경에 낡은 모니터와 냉장고 등 가전제품이 쌓여 있음. 프레이밍은 요구된 클로즈업 샷을 정확히 따름.",
        "entities": "마츠다(레퍼런스와 일치하는 노령의 남성)가 이전 샷과 동일한 스타일의 검은색 재킷을 입고 있음. 요구된 낡은 전화기 수화기가 등장함.",
        "hard_violations": [],
        "physics": "왼손이 수화기를 귀에 바짝 밀착시켜 단단히 쥐고 있으며, 손가락의 구조와 파지법이 물리적으로 자연스러움. 전화선 역시 아래로 자연스럽게 늘어져 있음."
       },
       {
        "label": "B",
        "direction": "시선이 카메라 렌즈를 정면으로 향하고 있어 지시문이 금지한 부자연스러운 정면 응시를 보여줌. 송화기는 턱 옆에 위치함.",
        "built_space": "어두운 상점 내부를 배경으로 낡은 전자제품 폐기물들이 배치되어 있음. 인물에 맞춘 클로즈업 샷임.",
        "entities": "마츠다(레퍼런스 인물과 일치)가 등장하며 검은색 재킷을 착용하고 있음. 낡은 전화기를 들고 있음.",
        "hard_violations": [],
        "physics": "오른손으로 전화기를 쥐고 귀에 대고 있으나, 수화기를 쥔 손가락 관절 형태가 약간 뭉개져 있어 다소 부자연스러움."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "지시된 '입꼬리를 올린' 표정과 자연스러운 시선 처리를 잘 구현했으며, 이전 샷의 조명 분위기와 배경의 디테일을 훌륭하게 유지했습니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "표정은 잘 묘사되었으나 카메라 렌즈를 정면으로 응시하여 '렌즈를 응시하지 말라'는 자연스러운 연기 지침을 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 화면 밖 왼쪽 아래를 자연스럽게 향하고 있으며 통화 중인 상황에 부합함. 전화기 송화기는 입을 가리지 않고 뺨 옆에 올바르게 위치함.",
        "built_space": "어두운 상점 내부로, 배경에 낡은 모니터와 냉장고 등 가전제품이 쌓여 있음. 프레이밍은 요구된 클로즈업 샷을 정확히 따름.",
        "entities": "마츠다(레퍼런스와 일치하는 노령의 남성)가 이전 샷과 동일한 스타일의 검은색 재킷을 입고 있음. 요구된 낡은 전화기 수화기가 등장함.",
        "hard_violations": [],
        "physics": "왼손이 수화기를 귀에 바짝 밀착시켜 단단히 쥐고 있으며, 손가락의 구조와 파지법이 물리적으로 자연스러움. 전화선 역시 아래로 자연스럽게 늘어져 있음."
       },
       {
        "label": "B",
        "direction": "시선이 카메라 렌즈를 정면으로 향하고 있어 지시문이 금지한 부자연스러운 정면 응시를 보여줌. 송화기는 턱 옆에 위치함.",
        "built_space": "어두운 상점 내부를 배경으로 낡은 전자제품 폐기물들이 배치되어 있음. 인물에 맞춘 클로즈업 샷임.",
        "entities": "마츠다(레퍼런스 인물과 일치)가 등장하며 검은색 재킷을 착용하고 있음. 낡은 전화기를 들고 있음.",
        "hard_violations": [],
        "physics": "오른손으로 전화기를 쥐고 귀에 대고 있으나, 수화기를 쥔 손가락 관절 형태가 약간 뭉개져 있어 다소 부자연스러움."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "얼굴 중심의 밀착 구도, 귀에 붙인 수화기와 드러난 미소, 어두운 상점 분위기가 핵심 지시를 가장 충실히 구현한다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "통화 동작과 인물 외형은 맞지만, 배경 가전이 지나치게 선명하고 실내가 밝으며 상단 조명까지 보여 어두운 상점 속 표정 중심 연출에서 멀어진다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 카메라 가까운 정면을 향하며 특정한 화면 속 대상을 보지는 않는다. 수화기의 위쪽은 귀에 밀착하고 아래 송화부는 입 왼쪽 아래에 놓여 미소를 가리지 않는다. 통화에 필요한 방향은 맞지만 렌즈 가까운 시선은 자연스러운 사적 순간이라는 지시에는 다소 덜 맞는다.",
        "built_space": "얼굴과 손이 화면 중심을 차지하고 양어깨 일부만 들어온다. 왼쪽에는 큰 밝은색 가전 한 대와 그 위에 쌓인 소형 기기들이, 오른쪽에는 회로가 드러난 기기 한 대와 겹친 폐가전들이 보인다. 낡고 복잡한 상점의 재질과 적치 상태가 이어지며 배경은 흐려져 있다. 원본의 고정 설비 수를 대조할 만큼 같은 벽면이 보이지는 않지만 명백한 중복 설비나 불가능한 반사는 없다.",
        "entities": "인물은 마츠다에 해당하는 노년 동아시아계 남성 한 명뿐이다. 일본 국적 자체는 외형으로 판별할 수 없으나, 짧은 회흑색 머리와 눈가 주름, 얼굴 윤곽이 인물 참조와 잘 맞는다. 검은 작업 재킷은 이전 장면의 옷을 따른다. 검은 낡은 유선 수화기 한 개와 나이 든 손이 보이고 입꼬리는 올라가 있다. 폐가전과 잡동사니도 존재한다. 컴퓨터 기사와 찰리는 이 클로즈업에 보이지 않으며, 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "손가락이 수화기 손잡이를 감싸 쥐고 귀 쪽으로 누른다. 손목은 재킷 소매와 자연스럽게 연결되고 수화기 선은 아래로 늘어진다. 머리는 목과 몸통에 정상적으로 지지되며 하체는 구도 밖이다. 배경 기기들은 다른 기기나 적치면 위에 놓여 있고, 지지 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "마츠다는 화면 왼쪽의 화면 밖 공간을 약간 내려다본다. 수화기의 청취부는 가까운 귀에 붙어 있고 송화부는 입 아래쪽에 있어 미소를 가리지 않는다. 손으로 수화기를 누르는 방향과 통화 자세는 자연스럽다.",
        "built_space": "얼굴을 비스듬히 보는 클로즈업이지만 어깨와 주변 공간의 비중이 크다. 왼쪽 상단에 창이 있고, 왼쪽 가전 더미 앞에는 나란한 브라운관 화면 두 대와 그 아래 밝은색 받침 가전이 보인다. 오른쪽에는 큰 밝은색 가전과 소형 화면 한 대가 쌓여 있다. 상단에는 밝게 켜진 선형 조명 일부가 보인다. 폐가전 상점이라는 장소 유형은 맞지만 배경이 선명하고 밝아, 어두운 분위기와 새 실용광원을 추가하지 말라는 조건에는 어긋난다. 불가능한 반사나 확실한 설비 중복은 보이지 않는다.",
        "entities": "노년 동아시아계 남성 한 명만 등장하며 회흑색 머리, 주름, 얼굴 생김새가 마츠다 참조와 대체로 일치한다. 검은 작업 재킷도 이전 장면과 부합한다. 손의 피부와 굵기는 인물에게 어울리고 낡은 검은 유선 수화기가 있다. 입꼬리가 살짝 올라가 있다. 배경에는 폐가전과 전화기 본체로 보이는 물건이 있으며 읽을 수 있는 글자는 없다. 찰리와 뉴스 화면은 구도 밖이므로 누락으로 보지 않는다.",
        "hard_violations": [],
        "physics": "손이 수화기 중앙을 확실히 쥐고 청취부를 귀에 붙인다. 손목과 팔은 소매 안으로 정상적으로 이어지고 전화선은 중력 방향으로 내려간다. 목과 어깨의 연결도 자연스럽다. 화면에 보이는 가전들은 받침 가전이나 폐기물 더미에 기대어 지지되며 공중에 떠 있는 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "얼굴 중심의 밀착 구도, 귀에 붙인 수화기와 드러난 미소, 어두운 상점 분위기가 핵심 지시를 가장 충실히 구현한다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "통화 동작과 인물 외형은 맞지만, 배경 가전이 지나치게 선명하고 실내가 밝으며 상단 조명까지 보여 어두운 상점 속 표정 중심 연출에서 멀어진다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "시선은 카메라 가까운 정면을 향하며 특정한 화면 속 대상을 보지는 않는다. 수화기의 위쪽은 귀에 밀착하고 아래 송화부는 입 왼쪽 아래에 놓여 미소를 가리지 않는다. 통화에 필요한 방향은 맞지만 렌즈 가까운 시선은 자연스러운 사적 순간이라는 지시에는 다소 덜 맞는다.",
        "built_space": "얼굴과 손이 화면 중심을 차지하고 양어깨 일부만 들어온다. 왼쪽에는 큰 밝은색 가전 한 대와 그 위에 쌓인 소형 기기들이, 오른쪽에는 회로가 드러난 기기 한 대와 겹친 폐가전들이 보인다. 낡고 복잡한 상점의 재질과 적치 상태가 이어지며 배경은 흐려져 있다. 원본의 고정 설비 수를 대조할 만큼 같은 벽면이 보이지는 않지만 명백한 중복 설비나 불가능한 반사는 없다.",
        "entities": "인물은 마츠다에 해당하는 노년 동아시아계 남성 한 명뿐이다. 일본 국적 자체는 외형으로 판별할 수 없으나, 짧은 회흑색 머리와 눈가 주름, 얼굴 윤곽이 인물 참조와 잘 맞는다. 검은 작업 재킷은 이전 장면의 옷을 따른다. 검은 낡은 유선 수화기 한 개와 나이 든 손이 보이고 입꼬리는 올라가 있다. 폐가전과 잡동사니도 존재한다. 컴퓨터 기사와 찰리는 이 클로즈업에 보이지 않으며, 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "손가락이 수화기 손잡이를 감싸 쥐고 귀 쪽으로 누른다. 손목은 재킷 소매와 자연스럽게 연결되고 수화기 선은 아래로 늘어진다. 머리는 목과 몸통에 정상적으로 지지되며 하체는 구도 밖이다. 배경 기기들은 다른 기기나 적치면 위에 놓여 있고, 지지 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "마츠다는 화면 왼쪽의 화면 밖 공간을 약간 내려다본다. 수화기의 청취부는 가까운 귀에 붙어 있고 송화부는 입 아래쪽에 있어 미소를 가리지 않는다. 손으로 수화기를 누르는 방향과 통화 자세는 자연스럽다.",
        "built_space": "얼굴을 비스듬히 보는 클로즈업이지만 어깨와 주변 공간의 비중이 크다. 왼쪽 상단에 창이 있고, 왼쪽 가전 더미 앞에는 나란한 브라운관 화면 두 대와 그 아래 밝은색 받침 가전이 보인다. 오른쪽에는 큰 밝은색 가전과 소형 화면 한 대가 쌓여 있다. 상단에는 밝게 켜진 선형 조명 일부가 보인다. 폐가전 상점이라는 장소 유형은 맞지만 배경이 선명하고 밝아, 어두운 분위기와 새 실용광원을 추가하지 말라는 조건에는 어긋난다. 불가능한 반사나 확실한 설비 중복은 보이지 않는다.",
        "entities": "노년 동아시아계 남성 한 명만 등장하며 회흑색 머리, 주름, 얼굴 생김새가 마츠다 참조와 대체로 일치한다. 검은 작업 재킷도 이전 장면과 부합한다. 손의 피부와 굵기는 인물에게 어울리고 낡은 검은 유선 수화기가 있다. 입꼬리가 살짝 올라가 있다. 배경에는 폐가전과 전화기 본체로 보이는 물건이 있으며 읽을 수 있는 글자는 없다. 찰리와 뉴스 화면은 구도 밖이므로 누락으로 보지 않는다.",
        "hard_violations": [],
        "physics": "손이 수화기 중앙을 확실히 쥐고 청취부를 귀에 붙인다. 손목과 팔은 소매 안으로 정상적으로 이어지고 전화선은 중력 방향으로 내려간다. 목과 어깨의 연결도 자연스럽다. 화면에 보이는 가전들은 받침 가전이나 폐기물 더미에 기대어 지지되며 공중에 떠 있는 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.778,
    "B": 1.667
   },
   "adjusted": {
    "A": 1.778,
    "B": 1.667
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1778,
   "B": 1667
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1778,
    "verdict_ko": "지시된 '입꼬리를 올린' 표정과 자연스러운 시선 처리를 잘 구현했으며, 이전 샷의 조명 분위기와 배경의 디테일을 훌륭하게 유지했습니다."
   },
   {
    "label": "B",
    "score": 1667,
    "verdict_ko": "표정은 잘 묘사되었으나 카메라 렌즈를 정면으로 응시하여 '렌즈를 응시하지 말라'는 자연스러운 연기 지침을 위반했습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 마츠다 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S15sh4_sel.png",
    "asset_id": "f1a62e80-19d1-4b03-bc16-0fe4eb47c9ef",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 마츠다: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1426352>",
    "asset_id": "36f5f867-079a-4d08-baa1-280e22734416",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-e74e-7f45-9fee-bb8b3b7d1add",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S15sh4"
  }
 },
 "S15sh10::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:37:27.421337+00:00",
  "fingerprint": "00e753a22afd3af4ffb77126353bd689a318e8411e9177a88c7001045f84894d",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S15sh10_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S15sh10_sel.png",
  "source_sha256": "5625789f958972d185d75ac8c2c65a3e0a77f6edc9ffd8afd1e731c3b4b6a1b5",
  "file": "S15sh10_cine.png",
  "staged_sha256": "4bcfe8695051b6bbda1e5f925ee18d7aa58fc4b3df297da49ff1bd4319de938e",
  "latency_ms": 11241
 },
 "S16sh3::signage": {
  "fp": "a72c33f111f761d2",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S16sh3": {
  "input_fingerprint": "d084f253b0115396",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 흙먼지가 씻겨나간 젖은 몸으로 양어깨를 치켜올린 채 즐거워하는 찰리의 상체.\n\nLOCATION (lock): In the open washing area directly in front of a container home in the refugee settlement. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Hose water stream (Striking Charlie and washing away the dirt) — Enters laterally from the off-screen hose position toward his torso; used as Connect the raised shoulders to the physical sensation without obscuring his face; Front of 라울's container home (Visible behind the washing action) — A partial exterior backdrop is retained beside the upper body; used as Locate the intimate action without widening to the observers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daylight with controlled highlights on the explicitly wet body and water stream, preserving a gentle rather than harsh mood.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A hose is spraying water outside Raul's container, leaving Charlie wet as the dirt washes off. His worn Ubik chest logo and blue-lit eyes remain; the previously established old coat and hat have no stated removal.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 흙먼지가 씻겨나간 젖은 몸으로 양어깨를 치켜올린 채 즐거워하는 찰리의 상체.\n\nLOCATION (lock): In the open washing area directly in front of a container home in the refugee settlement. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Hose water stream (Striking Charlie and washing away the dirt) — Enters laterally from the off-screen hose position toward his torso; used as Connect the raised shoulders to the physical sensation without obscuring his face; Front of 라울's container home (Visible behind the washing action) — A partial exterior backdrop is retained beside the upper body; used as Locate the intimate action without widening to the observers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daylight with controlled highlights on the explicitly wet body and water stream, preserving a gentle rather than harsh mood.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A hose is spraying water outside Raul's container, leaving Charlie wet as the dirt washes off. His worn Ubik chest logo and blue-lit eyes remain; the previously established old coat and hat have no stated removal.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 흙먼지가 씻겨나간 젖은 몸으로 양어깨를 치켜올린 채 즐거워하는 찰리의 상체.\n\nLOCATION (lock): In the open washing area directly in front of a container home in the refugee settlement. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Hose water stream (Striking Charlie and washing away the dirt) — Enters laterally from the off-screen hose position toward his torso; used as Connect the raised shoulders to the physical sensation without obscuring his face; Front of 라울's container home (Visible behind the washing action) — A partial exterior backdrop is retained beside the upper body; used as Locate the intimate action without widening to the observers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daylight with controlled highlights on the explicitly wet body and water stream, preserving a gentle rather than harsh mood.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A hose is spraying water outside Raul's container, leaving Charlie wet as the dirt washes off. His worn Ubik chest logo and blue-lit eyes remain; the previously established old coat and hat have no stated removal.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S16sh3__bgfirst_bg.png",
     "asset_id": "58f097e3-cb89-4739-96b1-20ba4de0ccc6",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S16sh3.png",
     "asset_id": "083f7cfb-ee18-4327-afb0-899eddfe6874",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L169B01.png",
     "asset_id": "a2e9a0db-ecb3-47e7-96eb-291c9f873ffb",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "찰리의 시선은 정면을 향함. 화면 우측 밖에서 날아온 물줄기가 가슴에 적중하며, 배경 우측 호스도 물을 뿜고 있음.",
    "built_space": "컨테이너 벽면, 창문, 주변 집기들이 배경을 이룸. 우측 배경에 참조 사진의 삼각대 호스가 그대로 유지됨.",
    "entities": "찰리는 파란 눈을 가졌으며, 필수 사항인 베이지 코트, 중절모, 가슴의 UBIK 로고를 모두 정확히 착용하고 있음.",
    "hard_violations": [
     "[gemini-pro] 중복된 설비 (공간에 하나뿐인 세척용 호스가 배경과 화면 밖 두 곳에 존재함)",
     "[gpt-high] 가슴의 'UBIK' 문자와 로고가 선명하게 읽혀, 읽을 수 있는 글자와 로고를 금지한 조건을 위반한다."
    ],
    "physics": "찰리는 두 발로 땅을 딛고 서 있음. 물줄기는 수압으로 날아오며, 배경의 호스는 삼각대가 지지함."
   },
   {
    "label": "B",
    "direction": "찰리는 정면을 향하며, 우측 노즐에서 시작된 물줄기가 가슴에 부딪힘.",
    "built_space": "컨테이너 벽면과 플라스틱 상자들이 배경에 보이나, 참조 사진의 삼각대 호스는 생략됨.",
    "entities": "찰리의 파란 눈은 묘사되었으나 필수 복장(코트, 모자)과 로고가 없음. 우측 가장자리에 프롬프트에 없는 사람의 손이 보임.",
    "hard_violations": [
     "[gemini-pro] 창조된 인물 (화면 우측 끝에 호스를 잡고 있는 출처 불명의 손 등장)"
    ],
    "physics": "찰리는 땅에 서서 양팔을 공중으로 들고 있음. 우측 끝의 호스는 사람의 손이 쥐어 지탱함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "필수 복장과 로고는 잘 유지했으나, 공간에 하나뿐인 세척용 호스가 배경과 화면 밖 두 곳에 중복 등장하는 치명적 위반이 있습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "필수 복장과 로고가 전면 누락되었으며, 우측 가장자리에 프롬프트에 명시되지 않은 사람의 손이 등장해 규정을 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 시선은 정면을 향함. 화면 우측 밖에서 날아온 물줄기가 가슴에 적중하며, 배경 우측 호스도 물을 뿜고 있음.",
        "built_space": "컨테이너 벽면, 창문, 주변 집기들이 배경을 이룸. 우측 배경에 참조 사진의 삼각대 호스가 그대로 유지됨.",
        "entities": "찰리는 파란 눈을 가졌으며, 필수 사항인 베이지 코트, 중절모, 가슴의 UBIK 로고를 모두 정확히 착용하고 있음.",
        "hard_violations": [
         "중복된 설비 (공간에 하나뿐인 세척용 호스가 배경과 화면 밖 두 곳에 존재함)"
        ],
        "physics": "찰리는 두 발로 땅을 딛고 서 있음. 물줄기는 수압으로 날아오며, 배경의 호스는 삼각대가 지지함."
       },
       {
        "label": "B",
        "direction": "찰리는 정면을 향하며, 우측 노즐에서 시작된 물줄기가 가슴에 부딪힘.",
        "built_space": "컨테이너 벽면과 플라스틱 상자들이 배경에 보이나, 참조 사진의 삼각대 호스는 생략됨.",
        "entities": "찰리의 파란 눈은 묘사되었으나 필수 복장(코트, 모자)과 로고가 없음. 우측 가장자리에 프롬프트에 없는 사람의 손이 보임.",
        "hard_violations": [
         "창조된 인물 (화면 우측 끝에 호스를 잡고 있는 출처 불명의 손 등장)"
        ],
        "physics": "찰리는 땅에 서서 양팔을 공중으로 들고 있음. 우측 끝의 호스는 사람의 손이 쥐어 지탱함."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "필수 복장과 로고는 잘 유지했으나, 공간에 하나뿐인 세척용 호스가 배경과 화면 밖 두 곳에 중복 등장하는 치명적 위반이 있습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "필수 복장과 로고가 전면 누락되었으며, 우측 가장자리에 프롬프트에 명시되지 않은 사람의 손이 등장해 규정을 위반했습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 시선은 정면을 향함. 화면 우측 밖에서 날아온 물줄기가 가슴에 적중하며, 배경 우측 호스도 물을 뿜고 있음.",
        "built_space": "컨테이너 벽면, 창문, 주변 집기들이 배경을 이룸. 우측 배경에 참조 사진의 삼각대 호스가 그대로 유지됨.",
        "entities": "찰리는 파란 눈을 가졌으며, 필수 사항인 베이지 코트, 중절모, 가슴의 UBIK 로고를 모두 정확히 착용하고 있음.",
        "hard_violations": [
         "중복된 설비 (공간에 하나뿐인 세척용 호스가 배경과 화면 밖 두 곳에 존재함)"
        ],
        "physics": "찰리는 두 발로 땅을 딛고 서 있음. 물줄기는 수압으로 날아오며, 배경의 호스는 삼각대가 지지함."
       },
       {
        "label": "B",
        "direction": "찰리는 정면을 향하며, 우측 노즐에서 시작된 물줄기가 가슴에 부딪힘.",
        "built_space": "컨테이너 벽면과 플라스틱 상자들이 배경에 보이나, 참조 사진의 삼각대 호스는 생략됨.",
        "entities": "찰리의 파란 눈은 묘사되었으나 필수 복장(코트, 모자)과 로고가 없음. 우측 가장자리에 프롬프트에 없는 사람의 손이 보임.",
        "hard_violations": [
         "창조된 인물 (화면 우측 끝에 호스를 잡고 있는 출처 불명의 손 등장)"
        ],
        "physics": "찰리는 땅에 서서 양팔을 공중으로 들고 있음. 우측 끝의 호스는 사람의 손이 쥐어 지탱함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "젖은 상체 중심의 미디엄 숏과 가슴에 맞는 측면 물줄기는 충실하지만, 어깨를 움츠리는 즐거움이 양팔을 높이 드는 동작으로 바뀌었고 모자와 외투도 빠졌다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "모자·외투와 젖은 장갑판은 유지했지만, 읽히는 가슴 문자가 금지 조건을 위반하고 양어깨를 치켜올리는 동작보다 손을 벌린 자세와 넓은 배경이 강조된다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴과 파란 눈은 화면 오른쪽 앞을 향한다. 물줄기는 오른쪽 화면 가장자리에서 왼쪽 위로 들어와 찰리의 가슴 장갑판에 실제로 부딪히며, 얼굴을 가리지 않는다. 양팔은 위로 들려 있고 주먹 일부는 상단에서 잘린다.",
        "built_space": "회색 골강판 컨테이너 외벽이 상체 뒤를 채운다. 왼쪽 어깨 뒤로 창문 하나가 일부 보이고, 오른쪽 아래에는 파란 통 두 개와 어두운 큰 통 하나, 흰 용기와 선반 일부, 급수대가 보인다. 출입문은 인물에 가려 확인할 수 없다. 참조의 외벽 재료와 세척 구역 설비가 이어지며, 불가능한 반사나 명백한 설비 중복은 보이지 않는다.",
        "entities": "찰리 한 명이 중심에 있으며 흰 각진 마스크, 샌드 베이지 장갑판, 육중한 팔과 노출된 기계 관절이 캐릭터 참조에 가깝다. 눈은 지시대로 파랗다. 기계형 외형이므로 인간의 나이·민족·성별은 판독할 수 없다. 머리와 몸통이 노출되어 있어 모자와 낡은 외투의 누락이 확인된다. 가슴에서 읽히는 문자나 로고는 보이지 않는다. 오른쪽 끝에는 물줄기 출구의 작은 갈색 부분이 있으나 별도 인물의 손이라고 확정할 만큼 드러나지는 않는다.",
        "hard_violations": [],
        "physics": "몸통은 화면 아래 골반으로 이어지며, 발이 잘렸다고 해서 공중에 떠 있는 모습은 아니다. 들어 올린 팔은 어깨와 팔꿈치 관절에 연결되어 있어 가능한 자세다. 물줄기는 가슴에서 비산하고 물방울은 장갑판을 따라 아래로 흐른다. 다만 물의 감촉에 양어깨를 움츠린 순간보다는 두 팔을 들어 환호하는 동작에 가깝다."
       },
       {
        "label": "B",
        "direction": "찰리의 얼굴은 화면 오른쪽 앞을 향하고 파란 눈은 위로 휜 형태로 즐거움을 나타낸다. 전경 물줄기는 오른쪽 아래 화면 밖에서 가슴 중앙으로 들어와 충돌한다. 양손은 허리 옆에서 손바닥을 위로 벌리고 있다. 뒤쪽 급수대에서도 별도의 물줄기가 오른쪽 땅을 향해 분사된다.",
        "built_space": "컨테이너 전면 대부분과 넓은 흙마당이 보인다. 창문은 왼쪽·중앙·오른쪽에 총 세 개가 보이고, 출입문은 인물 뒤에 대부분 가려져 있다. 왼쪽 화분 탁자 하나, 오른쪽 작은 선반 하나, 여러 저장 통, 급수대 하나와 가장자리 물탱크가 참조와 유사하게 배치된다. 장소는 잘 식별되지만 상체 옆에 외관 일부만 남기라는 지시보다 배경의 범위가 넓다.",
        "entities": "찰리 한 명의 베이지 장갑판, 흰 기계형 마스크, 무거운 팔과 파란 눈이 보인다. 인간의 나이·민족·성별은 외형상 판독할 수 없다. 낡은 모자와 열린 외투는 유지되어 있다. 가슴에는 'UBIK'이라는 문자가 명확히 읽히며, 이는 문자와 로고를 판독 불가능하게 하라는 최종 조건에 어긋난다. 추가 인물은 보이지 않는다.",
        "hard_violations": [
         "가슴의 'UBIK' 문자와 로고가 선명하게 읽혀, 읽을 수 있는 글자와 로고를 금지한 조건을 위반한다."
        ],
        "physics": "몸통은 아래쪽 골반과 다리로 이어지고, 벌린 팔과 손은 관절로 정상적으로 연결되어 있다. 모자는 머리에 얹혀 있고 외투는 어깨에서 아래로 늘어진다. 가슴에 맞은 물은 튀고 아래로 흘러 물리적으로 자연스럽다. 다만 양어깨를 뚜렷하게 치켜올리기보다는 팔꿈치를 굽혀 손을 벌린 자세로 보인다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "젖은 상체 중심의 미디엄 숏과 가슴에 맞는 측면 물줄기는 충실하지만, 어깨를 움츠리는 즐거움이 양팔을 높이 드는 동작으로 바뀌었고 모자와 외투도 빠졌다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "모자·외투와 젖은 장갑판은 유지했지만, 읽히는 가슴 문자가 금지 조건을 위반하고 양어깨를 치켜올리는 동작보다 손을 벌린 자세와 넓은 배경이 강조된다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴과 파란 눈은 화면 오른쪽 앞을 향한다. 물줄기는 오른쪽 화면 가장자리에서 왼쪽 위로 들어와 찰리의 가슴 장갑판에 실제로 부딪히며, 얼굴을 가리지 않는다. 양팔은 위로 들려 있고 주먹 일부는 상단에서 잘린다.",
        "built_space": "회색 골강판 컨테이너 외벽이 상체 뒤를 채운다. 왼쪽 어깨 뒤로 창문 하나가 일부 보이고, 오른쪽 아래에는 파란 통 두 개와 어두운 큰 통 하나, 흰 용기와 선반 일부, 급수대가 보인다. 출입문은 인물에 가려 확인할 수 없다. 참조의 외벽 재료와 세척 구역 설비가 이어지며, 불가능한 반사나 명백한 설비 중복은 보이지 않는다.",
        "entities": "찰리 한 명이 중심에 있으며 흰 각진 마스크, 샌드 베이지 장갑판, 육중한 팔과 노출된 기계 관절이 캐릭터 참조에 가깝다. 눈은 지시대로 파랗다. 기계형 외형이므로 인간의 나이·민족·성별은 판독할 수 없다. 머리와 몸통이 노출되어 있어 모자와 낡은 외투의 누락이 확인된다. 가슴에서 읽히는 문자나 로고는 보이지 않는다. 오른쪽 끝에는 물줄기 출구의 작은 갈색 부분이 있으나 별도 인물의 손이라고 확정할 만큼 드러나지는 않는다.",
        "hard_violations": [],
        "physics": "몸통은 화면 아래 골반으로 이어지며, 발이 잘렸다고 해서 공중에 떠 있는 모습은 아니다. 들어 올린 팔은 어깨와 팔꿈치 관절에 연결되어 있어 가능한 자세다. 물줄기는 가슴에서 비산하고 물방울은 장갑판을 따라 아래로 흐른다. 다만 물의 감촉에 양어깨를 움츠린 순간보다는 두 팔을 들어 환호하는 동작에 가깝다."
       },
       {
        "label": "A",
        "direction": "찰리의 얼굴은 화면 오른쪽 앞을 향하고 파란 눈은 위로 휜 형태로 즐거움을 나타낸다. 전경 물줄기는 오른쪽 아래 화면 밖에서 가슴 중앙으로 들어와 충돌한다. 양손은 허리 옆에서 손바닥을 위로 벌리고 있다. 뒤쪽 급수대에서도 별도의 물줄기가 오른쪽 땅을 향해 분사된다.",
        "built_space": "컨테이너 전면 대부분과 넓은 흙마당이 보인다. 창문은 왼쪽·중앙·오른쪽에 총 세 개가 보이고, 출입문은 인물 뒤에 대부분 가려져 있다. 왼쪽 화분 탁자 하나, 오른쪽 작은 선반 하나, 여러 저장 통, 급수대 하나와 가장자리 물탱크가 참조와 유사하게 배치된다. 장소는 잘 식별되지만 상체 옆에 외관 일부만 남기라는 지시보다 배경의 범위가 넓다.",
        "entities": "찰리 한 명의 베이지 장갑판, 흰 기계형 마스크, 무거운 팔과 파란 눈이 보인다. 인간의 나이·민족·성별은 외형상 판독할 수 없다. 낡은 모자와 열린 외투는 유지되어 있다. 가슴에는 'UBIK'이라는 문자가 명확히 읽히며, 이는 문자와 로고를 판독 불가능하게 하라는 최종 조건에 어긋난다. 추가 인물은 보이지 않는다.",
        "hard_violations": [
         "가슴의 'UBIK' 문자와 로고가 선명하게 읽혀, 읽을 수 있는 글자와 로고를 금지한 조건을 위반한다."
        ],
        "physics": "몸통은 아래쪽 골반과 다리로 이어지고, 벌린 팔과 손은 관절로 정상적으로 연결되어 있다. 모자는 머리에 얹혀 있고 외투는 어깨에서 아래로 늘어진다. 가슴에 맞은 물은 튀고 아래로 흘러 물리적으로 자연스럽다. 다만 양어깨를 뚜렷하게 치켜올리기보다는 팔꿈치를 굽혀 손을 벌린 자세로 보인다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.5,
    "B": 1.75
   },
   "adjusted": {
    "A": 1.25,
    "B": 1.5
   },
   "violations": {
    "A": [
     "[gemini-pro] 중복된 설비 (공간에 하나뿐인 세척용 호스가 배경과 화면 밖 두 곳에 존재함)",
     "[gpt-high] 가슴의 'UBIK' 문자와 로고가 선명하게 읽혀, 읽을 수 있는 글자와 로고를 금지한 조건을 위반한다."
    ],
    "B": [
     "[gemini-pro] 창조된 인물 (화면 우측 끝에 호스를 잡고 있는 출처 불명의 손 등장)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1250,
   "B": 1500
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1250,
    "verdict_ko": "필수 복장과 로고는 잘 유지했으나, 공간에 하나뿐인 세척용 호스가 배경과 화면 밖 두 곳에 중복 등장하는 치명적 위반이 있습니다.  ★위반: [gemini-pro] 중복된 설비 (공간에 하나뿐인 세척용 호스가 배경과 화면 밖 두 곳에 존재함) / [gpt-high] 가슴의 'UBIK' 문자와 로고가 선명하게 읽혀, 읽을 수 있는 글자와 로고를 금지한 조건을 위반한다."
   },
   {
    "label": "B",
    "score": 1500,
    "verdict_ko": "필수 복장과 로고가 전면 누락되었으며, 우측 가장자리에 프롬프트에 명시되지 않은 사람의 손이 등장해 규정을 위반했습니다.  ★위반: [gemini-pro] 창조된 인물 (화면 우측 끝에 호스를 잡고 있는 출처 불명의 손 등장)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L169B01.png",
    "asset_id": "a2e9a0db-ecb3-47e7-96eb-291c9f873ffb",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-e91d-72ae-a0d6-250a689ec8e4",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S16sh3__bgfirst_bg.png",
   "bg_asset_id": "58f097e3-cb89-4739-96b1-20ba4de0ccc6",
   "bg_record_key": "S16sh3::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S16sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:04:33.189338+00:00",
  "fingerprint": "6cb0b36cb5e2ed644d12d2aafb6b13919c8e6e35e9f94a9b31cfb3426d6e2662",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S16sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S16sh3_sel.png",
  "source_sha256": "12fa8d2269d479a29a77eb5e3037964d1a7b87ca5ff54b5d9cc33eb3b0b74d12",
  "file": "S16sh3_cine.png",
  "staged_sha256": "b0086895c4a0e9b7cf34686d0804136d12822bc15952e9a5099d85c0de6d50bd",
  "latency_ms": 17043
 },
 "S16sh6::signage": {
  "fp": "5b8a99dbccdef863",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S16sh6": {
  "input_fingerprint": "6a563362473eec82",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 한 손으로 턱을 괸 채 확신에 찬 눈빛으로 찰리 쪽을 주시하는 현우의 얼굴.\n\nLOCATION (lock): At the edge of the open ground outside the container home, a short distance from the robot-washing activity. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container-home exterior surroundings (Same exterior setting as the washing action) — Only soft, partial surroundings remain behind Hyunwoo; used as Maintain spatial continuity while preserving clear look room toward off-screen Charlie.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the neutral daylight of the exterior and restrained facial contrast, allowing the assured expression to carry the change in mood.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The washing area remains outside Raul's container, with the hose in use and Charlie's washed body wet. Charlie retains the worn chest logo, blue-lit eyes and established coat-and-hat disguise. 현우: He continues watching from a distance, wearing his outer shirt. His facial injuries and dog-bitten leg remain unhealed.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 한 손으로 턱을 괸 채 확신에 찬 눈빛으로 찰리 쪽을 주시하는 현우의 얼굴.\n\nLOCATION (lock): At the edge of the open ground outside the container home, a short distance from the robot-washing activity. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container-home exterior surroundings (Same exterior setting as the washing action) — Only soft, partial surroundings remain behind Hyunwoo; used as Maintain spatial continuity while preserving clear look room toward off-screen Charlie.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the neutral daylight of the exterior and restrained facial contrast, allowing the assured expression to carry the change in mood.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The washing area remains outside Raul's container, with the hose in use and Charlie's washed body wet. Charlie retains the worn chest logo, blue-lit eyes and established coat-and-hat disguise. 현우: He continues watching from a distance, wearing his outer shirt. His facial injuries and dog-bitten leg remain unhealed.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 한 손으로 턱을 괸 채 확신에 찬 눈빛으로 찰리 쪽을 주시하는 현우의 얼굴.\n\nLOCATION (lock): At the edge of the open ground outside the container home, a short distance from the robot-washing activity. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container-home exterior surroundings (Same exterior setting as the washing action) — Only soft, partial surroundings remain behind Hyunwoo; used as Maintain spatial continuity while preserving clear look room toward off-screen Charlie.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the neutral daylight of the exterior and restrained facial contrast, allowing the assured expression to carry the change in mood.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The washing area remains outside Raul's container, with the hose in use and Charlie's washed body wet. Charlie retains the worn chest logo, blue-lit eyes and established coat-and-hat disguise. 현우: He continues watching from a distance, wearing his outer shirt. His facial injuries and dog-bitten leg remain unhealed.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "화면 밖 우측(찰리가 있는 방향)을 확신에 찬 눈빛으로 주시함.",
    "built_space": "아웃포커싱된 푸른색 컨테이너 배경이 설정과 일치함.",
    "entities": "얼굴 상처가 있는 현우(레퍼런스와 일치)가 단독으로 등장함.",
    "hard_violations": [],
    "physics": "손이 턱을 자연스럽게 받치고 있으며 신체 하중이 안정적임."
   },
   {
    "label": "B",
    "direction": "화면 우측 전경에 위치한 인물을 향해 시선이 향함.",
    "built_space": "배경에 푸른색 컨테이너가 보임.",
    "entities": "현우와 함께 우측 전경에 지시문에 없는 여성 인물의 뒷모습이 포함됨.",
    "hard_violations": [
     "[gemini-pro] 지시문에 없는 추가 인물 등장",
     "[gpt-high] 현우만 보여야 하는 장면에 별도 인물의 머리와 상체를 오른쪽 전경에 추가했다."
    ],
    "physics": "손가락이 턱에 닿아 있으며 붕대를 감은 다리가 보이나 자세가 다소 경직됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "오프스크린을 향한 시선 처리와 턱을 괸 자세를 클로즈업으로 정확히 연출했으며, 추가 인물 없이 지시를 잘 따름."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "지시문에 명시되지 않은 인물이 프레임 안에 등장하여 단독 인물 조건 및 오프스크린 시선 조건을 위반함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "화면 밖 우측(찰리가 있는 방향)을 확신에 찬 눈빛으로 주시함.",
        "built_space": "아웃포커싱된 푸른색 컨테이너 배경이 설정과 일치함.",
        "entities": "얼굴 상처가 있는 현우(레퍼런스와 일치)가 단독으로 등장함.",
        "hard_violations": [],
        "physics": "손이 턱을 자연스럽게 받치고 있으며 신체 하중이 안정적임."
       },
       {
        "label": "B",
        "direction": "화면 우측 전경에 위치한 인물을 향해 시선이 향함.",
        "built_space": "배경에 푸른색 컨테이너가 보임.",
        "entities": "현우와 함께 우측 전경에 지시문에 없는 여성 인물의 뒷모습이 포함됨.",
        "hard_violations": [
         "지시문에 없는 추가 인물 등장"
        ],
        "physics": "손가락이 턱에 닿아 있으며 붕대를 감은 다리가 보이나 자세가 다소 경직됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "오프스크린을 향한 시선 처리와 턱을 괸 자세를 클로즈업으로 정확히 연출했으며, 추가 인물 없이 지시를 잘 따름."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "지시문에 명시되지 않은 인물이 프레임 안에 등장하여 단독 인물 조건 및 오프스크린 시선 조건을 위반함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "화면 밖 우측(찰리가 있는 방향)을 확신에 찬 눈빛으로 주시함.",
        "built_space": "아웃포커싱된 푸른색 컨테이너 배경이 설정과 일치함.",
        "entities": "얼굴 상처가 있는 현우(레퍼런스와 일치)가 단독으로 등장함.",
        "hard_violations": [],
        "physics": "손이 턱을 자연스럽게 받치고 있으며 신체 하중이 안정적임."
       },
       {
        "label": "B",
        "direction": "화면 우측 전경에 위치한 인물을 향해 시선이 향함.",
        "built_space": "배경에 푸른색 컨테이너가 보임.",
        "entities": "현우와 함께 우측 전경에 지시문에 없는 여성 인물의 뒷모습이 포함됨.",
        "hard_violations": [
         "지시문에 없는 추가 인물 등장"
        ],
        "physics": "손가락이 턱에 닿아 있으며 붕대를 감은 다리가 보이나 자세가 다소 경직됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "허용되지 않은 인물을 전경에 추가하고 그쪽을 바라보며, 얼굴 클로즈업 대신 무릎까지 포함한 넓은 구도로 핵심 연출을 위반한다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "현우만의 얼굴 클로즈업, 한 손으로 턱을 받친 자세, 화면 밖 찰리 방향의 시선과 컨테이너 외부의 연속성을 충실히 구현하나 확신보다는 생각에 잠긴 표정이 조금 강하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 얼굴과 눈은 화면 오른쪽 전경에 있는 검은 머리 인물을 향한다. 따라서 실제로 보이는 시선 대상은 화면 밖 찰리가 아니라 추가된 인물이다. 그 인물도 현우 쪽으로 고개를 돌리고 있다.",
        "built_space": "푸른 골강판 컨테이너 벽 한 면, 상단에 일부 잘린 창틀 하나, 하단 받침 블록과 흙바닥이 보인다. 현우는 지면 가까이 낮은 자세로 있고 오른쪽 전경의 추가 인물이 화면을 크게 차지한다. 재료와 낮 조명은 장소 참고와 유사하지만, 배경과 무릎까지 드러나는 구도는 요구한 얼굴 클로즈업이 아니다. 반사면이나 명백히 중복된 고정 설비는 없다.",
        "entities": "현우는 참고와 유사한 앳된 동아시아계 남성 얼굴과 검은 머리를 지녔으며, 겉셔츠와 어두운 속옷, 볼의 상처가 보인다. 무릎에는 붕대가 있지만 개에게 물린 상처인지는 판별할 수 없다. 오른쪽에는 밝은 옷을 입은 별도 인물의 머리와 상체가 추가되어 있다. 이는 허용된 현우도, 참고의 로봇 찰리도 아니다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "현우만 보여야 하는 장면에 별도 인물의 머리와 상체를 오른쪽 전경에 추가했다."
        ],
        "physics": "턱 아래에 한 손이 실제로 닿아 있고 손목과 팔이 자연스럽게 이어진다. 다른 팔은 굽힌 무릎 위에 놓여 지지를 받는다. 골반과 발의 접지점은 화면 밖이라 정확한 착석 방식은 확인할 수 없지만, 떠 있거나 불가능한 자세라는 증거는 없다."
       },
       {
        "label": "B",
        "direction": "현우의 고개와 두 눈은 카메라가 아니라 화면 오른쪽 밖의 한 지점을 향한다. 찰리 자체는 보이지 않으므로 대상의 신원을 직접 확인할 수는 없지만, 화면 밖 찰리를 주시한다는 연출에 맞으며 시선 앞 공간도 확보되어 있다.",
        "built_space": "흐린 배경에 푸른 골강판 벽 한 면과 흰 테두리 창 하나, 오른쪽 아래 어두운 통 약 네 개, 우측 끝의 밝은 용기와 녹색 상자가 놓인 작업대 일부가 보인다. 참고 장소의 벽 재질과 외부 세척 공간 구성이 이어진다. 현우의 얼굴과 어깨가 전경을 차지하고 주변은 부분적으로만 남아 클로즈업 요구에 맞는다. 설비 중복이나 불가능한 반사는 보이지 않는다.",
        "entities": "보이는 인물은 현우 한 명이다. 앳된 남성 얼굴, 헝클어진 검은 머리, 얼굴 윤곽과 이목구비가 인물 참고에 가깝다. 양 볼의 아물지 않은 찰과상과 낡은 겉셔츠, 짙은 속옷도 확인된다. 국적은 외형만으로 검증할 수 없다. 다리와 찰리 및 호스는 적절한 클로즈업 밖에 있어 상태를 판단할 수 없으며, 읽을 수 있는 글자도 없다.",
        "hard_violations": [],
        "physics": "한 손의 손바닥이 턱과 아래쪽 볼에 밀착해 머리를 받치고, 손목과 전완은 화면 아래로 자연스럽게 이어진다. 팔꿈치의 받침과 하체는 화면 밖이지만 보이는 접촉과 관절 구조는 현실적이다. 지지 없이 떠 있는 신체나 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "허용되지 않은 인물을 전경에 추가하고 그쪽을 바라보며, 얼굴 클로즈업 대신 무릎까지 포함한 넓은 구도로 핵심 연출을 위반한다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "현우만의 얼굴 클로즈업, 한 손으로 턱을 받친 자세, 화면 밖 찰리 방향의 시선과 컨테이너 외부의 연속성을 충실히 구현하나 확신보다는 생각에 잠긴 표정이 조금 강하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 얼굴과 눈은 화면 오른쪽 전경에 있는 검은 머리 인물을 향한다. 따라서 실제로 보이는 시선 대상은 화면 밖 찰리가 아니라 추가된 인물이다. 그 인물도 현우 쪽으로 고개를 돌리고 있다.",
        "built_space": "푸른 골강판 컨테이너 벽 한 면, 상단에 일부 잘린 창틀 하나, 하단 받침 블록과 흙바닥이 보인다. 현우는 지면 가까이 낮은 자세로 있고 오른쪽 전경의 추가 인물이 화면을 크게 차지한다. 재료와 낮 조명은 장소 참고와 유사하지만, 배경과 무릎까지 드러나는 구도는 요구한 얼굴 클로즈업이 아니다. 반사면이나 명백히 중복된 고정 설비는 없다.",
        "entities": "현우는 참고와 유사한 앳된 동아시아계 남성 얼굴과 검은 머리를 지녔으며, 겉셔츠와 어두운 속옷, 볼의 상처가 보인다. 무릎에는 붕대가 있지만 개에게 물린 상처인지는 판별할 수 없다. 오른쪽에는 밝은 옷을 입은 별도 인물의 머리와 상체가 추가되어 있다. 이는 허용된 현우도, 참고의 로봇 찰리도 아니다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "현우만 보여야 하는 장면에 별도 인물의 머리와 상체를 오른쪽 전경에 추가했다."
        ],
        "physics": "턱 아래에 한 손이 실제로 닿아 있고 손목과 팔이 자연스럽게 이어진다. 다른 팔은 굽힌 무릎 위에 놓여 지지를 받는다. 골반과 발의 접지점은 화면 밖이라 정확한 착석 방식은 확인할 수 없지만, 떠 있거나 불가능한 자세라는 증거는 없다."
       },
       {
        "label": "A",
        "direction": "현우의 고개와 두 눈은 카메라가 아니라 화면 오른쪽 밖의 한 지점을 향한다. 찰리 자체는 보이지 않으므로 대상의 신원을 직접 확인할 수는 없지만, 화면 밖 찰리를 주시한다는 연출에 맞으며 시선 앞 공간도 확보되어 있다.",
        "built_space": "흐린 배경에 푸른 골강판 벽 한 면과 흰 테두리 창 하나, 오른쪽 아래 어두운 통 약 네 개, 우측 끝의 밝은 용기와 녹색 상자가 놓인 작업대 일부가 보인다. 참고 장소의 벽 재질과 외부 세척 공간 구성이 이어진다. 현우의 얼굴과 어깨가 전경을 차지하고 주변은 부분적으로만 남아 클로즈업 요구에 맞는다. 설비 중복이나 불가능한 반사는 보이지 않는다.",
        "entities": "보이는 인물은 현우 한 명이다. 앳된 남성 얼굴, 헝클어진 검은 머리, 얼굴 윤곽과 이목구비가 인물 참고에 가깝다. 양 볼의 아물지 않은 찰과상과 낡은 겉셔츠, 짙은 속옷도 확인된다. 국적은 외형만으로 검증할 수 없다. 다리와 찰리 및 호스는 적절한 클로즈업 밖에 있어 상태를 판단할 수 없으며, 읽을 수 있는 글자도 없다.",
        "hard_violations": [],
        "physics": "한 손의 손바닥이 턱과 아래쪽 볼에 밀착해 머리를 받치고, 손목과 전완은 화면 아래로 자연스럽게 이어진다. 팔꿈치의 받침과 하체는 화면 밖이지만 보이는 접촉과 관절 구조는 현실적이다. 지지 없이 떠 있는 신체나 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.651
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.401
   },
   "violations": {
    "B": [
     "[gemini-pro] 지시문에 없는 추가 인물 등장",
     "[gpt-high] 현우만 보여야 하는 장면에 별도 인물의 머리와 상체를 오른쪽 전경에 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 401
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "오프스크린을 향한 시선 처리와 턱을 괸 자세를 클로즈업으로 정확히 연출했으며, 추가 인물 없이 지시를 잘 따름."
   },
   {
    "label": "B",
    "score": 401,
    "verdict_ko": "지시문에 명시되지 않은 인물이 프레임 안에 등장하여 단독 인물 조건 및 오프스크린 시선 조건을 위반함.  ★위반: [gemini-pro] 지시문에 없는 추가 인물 등장 / [gpt-high] 현우만 보여야 하는 장면에 별도 인물의 머리와 상체를 오른쪽 전경에 추가했다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S16sh3_sel.png",
    "asset_id": "ba2579d1-d28e-4a78-a375-d4a1ffd4382a",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-ec91-730d-bac5-6bf79e9da107",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S16sh3"
  }
 },
 "S16sh6::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:05:28.741163+00:00",
  "fingerprint": "54791e6672049a64c10bf38aed387ef118eb4c06c81a43843b401d2a4756cb47",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S16sh6_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S16sh6_sel.png",
  "source_sha256": "de42131d48c733e9184f2ccc373b496f920ea2eb59d7339a60366fa265547201",
  "file": "S16sh6_cine.png",
  "staged_sha256": "a16c54e13e69ec5c8f2a40fc0f1b119069ba90da1d539309fa9062261dbc1b57",
  "latency_ms": 10266
 },
 "S17sh1::signage": {
  "fp": "d571c1e172ce9549",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S17sh1": {
  "input_fingerprint": "50d9db1104f34188",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 상점 입구에 들어선 무장한 수하 1의 검은 실루엣 전신.\n\nLOCATION (lock): Just inside the junk shop entrance, silhouetted against the brighter daytime doorway with piled appliances farther inside. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Shop entrance (Being crossed by the entering subordinate) — Seen obliquely from inside, with the entrance boundaries surrounding his full figure; used as Frame the arrival and establish the route toward Matsuda without implying a door-leaf position; Silenced rifle (Carried by the subordinate, not firing) — Its outline remains adjacent to his body rather than aimed into the lens; used as Make the armed arrival legible without staging the later shot.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the described dark silhouette through controlled entrance-to-interior tonal separation without specifying an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The shop remains cluttered with appliances and scrap, with a television currently under repair. 수하 1: He has entered the shop and is looking around.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 수하 1 (한국인, 성인 남성, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 상점 입구에 들어선 무장한 수하 1의 검은 실루엣 전신.\n\nLOCATION (lock): Just inside the junk shop entrance, silhouetted against the brighter daytime doorway with piled appliances farther inside. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Shop entrance (Being crossed by the entering subordinate) — Seen obliquely from inside, with the entrance boundaries surrounding his full figure; used as Frame the arrival and establish the route toward Matsuda without implying a door-leaf position; Silenced rifle (Carried by the subordinate, not firing) — Its outline remains adjacent to his body rather than aimed into the lens; used as Make the armed arrival legible without staging the later shot.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the described dark silhouette through controlled entrance-to-interior tonal separation without specifying an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The shop remains cluttered with appliances and scrap, with a television currently under repair. 수하 1: He has entered the shop and is looking around.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 수하 1 (한국인, 성인 남성, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 상점 입구에 들어선 무장한 수하 1의 검은 실루엣 전신.\n\nLOCATION (lock): Just inside the junk shop entrance, silhouetted against the brighter daytime doorway with piled appliances farther inside. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Shop entrance (Being crossed by the entering subordinate) — Seen obliquely from inside, with the entrance boundaries surrounding his full figure; used as Frame the arrival and establish the route toward Matsuda without implying a door-leaf position; Silenced rifle (Carried by the subordinate, not firing) — Its outline remains adjacent to his body rather than aimed into the lens; used as Make the armed arrival legible without staging the later shot.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the described dark silhouette through controlled entrance-to-interior tonal separation without specifying an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The shop remains cluttered with appliances and scrap, with a television currently under repair. 수하 1: He has entered the shop and is looking around.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 수하 1 (한국인, 성인 남성, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "인물의 시선은 화면 왼쪽을 향하고 있으며, 오른손에 들린 총의 총구는 바닥을 향하고 있음.",
    "built_space": "정면에 밝은 야외 빛이 들어오는 상점 출입구가 위치함. 화면 왼쪽에는 레퍼런스 사진과 일치하는 세로줄 블라인드가 쳐진 창문과 그 아래로 브라운관 TV, 흰색 냉장고가 정확히 배치되어 있음.",
    "entities": "수하 1의 전신 실루엣이 묘사되었으며 얼굴 측면 윤곽이 일치함. 쥐고 있는 총기는 얇은 소음기와 특유의 개머리판 등 프롭 레퍼런스의 윤곽과 일치함.",
    "hard_violations": [],
    "physics": "오른발로 바닥을 디디며 상점으로 걸어 들어오는 자연스러운 보행 동작이며, 오른손으로 총기의 손잡이를 안정적으로 쥐고 있음."
   },
   {
    "label": "B",
    "direction": "인물의 시선은 화면 오른쪽을 향하고, 오른손에 들린 총구는 바닥 쪽을 향하고 있음.",
    "built_space": "정면에 유리 출입문이 배치되어 있으나, 레퍼런스에 존재했던 왼쪽 벽면의 창문 구조가 보이지 않으며 가전제품의 쌓인 형태와 위치도 이전 샷과 다름.",
    "entities": "수하 1의 전신 실루엣이 나타남. 그러나 들고 있는 무기가 지정된 권총형 프롭이 아닌 두꺼운 총열의 소총 형태로 묘사됨.",
    "hard_violations": [],
    "physics": "왼발로 바닥을 딛고 안으로 들어서는 보행 자세를 취하고 있으며, 오른손으로 무기를 지탱하고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "요구된 검은 실루엣의 전신 와이드 샷을 완벽하게 구현했으며, 이전 샷의 고물상 배경(왼쪽 창문 및 냉장고)과 소음기 달린 총기의 형태를 매우 충실하게 재현했습니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "실루엣 연출과 구도는 무난하나, 이전 샷에서 확정된 배경 구조(창문 부재 등)를 무시했으며 들고 있는 무기 역시 지정된 레퍼런스와 전혀 다른 소총 형태입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "인물의 시선은 화면 왼쪽을 향하고 있으며, 오른손에 들린 총의 총구는 바닥을 향하고 있음.",
        "built_space": "정면에 밝은 야외 빛이 들어오는 상점 출입구가 위치함. 화면 왼쪽에는 레퍼런스 사진과 일치하는 세로줄 블라인드가 쳐진 창문과 그 아래로 브라운관 TV, 흰색 냉장고가 정확히 배치되어 있음.",
        "entities": "수하 1의 전신 실루엣이 묘사되었으며 얼굴 측면 윤곽이 일치함. 쥐고 있는 총기는 얇은 소음기와 특유의 개머리판 등 프롭 레퍼런스의 윤곽과 일치함.",
        "hard_violations": [],
        "physics": "오른발로 바닥을 디디며 상점으로 걸어 들어오는 자연스러운 보행 동작이며, 오른손으로 총기의 손잡이를 안정적으로 쥐고 있음."
       },
       {
        "label": "B",
        "direction": "인물의 시선은 화면 오른쪽을 향하고, 오른손에 들린 총구는 바닥 쪽을 향하고 있음.",
        "built_space": "정면에 유리 출입문이 배치되어 있으나, 레퍼런스에 존재했던 왼쪽 벽면의 창문 구조가 보이지 않으며 가전제품의 쌓인 형태와 위치도 이전 샷과 다름.",
        "entities": "수하 1의 전신 실루엣이 나타남. 그러나 들고 있는 무기가 지정된 권총형 프롭이 아닌 두꺼운 총열의 소총 형태로 묘사됨.",
        "hard_violations": [],
        "physics": "왼발로 바닥을 딛고 안으로 들어서는 보행 자세를 취하고 있으며, 오른손으로 무기를 지탱하고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "요구된 검은 실루엣의 전신 와이드 샷을 완벽하게 구현했으며, 이전 샷의 고물상 배경(왼쪽 창문 및 냉장고)과 소음기 달린 총기의 형태를 매우 충실하게 재현했습니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "실루엣 연출과 구도는 무난하나, 이전 샷에서 확정된 배경 구조(창문 부재 등)를 무시했으며 들고 있는 무기 역시 지정된 레퍼런스와 전혀 다른 소총 형태입니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "인물의 시선은 화면 왼쪽을 향하고 있으며, 오른손에 들린 총의 총구는 바닥을 향하고 있음.",
        "built_space": "정면에 밝은 야외 빛이 들어오는 상점 출입구가 위치함. 화면 왼쪽에는 레퍼런스 사진과 일치하는 세로줄 블라인드가 쳐진 창문과 그 아래로 브라운관 TV, 흰색 냉장고가 정확히 배치되어 있음.",
        "entities": "수하 1의 전신 실루엣이 묘사되었으며 얼굴 측면 윤곽이 일치함. 쥐고 있는 총기는 얇은 소음기와 특유의 개머리판 등 프롭 레퍼런스의 윤곽과 일치함.",
        "hard_violations": [],
        "physics": "오른발로 바닥을 디디며 상점으로 걸어 들어오는 자연스러운 보행 동작이며, 오른손으로 총기의 손잡이를 안정적으로 쥐고 있음."
       },
       {
        "label": "B",
        "direction": "인물의 시선은 화면 오른쪽을 향하고, 오른손에 들린 총구는 바닥 쪽을 향하고 있음.",
        "built_space": "정면에 유리 출입문이 배치되어 있으나, 레퍼런스에 존재했던 왼쪽 벽면의 창문 구조가 보이지 않으며 가전제품의 쌓인 형태와 위치도 이전 샷과 다름.",
        "entities": "수하 1의 전신 실루엣이 나타남. 그러나 들고 있는 무기가 지정된 권총형 프롭이 아닌 두꺼운 총열의 소총 형태로 묘사됨.",
        "hard_violations": [],
        "physics": "왼발로 바닥을 딛고 안으로 들어서는 보행 자세를 취하고 있으며, 오른손으로 무기를 지탱하고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "전신 진입 동작과 아래로 내린 소음총은 충실하지만, 전술조끼 차림이 인물 참고의 정장과 다르고 입구의 사선 시점도 약하다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "밝은 출입구 안의 검은 전신 실루엣과 주변을 살피는 진입 순간, 낡은 가전 수리점의 연속성이 더 충실하나 정장과 참고 총기의 형태는 재현하지 못했다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "남자는 몸을 실내로 전진시키면서 고개와 시선을 화면 오른쪽 가전 더미 쪽으로 돌린다. 손에 든 총의 소음기는 몸 옆에서 아래쪽 바닥을 향하며 렌즈나 사람을 겨누지 않는다. 발사 징후는 없다.",
        "built_space": "실내 통로에서 출입구 한 곳을 바라본다. 출입구에는 위쪽 채광창 두 칸, 왼쪽 열린 문짝, 오른쪽 유리 구획이 보인다. 남자의 머리부터 신발까지 입구 경계 안에 들어온다. 왼쪽에는 화면이 드러난 브라운관 기기 약 네 대, 오른쪽에는 약 다섯 대와 금속 가전 받침들이 쌓여 있고 중앙 통로는 비어 있다. 낡은 벽과 폐가전의 재질은 참고와 유사하지만 출입구를 거의 정면으로 보아 요구된 사선 관찰은 약하다. 참고 사진에 출입구 전체가 없어 그 구조의 정확한 일치 여부는 확인할 수 없다.",
        "entities": "인물은 짧은 검은 머리의 한국인 성인 남성으로 읽히는 한 명뿐이며 이전 사진의 노인은 없다. 얼굴의 세부 동일성은 역광 때문에 제한적으로만 확인된다. 검은 전술조끼와 작업·전술복 차림은 참고의 남색 정장과 흰 셔츠와 다르다. 소음기가 붙은 긴 총 한 자루는 보이지만 참고의 권총형 몸체와 신축 개머리판을 가진 소형 총기와 형태가 다르다. 폐가전, 브라운관 텔레비전, 노출된 전자 부품은 있으나 어느 텔레비전이 현재 수리 중인지는 명확하지 않다. 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "화면 왼쪽 신발이 바닥을 딛고 체중을 지지하며 반대쪽 발은 무릎을 굽힌 채 앞으로 옮기는 중이다. 정상적인 보행으로 가능한 자세다. 총은 내려온 손이 잡고 있어 공중에 뜨지 않는다. 가전과 부품은 바닥, 받침 가전, 선반 또는 더미에 기대어 지지된다. 인물의 옷과 신체에는 입체적인 음영이 남아 있어 평면 검은 도형처럼 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "남자는 실내로 걸어 들어오면서 고개를 화면 왼쪽의 텔레비전과 폐가전 쪽으로 돌려 살핀다. 총은 몸의 화면 왼쪽에 붙여 들고 있으며 긴 소음기 끝은 앞쪽 바닥으로 향한다. 렌즈를 조준하거나 발사하는 모습이 아니다.",
        "built_space": "출입구 한 곳과 그 왼쪽의 두 칸짜리 큰 창, 오른쪽 유리 구획이 보인다. 남자의 전신은 밝은 출입구 경계에 둘러싸여 있다. 왼쪽 흰 가전 받침 위에 텔레비전 한 대, 오른쪽 흰 가전 위에 텔레비전 한 대와 전화기 한 대가 놓여 있다. 왼쪽에는 의자 한 개와 기울어진 브라운관 기기, 통로 가장자리에는 추가 텔레비전들이 있다. 오른쪽 위에는 창문형 냉방기 한 대가 보인다. 흰 가전, 낡은 브라운관, 전화기와 창의 조합이 참고 장소의 특징을 더 잘 이어 간다. 다만 입구를 보는 각도는 거의 정면이어서 사선 시점 요구는 충분히 강하지 않다.",
        "entities": "짧은 검은 머리의 한국인 성인 남성으로 읽히는 한 명만 등장한다. 참고와 유사한 머리 윤곽은 보이지만 어두운 얼굴로 정확한 동일 인물 여부를 단정하기 어렵다. 검은 상의와 바지에는 참고의 흰 셔츠와 정장 라펠이 드러나지 않는다. 몸 옆의 소음총은 요구된 무장 상태를 보여 주지만 참고 총기보다 길고 일반적인 소총 형태다. 왼쪽 텔레비전 앞의 노출 부품은 수리 작업의 흔적으로 읽힌다. 이전 사진의 노인이나 추가 인물, 판독 가능한 글자는 없다.",
        "hard_violations": [],
        "physics": "화면 왼쪽 발이 실내 바닥을 딛고 있고 다른 발은 뒤꿈치를 들어 다음 걸음을 옮긴다. 체중 이동과 착지 경로가 자연스럽다. 총은 손으로 잡고 아래로 늘어뜨려 안정적으로 지지된다. 텔레비전은 가전 상판이나 바닥에 놓여 있고 기울어진 기기는 의자와 주변 더미에 기대어 있다. 지지 없이 떠 있는 인물이나 물체는 보이지 않는다. 역광 실루엣 안에서도 얼굴과 옷의 부피가 남는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "전신 진입 동작과 아래로 내린 소음총은 충실하지만, 전술조끼 차림이 인물 참고의 정장과 다르고 입구의 사선 시점도 약하다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "밝은 출입구 안의 검은 전신 실루엣과 주변을 살피는 진입 순간, 낡은 가전 수리점의 연속성이 더 충실하나 정장과 참고 총기의 형태는 재현하지 못했다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "남자는 몸을 실내로 전진시키면서 고개와 시선을 화면 오른쪽 가전 더미 쪽으로 돌린다. 손에 든 총의 소음기는 몸 옆에서 아래쪽 바닥을 향하며 렌즈나 사람을 겨누지 않는다. 발사 징후는 없다.",
        "built_space": "실내 통로에서 출입구 한 곳을 바라본다. 출입구에는 위쪽 채광창 두 칸, 왼쪽 열린 문짝, 오른쪽 유리 구획이 보인다. 남자의 머리부터 신발까지 입구 경계 안에 들어온다. 왼쪽에는 화면이 드러난 브라운관 기기 약 네 대, 오른쪽에는 약 다섯 대와 금속 가전 받침들이 쌓여 있고 중앙 통로는 비어 있다. 낡은 벽과 폐가전의 재질은 참고와 유사하지만 출입구를 거의 정면으로 보아 요구된 사선 관찰은 약하다. 참고 사진에 출입구 전체가 없어 그 구조의 정확한 일치 여부는 확인할 수 없다.",
        "entities": "인물은 짧은 검은 머리의 한국인 성인 남성으로 읽히는 한 명뿐이며 이전 사진의 노인은 없다. 얼굴의 세부 동일성은 역광 때문에 제한적으로만 확인된다. 검은 전술조끼와 작업·전술복 차림은 참고의 남색 정장과 흰 셔츠와 다르다. 소음기가 붙은 긴 총 한 자루는 보이지만 참고의 권총형 몸체와 신축 개머리판을 가진 소형 총기와 형태가 다르다. 폐가전, 브라운관 텔레비전, 노출된 전자 부품은 있으나 어느 텔레비전이 현재 수리 중인지는 명확하지 않다. 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "화면 왼쪽 신발이 바닥을 딛고 체중을 지지하며 반대쪽 발은 무릎을 굽힌 채 앞으로 옮기는 중이다. 정상적인 보행으로 가능한 자세다. 총은 내려온 손이 잡고 있어 공중에 뜨지 않는다. 가전과 부품은 바닥, 받침 가전, 선반 또는 더미에 기대어 지지된다. 인물의 옷과 신체에는 입체적인 음영이 남아 있어 평면 검은 도형처럼 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "남자는 실내로 걸어 들어오면서 고개를 화면 왼쪽의 텔레비전과 폐가전 쪽으로 돌려 살핀다. 총은 몸의 화면 왼쪽에 붙여 들고 있으며 긴 소음기 끝은 앞쪽 바닥으로 향한다. 렌즈를 조준하거나 발사하는 모습이 아니다.",
        "built_space": "출입구 한 곳과 그 왼쪽의 두 칸짜리 큰 창, 오른쪽 유리 구획이 보인다. 남자의 전신은 밝은 출입구 경계에 둘러싸여 있다. 왼쪽 흰 가전 받침 위에 텔레비전 한 대, 오른쪽 흰 가전 위에 텔레비전 한 대와 전화기 한 대가 놓여 있다. 왼쪽에는 의자 한 개와 기울어진 브라운관 기기, 통로 가장자리에는 추가 텔레비전들이 있다. 오른쪽 위에는 창문형 냉방기 한 대가 보인다. 흰 가전, 낡은 브라운관, 전화기와 창의 조합이 참고 장소의 특징을 더 잘 이어 간다. 다만 입구를 보는 각도는 거의 정면이어서 사선 시점 요구는 충분히 강하지 않다.",
        "entities": "짧은 검은 머리의 한국인 성인 남성으로 읽히는 한 명만 등장한다. 참고와 유사한 머리 윤곽은 보이지만 어두운 얼굴로 정확한 동일 인물 여부를 단정하기 어렵다. 검은 상의와 바지에는 참고의 흰 셔츠와 정장 라펠이 드러나지 않는다. 몸 옆의 소음총은 요구된 무장 상태를 보여 주지만 참고 총기보다 길고 일반적인 소총 형태다. 왼쪽 텔레비전 앞의 노출 부품은 수리 작업의 흔적으로 읽힌다. 이전 사진의 노인이나 추가 인물, 판독 가능한 글자는 없다.",
        "hard_violations": [],
        "physics": "화면 왼쪽 발이 실내 바닥을 딛고 있고 다른 발은 뒤꿈치를 들어 다음 걸음을 옮긴다. 체중 이동과 착지 경로가 자연스럽다. 총은 손으로 잡고 아래로 늘어뜨려 안정적으로 지지된다. 텔레비전은 가전 상판이나 바닥에 놓여 있고 기울어진 기기는 의자와 주변 더미에 기대어 있다. 지지 없이 떠 있는 인물이나 물체는 보이지 않는다. 역광 실루엣 안에서도 얼굴과 옷의 부피가 남는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.431
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.431
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1431
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "요구된 검은 실루엣의 전신 와이드 샷을 완벽하게 구현했으며, 이전 샷의 고물상 배경(왼쪽 창문 및 냉장고)과 소음기 달린 총기의 형태를 매우 충실하게 재현했습니다."
   },
   {
    "label": "B",
    "score": 1431,
    "verdict_ko": "실루엣 연출과 구도는 무난하나, 이전 샷에서 확정된 배경 구조(창문 부재 등)를 무시했으며 들고 있는 무기 역시 지정된 레퍼런스와 전혀 다른 소총 형태입니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S15sh10_sel.png",
    "asset_id": "797aadcb-72c2-43e0-b382-f0bbeda6fe0f",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 수하 1: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1192369>",
    "asset_id": "a969dd53-900f-4bf4-b80a-e38b77b03fbf",
    "role": "character_ref"
   },
   {
    "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
    "path": "<bytes:842741>",
    "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
    "role": "prop_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-ee55-7725-b689-a99301ea0586",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S15sh10"
  }
 },
 "S17sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:40:39.413777+00:00",
  "fingerprint": "74e0350f441b30684090a7332c060d8993e92986a7daf4a04abd4a7297ac3bf1",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S17sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S17sh1_sel.png",
  "source_sha256": "f5684be85a0b6e5484b3f96e2aa660856e9220458b6fee462b67a0359f7380fa",
  "file": "S17sh1_cine.png",
  "staged_sha256": "c93554e01ecb3002152f9b1c42df81ebd34de0875b2c62310004a9fadcb9d374",
  "latency_ms": 12900
 },
 "S17sh7::signage": {
  "fp": "695a031b3956ab20",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S17sh7": {
  "input_fingerprint": "5b4fda8f13930683",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쓰러진 마츠다를 등진 채 무전기를 입가에 바짝 댄 수하 1의 차가운 옆얼굴.\n\nLOCATION (lock): Beside the repair desk inside the junk shop, with the fallen owner behind it in subdued daytime light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Desk supporting the collapsed Matsuda, behind the subordinate's back in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Desk beneath Matsuda (Supporting Matsuda's collapsed upper body) — The desktop is visible obliquely in the right background behind the subordinate; used as Provide an unmistakable shared-space anchor for the aftermath; Radio (Held at the subordinate's mouth during his report) — Seen from the side, below the nose and beside the lips; used as Identify the reporting action without concealing the cold profile.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained shop ambience with enough tonal separation to distinguish the speaking profile from Matsuda's collapsed body.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 수하 1 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Matsuda is slumped forward over the desk, his collapsed upper body resting on the desktop after being shot. The source does not specify his head's turn, the placement of his arms, or the position of his lower body.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The television repair and surrounding appliance-and-scrap clutter remain in place. 수하 1: He remains inside the shop, holding and using a radio.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 수하 1 right now, so 수하 1's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 수하 1: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 수하 1 (한국인, 성인 남성, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쓰러진 마츠다를 등진 채 무전기를 입가에 바짝 댄 수하 1의 차가운 옆얼굴.\n\nLOCATION (lock): Beside the repair desk inside the junk shop, with the fallen owner behind it in subdued daytime light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Desk supporting the collapsed Matsuda, behind the subordinate's back in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Desk beneath Matsuda (Supporting Matsuda's collapsed upper body) — The desktop is visible obliquely in the right background behind the subordinate; used as Provide an unmistakable shared-space anchor for the aftermath; Radio (Held at the subordinate's mouth during his report) — Seen from the side, below the nose and beside the lips; used as Identify the reporting action without concealing the cold profile.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained shop ambience with enough tonal separation to distinguish the speaking profile from Matsuda's collapsed body.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 수하 1 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Matsuda is slumped forward over the desk, his collapsed upper body resting on the desktop after being shot. The source does not specify his head's turn, the placement of his arms, or the position of his lower body.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The television repair and surrounding appliance-and-scrap clutter remain in place. 수하 1: He remains inside the shop, holding and using a radio.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 수하 1 right now, so 수하 1's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 수하 1: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 수하 1 (한국인, 성인 남성, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쓰러진 마츠다를 등진 채 무전기를 입가에 바짝 댄 수하 1의 차가운 옆얼굴.\n\nLOCATION (lock): Beside the repair desk inside the junk shop, with the fallen owner behind it in subdued daytime light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Desk supporting the collapsed Matsuda, behind the subordinate's back in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Desk beneath Matsuda (Supporting Matsuda's collapsed upper body) — The desktop is visible obliquely in the right background behind the subordinate; used as Provide an unmistakable shared-space anchor for the aftermath; Radio (Held at the subordinate's mouth during his report) — Seen from the side, below the nose and beside the lips; used as Identify the reporting action without concealing the cold profile.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained shop ambience with enough tonal separation to distinguish the speaking profile from Matsuda's collapsed body.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 수하 1 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Matsuda is slumped forward over the desk, his collapsed upper body resting on the desktop after being shot. The source does not specify his head's turn, the placement of his arms, or the position of his lower body.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The television repair and surrounding appliance-and-scrap clutter remain in place. 수하 1: He remains inside the shop, holding and using a radio.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 수하 1 right now, so 수하 1's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 수하 1: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 수하 1 (한국인, 성인 남성, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "수하 1은 앞을 응시하고, 무전기를 입가 주변에 들고 있음.",
    "built_space": "구형 모니터와 TV가 쌓인 고물상 배경이나, 맥락에 맞지 않는 깔끔한 사무용 나무 책상과 의자가 배치됨.",
    "entities": "수하 1의 얼굴은 일치하나 이전 컷에 없던 전술 조끼와 전술 장갑을 착용함. 마츠다는 책상에 엎드려 있음.",
    "hard_violations": [
     "[gemini-pro] 이전 컷과 일치해야 하는 의상 지침 위반 (레퍼런스에 없는 전술 조끼와 장갑 추가)"
    ],
    "physics": "무전기를 쥔 손의 파지법이나 책상에 엎드려 지탱된 마츠다의 자세는 물리적으로 타당함."
   },
   {
    "label": "B",
    "direction": "수하 1은 앞을 향해 시선을 두고 있으며, 무전기를 입가에 바짝 대고 있음.",
    "built_space": "고물상 내부로 뒤편에 구형 TV들이 쌓여 있고, 공구와 부품이 널브러진 작업용 수리 책상이 배경에 배치됨.",
    "entities": "수하 1의 얼굴과 평범한 검은 의상(맨손 포함)이 이전 컷 레퍼런스와 일치함. 쓰러진 마츠다와 무전기도 명확히 묘사됨.",
    "hard_violations": [],
    "physics": "수하 1의 손이 무전기를 자연스럽게 쥐고 있으며, 마츠다의 상체는 책상 위에 체중이 실려 늘어져 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "지정된 클로즈업 구도를 잘 따랐으며, 인물의 의상과 수리용 책상의 디테일이 레퍼런스 및 배경 설정과 훌륭하게 일치합니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "레퍼런스에 없는 전술 장비(조끼, 장갑)를 착용하여 의상 고정 지침을 위반했고, 고물상 수리대 대신 뜬금없는 사무용 책상이 배치되어 몰입을 방해합니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "수하 1은 앞을 향해 시선을 두고 있으며, 무전기를 입가에 바짝 대고 있음.",
        "built_space": "고물상 내부로 뒤편에 구형 TV들이 쌓여 있고, 공구와 부품이 널브러진 작업용 수리 책상이 배경에 배치됨.",
        "entities": "수하 1의 얼굴과 평범한 검은 의상(맨손 포함)이 이전 컷 레퍼런스와 일치함. 쓰러진 마츠다와 무전기도 명확히 묘사됨.",
        "hard_violations": [],
        "physics": "수하 1의 손이 무전기를 자연스럽게 쥐고 있으며, 마츠다의 상체는 책상 위에 체중이 실려 늘어져 있음."
       },
       {
        "label": "A",
        "direction": "수하 1은 앞을 응시하고, 무전기를 입가 주변에 들고 있음.",
        "built_space": "구형 모니터와 TV가 쌓인 고물상 배경이나, 맥락에 맞지 않는 깔끔한 사무용 나무 책상과 의자가 배치됨.",
        "entities": "수하 1의 얼굴은 일치하나 이전 컷에 없던 전술 조끼와 전술 장갑을 착용함. 마츠다는 책상에 엎드려 있음.",
        "hard_violations": [
         "이전 컷과 일치해야 하는 의상 지침 위반 (레퍼런스에 없는 전술 조끼와 장갑 추가)"
        ],
        "physics": "무전기를 쥔 손의 파지법이나 책상에 엎드려 지탱된 마츠다의 자세는 물리적으로 타당함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "지정된 클로즈업 구도를 잘 따랐으며, 인물의 의상과 수리용 책상의 디테일이 레퍼런스 및 배경 설정과 훌륭하게 일치합니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "레퍼런스에 없는 전술 장비(조끼, 장갑)를 착용하여 의상 고정 지침을 위반했고, 고물상 수리대 대신 뜬금없는 사무용 책상이 배치되어 몰입을 방해합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "수하 1은 앞을 향해 시선을 두고 있으며, 무전기를 입가에 바짝 대고 있음.",
        "built_space": "고물상 내부로 뒤편에 구형 TV들이 쌓여 있고, 공구와 부품이 널브러진 작업용 수리 책상이 배경에 배치됨.",
        "entities": "수하 1의 얼굴과 평범한 검은 의상(맨손 포함)이 이전 컷 레퍼런스와 일치함. 쓰러진 마츠다와 무전기도 명확히 묘사됨.",
        "hard_violations": [],
        "physics": "수하 1의 손이 무전기를 자연스럽게 쥐고 있으며, 마츠다의 상체는 책상 위에 체중이 실려 늘어져 있음."
       },
       {
        "label": "A",
        "direction": "수하 1은 앞을 응시하고, 무전기를 입가 주변에 들고 있음.",
        "built_space": "구형 모니터와 TV가 쌓인 고물상 배경이나, 맥락에 맞지 않는 깔끔한 사무용 나무 책상과 의자가 배치됨.",
        "entities": "수하 1의 얼굴은 일치하나 이전 컷에 없던 전술 조끼와 전술 장갑을 착용함. 마츠다는 책상에 엎드려 있음.",
        "hard_violations": [
         "이전 컷과 일치해야 하는 의상 지침 위반 (레퍼런스에 없는 전술 조끼와 장갑 추가)"
        ],
        "physics": "무전기를 쥔 손의 파지법이나 책상에 엎드려 지탱된 마츠다의 자세는 물리적으로 타당함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "상체 전체가 책상에 엎어진 자세는 미흡하지만, 차가운 옆얼굴과 입가의 무전기를 강조한 정확한 클로즈업이 B보다 우선하며 맨손과 검은 옷의 연속성도 낫다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "마츠다의 엎어진 자세와 책상 지지는 더 정확하지만, 가슴과 책상 전면까지 보이는 넓은 구도가 지정된 클로즈업을 벗어나고 장갑도 이전 모습과 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "수하의 얼굴과 시선은 화면 오른쪽을 향하며 카메라나 마츠다를 직접 바라보지 않는다. 무전기 상단은 코 아래 입술 바로 옆에 있고 안테나는 위로 향한다. 마츠다는 수하보다 깊은 오른쪽 배경에 있으나, 수하가 그에게 등을 완전히 돌렸는지는 이 측면 각도만으로 명확하지 않다.",
        "built_space": "수하의 얼굴과 손이 전경을 크게 차지하고, 수리 책상 하나의 상판이 오른쪽 배경에 비스듬히 보인다. 뒤에는 브라운관 텔레비전 두 대, 오른쪽 부품 서랍장 하나, 상판의 공구와 부품이 보인다. 낡은 수리점의 재질과 낮빛은 이어지지만, 이전 사진의 출입문과 가전 배치를 그대로 확인할 수 있는 구도는 아니다.",
        "entities": "짧은 검은 머리의 성인 동아시아계 남성은 참조 인물과 얼굴 및 머리 형태가 대체로 닮았고, 이전 장면처럼 검은 옷과 맨손이 보인다. 그 손이 휴대용 무전기를 쥐고 있다. 배경에는 회색 머리와 회갈색 작업복의 고령 남성이 마츠다로 제시된다. 마츠다의 외모를 고정할 참조는 없으며 추가 인물이나 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "무전기는 수하의 손가락과 손바닥으로 지지된다. 마츠다의 옆머리는 상판에 닿아 있고 몸통은 책상 가장자리 바깥으로 아래쪽에 이어진다. 공중에 떠 있는 몸은 아니지만, 상체 전체가 상판에 앞으로 엎어져 받쳐진 정해진 자세보다는 머리를 책상에 기댄 자세에 가깝다. 보이는 공구와 부품은 상판에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "수하의 얼굴과 시선은 화면 오른쪽을 향한다. 무전기는 입술 가까이 세워져 있고 안테나는 오른쪽 위로 기울어져 있다. 스피커와 버튼 면이 카메라 쪽으로 상당 부분 노출되지만 입 역시 스피커 가까이에 있다. 마츠다는 오른쪽 뒤에 있으며, 수하가 마츠다를 바라보는 모습은 아니다.",
        "built_space": "왼쪽 전경에 수하의 머리부터 가슴까지, 오른쪽에는 책상 하나의 상판과 전면이 크게 보인다. 마츠다는 책상 뒤 의자에 앉아 앞으로 엎어져 있다. 배경에는 브라운관 텔레비전 다섯 대, 큰 창 여러 면, 해체된 가전과 금속 외함이 보인다. 수리점의 분위기는 유지하지만 창과 텔레비전 배열의 연속성은 확실하지 않으며, 지정된 클로즈업보다 공간을 넓게 보여 준다.",
        "entities": "수하는 짧은 검은 머리와 검은 옷을 입은 성인 동아시아계 남성으로, 얼굴은 참조와 대체로 유사하다. 다만 무전기를 쥔 손에 손가락이 드러나는 장갑이 추가되어 이전 사진의 맨손과 다르다. 검은 머리와 어두운 작업복의 남성이 마츠다로 제시되며, 외모 자체는 주어진 마츠다 설정에 어긋나지 않는다. 추가 인물이나 확실히 판독되는 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "무전기는 장갑 낀 손으로 확실하게 지지된다. 마츠다는 의자에 하체를 두고 앞으로 접혀 있으며 상체와 팔이 책상에 얹혀 있고, 머리도 상판과 팔 쪽으로 떨어져 있다. 들린 손이나 공중에 떠 있는 신체 부위는 보이지 않아 정해진 무기력한 엎드림에 A보다 가깝다. 가전과 부품도 책상이나 다른 가전 위에 놓여 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "상체 전체가 책상에 엎어진 자세는 미흡하지만, 차가운 옆얼굴과 입가의 무전기를 강조한 정확한 클로즈업이 B보다 우선하며 맨손과 검은 옷의 연속성도 낫다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "마츠다의 엎어진 자세와 책상 지지는 더 정확하지만, 가슴과 책상 전면까지 보이는 넓은 구도가 지정된 클로즈업을 벗어나고 장갑도 이전 모습과 다르다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "수하의 얼굴과 시선은 화면 오른쪽을 향하며 카메라나 마츠다를 직접 바라보지 않는다. 무전기 상단은 코 아래 입술 바로 옆에 있고 안테나는 위로 향한다. 마츠다는 수하보다 깊은 오른쪽 배경에 있으나, 수하가 그에게 등을 완전히 돌렸는지는 이 측면 각도만으로 명확하지 않다.",
        "built_space": "수하의 얼굴과 손이 전경을 크게 차지하고, 수리 책상 하나의 상판이 오른쪽 배경에 비스듬히 보인다. 뒤에는 브라운관 텔레비전 두 대, 오른쪽 부품 서랍장 하나, 상판의 공구와 부품이 보인다. 낡은 수리점의 재질과 낮빛은 이어지지만, 이전 사진의 출입문과 가전 배치를 그대로 확인할 수 있는 구도는 아니다.",
        "entities": "짧은 검은 머리의 성인 동아시아계 남성은 참조 인물과 얼굴 및 머리 형태가 대체로 닮았고, 이전 장면처럼 검은 옷과 맨손이 보인다. 그 손이 휴대용 무전기를 쥐고 있다. 배경에는 회색 머리와 회갈색 작업복의 고령 남성이 마츠다로 제시된다. 마츠다의 외모를 고정할 참조는 없으며 추가 인물이나 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "무전기는 수하의 손가락과 손바닥으로 지지된다. 마츠다의 옆머리는 상판에 닿아 있고 몸통은 책상 가장자리 바깥으로 아래쪽에 이어진다. 공중에 떠 있는 몸은 아니지만, 상체 전체가 상판에 앞으로 엎어져 받쳐진 정해진 자세보다는 머리를 책상에 기댄 자세에 가깝다. 보이는 공구와 부품은 상판에 놓여 있다."
       },
       {
        "label": "A",
        "direction": "수하의 얼굴과 시선은 화면 오른쪽을 향한다. 무전기는 입술 가까이 세워져 있고 안테나는 오른쪽 위로 기울어져 있다. 스피커와 버튼 면이 카메라 쪽으로 상당 부분 노출되지만 입 역시 스피커 가까이에 있다. 마츠다는 오른쪽 뒤에 있으며, 수하가 마츠다를 바라보는 모습은 아니다.",
        "built_space": "왼쪽 전경에 수하의 머리부터 가슴까지, 오른쪽에는 책상 하나의 상판과 전면이 크게 보인다. 마츠다는 책상 뒤 의자에 앉아 앞으로 엎어져 있다. 배경에는 브라운관 텔레비전 다섯 대, 큰 창 여러 면, 해체된 가전과 금속 외함이 보인다. 수리점의 분위기는 유지하지만 창과 텔레비전 배열의 연속성은 확실하지 않으며, 지정된 클로즈업보다 공간을 넓게 보여 준다.",
        "entities": "수하는 짧은 검은 머리와 검은 옷을 입은 성인 동아시아계 남성으로, 얼굴은 참조와 대체로 유사하다. 다만 무전기를 쥔 손에 손가락이 드러나는 장갑이 추가되어 이전 사진의 맨손과 다르다. 검은 머리와 어두운 작업복의 남성이 마츠다로 제시되며, 외모 자체는 주어진 마츠다 설정에 어긋나지 않는다. 추가 인물이나 확실히 판독되는 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "무전기는 장갑 낀 손으로 확실하게 지지된다. 마츠다는 의자에 하체를 두고 앞으로 접혀 있으며 상체와 팔이 책상에 얹혀 있고, 머리도 상판과 팔 쪽으로 떨어져 있다. 들린 손이나 공중에 떠 있는 신체 부위는 보이지 않아 정해진 무기력한 엎드림에 A보다 가깝다. 가전과 부품도 책상이나 다른 가전 위에 놓여 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.232,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.982,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 이전 컷과 일치해야 하는 의상 지침 위반 (레퍼런스에 없는 전술 조끼와 장갑 추가)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 982
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "지정된 클로즈업 구도를 잘 따랐으며, 인물의 의상과 수리용 책상의 디테일이 레퍼런스 및 배경 설정과 훌륭하게 일치합니다."
   },
   {
    "label": "A",
    "score": 982,
    "verdict_ko": "레퍼런스에 없는 전술 장비(조끼, 장갑)를 착용하여 의상 고정 지침을 위반했고, 고물상 수리대 대신 뜬금없는 사무용 책상이 배치되어 몰입을 방해합니다.  ★위반: [gemini-pro] 이전 컷과 일치해야 하는 의상 지침 위반 (레퍼런스에 없는 전술 조끼와 장갑 추가)"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 수하 1 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S17sh1_sel.png",
    "asset_id": "a3f8b229-6861-4f48-9fe0-7d6f9ff5bcbc",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 수하 1: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1192369>",
    "asset_id": "a969dd53-900f-4bf4-b80a-e38b77b03fbf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-f01d-79c2-80a5-b2ef5c178c26",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S17sh1"
  }
 },
 "S17sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:41:54.181278+00:00",
  "fingerprint": "aff1252aba3b2e39f82baadbe1e538b8e359f557e2241b6dd0c5b4ca015572e9",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S17sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S17sh7_sel.png",
  "source_sha256": "f83573c913e8b816062e5f57dfd6ecce39679a57f158cc7b2aab9d62c8592cf1",
  "file": "S17sh7_cine.png",
  "staged_sha256": "7dc8dde2286083c583c861361a643f0880b8abd96b846ce424ee0e7b6875232d",
  "latency_ms": 11590
 },
 "S18sh1::signage": {
  "fp": "9ffc554368356fea",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::a44191158dcd7174": {
  "subjects": [],
  "subject_text": "박철진의 민병대 사무실\n책상과 의자, 스피커폰이 놓인 사무실. 책상 맞은편에 빈 공간이 있고 그 뒤로 벽면이 드러나 있다.",
  "identity": "canonical",
  "scope_id": "L171",
  "scope_role": "location_interior",
  "scope_sha": "ecbb2eee6c55995e"
 },
 "S18sh1::bgfirst_bg": {
  "input_fingerprint": "4944a65c5e16b663",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 책상 위의 스피커폰을 향해 몸을 기울인 채 능글맞게 웃고 있는 박철진의 상체.\n\nLOCATION (lock): At the desk inside a militia commander's office, beside a speakerphone in ordinary daytime interior light.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Desk (Supporting the speakerphone during the call) — The near edge runs obliquely across the bottom of the composition; used as Connect the leaning torso to the call while preserving a modest foreground footprint; Speakerphone (In use for the ongoing call) — Seen obliquely from above on the desktop beneath Park's face; used as Supply the visible focus of his speech and downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral office ambience and controlled tonal contrast, keeping the smile and speakerphone readable without introducing a specific fixture or color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 책상 위의 스피커폰을 향해 몸을 기울인 채 능글맞게 웃고 있는 박철진의 상체.\n\nLOCATION (lock): At the desk inside a militia commander's office, beside a speakerphone in ordinary daytime interior light.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Desk (Supporting the speakerphone during the call) — The near edge runs obliquely across the bottom of the composition; used as Connect the leaning torso to the call while preserving a modest foreground footprint; Speakerphone (In use for the ongoing call) — Seen obliquely from above on the desktop beneath Park's face; used as Supply the visible focus of his speech and downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral office ambience and controlled tonal contrast, keeping the smile and speakerphone readable without introducing a specific fixture or color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S18sh1__bgfirst_bg.png",
  "asset_id": "5a3734d6-9057-4f50-afa3-4307c7c976a2",
  "input_asset_ids": [
   "86ca7f7a-7c3a-4769-b837-fca22d49aac1",
   "64c1729a-dc88-4c0d-aef7-ac2fc3c15d65"
  ]
 },
 "S18sh1": {
  "input_fingerprint": "df00f0d5d4ed82fa",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 책상 위의 스피커폰을 향해 몸을 기울인 채 능글맞게 웃고 있는 박철진의 상체.\n\nLOCATION (lock): At the desk inside a militia commander's office, beside a speakerphone in ordinary daytime interior light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Desk (Supporting the speakerphone during the call) — The near edge runs obliquely across the bottom of the composition; used as Connect the leaning torso to the call while preserving a modest foreground footprint; Speakerphone (In use for the ongoing call) — Seen obliquely from above on the desktop beneath Park's face; used as Supply the visible focus of his speech and downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral office ambience and controlled tonal contrast, keeping the smile and speakerphone readable without introducing a specific fixture or color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The speakerphone is in use in the militia office. 박철진: He has a knife in hand, playing with it during the speakerphone conversation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 책상 위의 스피커폰을 향해 몸을 기울인 채 능글맞게 웃고 있는 박철진의 상체.\n\nLOCATION (lock): At the desk inside a militia commander's office, beside a speakerphone in ordinary daytime interior light. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Desk (Supporting the speakerphone during the call) — The near edge runs obliquely across the bottom of the composition; used as Connect the leaning torso to the call while preserving a modest foreground footprint; Speakerphone (In use for the ongoing call) — Seen obliquely from above on the desktop beneath Park's face; used as Supply the visible focus of his speech and downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral office ambience and controlled tonal contrast, keeping the smile and speakerphone readable without introducing a specific fixture or color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The speakerphone is in use in the militia office. 박철진: He has a knife in hand, playing with it during the speakerphone conversation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 책상 위의 스피커폰을 향해 몸을 기울인 채 능글맞게 웃고 있는 박철진의 상체.\n\nLOCATION (lock): At the desk inside a militia commander's office, beside a speakerphone in ordinary daytime interior light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Desk (Supporting the speakerphone during the call) — The near edge runs obliquely across the bottom of the composition; used as Connect the leaning torso to the call while preserving a modest foreground footprint; Speakerphone (In use for the ongoing call) — Seen obliquely from above on the desktop beneath Park's face; used as Supply the visible focus of his speech and downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral office ambience and controlled tonal contrast, keeping the smile and speakerphone readable without introducing a specific fixture or color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The speakerphone is in use in the militia office. 박철진: He has a knife in hand, playing with it during the speakerphone conversation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S18sh1__bgfirst_bg.png",
     "asset_id": "5a3734d6-9057-4f50-afa3-4307c7c976a2",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S18sh1.png",
     "asset_id": "86ca7f7a-7c3a-4769-b837-fca22d49aac1",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1401722>",
     "asset_id": "fee7383c-fb61-4b3a-ba7c-79f2555de00b",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L171B01.png",
     "asset_id": "64c1729a-dc88-4c0d-aef7-ac2fc3c15d65",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1401722>",
     "asset_id": "fee7383c-fb61-4b3a-ba7c-79f2555de00b",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "시선은 책상 위의 스피커폰을 향해 아래로 고정되어 있음.",
    "built_space": "창문, 파일 캐비닛, 선반 등 참조 이미지의 사무실 구조와 가구 배치가 일치하며, 인물은 책상 뒤에 위치함.",
    "entities": "박철진(참조 이미지의 얼굴, 헤어스타일, 네이비 정장과 넥타이 일치), 스피커폰, 단검 모두 명확히 존재함.",
    "hard_violations": [],
    "physics": "오른손으로 단검을 안정적으로 쥐고 있으며, 양팔은 책상에 지지된 채 몸을 기울인 자세가 물리적으로 자연스러움."
   },
   {
    "label": "B",
    "direction": "시선은 책상 위의 거대한 스피커폰을 향해 있음.",
    "built_space": "사무실 내부가 보이나 배경 요소가 다소 생략되었고, 인물은 책상 위에 몸을 기대고 있음.",
    "entities": "박철진의 얼굴과 헤어스타일은 일치하나 의상(갈색 재킷)이 참조 이미지와 불일치함. 단검은 존재하나 스피커폰이 비정상적으로 큼.",
    "hard_violations": [
     "[gemini-pro] 스피커폰의 크기가 인물이나 책상 비율에 비해 물리적으로 불가능할 정도로 거대하게 왜곡됨."
    ],
    "physics": "오른손으로 단검을 쥐고 있고 몸의 하중은 책상에 지지되어 있으나, 스피커폰의 스케일 오류로 인해 장면의 물리적 현실성이 무너짐."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 정장 의상과 사무실 배경을 충실히 반영하고, 스피커폰을 향해 기울인 자세와 단검을 든 동작을 자연스럽게 연출함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "캐릭터 레퍼런스의 의상(정장)을 따르지 않았으며, 스피커폰의 크기가 사람에 비해 물리적으로 비정상적으로 크게 왜곡됨."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 책상 위의 스피커폰을 향해 아래로 고정되어 있음.",
        "built_space": "창문, 파일 캐비닛, 선반 등 참조 이미지의 사무실 구조와 가구 배치가 일치하며, 인물은 책상 뒤에 위치함.",
        "entities": "박철진(참조 이미지의 얼굴, 헤어스타일, 네이비 정장과 넥타이 일치), 스피커폰, 단검 모두 명확히 존재함.",
        "hard_violations": [],
        "physics": "오른손으로 단검을 안정적으로 쥐고 있으며, 양팔은 책상에 지지된 채 몸을 기울인 자세가 물리적으로 자연스러움."
       },
       {
        "label": "B",
        "direction": "시선은 책상 위의 거대한 스피커폰을 향해 있음.",
        "built_space": "사무실 내부가 보이나 배경 요소가 다소 생략되었고, 인물은 책상 위에 몸을 기대고 있음.",
        "entities": "박철진의 얼굴과 헤어스타일은 일치하나 의상(갈색 재킷)이 참조 이미지와 불일치함. 단검은 존재하나 스피커폰이 비정상적으로 큼.",
        "hard_violations": [
         "스피커폰의 크기가 인물이나 책상 비율에 비해 물리적으로 불가능할 정도로 거대하게 왜곡됨."
        ],
        "physics": "오른손으로 단검을 쥐고 있고 몸의 하중은 책상에 지지되어 있으나, 스피커폰의 스케일 오류로 인해 장면의 물리적 현실성이 무너짐."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 정장 의상과 사무실 배경을 충실히 반영하고, 스피커폰을 향해 기울인 자세와 단검을 든 동작을 자연스럽게 연출함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "캐릭터 레퍼런스의 의상(정장)을 따르지 않았으며, 스피커폰의 크기가 사람에 비해 물리적으로 비정상적으로 크게 왜곡됨."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 책상 위의 스피커폰을 향해 아래로 고정되어 있음.",
        "built_space": "창문, 파일 캐비닛, 선반 등 참조 이미지의 사무실 구조와 가구 배치가 일치하며, 인물은 책상 뒤에 위치함.",
        "entities": "박철진(참조 이미지의 얼굴, 헤어스타일, 네이비 정장과 넥타이 일치), 스피커폰, 단검 모두 명확히 존재함.",
        "hard_violations": [],
        "physics": "오른손으로 단검을 안정적으로 쥐고 있으며, 양팔은 책상에 지지된 채 몸을 기울인 자세가 물리적으로 자연스러움."
       },
       {
        "label": "B",
        "direction": "시선은 책상 위의 거대한 스피커폰을 향해 있음.",
        "built_space": "사무실 내부가 보이나 배경 요소가 다소 생략되었고, 인물은 책상 위에 몸을 기대고 있음.",
        "entities": "박철진의 얼굴과 헤어스타일은 일치하나 의상(갈색 재킷)이 참조 이미지와 불일치함. 단검은 존재하나 스피커폰이 비정상적으로 큼.",
        "hard_violations": [
         "스피커폰의 크기가 인물이나 책상 비율에 비해 물리적으로 불가능할 정도로 거대하게 왜곡됨."
        ],
        "physics": "오른손으로 단검을 쥐고 있고 몸의 하중은 책상에 지지되어 있으나, 스피커폰의 스케일 오류로 인해 장면의 물리적 현실성이 무너짐."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "스피커폰 쪽으로 상체를 깊이 기울인 능글맞은 웃음과 기기의 사용 방향이 충실하지만, 야전 재킷은 인물 참조의 정장과 다르다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "정장과 인물 인상, 아래로 향한 시선은 잘 맞지만, 스피커폰 조작면이 사용자 대신 카메라를 향해 명시된 소품 방향을 어긴다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴과 상체가 책상 위 스피커폰을 향해 앞으로 숙여져 있고, 시선도 화면 오른쪽 아래의 기기 쪽으로 내려간다. 오른손의 칼끝은 화면 왼쪽 빈 공간을 향하며 사람을 겨누지 않는다. 스피커폰은 카메라에 뒷면을 보이고 사용자 쪽에 조작면이 놓이는 방향이다.",
        "built_space": "목재 책상 하나가 전경을 차지하고 그 위에 스피커폰 하나가 있다. 왼쪽 배경에 지도 하나, 높은 녹색 서류함 하나, 낮은 회색 수납장 일부와 화분이 보이며, 뒤벽에는 걸린 칼 하나가 있다. 흰색·회청색 이중 도색 벽과 왼쪽에서 들어오는 낮빛은 장소 참조와 부합한다. 박철진은 책상 뒤에서 몸을 내밀고 있다. 상체 중심의 미디엄 구도이지만 책상 가까운 모서리의 사선은 약하고 전경 상판 비중이 다소 크다.",
        "entities": "보이는 사람은 박철진에 해당하는 중년 동아시아계 남성 한 명뿐이다. 짧은 검은 머리, 얼굴 윤곽과 작은 귀걸이는 인물 참조와 대체로 일치한다. 다만 남색 정장·흰 셔츠·줄무늬 넥타이 대신 회갈색 야전 재킷과 어두운 셔츠를 입었다. 오른손에는 실제 접이식 칼이 있고 책상에는 검은 회의용 스피커폰이 있다. 웃음은 능글맞은 표정으로 읽히며, 뚜렷하게 읽히는 문구나 오버레이는 없다.",
        "hard_violations": [],
        "physics": "왼쪽 팔꿈치와 전완이 책상에 닿아 앞으로 기울인 상체를 받치고, 오른쪽 전완과 손도 상판 가까이 놓여 있다. 칼 손잡이는 오른손 손가락으로 잡혀 있으며 스피커폰과 서류, 필기구 통은 상판에 받쳐져 있다. 보이는 범위에서 무지지 부유나 불가능한 관절 배치는 없다."
       },
       {
        "label": "B",
        "direction": "박철진의 얼굴과 시선은 화면 왼쪽 아래의 스피커폰 쪽으로 향한다. 오른손으로 든 칼은 왼쪽 위를 향하고 별도의 공격 대상은 없다. 그러나 스피커폰의 키패드와 경사진 조작면이 사용자에게서 멀어지는 카메라 쪽을 향하므로, 소품의 기능면을 사용자의 눈 쪽에 두라는 지시와 맞지 않는다.",
        "built_space": "목재 책상 하나, 회의용 스피커폰 하나, 박철진 뒤의 사무용 의자 하나가 보인다. 왼쪽에는 큰 창, 낮은 수납 선반, 회색 서류함과 그 뒤의 녹색 서류함이 있고, 오른쪽에는 게시판 하나가 있다. 뒤벽의 걸린 칼 하나와 이중 도색 벽도 유지된다. 참조와 같은 사무실로 읽히지만 회색 서류함의 높이와 배치는 달라졌다. 인물은 책상 뒤 의자에 앉아 있으며 상체 미디엄 구도와 하단의 비스듬한 책상 모서리는 요구에 가깝다.",
        "entities": "등장인물은 중년 동아시아계 남성 한 명이다. 짧은 검은 머리, 얼굴 형태, 귀걸이, 남색 정장, 흰 셔츠와 사선 줄무늬 넥타이가 인물 참조에 잘 맞는다. 오른손에 접이식 칼을 들고 있고 책상 위에는 회의용 스피커폰이 있다. 입꼬리를 올린 웃음은 능글맞게 읽힌다. 별도의 인물이나 자막은 없다.",
        "hard_violations": [],
        "physics": "의자 등받이와 팔걸이가 보이며 앉아서 앞으로 숙인 자세가 자연스럽다. 왼손과 전완은 책상에 놓여 있고, 들어 올린 오른손이 칼 손잡이를 확실히 쥐고 있다. 스피커폰, 서류 더미, 필기구 통과 상자는 모두 책상 위에 놓여 있다. 떠 있는 물체나 물리적으로 불가능한 신체 자세는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "스피커폰 쪽으로 상체를 깊이 기울인 능글맞은 웃음과 기기의 사용 방향이 충실하지만, 야전 재킷은 인물 참조의 정장과 다르다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "정장과 인물 인상, 아래로 향한 시선은 잘 맞지만, 스피커폰 조작면이 사용자 대신 카메라를 향해 명시된 소품 방향을 어긴다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴과 상체가 책상 위 스피커폰을 향해 앞으로 숙여져 있고, 시선도 화면 오른쪽 아래의 기기 쪽으로 내려간다. 오른손의 칼끝은 화면 왼쪽 빈 공간을 향하며 사람을 겨누지 않는다. 스피커폰은 카메라에 뒷면을 보이고 사용자 쪽에 조작면이 놓이는 방향이다.",
        "built_space": "목재 책상 하나가 전경을 차지하고 그 위에 스피커폰 하나가 있다. 왼쪽 배경에 지도 하나, 높은 녹색 서류함 하나, 낮은 회색 수납장 일부와 화분이 보이며, 뒤벽에는 걸린 칼 하나가 있다. 흰색·회청색 이중 도색 벽과 왼쪽에서 들어오는 낮빛은 장소 참조와 부합한다. 박철진은 책상 뒤에서 몸을 내밀고 있다. 상체 중심의 미디엄 구도이지만 책상 가까운 모서리의 사선은 약하고 전경 상판 비중이 다소 크다.",
        "entities": "보이는 사람은 박철진에 해당하는 중년 동아시아계 남성 한 명뿐이다. 짧은 검은 머리, 얼굴 윤곽과 작은 귀걸이는 인물 참조와 대체로 일치한다. 다만 남색 정장·흰 셔츠·줄무늬 넥타이 대신 회갈색 야전 재킷과 어두운 셔츠를 입었다. 오른손에는 실제 접이식 칼이 있고 책상에는 검은 회의용 스피커폰이 있다. 웃음은 능글맞은 표정으로 읽히며, 뚜렷하게 읽히는 문구나 오버레이는 없다.",
        "hard_violations": [],
        "physics": "왼쪽 팔꿈치와 전완이 책상에 닿아 앞으로 기울인 상체를 받치고, 오른쪽 전완과 손도 상판 가까이 놓여 있다. 칼 손잡이는 오른손 손가락으로 잡혀 있으며 스피커폰과 서류, 필기구 통은 상판에 받쳐져 있다. 보이는 범위에서 무지지 부유나 불가능한 관절 배치는 없다."
       },
       {
        "label": "A",
        "direction": "박철진의 얼굴과 시선은 화면 왼쪽 아래의 스피커폰 쪽으로 향한다. 오른손으로 든 칼은 왼쪽 위를 향하고 별도의 공격 대상은 없다. 그러나 스피커폰의 키패드와 경사진 조작면이 사용자에게서 멀어지는 카메라 쪽을 향하므로, 소품의 기능면을 사용자의 눈 쪽에 두라는 지시와 맞지 않는다.",
        "built_space": "목재 책상 하나, 회의용 스피커폰 하나, 박철진 뒤의 사무용 의자 하나가 보인다. 왼쪽에는 큰 창, 낮은 수납 선반, 회색 서류함과 그 뒤의 녹색 서류함이 있고, 오른쪽에는 게시판 하나가 있다. 뒤벽의 걸린 칼 하나와 이중 도색 벽도 유지된다. 참조와 같은 사무실로 읽히지만 회색 서류함의 높이와 배치는 달라졌다. 인물은 책상 뒤 의자에 앉아 있으며 상체 미디엄 구도와 하단의 비스듬한 책상 모서리는 요구에 가깝다.",
        "entities": "등장인물은 중년 동아시아계 남성 한 명이다. 짧은 검은 머리, 얼굴 형태, 귀걸이, 남색 정장, 흰 셔츠와 사선 줄무늬 넥타이가 인물 참조에 잘 맞는다. 오른손에 접이식 칼을 들고 있고 책상 위에는 회의용 스피커폰이 있다. 입꼬리를 올린 웃음은 능글맞게 읽힌다. 별도의 인물이나 자막은 없다.",
        "hard_violations": [],
        "physics": "의자 등받이와 팔걸이가 보이며 앉아서 앞으로 숙인 자세가 자연스럽다. 왼손과 전완은 책상에 놓여 있고, 들어 올린 오른손이 칼 손잡이를 확실히 쥐고 있다. 스피커폰, 서류 더미, 필기구 통과 상자는 모두 책상 위에 놓여 있다. 떠 있는 물체나 물리적으로 불가능한 신체 자세는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.75,
    "B": 1.429
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.179
   },
   "violations": {
    "B": [
     "[gemini-pro] 스피커폰의 크기가 인물이나 책상 비율에 비해 물리적으로 불가능할 정도로 거대하게 왜곡됨."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1750,
   "B": 1179
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "지정된 정장 의상과 사무실 배경을 충실히 반영하고, 스피커폰을 향해 기울인 자세와 단검을 든 동작을 자연스럽게 연출함."
   },
   {
    "label": "B",
    "score": 1179,
    "verdict_ko": "캐릭터 레퍼런스의 의상(정장)을 따르지 않았으며, 스피커폰의 크기가 사람에 비해 물리적으로 비정상적으로 크게 왜곡됨.  ★위반: [gemini-pro] 스피커폰의 크기가 인물이나 책상 비율에 비해 물리적으로 불가능할 정도로 거대하게 왜곡됨."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L171B01.png",
    "asset_id": "64c1729a-dc88-4c0d-aef7-ac2fc3c15d65",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1401722>",
    "asset_id": "fee7383c-fb61-4b3a-ba7c-79f2555de00b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-f1db-70dd-acd7-3ff7a4f21269",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S18sh1__bgfirst_bg.png",
   "bg_asset_id": "5a3734d6-9057-4f50-afa3-4307c7c976a2",
   "bg_record_key": "S18sh1::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S18sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:43:32.391318+00:00",
  "fingerprint": "c68043aeca5374ae00d70fcf9d56e519bdcf78c20d4e7f1059ff10d45e44ebaf",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S18sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S18sh1_sel.png",
  "source_sha256": "48d4a2bb10be99b2eeb8824c6657e28745272cbbb6b65bfb6bece18e65f34a23",
  "file": "S18sh1_cine.png",
  "staged_sha256": "343037171e44a1cf84a8258b006c15c77583938906bc3d2155d34236aa38d981",
  "latency_ms": 13398
 },
 "S18sh5::signage": {
  "fp": "fa92ba9eec020c5a",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S18sh5": {
  "input_fingerprint": "f9f2fca11bf9390a",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 벽에 박힌 단검 바로 옆에서 몸을 움츠린 채 두 눈을 질끈 감은 윤성찬의 붉어진 얼굴.\n\nLOCATION (lock): Inside the militia commander's office, directly beside the wall struck by a thrown dagger, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Wall section with the embedded dagger immediately beside and behind 윤성찬's head in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Wall behind 윤성찬 (The thrown dagger is embedded beside him) — An adjacent section is visible behind the right side of his head; used as Make the near miss spatially legible without placing the weapon in front of his face; Embedded dagger (Lodged in the wall after being thrown) — The exposed handle projects from the wall and is seen obliquely beside, not overlapping, his head; used as Provide a small but readable cause for his recoil.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain neutral office illumination with restrained contrast, preserving the explicitly flushed skin without turning it into a colored lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The thrown knife remains lodged in the office wall; the speakerphone remains in place. 윤성찬: He remains standing with his handkerchief, his face flushed from the fright.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 벽에 박힌 단검 바로 옆에서 몸을 움츠린 채 두 눈을 질끈 감은 윤성찬의 붉어진 얼굴.\n\nLOCATION (lock): Inside the militia commander's office, directly beside the wall struck by a thrown dagger, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Wall section with the embedded dagger immediately beside and behind 윤성찬's head in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Wall behind 윤성찬 (The thrown dagger is embedded beside him) — An adjacent section is visible behind the right side of his head; used as Make the near miss spatially legible without placing the weapon in front of his face; Embedded dagger (Lodged in the wall after being thrown) — The exposed handle projects from the wall and is seen obliquely beside, not overlapping, his head; used as Provide a small but readable cause for his recoil.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain neutral office illumination with restrained contrast, preserving the explicitly flushed skin without turning it into a colored lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The thrown knife remains lodged in the office wall; the speakerphone remains in place. 윤성찬: He remains standing with his handkerchief, his face flushed from the fright.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 벽에 박힌 단검 바로 옆에서 몸을 움츠린 채 두 눈을 질끈 감은 윤성찬의 붉어진 얼굴.\n\nLOCATION (lock): Inside the militia commander's office, directly beside the wall struck by a thrown dagger, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Wall section with the embedded dagger immediately beside and behind 윤성찬's head in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Wall behind 윤성찬 (The thrown dagger is embedded beside him) — An adjacent section is visible behind the right side of his head; used as Make the near miss spatially legible without placing the weapon in front of his face; Embedded dagger (Lodged in the wall after being thrown) — The exposed handle projects from the wall and is seen obliquely beside, not overlapping, his head; used as Provide a small but readable cause for his recoil.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain neutral office illumination with restrained contrast, preserving the explicitly flushed skin without turning it into a colored lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The thrown knife remains lodged in the office wall; the speakerphone remains in place. 윤성찬: He remains standing with his handkerchief, his face flushed from the fright.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "인물은 고개를 숙이고 눈을 감은 채 움츠림. 단검의 자루는 화면 좌측을 향해 튀어나와 있음.",
    "built_space": "이전 샷에서 확립된 투톤 컬러(아래가 더 어두운 회색) 벽면이 배경에 정확히 유지됨. 좌측 가장자리에 캐비닛 일부가 보임.",
    "entities": "윤성찬: 붉어진 얼굴, 질끈 감은 눈, 주름진 노안 등 레퍼런스 및 텍스트와 일치. 손수건을 들고 있음. 단검: 프롬프트대로 인물 우측 벽에 박혀 있음.",
    "hard_violations": [],
    "physics": "단검은 벽면에 박혀 고정됨. 손수건은 인물의 손에 쥐어져 있음."
   },
   {
    "label": "B",
    "direction": "인물은 눈을 질끈 감고 몸을 움츠림. 단검의 자루는 우측 상단을 향하고 있음.",
    "built_space": "이전 샷과 달리 투톤 분할이 없는 단색의 밝은 회색 벽면으로 렌더링됨. 우측 배경에 캐비닛이 위치함.",
    "entities": "윤성찬: 붉어진 얼굴, 감은 눈, 손수건, 레퍼런스와 흡사한 질감의 재킷 착용. 단검: 이전 샷의 스위치블레이드 형태와 매우 유사하나 위치가 인물 좌측으로 어긋남.",
    "hard_violations": [],
    "physics": "단검의 날이 벽면에 박혀 지지됨. 손수건은 손에 꽉 쥐어져 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "단검을 인물 우측 배경에 배치하라는 프레이밍 지시와 이전 샷의 투톤 벽면(Location Lock)을 정확히 반영하여 가장 충실한 결과물입니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "인물의 복장과 단검의 디테일은 우수하나, 단검이 좌측에 배치되었고 이전 샷의 투톤 벽면 설정이 누락되어 프레이밍 및 장소 유지 규칙을 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "인물은 고개를 숙이고 눈을 감은 채 움츠림. 단검의 자루는 화면 좌측을 향해 튀어나와 있음.",
        "built_space": "이전 샷에서 확립된 투톤 컬러(아래가 더 어두운 회색) 벽면이 배경에 정확히 유지됨. 좌측 가장자리에 캐비닛 일부가 보임.",
        "entities": "윤성찬: 붉어진 얼굴, 질끈 감은 눈, 주름진 노안 등 레퍼런스 및 텍스트와 일치. 손수건을 들고 있음. 단검: 프롬프트대로 인물 우측 벽에 박혀 있음.",
        "hard_violations": [],
        "physics": "단검은 벽면에 박혀 고정됨. 손수건은 인물의 손에 쥐어져 있음."
       },
       {
        "label": "B",
        "direction": "인물은 눈을 질끈 감고 몸을 움츠림. 단검의 자루는 우측 상단을 향하고 있음.",
        "built_space": "이전 샷과 달리 투톤 분할이 없는 단색의 밝은 회색 벽면으로 렌더링됨. 우측 배경에 캐비닛이 위치함.",
        "entities": "윤성찬: 붉어진 얼굴, 감은 눈, 손수건, 레퍼런스와 흡사한 질감의 재킷 착용. 단검: 이전 샷의 스위치블레이드 형태와 매우 유사하나 위치가 인물 좌측으로 어긋남.",
        "hard_violations": [],
        "physics": "단검의 날이 벽면에 박혀 지지됨. 손수건은 손에 꽉 쥐어져 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "단검을 인물 우측 배경에 배치하라는 프레이밍 지시와 이전 샷의 투톤 벽면(Location Lock)을 정확히 반영하여 가장 충실한 결과물입니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "인물의 복장과 단검의 디테일은 우수하나, 단검이 좌측에 배치되었고 이전 샷의 투톤 벽면 설정이 누락되어 프레이밍 및 장소 유지 규칙을 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "인물은 고개를 숙이고 눈을 감은 채 움츠림. 단검의 자루는 화면 좌측을 향해 튀어나와 있음.",
        "built_space": "이전 샷에서 확립된 투톤 컬러(아래가 더 어두운 회색) 벽면이 배경에 정확히 유지됨. 좌측 가장자리에 캐비닛 일부가 보임.",
        "entities": "윤성찬: 붉어진 얼굴, 질끈 감은 눈, 주름진 노안 등 레퍼런스 및 텍스트와 일치. 손수건을 들고 있음. 단검: 프롬프트대로 인물 우측 벽에 박혀 있음.",
        "hard_violations": [],
        "physics": "단검은 벽면에 박혀 고정됨. 손수건은 인물의 손에 쥐어져 있음."
       },
       {
        "label": "B",
        "direction": "인물은 눈을 질끈 감고 몸을 움츠림. 단검의 자루는 우측 상단을 향하고 있음.",
        "built_space": "이전 샷과 달리 투톤 분할이 없는 단색의 밝은 회색 벽면으로 렌더링됨. 우측 배경에 캐비닛이 위치함.",
        "entities": "윤성찬: 붉어진 얼굴, 감은 눈, 손수건, 레퍼런스와 흡사한 질감의 재킷 착용. 단검: 이전 샷의 스위치블레이드 형태와 매우 유사하나 위치가 인물 좌측으로 어긋남.",
        "hard_violations": [],
        "physics": "단검의 날이 벽면에 박혀 지지됨. 손수건은 손에 꽉 쥐어져 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "머리 오른쪽 뒤의 단검, 두 눈을 질끈 감고 움츠린 붉은 얼굴, 사무실의 투톤 벽을 가장 충실히 구현했으나 안경이 없고 재킷과 머리 모양은 참조와 다릅니다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "붉어진 얼굴과 움츠린 동작은 맞지만 단검이 지정된 화면 중간 오른쪽이 아닌 왼쪽에 크게 배치되며, 참조의 안경도 빠졌습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 눈을 감고 얼굴을 화면 오른쪽으로 돌려 칼에서 피하는 모습입니다. 단검의 칼끝은 왼쪽 벽 안으로 들어가고 손잡이는 오른쪽 위로 돌출되어 머리와 겹치지 않습니다. 다만 단검과 박힌 벽은 지정된 머리 오른쪽 뒤가 아니라 화면 왼쪽에 있습니다.",
        "built_space": "왼쪽에 마모된 밝은 벽 한 면, 오른쪽 배경에 금속 서류장 한 개와 창 일부, 수납 선반 일부가 보입니다. 인물은 벽 바로 앞에 어깨를 붙이고 있습니다. 사무실 재료와 낮의 조명은 유사하지만 참조의 회청색 하단 벽은 보이지 않아 장소 연속성이 덜 분명합니다. 단검 한 개가 비교적 크게 드러납니다. 스피커폰은 근접 촬영 범위 밖입니다.",
        "entities": "고령의 한국인 남성으로 보이는 인물 한 명이며, 회색 머리와 깊은 이마·눈가 주름, 붉어진 피부가 요구와 부합합니다. 짙은 재킷과 남색 셔츠는 인물 참조에 가깝지만 참조의 금속테 안경이 없습니다. 손에는 흰 손수건이 있고 벽에는 단검 한 개가 있습니다. 추가 인물이나 읽을 수 있는 글자는 보이지 않습니다.",
        "hard_violations": [],
        "physics": "칼날 끝이 벽에 박힌 접점과 벽 손상이 보여 단검은 벽에 의해 지지됩니다. 손수건은 손가락으로 쥐고 있습니다. 어깨를 올리고 목을 움츠린 자세는 놀라 피하는 동작으로 가능합니다. 하체는 프레임 밖이므로 발의 지지는 확인할 수 없지만 몸이 떠 있는 모습은 아닙니다."
       },
       {
        "label": "B",
        "direction": "두 눈을 강하게 감고 고개를 왼쪽 아래로 숙여 오른쪽 뒤의 단검에서 몸을 피합니다. 단검의 칼끝은 화면 오른쪽 벽 안에 박혀 있고 손잡이는 왼쪽으로 비스듬히 돌출됩니다. 칼은 얼굴과 겹치지 않으며 머리 바로 오른쪽 뒤에 있어 아슬아슬하게 빗나간 관계가 분명합니다.",
        "built_space": "밝은 상단과 회청색 하단으로 나뉜 벽 한 면이 배경을 차지하고, 왼쪽 끝에 서류장 한 개의 일부가 보입니다. 벽의 도장과 마모는 이전 사무실의 재료에 부합합니다. 인물은 벽 바로 앞에 있고, 단검 한 개는 화면 중간 오른쪽의 벽에 박혀 있습니다. 얼굴 중심의 근접 구도이며 스피커폰은 촬영 범위 밖입니다.",
        "entities": "고령의 한국인 남성으로 보이는 인물 한 명이 깊은 주름과 뚜렷하게 붉어진 얼굴을 보입니다. 참조와 유사한 회색 머리와 얼굴 특징이 있으나 머리가 더 흐트러졌고 금속테 안경이 없습니다. 재킷도 참조보다 밝은 회색이며 조직감이 다릅니다. 어두운 셔츠, 손에 쥔 손수건, 벽에 박힌 단검이 보이고 추가 인물이나 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "단검의 칼날이 손상된 벽에 들어가 있어 지지점이 명확합니다. 손수건은 가슴 앞에서 손가락으로 움켜쥐고 있습니다. 올라간 어깨와 숙인 목, 앞으로 굽힌 상체는 겁에 질려 움츠리는 동작으로 자연스럽습니다. 하체는 잘려 있어 서 있는 발은 확인되지 않지만 공중 부양이나 지지 없는 물체는 보이지 않습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "머리 오른쪽 뒤의 단검, 두 눈을 질끈 감고 움츠린 붉은 얼굴, 사무실의 투톤 벽을 가장 충실히 구현했으나 안경이 없고 재킷과 머리 모양은 참조와 다릅니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "붉어진 얼굴과 움츠린 동작은 맞지만 단검이 지정된 화면 중간 오른쪽이 아닌 왼쪽에 크게 배치되며, 참조의 안경도 빠졌습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "두 눈을 감고 얼굴을 화면 오른쪽으로 돌려 칼에서 피하는 모습입니다. 단검의 칼끝은 왼쪽 벽 안으로 들어가고 손잡이는 오른쪽 위로 돌출되어 머리와 겹치지 않습니다. 다만 단검과 박힌 벽은 지정된 머리 오른쪽 뒤가 아니라 화면 왼쪽에 있습니다.",
        "built_space": "왼쪽에 마모된 밝은 벽 한 면, 오른쪽 배경에 금속 서류장 한 개와 창 일부, 수납 선반 일부가 보입니다. 인물은 벽 바로 앞에 어깨를 붙이고 있습니다. 사무실 재료와 낮의 조명은 유사하지만 참조의 회청색 하단 벽은 보이지 않아 장소 연속성이 덜 분명합니다. 단검 한 개가 비교적 크게 드러납니다. 스피커폰은 근접 촬영 범위 밖입니다.",
        "entities": "고령의 한국인 남성으로 보이는 인물 한 명이며, 회색 머리와 깊은 이마·눈가 주름, 붉어진 피부가 요구와 부합합니다. 짙은 재킷과 남색 셔츠는 인물 참조에 가깝지만 참조의 금속테 안경이 없습니다. 손에는 흰 손수건이 있고 벽에는 단검 한 개가 있습니다. 추가 인물이나 읽을 수 있는 글자는 보이지 않습니다.",
        "hard_violations": [],
        "physics": "칼날 끝이 벽에 박힌 접점과 벽 손상이 보여 단검은 벽에 의해 지지됩니다. 손수건은 손가락으로 쥐고 있습니다. 어깨를 올리고 목을 움츠린 자세는 놀라 피하는 동작으로 가능합니다. 하체는 프레임 밖이므로 발의 지지는 확인할 수 없지만 몸이 떠 있는 모습은 아닙니다."
       },
       {
        "label": "A",
        "direction": "두 눈을 강하게 감고 고개를 왼쪽 아래로 숙여 오른쪽 뒤의 단검에서 몸을 피합니다. 단검의 칼끝은 화면 오른쪽 벽 안에 박혀 있고 손잡이는 왼쪽으로 비스듬히 돌출됩니다. 칼은 얼굴과 겹치지 않으며 머리 바로 오른쪽 뒤에 있어 아슬아슬하게 빗나간 관계가 분명합니다.",
        "built_space": "밝은 상단과 회청색 하단으로 나뉜 벽 한 면이 배경을 차지하고, 왼쪽 끝에 서류장 한 개의 일부가 보입니다. 벽의 도장과 마모는 이전 사무실의 재료에 부합합니다. 인물은 벽 바로 앞에 있고, 단검 한 개는 화면 중간 오른쪽의 벽에 박혀 있습니다. 얼굴 중심의 근접 구도이며 스피커폰은 촬영 범위 밖입니다.",
        "entities": "고령의 한국인 남성으로 보이는 인물 한 명이 깊은 주름과 뚜렷하게 붉어진 얼굴을 보입니다. 참조와 유사한 회색 머리와 얼굴 특징이 있으나 머리가 더 흐트러졌고 금속테 안경이 없습니다. 재킷도 참조보다 밝은 회색이며 조직감이 다릅니다. 어두운 셔츠, 손에 쥔 손수건, 벽에 박힌 단검이 보이고 추가 인물이나 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "단검의 칼날이 손상된 벽에 들어가 있어 지지점이 명확합니다. 손수건은 가슴 앞에서 손가락으로 움켜쥐고 있습니다. 올라간 어깨와 숙인 목, 앞으로 굽힌 상체는 겁에 질려 움츠리는 동작으로 자연스럽습니다. 하체는 잘려 있어 서 있는 발은 확인되지 않지만 공중 부양이나 지지 없는 물체는 보이지 않습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.321
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.321
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1321
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "단검을 인물 우측 배경에 배치하라는 프레이밍 지시와 이전 샷의 투톤 벽면(Location Lock)을 정확히 반영하여 가장 충실한 결과물입니다."
   },
   {
    "label": "B",
    "score": 1321,
    "verdict_ko": "인물의 복장과 단검의 디테일은 우수하나, 단검이 좌측에 배치되었고 이전 샷의 투톤 벽면 설정이 누락되어 프레이밍 및 장소 유지 규칙을 위반했습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S18sh1_sel.png",
    "asset_id": "0bc31807-2f1a-45e8-b5a6-feafe863b6fd",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 윤성찬: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1453233>",
    "asset_id": "04d34665-3829-49fb-a6d2-25b9fd051d63",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-f53f-7df7-bafe-407d1989d262",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S18sh1"
  }
 },
 "S18sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:44:47.743209+00:00",
  "fingerprint": "dfb2100d24954e8ed44243c60304eb0be54e5245b4dc41a9b7d60beca9cf82fc",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S18sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S18sh5_sel.png",
  "source_sha256": "aedae02a7c12686c12465cc937204df19598a5d6db7847f5c1efd5fabf5eed68",
  "file": "S18sh5_cine.png",
  "staged_sha256": "fee5517eb0ba63579d9c5fe6ff83ad69669195f8f2fb4d9f59676e46dd395a82",
  "latency_ms": 11407
 },
 "S18sh6::signage": {
  "fp": "fbada9e8e6a451d6",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S18sh6": {
  "input_fingerprint": "da39fa874bc28292",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 겁먹은 윤성찬을 향해 허리를 90도로 굽힌 채 고개를 들어 입꼬리를 길게 찢어 웃고 있는 박철진의 자세.\n\nLOCATION (lock): In the open space between the commander's desk and the visitor inside the militia office, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office interior (Only a limited portion remains visible around the bowed figure); used as Preserve spatial context around the bend of the torso without adding foreground obstructions.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the office illumination neutral and the facial contrast controlled, allowing the threatening smile to remain plainly visible without an expressionistic lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The knife remains lodged in the wall, and the speakerphone remains in the office. 박철진: He no longer holds the knife and bends into a deep, ninety-degree bow.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 겁먹은 윤성찬을 향해 허리를 90도로 굽힌 채 고개를 들어 입꼬리를 길게 찢어 웃고 있는 박철진의 자세.\n\nLOCATION (lock): In the open space between the commander's desk and the visitor inside the militia office, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office interior (Only a limited portion remains visible around the bowed figure); used as Preserve spatial context around the bend of the torso without adding foreground obstructions.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the office illumination neutral and the facial contrast controlled, allowing the threatening smile to remain plainly visible without an expressionistic lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The knife remains lodged in the wall, and the speakerphone remains in the office. 박철진: He no longer holds the knife and bends into a deep, ninety-degree bow.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 겁먹은 윤성찬을 향해 허리를 90도로 굽힌 채 고개를 들어 입꼬리를 길게 찢어 웃고 있는 박철진의 자세.\n\nLOCATION (lock): In the open space between the commander's desk and the visitor inside the militia office, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office interior (Only a limited portion remains visible around the bowed figure); used as Preserve spatial context around the bend of the torso without adding foreground obstructions.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the office illumination neutral and the facial contrast controlled, allowing the threatening smile to remain plainly visible without an expressionistic lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The knife remains lodged in the wall, and the speakerphone remains in the office. 박철진: He no longer holds the knife and bends into a deep, ninety-degree bow.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "박철진이 화면 좌측 전경에 있는 노인을 향해 시선을 고정하고 있음.",
    "built_space": "사무실 내부. 좌측에 투톤 벽면과 박힌 칼이 있고, 우측에 책상과 명패가 배치됨.",
    "entities": "박철진은 레퍼런스와 일치하나, 전경의 인물은 이전 샷의 노인(회색 머리와 양복)을 그대로 복사함. 우측 명패에 읽을 수 있는 문자 형태가 보임.",
    "hard_violations": [
     "[gemini-pro] 이전 샷의 인물 외형(회색 머리의 노인)을 그대로 차용함",
     "[gemini-pro] 명패에 읽을 수 있는 텍스트 노출",
     "[gpt-high] 허용 인물 목록에 없는 노인을 왼쪽 전경에 추가했으며, 이전 장면 인물의 외형과 옷을 재사용하지 말라는 지시에도 어긋난다.",
     "[gpt-high] 오른쪽 책상 명패에 읽을 수 있는 한글을 노출해 문자 금지 지시를 위반했다."
    ],
    "physics": "박철진은 상체를 약간만 숙인 상태로 두 다리로 지탱하여 서 있음."
   },
   {
    "label": "B",
    "direction": "박철진이 화면 우측 전경에 있는 인물을 향해 고개를 들고 시선을 맞춤.",
    "built_space": "사무실 내부. 좌측 벽면에 칼이 박혀 있고, 배경에 책상과 스피커폰이 배치됨.",
    "entities": "박철진은 레퍼런스의 얼굴 및 복장과 일치함. 전경 인물은 검은 머리로 이전 샷 인물과 다름. 책상 위에 지정된 스피커폰이 있음. 텍스트 노출 없음.",
    "hard_violations": [
     "[gpt-high] 박철진만 허용된 인물 목록과 달리 오른쪽 전경에 별도의 남자 몸을 추가했다."
    ],
    "physics": "박철진은 허리를 90도로 깊게 숙인 상태에서 상체를 다리로 안정적으로 지탱하고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "이전 샷의 인물을 그대로 사용하지 말라는 제약을 어겼으며, 명패에 읽을 수 있는 문자가 노출되어 심각한 규정 위반임."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "90도로 허리를 굽힌 채 고개를 들고 웃는 지정된 포즈를 훌륭히 소화했으며, 스피커폰과 배경 설정 등 모든 제약을 충실히 준수함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진이 화면 좌측 전경에 있는 노인을 향해 시선을 고정하고 있음.",
        "built_space": "사무실 내부. 좌측에 투톤 벽면과 박힌 칼이 있고, 우측에 책상과 명패가 배치됨.",
        "entities": "박철진은 레퍼런스와 일치하나, 전경의 인물은 이전 샷의 노인(회색 머리와 양복)을 그대로 복사함. 우측 명패에 읽을 수 있는 문자 형태가 보임.",
        "hard_violations": [
         "이전 샷의 인물 외형(회색 머리의 노인)을 그대로 차용함",
         "명패에 읽을 수 있는 텍스트 노출"
        ],
        "physics": "박철진은 상체를 약간만 숙인 상태로 두 다리로 지탱하여 서 있음."
       },
       {
        "label": "B",
        "direction": "박철진이 화면 우측 전경에 있는 인물을 향해 고개를 들고 시선을 맞춤.",
        "built_space": "사무실 내부. 좌측 벽면에 칼이 박혀 있고, 배경에 책상과 스피커폰이 배치됨.",
        "entities": "박철진은 레퍼런스의 얼굴 및 복장과 일치함. 전경 인물은 검은 머리로 이전 샷 인물과 다름. 책상 위에 지정된 스피커폰이 있음. 텍스트 노출 없음.",
        "hard_violations": [],
        "physics": "박철진은 허리를 90도로 깊게 숙인 상태에서 상체를 다리로 안정적으로 지탱하고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "이전 샷의 인물을 그대로 사용하지 말라는 제약을 어겼으며, 명패에 읽을 수 있는 문자가 노출되어 심각한 규정 위반임."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "90도로 허리를 굽힌 채 고개를 들고 웃는 지정된 포즈를 훌륭히 소화했으며, 스피커폰과 배경 설정 등 모든 제약을 충실히 준수함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "박철진이 화면 좌측 전경에 있는 노인을 향해 시선을 고정하고 있음.",
        "built_space": "사무실 내부. 좌측에 투톤 벽면과 박힌 칼이 있고, 우측에 책상과 명패가 배치됨.",
        "entities": "박철진은 레퍼런스와 일치하나, 전경의 인물은 이전 샷의 노인(회색 머리와 양복)을 그대로 복사함. 우측 명패에 읽을 수 있는 문자 형태가 보임.",
        "hard_violations": [
         "이전 샷의 인물 외형(회색 머리의 노인)을 그대로 차용함",
         "명패에 읽을 수 있는 텍스트 노출"
        ],
        "physics": "박철진은 상체를 약간만 숙인 상태로 두 다리로 지탱하여 서 있음."
       },
       {
        "label": "B",
        "direction": "박철진이 화면 우측 전경에 있는 인물을 향해 고개를 들고 시선을 맞춤.",
        "built_space": "사무실 내부. 좌측 벽면에 칼이 박혀 있고, 배경에 책상과 스피커폰이 배치됨.",
        "entities": "박철진은 레퍼런스의 얼굴 및 복장과 일치함. 전경 인물은 검은 머리로 이전 샷 인물과 다름. 책상 위에 지정된 스피커폰이 있음. 텍스트 노출 없음.",
        "hard_violations": [],
        "physics": "박철진은 허리를 90도로 깊게 숙인 상태에서 상체를 다리로 안정적으로 지탱하고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "대상을 향한 시선과 웃음, 인물 외형은 비교적 충실하지만, 허용되지 않은 전경 인물을 추가했고 허리도 명확한 90도 굽힘에 미달한다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "상대를 향한 위협적인 웃음은 구현했지만 이전 장면의 노인을 전경에 재등장시키고 읽을 수 있는 명패까지 넣었으며, 90도 굽힘도 불충분하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진은 몸을 화면 오른쪽으로 숙이고 고개를 들어 오른쪽 전경 남자의 얼굴을 바라보며 이를 드러내 웃는다. 시선의 대상은 명확하지만, 그 남자가 윤성찬인지는 확인할 수 없다. 왼쪽 벽의 칼은 칼끝이 벽 안으로 들어가고 손잡이가 오른쪽으로 돌출되어 있다.",
        "built_space": "흰색 상부와 회색 하부로 나뉜 낡은 벽은 장소 참조와 유사하다. 뒤쪽 사무용 책상 하나와 검은 의자 하나, 오른쪽 앞 탁자 하나, 왼쪽 아래 의자 일부, 뒤쪽 수납장 하나, 오른쪽 문 하나, 왼쪽 벽 게시물 두 개가 보인다. 칼은 왼쪽 벽에 하나, 스피커폰은 앞 탁자 위에 하나 있다. 박철진은 가구 사이에 서 있지만, 오른쪽 전경 인물이 화면을 가려 전경 방해물 없는 구도 지시를 어긴다.",
        "entities": "주인공은 중년 동아시아계 남성으로 보이며, 짧게 빗어 넘긴 검은 머리, 귀걸이, 남색 정장, 흰 셔츠와 사선무늬 넥타이가 박철진 참조에 대체로 부합한다. 추가로 검은 머리와 회색 계열 옷을 입은 남자의 뒤통수와 어깨가 보이는데, 허용 인물 목록에 없는 인물이다. 벽에 꽂힌 칼과 사무실 스피커폰은 모두 보이고, 박철진은 칼을 들고 있지 않다. 뚜렷하게 읽히는 글자는 없다.",
        "hard_violations": [
         "박철진만 허용된 인물 목록과 달리 오른쪽 전경에 별도의 남자 몸을 추가했다."
        ],
        "physics": "박철진의 골반과 하체는 화면 아래로 이어지고 상체를 앞으로 기울인 서 있는 자세로 읽힌다. 발은 잘렸지만 공중에 떠 있는 모습은 아니다. 다만 상체가 수평에 가까워질 만큼 허리에서 90도로 접힌 모습은 아니고 비스듬히 숙인 자세다. 스피커폰은 탁자에 놓여 있고 칼은 벽에 박혀 지지된다. 웃음과 목의 들림은 정상적인 인체 연기로 가능하다."
       },
       {
        "label": "B",
        "direction": "박철진의 얼굴과 눈은 왼쪽 전경의 회색 머리 노인을 향한다. 고개를 든 채 입꼬리를 올리고 이를 드러내 웃어 상대를 겨냥한 표정은 분명하다. 노인 역시 박철진 쪽으로 얼굴을 돌리고 있다. 뒤쪽 칼은 칼끝이 벽에 박히고 손잡이가 오른쪽으로 나와 있다.",
        "built_space": "낡은 흰색·회색 이색 벽과 모서리가 배경을 이룬다. 오른쪽에는 책상 하나, 그 뒤 낮은 서랍장 또는 수납가구 하나, 벽 액자 일부 하나가 보인다. 책상 주변에는 서류 정리함, 파일 묶음, 필기구 통 하나와 검은 명패 하나가 놓여 있다. 벽의 칼은 하나다. 박철진은 책상 앞 공간에 있지만 왼쪽 노인이 상당한 전경을 차지해 방해물 없는 제한된 배경 구도에서 벗어난다. 스피커폰은 이 프레임에서 확인되지 않는다.",
        "entities": "박철진은 중년 동아시아계 남성으로 보이고 검은 머리, 귀걸이, 남색 정장, 흰 셔츠와 사선무늬 넥타이는 참조와 대체로 일치한다. 왼쪽에는 이전 장면의 인물과 유사한 회색 머리·회색 재킷의 노인이 추가되어 있다. 벽에 박힌 칼은 유지되며 박철진의 두 손은 비어 있다. 오른쪽 검은 명패에는 식별 가능한 한글이 보인다.",
        "hard_violations": [
         "허용 인물 목록에 없는 노인을 왼쪽 전경에 추가했으며, 이전 장면 인물의 외형과 옷을 재사용하지 말라는 지시에도 어긋난다.",
         "오른쪽 책상 명패에 읽을 수 있는 한글을 노출해 문자 금지 지시를 위반했다."
        ],
        "physics": "박철진은 하체 위에서 상체를 앞으로 숙이고 두 손을 배 앞에서 맞잡고 있다. 손의 접촉과 몸의 연결은 자연스럽고, 잘린 하체가 지면 지지를 이어가는 자세라 부유 문제는 없다. 그러나 등과 몸통이 뚜렷하게 수평을 이루는 90도 인사보다는 얕게 기울어진 자세다. 칼은 벽에 박혀 있고 책상 물건들은 가구 표면에 놓여 있다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "대상을 향한 시선과 웃음, 인물 외형은 비교적 충실하지만, 허용되지 않은 전경 인물을 추가했고 허리도 명확한 90도 굽힘에 미달한다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "상대를 향한 위협적인 웃음은 구현했지만 이전 장면의 노인을 전경에 재등장시키고 읽을 수 있는 명패까지 넣었으며, 90도 굽힘도 불충분하다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "박철진은 몸을 화면 오른쪽으로 숙이고 고개를 들어 오른쪽 전경 남자의 얼굴을 바라보며 이를 드러내 웃는다. 시선의 대상은 명확하지만, 그 남자가 윤성찬인지는 확인할 수 없다. 왼쪽 벽의 칼은 칼끝이 벽 안으로 들어가고 손잡이가 오른쪽으로 돌출되어 있다.",
        "built_space": "흰색 상부와 회색 하부로 나뉜 낡은 벽은 장소 참조와 유사하다. 뒤쪽 사무용 책상 하나와 검은 의자 하나, 오른쪽 앞 탁자 하나, 왼쪽 아래 의자 일부, 뒤쪽 수납장 하나, 오른쪽 문 하나, 왼쪽 벽 게시물 두 개가 보인다. 칼은 왼쪽 벽에 하나, 스피커폰은 앞 탁자 위에 하나 있다. 박철진은 가구 사이에 서 있지만, 오른쪽 전경 인물이 화면을 가려 전경 방해물 없는 구도 지시를 어긴다.",
        "entities": "주인공은 중년 동아시아계 남성으로 보이며, 짧게 빗어 넘긴 검은 머리, 귀걸이, 남색 정장, 흰 셔츠와 사선무늬 넥타이가 박철진 참조에 대체로 부합한다. 추가로 검은 머리와 회색 계열 옷을 입은 남자의 뒤통수와 어깨가 보이는데, 허용 인물 목록에 없는 인물이다. 벽에 꽂힌 칼과 사무실 스피커폰은 모두 보이고, 박철진은 칼을 들고 있지 않다. 뚜렷하게 읽히는 글자는 없다.",
        "hard_violations": [
         "박철진만 허용된 인물 목록과 달리 오른쪽 전경에 별도의 남자 몸을 추가했다."
        ],
        "physics": "박철진의 골반과 하체는 화면 아래로 이어지고 상체를 앞으로 기울인 서 있는 자세로 읽힌다. 발은 잘렸지만 공중에 떠 있는 모습은 아니다. 다만 상체가 수평에 가까워질 만큼 허리에서 90도로 접힌 모습은 아니고 비스듬히 숙인 자세다. 스피커폰은 탁자에 놓여 있고 칼은 벽에 박혀 지지된다. 웃음과 목의 들림은 정상적인 인체 연기로 가능하다."
       },
       {
        "label": "A",
        "direction": "박철진의 얼굴과 눈은 왼쪽 전경의 회색 머리 노인을 향한다. 고개를 든 채 입꼬리를 올리고 이를 드러내 웃어 상대를 겨냥한 표정은 분명하다. 노인 역시 박철진 쪽으로 얼굴을 돌리고 있다. 뒤쪽 칼은 칼끝이 벽에 박히고 손잡이가 오른쪽으로 나와 있다.",
        "built_space": "낡은 흰색·회색 이색 벽과 모서리가 배경을 이룬다. 오른쪽에는 책상 하나, 그 뒤 낮은 서랍장 또는 수납가구 하나, 벽 액자 일부 하나가 보인다. 책상 주변에는 서류 정리함, 파일 묶음, 필기구 통 하나와 검은 명패 하나가 놓여 있다. 벽의 칼은 하나다. 박철진은 책상 앞 공간에 있지만 왼쪽 노인이 상당한 전경을 차지해 방해물 없는 제한된 배경 구도에서 벗어난다. 스피커폰은 이 프레임에서 확인되지 않는다.",
        "entities": "박철진은 중년 동아시아계 남성으로 보이고 검은 머리, 귀걸이, 남색 정장, 흰 셔츠와 사선무늬 넥타이는 참조와 대체로 일치한다. 왼쪽에는 이전 장면의 인물과 유사한 회색 머리·회색 재킷의 노인이 추가되어 있다. 벽에 박힌 칼은 유지되며 박철진의 두 손은 비어 있다. 오른쪽 검은 명패에는 식별 가능한 한글이 보인다.",
        "hard_violations": [
         "허용 인물 목록에 없는 노인을 왼쪽 전경에 추가했으며, 이전 장면 인물의 외형과 옷을 재사용하지 말라는 지시에도 어긋난다.",
         "오른쪽 책상 명패에 읽을 수 있는 한글을 노출해 문자 금지 지시를 위반했다."
        ],
        "physics": "박철진은 하체 위에서 상체를 앞으로 숙이고 두 손을 배 앞에서 맞잡고 있다. 손의 접촉과 몸의 연결은 자연스럽고, 잘린 하체가 지면 지지를 이어가는 자세라 부유 문제는 없다. 그러나 등과 몸통이 뚜렷하게 수평을 이루는 90도 인사보다는 얕게 기울어진 자세다. 칼은 벽에 박혀 있고 책상 물건들은 가구 표면에 놓여 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.095,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.845,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 이전 샷의 인물 외형(회색 머리의 노인)을 그대로 차용함",
     "[gemini-pro] 명패에 읽을 수 있는 텍스트 노출",
     "[gpt-high] 허용 인물 목록에 없는 노인을 왼쪽 전경에 추가했으며, 이전 장면 인물의 외형과 옷을 재사용하지 말라는 지시에도 어긋난다.",
     "[gpt-high] 오른쪽 책상 명패에 읽을 수 있는 한글을 노출해 문자 금지 지시를 위반했다."
    ],
    "B": [
     "[gpt-high] 박철진만 허용된 인물 목록과 달리 오른쪽 전경에 별도의 남자 몸을 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "A": 845,
   "B": 1750
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 845,
    "verdict_ko": "이전 샷의 인물을 그대로 사용하지 말라는 제약을 어겼으며, 명패에 읽을 수 있는 문자가 노출되어 심각한 규정 위반임.  ★위반: [gemini-pro] 이전 샷의 인물 외형(회색 머리의 노인)을 그대로 차용함 / [gemini-pro] 명패에 읽을 수 있는 텍스트 노출 / [gpt-high] 허용 인물 목록에 없는 노인을 왼쪽 전경에 추가했으며, 이전 장면 인물의 외형과 옷을 재사용하지 말라는 지시에도 어긋난다. / [gpt-high] 오른쪽 책상 명패에 읽을 수 있는 한글을 노출해 문자 금지 지시를 위반했다."
   },
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "90도로 허리를 굽힌 채 고개를 들고 웃는 지정된 포즈를 훌륭히 소화했으며, 스피커폰과 배경 설정 등 모든 제약을 충실히 준수함.  ★위반: [gpt-high] 박철진만 허용된 인물 목록과 달리 오른쪽 전경에 별도의 남자 몸을 추가했다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S18sh5_sel.png",
    "asset_id": "18293ecc-a3be-411c-bee8-12e3e4cb279a",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1401722>",
    "asset_id": "fee7383c-fb61-4b3a-ba7c-79f2555de00b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-f708-7f86-a9d5-a08abcca9b50",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S18sh5"
  }
 },
 "S18sh6::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:46:09.242094+00:00",
  "fingerprint": "e35fd0582fceb1d9559045d058a72551198b469cf498f2446ab8a38580316705",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S18sh6_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S18sh6_sel.png",
  "source_sha256": "7e19de5ef41e6bee2fc7785ca172d3f825b0eb9e2dcbcae9bd56b14d5a0a3c78",
  "file": "S18sh6_cine.png",
  "staged_sha256": "3ccb4ecfe1548b699ce4de3588c04c99fa3d9555f4e89c2858af863fe0e1d180",
  "latency_ms": 9373
 },
 "S19sh3::confined_fp_apt": {
  "applies": true,
  "reason_ko": "이 숏은 차량 내부라는 제한된 공간에서 진행되며, 뒷좌석에 앉은 인물이 앞좌석을 향해 몸을 기울이는 구체적인 행동을 묘사하고 있습니다. 따라서 앞뒤 좌석의 공간적 배치와 인물의 위치 관계를 정확하게 표현하지 않으면 상황을 이해하기 어려우므로 평면도 레이아웃 보조가 필요합니다.",
  "input_fingerprint": "5244c6375b6673fe"
 },
 "S19sh3::signage": {
  "fp": "a689f1a4fcea29e3",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "confinedfp::3a46c2692d54": {
  "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/confinedfp_base_3a46c2692d54.png",
  "place_text": "Inside the rear passenger compartment of a car, immediately behind the front seats, with daylight entering through the windows.",
  "input_fingerprint": "4fdd5d2a3a8abf12"
 },
 "S19sh3::confined_fp": {
  "reads": {
   "controls": "A steering wheel is located at the front-left driver's seat.",
   "mirrors": "No mirrors are present in the diagram.",
   "camera": "The camera is positioned between the front driver and passenger seats, pointing directly backward toward the rear center seat.",
   "occupants": "Yoon Sung-chan (윤성찬) is seated in the Rear Center position. All other seats are empty."
  },
  "mismatches": [],
  "scene_description_en": "The camera is positioned between the front seats, looking straight back into the rear cabin. On the near left edge of the screen, the inner side of the empty driver's seat frames the view. On the near right edge, the inner side of the empty front passenger seat provides the opposite frame. In the center of the frame sits Yoon Sung-chan, occupying the rear center seat and facing forward toward the camera. In the background, flanking him on the far left and far right, are the empty rear left and rear right seats.",
  "fixed": true,
  "input_fingerprint": "9a6fa3d50bd62343"
 },
 "era_assess::1dd1249d0430d782": {
  "subjects": [],
  "subject_text": "윤성찬의 자동차 내부\n앞좌석과 뒷좌석이 구분된 고급 세단 실내. 운전대와 계기판이 앞쪽에 있고 뒷좌석 옆으로 창문과 도어 내장재가 이어진다.",
  "identity": "canonical",
  "scope_id": "L172",
  "scope_role": "location_interior",
  "scope_sha": "e79f96e1a718ce2b"
 },
 "S19sh3": {
  "input_fingerprint": "5118b9d690c50c97",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 주먹을 꽉 쥔 채 앞좌석을 향해 상체를 바짝 내민 윤성찬의 격양된 얼굴.\n\nLOCATION (lock): Inside the rear passenger compartment of a car, immediately behind the front seats, with daylight entering through the windows. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Front seats (Their inner edges remain visible beside the camera's viewing gap) — The inner sides and partial rear faces bracket the view toward the rear passenger; used as Create a narrow spatial channel around the face and fist without obscuring either.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daytime ambient illumination within the car, retaining readable facial tension and restrained contrast without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): 윤성찬: He is seated inside the car and retains his handkerchief.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera is positioned between the front seats, looking straight back into the rear cabin. On the near left edge of the screen, the inner side of the empty driver's seat frames the view. On the near right edge, the inner side of the empty front passenger seat provides the opposite frame. In the center of the frame sits Yoon Sung-chan, occupying the rear center seat and facing forward toward the camera. In the background, flanking him on the far left and far right, are the empty rear left and rear right seats.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 주먹을 꽉 쥔 채 앞좌석을 향해 상체를 바짝 내민 윤성찬의 격양된 얼굴.\n\nLOCATION (lock): Inside the rear passenger compartment of a car, immediately behind the front seats, with daylight entering through the windows. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daytime ambient illumination within the car, retaining readable facial tension and restrained contrast without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): 윤성찬: He is seated inside the car and retains his handkerchief.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera is positioned between the front seats, looking straight back into the rear cabin. On the near left edge of the screen, the inner side of the empty driver's seat frames the view. On the near right edge, the inner side of the empty front passenger seat provides the opposite frame. In the center of the frame sits Yoon Sung-chan, occupying the rear center seat and facing forward toward the camera. In the background, flanking him on the far left and far right, are the empty rear left and rear right seats.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 주먹을 꽉 쥔 채 앞좌석을 향해 상체를 바짝 내민 윤성찬의 격양된 얼굴.\n\nLOCATION (lock): Inside the rear passenger compartment of a car, immediately behind the front seats, with daylight entering through the windows. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daytime ambient illumination within the car, retaining readable facial tension and restrained contrast without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): 윤성찬: He is seated inside the car and retains his handkerchief.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S19sh3_confinedfp.png",
     "asset_id": null,
     "role": null
    },
    {
     "label": "윤성찬",
     "path": "<bytes:1453233>",
     "asset_id": "04d34665-3829-49fb-a6d2-25b9fd051d63",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S19sh3_confinedfp.png",
     "asset_id": null,
     "role": null
    },
    {
     "label": "윤성찬",
     "path": "<bytes:1453233>",
     "asset_id": "04d34665-3829-49fb-a6d2-25b9fd051d63",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라를 향해 상체를 숙이고 양 주먹을 꽉 쥔 상태로 정면을 응시함.",
    "built_space": "차량 뒷좌석 중앙. 카메라가 앞좌석 사이에서 뒤를 향하며, 양옆에 앞좌석 등받이 뒷면이 정상적으로 배치됨.",
    "entities": "윤성찬. 참조 이미지의 안경, 헤어스타일, 주름, 의상이 정확히 일치함.",
    "hard_violations": [],
    "physics": "뒷좌석에 안정적으로 앉아 다리와 골반으로 체중을 지탱하며 자연스러운 자세를 유지함."
   },
   {
    "label": "B",
    "direction": "상체를 크게 숙이고 한쪽 주먹을 쥔 채 화면 우측을 빗겨 응시함.",
    "built_space": "차량 내부이나, 전경에 위치한 두 앞좌석이 등받이 뒷면이 아닌 탑승자가 앉는 정면(굴곡 및 측면 지지대)으로 잘못 렌더링되어 공간 방향성이 붕괴됨.",
    "entities": "윤성찬. 의상과 주름은 일치하나 참조 이미지의 안경이 누락됨.",
    "hard_violations": [
     "[gemini-pro] physically impossible staging (카메라 위치 및 방향상 앞좌석의 뒷면이 보여야 하나 시트의 정면이 렌더링됨)",
     "[gpt-high] 후방을 보는 지정 카메라에 앞좌석의 뒷면이 아닌 앞쪽 착석면이 노출되어, 고정 좌석의 방향과 카메라 위치 관계가 배치 지시와 모순된다."
    ],
    "physics": "좌석에 체중을 싣고 상체를 과도하게 앞으로 기울인 상태로 지탱됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "공간 배치와 인물 외형(안경 포함)을 완벽히 구현했으나, 지시된 클로즈업보다 다소 넓은 화각으로 렌더링됨."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "전경의 앞좌석이 뒷면이 아닌 정면으로 렌더링되는 치명적인 공간 구조 오류(Hard Violation)가 발생했으며 인물의 안경이 누락됨."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라를 향해 상체를 숙이고 양 주먹을 꽉 쥔 상태로 정면을 응시함.",
        "built_space": "차량 뒷좌석 중앙. 카메라가 앞좌석 사이에서 뒤를 향하며, 양옆에 앞좌석 등받이 뒷면이 정상적으로 배치됨.",
        "entities": "윤성찬. 참조 이미지의 안경, 헤어스타일, 주름, 의상이 정확히 일치함.",
        "hard_violations": [],
        "physics": "뒷좌석에 안정적으로 앉아 다리와 골반으로 체중을 지탱하며 자연스러운 자세를 유지함."
       },
       {
        "label": "B",
        "direction": "상체를 크게 숙이고 한쪽 주먹을 쥔 채 화면 우측을 빗겨 응시함.",
        "built_space": "차량 내부이나, 전경에 위치한 두 앞좌석이 등받이 뒷면이 아닌 탑승자가 앉는 정면(굴곡 및 측면 지지대)으로 잘못 렌더링되어 공간 방향성이 붕괴됨.",
        "entities": "윤성찬. 의상과 주름은 일치하나 참조 이미지의 안경이 누락됨.",
        "hard_violations": [
         "physically impossible staging (카메라 위치 및 방향상 앞좌석의 뒷면이 보여야 하나 시트의 정면이 렌더링됨)"
        ],
        "physics": "좌석에 체중을 싣고 상체를 과도하게 앞으로 기울인 상태로 지탱됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "공간 배치와 인물 외형(안경 포함)을 완벽히 구현했으나, 지시된 클로즈업보다 다소 넓은 화각으로 렌더링됨."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "전경의 앞좌석이 뒷면이 아닌 정면으로 렌더링되는 치명적인 공간 구조 오류(Hard Violation)가 발생했으며 인물의 안경이 누락됨."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라를 향해 상체를 숙이고 양 주먹을 꽉 쥔 상태로 정면을 응시함.",
        "built_space": "차량 뒷좌석 중앙. 카메라가 앞좌석 사이에서 뒤를 향하며, 양옆에 앞좌석 등받이 뒷면이 정상적으로 배치됨.",
        "entities": "윤성찬. 참조 이미지의 안경, 헤어스타일, 주름, 의상이 정확히 일치함.",
        "hard_violations": [],
        "physics": "뒷좌석에 안정적으로 앉아 다리와 골반으로 체중을 지탱하며 자연스러운 자세를 유지함."
       },
       {
        "label": "B",
        "direction": "상체를 크게 숙이고 한쪽 주먹을 쥔 채 화면 우측을 빗겨 응시함.",
        "built_space": "차량 내부이나, 전경에 위치한 두 앞좌석이 등받이 뒷면이 아닌 탑승자가 앉는 정면(굴곡 및 측면 지지대)으로 잘못 렌더링되어 공간 방향성이 붕괴됨.",
        "entities": "윤성찬. 의상과 주름은 일치하나 참조 이미지의 안경이 누락됨.",
        "hard_violations": [
         "physically impossible staging (카메라 위치 및 방향상 앞좌석의 뒷면이 보여야 하나 시트의 정면이 렌더링됨)"
        ],
        "physics": "좌석에 체중을 싣고 상체를 과도하게 앞으로 기울인 상태로 지탱됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "얼굴과 주먹의 근접 구도 및 전방으로 내민 동작은 좋지만, 앞좌석의 뒤가 아니라 착석면을 보여 지정된 카메라·좌석 관계를 어기며 참조의 안경도 빠졌다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "앞좌석 사이에서 뒷좌석 중앙의 윤성찬을 보는 공간 관계와 인물 정체성은 맞지만, 허벅지까지 보이는 넓은 구도라 격양된 얼굴의 클로즈업 지시를 충족하지 못한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴과 상체를 두 앞좌석 사이로 깊게 내밀고, 눈은 카메라보다 약간 위쪽인 화면 오른쪽 전방을 향한다. 쥔 주먹도 앞쪽으로 나와 있다. 앞좌석 쪽으로 향하는 동작은 읽히며, 시선이 닿는 사람은 화면에 없다.",
        "built_space": "좌우에 앞좌석 등받이와 머리받침이 각각 하나씩 있고 그 사이에 인물이 있다. 그러나 카메라를 향한 면은 안쪽 가장자리와 부분적인 뒷면이 아니라, 천으로 된 중앙 쿠션과 양옆 볼스터가 드러난 좌석의 착석면이다. 지정된 앞좌석 사이 카메라에서 후방을 보는 배치와 맞지 않는다. 양옆 창과 뒤쪽 창 일부가 보이며 반사는 없다.",
        "entities": "노년의 한국인 남성 한 명으로 표현되며, 회색으로 넘긴 머리와 깊은 이마·눈가 주름, 짙은 재킷과 남색 셔츠는 참조에 대체로 맞는다. 참조의 금속테 안경은 없다. 얼굴에는 긴장이 있으나 입과 눈의 연기는 분노보다 다급한 호소에도 가깝다. 주먹 하나가 보이고 손수건은 보이지 않으며, 나머지 손과 보관 위치는 화면 밖이다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "후방을 보는 지정 카메라에 앞좌석의 뒷면이 아닌 앞쪽 착석면이 노출되어, 고정 좌석의 방향과 카메라 위치 관계가 배치 지시와 모순된다."
        ],
        "physics": "주먹은 손목과 팔에 연결되어 있고, 앞으로 기울인 상체도 몸통과 자연스럽게 이어진다. 골반과 다리는 잘려 있어 좌석 접촉을 직접 확인할 수 없지만, 보이는 부분에 부유나 불가능한 관절 자세는 없다."
       },
       {
        "label": "B",
        "direction": "윤성찬의 시선은 앞좌석 사이 카메라 쪽을 향하고 상체도 그 방향으로 기울어 있다. 두 주먹을 몸 앞에 쥐고 있어 전방을 향한 긴장은 읽히지만, 얼굴이 앞좌석 틈까지 바짝 들어온 정도는 약하다.",
        "built_space": "전경 양쪽에 앞좌석 두 개의 뒷면과 머리받침이 보이고, 그 사이로 뒷좌석 중앙에 앉은 인물이 보인다. 뒤에는 벤치형 좌석과 좌우 머리받침 두 개, 좌우 안전벨트가 있으며 중앙 머리받침 위치는 인물에 가려져 있다. 양옆 창과 뒤창, 천장 손잡이 두 개가 보인다. 좌석 방향과 카메라의 후방 시선은 맞지만, 통로가 넓어 얼굴과 주먹을 좁게 감싸는 클로즈업은 아니다.",
        "entities": "노년의 한국인 남성 한 명이며, 회색 머리, 금속테 안경, 깊은 얼굴 주름, 짙은 회색 재킷과 남색 셔츠가 참조와 잘 맞는다. 미간과 입에 긴장이 있고 두 손 모두 주먹을 쥐었다. 손수건은 두 손에 보이지 않지만 주머니 속 보유 여부는 판단할 수 없다. 추가 인물이나 명확히 읽히는 글자는 없다.",
        "hard_violations": [],
        "physics": "골반과 허벅지가 뒷좌석 방석에 놓여 체중을 지지하고, 그 위에서 상체를 앞으로 숙인 자세다. 두 주먹은 굽힌 팔과 손목으로 지지된다. 발은 화면 밖이지만 앉은 자세의 지지는 충분히 보이며 부유하는 신체나 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "얼굴과 주먹의 근접 구도 및 전방으로 내민 동작은 좋지만, 앞좌석의 뒤가 아니라 착석면을 보여 지정된 카메라·좌석 관계를 어기며 참조의 안경도 빠졌다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "앞좌석 사이에서 뒷좌석 중앙의 윤성찬을 보는 공간 관계와 인물 정체성은 맞지만, 허벅지까지 보이는 넓은 구도라 격양된 얼굴의 클로즈업 지시를 충족하지 못한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴과 상체를 두 앞좌석 사이로 깊게 내밀고, 눈은 카메라보다 약간 위쪽인 화면 오른쪽 전방을 향한다. 쥔 주먹도 앞쪽으로 나와 있다. 앞좌석 쪽으로 향하는 동작은 읽히며, 시선이 닿는 사람은 화면에 없다.",
        "built_space": "좌우에 앞좌석 등받이와 머리받침이 각각 하나씩 있고 그 사이에 인물이 있다. 그러나 카메라를 향한 면은 안쪽 가장자리와 부분적인 뒷면이 아니라, 천으로 된 중앙 쿠션과 양옆 볼스터가 드러난 좌석의 착석면이다. 지정된 앞좌석 사이 카메라에서 후방을 보는 배치와 맞지 않는다. 양옆 창과 뒤쪽 창 일부가 보이며 반사는 없다.",
        "entities": "노년의 한국인 남성 한 명으로 표현되며, 회색으로 넘긴 머리와 깊은 이마·눈가 주름, 짙은 재킷과 남색 셔츠는 참조에 대체로 맞는다. 참조의 금속테 안경은 없다. 얼굴에는 긴장이 있으나 입과 눈의 연기는 분노보다 다급한 호소에도 가깝다. 주먹 하나가 보이고 손수건은 보이지 않으며, 나머지 손과 보관 위치는 화면 밖이다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "후방을 보는 지정 카메라에 앞좌석의 뒷면이 아닌 앞쪽 착석면이 노출되어, 고정 좌석의 방향과 카메라 위치 관계가 배치 지시와 모순된다."
        ],
        "physics": "주먹은 손목과 팔에 연결되어 있고, 앞으로 기울인 상체도 몸통과 자연스럽게 이어진다. 골반과 다리는 잘려 있어 좌석 접촉을 직접 확인할 수 없지만, 보이는 부분에 부유나 불가능한 관절 자세는 없다."
       },
       {
        "label": "A",
        "direction": "윤성찬의 시선은 앞좌석 사이 카메라 쪽을 향하고 상체도 그 방향으로 기울어 있다. 두 주먹을 몸 앞에 쥐고 있어 전방을 향한 긴장은 읽히지만, 얼굴이 앞좌석 틈까지 바짝 들어온 정도는 약하다.",
        "built_space": "전경 양쪽에 앞좌석 두 개의 뒷면과 머리받침이 보이고, 그 사이로 뒷좌석 중앙에 앉은 인물이 보인다. 뒤에는 벤치형 좌석과 좌우 머리받침 두 개, 좌우 안전벨트가 있으며 중앙 머리받침 위치는 인물에 가려져 있다. 양옆 창과 뒤창, 천장 손잡이 두 개가 보인다. 좌석 방향과 카메라의 후방 시선은 맞지만, 통로가 넓어 얼굴과 주먹을 좁게 감싸는 클로즈업은 아니다.",
        "entities": "노년의 한국인 남성 한 명이며, 회색 머리, 금속테 안경, 깊은 얼굴 주름, 짙은 회색 재킷과 남색 셔츠가 참조와 잘 맞는다. 미간과 입에 긴장이 있고 두 손 모두 주먹을 쥐었다. 손수건은 두 손에 보이지 않지만 주머니 속 보유 여부는 판단할 수 없다. 추가 인물이나 명확히 읽히는 글자는 없다.",
        "hard_violations": [],
        "physics": "골반과 허벅지가 뒷좌석 방석에 놓여 체중을 지지하고, 그 위에서 상체를 앞으로 숙인 자세다. 두 주먹은 굽힌 팔과 손목으로 지지된다. 발은 화면 밖이지만 앉은 자세의 지지는 충분히 보이며 부유하는 신체나 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.929
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.679
   },
   "violations": {
    "B": [
     "[gemini-pro] physically impossible staging (카메라 위치 및 방향상 앞좌석의 뒷면이 보여야 하나 시트의 정면이 렌더링됨)",
     "[gpt-high] 후방을 보는 지정 카메라에 앞좌석의 뒷면이 아닌 앞쪽 착석면이 노출되어, 고정 좌석의 방향과 카메라 위치 관계가 배치 지시와 모순된다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 679
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "공간 배치와 인물 외형(안경 포함)을 완벽히 구현했으나, 지시된 클로즈업보다 다소 넓은 화각으로 렌더링됨."
   },
   {
    "label": "B",
    "score": 679,
    "verdict_ko": "전경의 앞좌석이 뒷면이 아닌 정면으로 렌더링되는 치명적인 공간 구조 오류(Hard Violation)가 발생했으며 인물의 안경이 누락됨.  ★위반: [gemini-pro] physically impossible staging (카메라 위치 및 방향상 앞좌석의 뒷면이 보여야 하나 시트의 정면이 렌더링됨) / [gpt-high] 후방을 보는 지정 카메라에 앞좌석의 뒷면이 아닌 앞쪽 착석면이 노출되어, 고정 좌석의 방향과 카메라 위치 관계가 배치 지시와 모순된다."
   }
  ],
  "refs": [
   {
    "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S19sh3_confinedfp.png",
    "asset_id": null,
    "role": null
   },
   {
    "label": "윤성찬",
    "path": "<bytes:1453233>",
    "asset_id": "04d34665-3829-49fb-a6d2-25b9fd051d63",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-f8d8-7e10-8cfb-78b77b289960",
  "confined_fp": {
   "base_key": "confinedfp::3a46c2692d54",
   "apt_reason": "이 숏은 차량 내부라는 제한된 공간에서 진행되며, 뒷좌석에 앉은 인물이 앞좌석을 향해 몸을 기울이는 구체적인 행동을 묘사하고 있습니다. 따라서 앞뒤 좌석의 공간적 배치와 인물의 위치 관계를 정확하게 표현하지 않으면 상황을 이해하기 어려우므로 평면도 레이아웃 보조가 필요합니다.",
   "fixed": true,
   "mismatches": []
  },
  "ref_mode": "confined_fp: 도면+장면설명+엔티티",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S19sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:49:30.450638+00:00",
  "fingerprint": "3c3a4be730f1f0b0cc1b1e889df1d70932b061785af7b77f37c01ba9022fa7a9",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S19sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S19sh3_sel.png",
  "source_sha256": "fc3644a1c1d00f46a5601b82fbc92d4e9ad44cc4fcf09b2c7939a1e1550ca11b",
  "file": "S19sh3_cine.png",
  "staged_sha256": "db481afe7937ccae13e877c9796815982a4b9e3da3d7d3f198af7435dee7c4c6",
  "latency_ms": 11745
 },
 "S20sh3::signage": {
  "fp": "f1d679db43b52ec9",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S20sh3": {
  "input_fingerprint": "e152c8e8a3440921",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 컨테이너 모퉁이 뒤에 바짝 숨은 채 눈이 커진 현우와 페드로의 얼굴.\n\nLOCATION (lock): Behind a container corner near the searched family home, along an outdoor lane in the refugee settlement. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container corner (Concealing both men from the search beyond it) — The concealed-side face ends at a narrow vertical edge beside their sightline; used as Make the hiding position immediately readable while leaving both faces unobstructed.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain neutral daytime ambient light and controlled contrast, keeping both startled faces legible without inventing atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The searched container is the home marked 7-31. 현우: He is hiding near his container, wearing his outer shirt. His facial injuries and dog-bitten leg remain unhealed. 페드로: He is hiding near the container.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 컨테이너 모퉁이 뒤에 바짝 숨은 채 눈이 커진 현우와 페드로의 얼굴.\n\nLOCATION (lock): Behind a container corner near the searched family home, along an outdoor lane in the refugee settlement. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container corner (Concealing both men from the search beyond it) — The concealed-side face ends at a narrow vertical edge beside their sightline; used as Make the hiding position immediately readable while leaving both faces unobstructed.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain neutral daytime ambient light and controlled contrast, keeping both startled faces legible without inventing atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The searched container is the home marked 7-31. 현우: He is hiding near his container, wearing his outer shirt. His facial injuries and dog-bitten leg remain unhealed. 페드로: He is hiding near the container.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 컨테이너 모퉁이 뒤에 바짝 숨은 채 눈이 커진 현우와 페드로의 얼굴.\n\nLOCATION (lock): Behind a container corner near the searched family home, along an outdoor lane in the refugee settlement. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container corner (Concealing both men from the search beyond it) — The concealed-side face ends at a narrow vertical edge beside their sightline; used as Make the hiding position immediately readable while leaving both faces unobstructed.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain neutral daytime ambient light and controlled contrast, keeping both startled faces legible without inventing atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The searched container is the home marked 7-31. 현우: He is hiding near his container, wearing his outer shirt. His facial injuries and dog-bitten leg remain unhealed. 페드로: He is hiding near the container.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "두 인물의 시선이 화면 오른쪽 밖의 대상을 향하고 있음.",
    "built_space": "하나의 모퉁이가 아닌 두 컨테이너 사이의 좁은 틈새에 위치해 있으며, 배경으로 골목길이 보임. 왼쪽 컨테이너 표면에 비현실적인 검은 조각이 붙어 있음.",
    "entities": "현우와 페드로의 외양과 복장, 현우의 얼굴 상처가 레퍼런스와 일치함. 그러나 배경 우측에 프롬프트에서 금지한 추가 인물들이 다수 존재함.",
    "hard_violations": [
     "[gemini-pro] 프롬프트에 명시되지 않은 추가 인물들이 배경에 등장함",
     "[gemini-pro] 왼쪽 컨테이너 벽면에 물리적으로 불가능한 평면적 그래픽 아티팩트(검은색 덩어리)가 누출됨",
     "[gemini-pro] 컨테이너 모퉁이가 아닌 두 벽체 사이의 틈으로 공간이 잘못 연출됨",
     "[gpt-high] 오른쪽 배경 통로에 현우와 페드로 이외의 인물이 최소 두 명 보여, 두 사람 외에는 누구도 등장시키지 말라는 명시적 제한을 위반한다."
    ],
    "physics": "오른쪽 벽에 짚은 손과 몸의 밀착으로 자세가 지탱되고 있으나 짚은 손의 해부학적 형태가 어색함."
   },
   {
    "label": "B",
    "direction": "두 인물의 시선이 화면 왼쪽(인물들의 오른쪽) 밖의 특정 대상을 뚜렷하게 주시하고 있음.",
    "built_space": "레퍼런스와 일치하는 파란색 컨테이너 모퉁이 뒤에 올바르게 위치해 있으며, 뒤쪽으로 야외 골목의 원근감이 적절히 표현됨.",
    "entities": "현우와 페드로의 얼굴, 복장, 앳된 외모가 매우 정확하게 묘사되었으며 현우의 얼굴 상처도 잘 나타남. 추가 인물은 없음.",
    "hard_violations": [],
    "physics": "모퉁이 벽에 몸을 바짝 기대어 지탱하고 있으며, 아래쪽 가장자리에 손가락 일부가 자연스럽게 닿아 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "레퍼런스와 인물은 유사하나, 배경에 프롬프트에 없는 추가 인물들이 등장하고 컨테이너 표면에 기형적인 검은색 그래픽 아티팩트가 있어 실격입니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "컨테이너 모퉁이 뒤에 숨어 놀란 두 사람의 클로즈업 샷을 프롬프트의 지시대로 매우 정확하고 자연스럽게 구현했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 인물의 시선이 화면 오른쪽 밖의 대상을 향하고 있음.",
        "built_space": "하나의 모퉁이가 아닌 두 컨테이너 사이의 좁은 틈새에 위치해 있으며, 배경으로 골목길이 보임. 왼쪽 컨테이너 표면에 비현실적인 검은 조각이 붙어 있음.",
        "entities": "현우와 페드로의 외양과 복장, 현우의 얼굴 상처가 레퍼런스와 일치함. 그러나 배경 우측에 프롬프트에서 금지한 추가 인물들이 다수 존재함.",
        "hard_violations": [
         "프롬프트에 명시되지 않은 추가 인물들이 배경에 등장함",
         "왼쪽 컨테이너 벽면에 물리적으로 불가능한 평면적 그래픽 아티팩트(검은색 덩어리)가 누출됨",
         "컨테이너 모퉁이가 아닌 두 벽체 사이의 틈으로 공간이 잘못 연출됨"
        ],
        "physics": "오른쪽 벽에 짚은 손과 몸의 밀착으로 자세가 지탱되고 있으나 짚은 손의 해부학적 형태가 어색함."
       },
       {
        "label": "B",
        "direction": "두 인물의 시선이 화면 왼쪽(인물들의 오른쪽) 밖의 특정 대상을 뚜렷하게 주시하고 있음.",
        "built_space": "레퍼런스와 일치하는 파란색 컨테이너 모퉁이 뒤에 올바르게 위치해 있으며, 뒤쪽으로 야외 골목의 원근감이 적절히 표현됨.",
        "entities": "현우와 페드로의 얼굴, 복장, 앳된 외모가 매우 정확하게 묘사되었으며 현우의 얼굴 상처도 잘 나타남. 추가 인물은 없음.",
        "hard_violations": [],
        "physics": "모퉁이 벽에 몸을 바짝 기대어 지탱하고 있으며, 아래쪽 가장자리에 손가락 일부가 자연스럽게 닿아 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "레퍼런스와 인물은 유사하나, 배경에 프롬프트에 없는 추가 인물들이 등장하고 컨테이너 표면에 기형적인 검은색 그래픽 아티팩트가 있어 실격입니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "컨테이너 모퉁이 뒤에 숨어 놀란 두 사람의 클로즈업 샷을 프롬프트의 지시대로 매우 정확하고 자연스럽게 구현했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "두 인물의 시선이 화면 오른쪽 밖의 대상을 향하고 있음.",
        "built_space": "하나의 모퉁이가 아닌 두 컨테이너 사이의 좁은 틈새에 위치해 있으며, 배경으로 골목길이 보임. 왼쪽 컨테이너 표면에 비현실적인 검은 조각이 붙어 있음.",
        "entities": "현우와 페드로의 외양과 복장, 현우의 얼굴 상처가 레퍼런스와 일치함. 그러나 배경 우측에 프롬프트에서 금지한 추가 인물들이 다수 존재함.",
        "hard_violations": [
         "프롬프트에 명시되지 않은 추가 인물들이 배경에 등장함",
         "왼쪽 컨테이너 벽면에 물리적으로 불가능한 평면적 그래픽 아티팩트(검은색 덩어리)가 누출됨",
         "컨테이너 모퉁이가 아닌 두 벽체 사이의 틈으로 공간이 잘못 연출됨"
        ],
        "physics": "오른쪽 벽에 짚은 손과 몸의 밀착으로 자세가 지탱되고 있으나 짚은 손의 해부학적 형태가 어색함."
       },
       {
        "label": "B",
        "direction": "두 인물의 시선이 화면 왼쪽(인물들의 오른쪽) 밖의 특정 대상을 뚜렷하게 주시하고 있음.",
        "built_space": "레퍼런스와 일치하는 파란색 컨테이너 모퉁이 뒤에 올바르게 위치해 있으며, 뒤쪽으로 야외 골목의 원근감이 적절히 표현됨.",
        "entities": "현우와 페드로의 얼굴, 복장, 앳된 외모가 매우 정확하게 묘사되었으며 현우의 얼굴 상처도 잘 나타남. 추가 인물은 없음.",
        "hard_violations": [],
        "physics": "모퉁이 벽에 몸을 바짝 기대어 지탱하고 있으며, 아래쪽 가장자리에 손가락 일부가 자연스럽게 닿아 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "두 얼굴을 가리지 않는 모퉁이 밀착 클로즈업과 같은 방향의 경계 시선이 요구에 충실하며, 놀라 눈이 커진 표정은 조금 더 강해도 좋다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "금지된 배경 인물이 추가되었고, 모퉁이 대신 좁은 틈에 숨은 구도에서 페드로의 얼굴까지 가려져 핵심 연출을 놓쳤다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우와 페드로 모두 화면 왼쪽, 컨테이너 수직 모서리 너머의 화면 밖 공간을 바라본다. 수색 대상이나 수색자는 보이지 않지만 두 사람이 같은 방향을 경계하는 관계는 명확하다. 렌즈를 응시하지 않으며 무기나 지향성 소품은 없다.",
        "built_space": "화면 왼쪽에 녹슨 청회색 골판 금속 벽 한 면과 그 끝의 좁은 수직 모서리 하나가 있다. 현우가 모서리에 밀착하고 페드로가 바로 뒤에서 어깨 너머로 내다보며, 두 얼굴은 모두 가려지지 않는다. 오른쪽 배경에는 흐릿한 외부 통로와 컨테이너 벽이 보인다. 참조의 낡은 금속 컨테이너 재질과 이어지며, 문·창·설비가 잘려 나간 클로즈업이므로 그 수량은 확인할 수 없다.",
        "entities": "젊은 남성 두 명만 보인다. 현우의 동아시아계 얼굴, 헝클어진 검은 머리, 마른 체격과 회녹색 겉옷·짙은 속옷이 참조에 부합하며 얼굴에는 낫지 않은 긁힌 상처가 있다. 페드로는 참조와 유사한 갈색 피부톤, 짙은 머리와 앳된 얼굴, 남색 상의를 갖췄다. 정확한 나이와 국적은 영상만으로 확정할 수 없지만 두 사람의 외관은 설정과 부합한다. 다리 상처와 집 번호는 구도 밖이며 읽을 수 있는 글자는 없다. 자연스러운 낮 조명이다.",
        "hard_violations": [],
        "physics": "두 사람의 상체는 세운 상태에서 모서리 쪽으로 조금 기울어 있고 하체는 화면 밖이다. 현우의 손 일부가 아래쪽 모서리에 닿아 있다. 발 접지는 확인되지 않지만 공중에 뜬 자세가 아니며, 서로 바짝 붙어 숨는 동작으로 가능한 자세다. 떠 있는 물체나 불가능한 관절은 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "현우는 화면 왼쪽을 보고 페드로는 오른쪽 수직 금속 부재 너머를 본다. 두 사람의 경계 방향이 갈라져 있다. 오른쪽 먼 통로에 사람들이 보이지만 현우의 시선은 그쪽으로 향하지 않는다.",
        "built_space": "왼쪽 금속 벽 끝과 오른쪽의 넓은 수직 금속 부재 사이 좁은 틈에 두 사람이 끼어 있으며, 뒤쪽에도 골판 벽이 보인다. 하나의 모퉁이 뒤에 붙어 숨었다기보다 문틈이나 벽 사이 개구부에 들어간 구도다. 오른쪽 부재가 페드로 얼굴 일부를 가린다. 왼쪽 벽에는 검은 덮개 같은 면 하나, 오른쪽 통로에는 낮은 설비함 하나가 보인다. 금속의 색과 녹은 참조와 유사하지만 얼굴보다 상체와 구조물이 많이 포함된다.",
        "entities": "현우와 페드로로 보이는 젊은 남성 두 명 외에 오른쪽 배경 통로에 추가 인물들이 보인다. 현우의 검은 머리, 회녹색 겉옷과 얼굴 상처는 설정에 맞는다. 페드로의 피부톤과 짙은 상의는 대체로 맞지만 머리는 참조보다 곱슬기가 강하고 얼굴 일부가 가려져 동일성을 확인하기 어렵다. 읽을 수 있는 글자는 없고 낮 장면이다. 다리와 집 번호는 보이지 않는다.",
        "hard_violations": [
         "오른쪽 배경 통로에 현우와 페드로 이외의 인물이 최소 두 명 보여, 두 사람 외에는 누구도 등장시키지 말라는 명시적 제한을 위반한다."
        ],
        "physics": "현우는 팔꿈치를 굽혀 오른쪽 수직 금속 부재에 손바닥을 대고 있으며 손과 표면의 접촉은 자연스럽다. 페드로는 뒤에 몸을 세우고 붙어 있다. 두 사람의 하체는 잘렸으나 떠 있다는 징후는 없고, 보이는 자세는 물리적으로 가능하다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "두 얼굴을 가리지 않는 모퉁이 밀착 클로즈업과 같은 방향의 경계 시선이 요구에 충실하며, 놀라 눈이 커진 표정은 조금 더 강해도 좋다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "금지된 배경 인물이 추가되었고, 모퉁이 대신 좁은 틈에 숨은 구도에서 페드로의 얼굴까지 가려져 핵심 연출을 놓쳤다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우와 페드로 모두 화면 왼쪽, 컨테이너 수직 모서리 너머의 화면 밖 공간을 바라본다. 수색 대상이나 수색자는 보이지 않지만 두 사람이 같은 방향을 경계하는 관계는 명확하다. 렌즈를 응시하지 않으며 무기나 지향성 소품은 없다.",
        "built_space": "화면 왼쪽에 녹슨 청회색 골판 금속 벽 한 면과 그 끝의 좁은 수직 모서리 하나가 있다. 현우가 모서리에 밀착하고 페드로가 바로 뒤에서 어깨 너머로 내다보며, 두 얼굴은 모두 가려지지 않는다. 오른쪽 배경에는 흐릿한 외부 통로와 컨테이너 벽이 보인다. 참조의 낡은 금속 컨테이너 재질과 이어지며, 문·창·설비가 잘려 나간 클로즈업이므로 그 수량은 확인할 수 없다.",
        "entities": "젊은 남성 두 명만 보인다. 현우의 동아시아계 얼굴, 헝클어진 검은 머리, 마른 체격과 회녹색 겉옷·짙은 속옷이 참조에 부합하며 얼굴에는 낫지 않은 긁힌 상처가 있다. 페드로는 참조와 유사한 갈색 피부톤, 짙은 머리와 앳된 얼굴, 남색 상의를 갖췄다. 정확한 나이와 국적은 영상만으로 확정할 수 없지만 두 사람의 외관은 설정과 부합한다. 다리 상처와 집 번호는 구도 밖이며 읽을 수 있는 글자는 없다. 자연스러운 낮 조명이다.",
        "hard_violations": [],
        "physics": "두 사람의 상체는 세운 상태에서 모서리 쪽으로 조금 기울어 있고 하체는 화면 밖이다. 현우의 손 일부가 아래쪽 모서리에 닿아 있다. 발 접지는 확인되지 않지만 공중에 뜬 자세가 아니며, 서로 바짝 붙어 숨는 동작으로 가능한 자세다. 떠 있는 물체나 불가능한 관절은 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "현우는 화면 왼쪽을 보고 페드로는 오른쪽 수직 금속 부재 너머를 본다. 두 사람의 경계 방향이 갈라져 있다. 오른쪽 먼 통로에 사람들이 보이지만 현우의 시선은 그쪽으로 향하지 않는다.",
        "built_space": "왼쪽 금속 벽 끝과 오른쪽의 넓은 수직 금속 부재 사이 좁은 틈에 두 사람이 끼어 있으며, 뒤쪽에도 골판 벽이 보인다. 하나의 모퉁이 뒤에 붙어 숨었다기보다 문틈이나 벽 사이 개구부에 들어간 구도다. 오른쪽 부재가 페드로 얼굴 일부를 가린다. 왼쪽 벽에는 검은 덮개 같은 면 하나, 오른쪽 통로에는 낮은 설비함 하나가 보인다. 금속의 색과 녹은 참조와 유사하지만 얼굴보다 상체와 구조물이 많이 포함된다.",
        "entities": "현우와 페드로로 보이는 젊은 남성 두 명 외에 오른쪽 배경 통로에 추가 인물들이 보인다. 현우의 검은 머리, 회녹색 겉옷과 얼굴 상처는 설정에 맞는다. 페드로의 피부톤과 짙은 상의는 대체로 맞지만 머리는 참조보다 곱슬기가 강하고 얼굴 일부가 가려져 동일성을 확인하기 어렵다. 읽을 수 있는 글자는 없고 낮 장면이다. 다리와 집 번호는 보이지 않는다.",
        "hard_violations": [
         "오른쪽 배경 통로에 현우와 페드로 이외의 인물이 최소 두 명 보여, 두 사람 외에는 누구도 등장시키지 말라는 명시적 제한을 위반한다."
        ],
        "physics": "현우는 팔꿈치를 굽혀 오른쪽 수직 금속 부재에 손바닥을 대고 있으며 손과 표면의 접촉은 자연스럽다. 페드로는 뒤에 몸을 세우고 붙어 있다. 두 사람의 하체는 잘렸으나 떠 있다는 징후는 없고, 보이는 자세는 물리적으로 가능하다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.651,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.401,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 프롬프트에 명시되지 않은 추가 인물들이 배경에 등장함",
     "[gemini-pro] 왼쪽 컨테이너 벽면에 물리적으로 불가능한 평면적 그래픽 아티팩트(검은색 덩어리)가 누출됨",
     "[gemini-pro] 컨테이너 모퉁이가 아닌 두 벽체 사이의 틈으로 공간이 잘못 연출됨",
     "[gpt-high] 오른쪽 배경 통로에 현우와 페드로 이외의 인물이 최소 두 명 보여, 두 사람 외에는 누구도 등장시키지 말라는 명시적 제한을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "A": 401,
   "B": 2000
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 401,
    "verdict_ko": "레퍼런스와 인물은 유사하나, 배경에 프롬프트에 없는 추가 인물들이 등장하고 컨테이너 표면에 기형적인 검은색 그래픽 아티팩트가 있어 실격입니다.  ★위반: [gemini-pro] 프롬프트에 명시되지 않은 추가 인물들이 배경에 등장함 / [gemini-pro] 왼쪽 컨테이너 벽면에 물리적으로 불가능한 평면적 그래픽 아티팩트(검은색 덩어리)가 누출됨 / [gemini-pro] 컨테이너 모퉁이가 아닌 두 벽체 사이의 틈으로 공간이 잘못 연출됨 / [gpt-high] 오른쪽 배경 통로에 현우와 페드로 이외의 인물이 최소 두 명 보여, 두 사람 외에는 누구도 등장시키지 말라는 명시적 제한을 위반한다."
   },
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "컨테이너 모퉁이 뒤에 숨어 놀란 두 사람의 클로즈업 샷을 프롬프트의 지시대로 매우 정확하고 자연스럽게 구현했습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S11sh4_sel.png",
    "asset_id": "9977924e-b912-46a3-8a02-f3458bc73a59",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 페드로: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1278830>",
    "asset_id": "b09df655-64d4-4db4-a1b3-2f0bb5d29c95",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-fc2b-7606-9665-c0753a116ee7",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S11sh4"
  },
  "lane_policy": "ab_select_bypass:prev"
 },
 "S20sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:50:30.485800+00:00",
  "fingerprint": "e8d814dc9b8826925bb9d3843bea3530f18db0685d36df3083a308154cdb2e66",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S20sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S20sh3_sel.png",
  "source_sha256": "b067a9939f30a6dac520a6ff89aff2868c53abf40bafb805bf209f57d2a20c5c",
  "file": "S20sh3_cine.png",
  "staged_sha256": "3abb1461f51635e5f225bfceeb858681d9c1e4b85a5bca07689a1953ede2cc96",
  "latency_ms": 10476
 },
 "S20sh6::signage": {
  "fp": "9c8a18a1fb831a09",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S20sh6": {
  "input_fingerprint": "c1cadf611b176734",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 양손이 경찰(현우 체포·호송)의 가슴에 닿아 거칠게 뒤로 밀어내는 힘이 실린 페드로의 미드액션 자세.\n\nLOCATION (lock): On the outdoor container-lined lane near the searched home, where police block the route toward the settlement market. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Refugee-settlement passage (The officer blocks the route as the shove begins); used as Leave a modest visible strip around the bodies so the backward displacement reads within a real passage.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the same neutral daytime illumination and restrained tonal contrast as the concealment beat, with no dramatic lighting shift at contact.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 페드로 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Container 7-31 remains the site of the militia search. 페드로: He is at the blocked escape route, with his arms extended in a forceful push.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 양손이 경찰(현우 체포·호송)의 가슴에 닿아 거칠게 뒤로 밀어내는 힘이 실린 페드로의 미드액션 자세.\n\nLOCATION (lock): On the outdoor container-lined lane near the searched home, where police block the route toward the settlement market. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Refugee-settlement passage (The officer blocks the route as the shove begins); used as Leave a modest visible strip around the bodies so the backward displacement reads within a real passage.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the same neutral daytime illumination and restrained tonal contrast as the concealment beat, with no dramatic lighting shift at contact.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 페드로 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Container 7-31 remains the site of the militia search. 페드로: He is at the blocked escape route, with his arms extended in a forceful push.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 양손이 경찰(현우 체포·호송)의 가슴에 닿아 거칠게 뒤로 밀어내는 힘이 실린 페드로의 미드액션 자세.\n\nLOCATION (lock): On the outdoor container-lined lane near the searched home, where police block the route toward the settlement market. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Refugee-settlement passage (The officer blocks the route as the shove begins); used as Leave a modest visible strip around the bodies so the backward displacement reads within a real passage.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the same neutral daytime illumination and restrained tonal contrast as the concealment beat, with no dramatic lighting shift at contact.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 페드로 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Container 7-31 remains the site of the militia search. 페드로: He is at the blocked escape route, with his arms extended in a forceful push.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "페드로가 시선을 앞의 경찰에게 고정한 채 양손으로 옷깃과 가슴을 잡고 밀쳐내고 있습니다.",
    "built_space": "레퍼런스와 동일한 형태의 컨테이너 야외 통로가 보입니다.",
    "entities": "페드로의 외형은 잘 반영되었습니다. 그러나 경찰 역할에 이전 컷 레퍼런스에 등장한 인물(현우)의 얼굴을 금지 지침을 어기고 그대로 덮어씌웠습니다.",
    "hard_violations": [
     "[gemini-pro] 이전 컷의 인물 얼굴을 다른 캐릭터(경찰)에 그대로 사용 (금지된 인물 복사)"
    ],
    "physics": "밀려나는 경찰이 뒤로 크게 기울어지며 발이 끌리고 먼지가 일어나는 등, 거칠게 밀어내는 힘과 지지 상태가 물리적으로 잘 표현되었습니다."
   },
   {
    "label": "B",
    "direction": "페드로의 시선과 양손이 목표인 경찰의 가슴을 향해 정확히 위치하여 밀고 있습니다.",
    "built_space": "컨테이너가 늘어선 야외 통로가 레퍼런스와 유사하게 구현되었습니다.",
    "entities": "페드로의 외형(얼굴, 머리, 옷)은 레퍼런스와 일치합니다. 그러나 샷 텍스트에 명시되지 않은 다수의 인물(경찰과 시민들)이 배경에 나타납니다.",
    "hard_violations": [
     "[gemini-pro] 지문(Shot Text)에 존재하지 않는 다수의 사람들을 배경에 임의로 생성 (발명된 인물)",
     "[gemini-pro] 화면 내 읽을 수 있는 텍스트('POLICE', 이름표) 노출",
     "[gpt-high] 샷 텍스트에 없는 다수의 경찰과 주민을 배경에 추가했다.",
     "[gpt-high] 경찰 소매에 읽을 수 있는 영문 표기를 노출하여 읽히는 글자 금지 조건을 위반했다."
    ],
    "physics": "페드로가 경찰을 두 손으로 밀고 있으며, 두 사람 모두 발을 땅에 딛고 서 있는 형태로 무게 중심이 유지되고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "미디엄 샷 프레이밍과 주인공의 외형은 잘 맞췄으나, 지문에 없는 다수의 군중을 임의로 추가하고 옷에 텍스트가 노출되어 지침을 위반했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "이전 컷에 나온 다른 인물의 얼굴을 사용하지 말라는 강력한 지침을 어기고 경찰에 적용했으며, 요청된 미디엄 샷 대신 전신 샷을 렌더링했습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "페드로의 시선과 양손이 목표인 경찰의 가슴을 향해 정확히 위치하여 밀고 있습니다.",
        "built_space": "컨테이너가 늘어선 야외 통로가 레퍼런스와 유사하게 구현되었습니다.",
        "entities": "페드로의 외형(얼굴, 머리, 옷)은 레퍼런스와 일치합니다. 그러나 샷 텍스트에 명시되지 않은 다수의 인물(경찰과 시민들)이 배경에 나타납니다.",
        "hard_violations": [
         "지문(Shot Text)에 존재하지 않는 다수의 사람들을 배경에 임의로 생성 (발명된 인물)",
         "화면 내 읽을 수 있는 텍스트('POLICE', 이름표) 노출"
        ],
        "physics": "페드로가 경찰을 두 손으로 밀고 있으며, 두 사람 모두 발을 땅에 딛고 서 있는 형태로 무게 중심이 유지되고 있습니다."
       },
       {
        "label": "A",
        "direction": "페드로가 시선을 앞의 경찰에게 고정한 채 양손으로 옷깃과 가슴을 잡고 밀쳐내고 있습니다.",
        "built_space": "레퍼런스와 동일한 형태의 컨테이너 야외 통로가 보입니다.",
        "entities": "페드로의 외형은 잘 반영되었습니다. 그러나 경찰 역할에 이전 컷 레퍼런스에 등장한 인물(현우)의 얼굴을 금지 지침을 어기고 그대로 덮어씌웠습니다.",
        "hard_violations": [
         "이전 컷의 인물 얼굴을 다른 캐릭터(경찰)에 그대로 사용 (금지된 인물 복사)"
        ],
        "physics": "밀려나는 경찰이 뒤로 크게 기울어지며 발이 끌리고 먼지가 일어나는 등, 거칠게 밀어내는 힘과 지지 상태가 물리적으로 잘 표현되었습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "미디엄 샷 프레이밍과 주인공의 외형은 잘 맞췄으나, 지문에 없는 다수의 군중을 임의로 추가하고 옷에 텍스트가 노출되어 지침을 위반했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "이전 컷에 나온 다른 인물의 얼굴을 사용하지 말라는 강력한 지침을 어기고 경찰에 적용했으며, 요청된 미디엄 샷 대신 전신 샷을 렌더링했습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "페드로의 시선과 양손이 목표인 경찰의 가슴을 향해 정확히 위치하여 밀고 있습니다.",
        "built_space": "컨테이너가 늘어선 야외 통로가 레퍼런스와 유사하게 구현되었습니다.",
        "entities": "페드로의 외형(얼굴, 머리, 옷)은 레퍼런스와 일치합니다. 그러나 샷 텍스트에 명시되지 않은 다수의 인물(경찰과 시민들)이 배경에 나타납니다.",
        "hard_violations": [
         "지문(Shot Text)에 존재하지 않는 다수의 사람들을 배경에 임의로 생성 (발명된 인물)",
         "화면 내 읽을 수 있는 텍스트('POLICE', 이름표) 노출"
        ],
        "physics": "페드로가 경찰을 두 손으로 밀고 있으며, 두 사람 모두 발을 땅에 딛고 서 있는 형태로 무게 중심이 유지되고 있습니다."
       },
       {
        "label": "A",
        "direction": "페드로가 시선을 앞의 경찰에게 고정한 채 양손으로 옷깃과 가슴을 잡고 밀쳐내고 있습니다.",
        "built_space": "레퍼런스와 동일한 형태의 컨테이너 야외 통로가 보입니다.",
        "entities": "페드로의 외형은 잘 반영되었습니다. 그러나 경찰 역할에 이전 컷 레퍼런스에 등장한 인물(현우)의 얼굴을 금지 지침을 어기고 그대로 덮어씌웠습니다.",
        "hard_violations": [
         "이전 컷의 인물 얼굴을 다른 캐릭터(경찰)에 그대로 사용 (금지된 인물 복사)"
        ],
        "physics": "밀려나는 경찰이 뒤로 크게 기울어지며 발이 끌리고 먼지가 일어나는 등, 거칠게 밀어내는 힘과 지지 상태가 물리적으로 잘 표현되었습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "양손으로 경찰 가슴을 미는 접촉과 미디엄 구도는 가깝지만, 배경에 다수의 인물을 추가하고 읽히는 제복 문자를 노출하여 실격이다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "두 사람만으로 지면에 지지된 밀침과 경찰의 후방 이동을 구현했지만, 전신으로 넓힌 구도와 가슴을 밀기보다 조끼를 움켜쥔 손이 지시에서 벗어난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "오른쪽 페드로는 왼쪽 경찰의 얼굴을 바라보며 양팔을 뻗고, 두 손바닥을 경찰의 윗가슴에 댄다. 경찰도 페드로를 바라본다. 힘은 화면 왼쪽 컨테이너 벽 방향으로 향하며, 경찰이 통로 뒤쪽으로 물러나기보다는 옆 벽에 밀리는 배치다. 배경 인물들은 대체로 두 사람 쪽을 향한다. 겨누는 무기는 없다.",
        "built_space": "녹슨 청회색 컨테이너가 통로 양쪽에 늘어서 있고, 왼쪽에 큰 창살 창 하나, 오른쪽 가까이에 창살 창 하나와 뒤쪽의 작은 개구부들이 보인다. 바닥에는 노란 선 하나가 길게 이어진다. 경찰은 왼쪽 벽 바로 앞, 페드로는 통로 중앙에 있다. 재료와 노후감은 참고와 비슷하지만, 인물 뒤 통로와 군중을 상당히 많이 노출한다. 참고에서 확인되지 않는 창과 간판이 추가되어 정확히 같은 장소인지는 확인하기 어렵다.",
        "entities": "전경에는 남색 반소매 차림의 젊은 페드로와 동아시아계 성인 남성 경찰이 있다. 페드로의 짙은 머리, 피부색, 얼굴과 체격은 참고에 대체로 부합한다. 경찰은 청록색 제복과 검은 장구 조끼를 착용한다. 그러나 뒤에 경찰 및 주민으로 보이는 인물이 최소 열 명 더 있어 두 사람의 장면이라는 조건을 위반한다. 경찰 소매의 영문 경찰 표기는 읽을 수 있고 가슴에도 문자가 보인다.",
        "hard_violations": [
         "샷 텍스트에 없는 다수의 경찰과 주민을 배경에 추가했다.",
         "경찰 소매에 읽을 수 있는 영문 표기를 노출하여 읽히는 글자 금지 조건을 위반했다."
        ],
        "physics": "페드로의 양손은 경찰 조끼에 실제로 닿고, 앞으로 기울어진 몸통과 뻗은 팔이 미는 힘을 전달한다. 경찰은 벽에 등과 어깨가 가까이 붙어 있으며 뒤로 젖혀진 자세다. 두 사람의 발은 프레임 밖이므로 접지는 확인할 수 없지만, 보이는 신체에 공중 부양이나 불가능한 관절은 없다. 배경 인물들은 통로 바닥에 서 있다."
       },
       {
        "label": "B",
        "direction": "오른쪽 페드로는 왼쪽 경찰의 얼굴을 바라보고 양팔을 경찰 윗가슴으로 뻗는다. 경찰은 페드로 쪽을 보며 몸통이 왼쪽 뒤로 젖혀진다. 양손은 가슴 위 조끼와 옷깃 부근에 닿지만, 펼친 손바닥으로 밀기보다 천을 움켜쥔 모습이다. 힘과 후퇴 방향은 서로 연결되지만 통로의 길이 방향보다는 왼쪽 벽 방향이다.",
        "built_space": "양쪽에 녹슨 청회색 컨테이너 벽이 있고, 왼쪽 전경에는 문 잠금봉과 경첩, 그 뒤에는 철망 구획 하나가 보인다. 통로 끝에는 가로 구조물과 작은 개구부가 있다. 경찰은 왼쪽 컨테이너 문 앞, 페드로는 통로를 가로질러 다리를 벌리고 있다. 재료와 낮의 조명은 참고에 대체로 맞지만 철망과 문 설비는 참고에서 확인되지 않는다. 두 사람의 발과 넓은 바닥까지 담아 지정된 미디엄 숏보다 확실히 넓다.",
        "entities": "보이는 사람은 페드로와 경찰 두 명뿐이다. 페드로는 참고에 가까운 짙은 곱슬머리와 앳된 얼굴, 남색 반소매를 유지한다. 경찰은 동아시아계 젊은 성인 남성으로 검은 제복과 장구 조끼를 착용한다. 이전 사진의 다른 인물이 입었던 녹색 겉옷이나 얼굴 상처를 옮겨오지는 않았다. 조끼에 작은 표식은 있으나 문구가 명확히 읽히지는 않는다.",
        "hard_violations": [],
        "physics": "페드로는 오른쪽으로 뻗은 뒤쪽 발을 바닥에 확실히 붙이고 앞쪽 무릎을 굽혀 체중을 밀어 넣는다. 양손은 경찰의 조끼를 잡고 있어 접촉이 유지된다. 경찰은 한쪽 부츠가 몸 아래 바닥에 남아 있고 다른 쪽 부츠는 앞쪽으로 뻗어 지면을 스치며, 그 주변에 먼지가 인다. 몸통의 후방 기울기는 밀침에 따른 균형 상실로 설명 가능하며, 아무 지지 없이 떠 있는 몸은 아니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "양손으로 경찰 가슴을 미는 접촉과 미디엄 구도는 가깝지만, 배경에 다수의 인물을 추가하고 읽히는 제복 문자를 노출하여 실격이다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "두 사람만으로 지면에 지지된 밀침과 경찰의 후방 이동을 구현했지만, 전신으로 넓힌 구도와 가슴을 밀기보다 조끼를 움켜쥔 손이 지시에서 벗어난다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "오른쪽 페드로는 왼쪽 경찰의 얼굴을 바라보며 양팔을 뻗고, 두 손바닥을 경찰의 윗가슴에 댄다. 경찰도 페드로를 바라본다. 힘은 화면 왼쪽 컨테이너 벽 방향으로 향하며, 경찰이 통로 뒤쪽으로 물러나기보다는 옆 벽에 밀리는 배치다. 배경 인물들은 대체로 두 사람 쪽을 향한다. 겨누는 무기는 없다.",
        "built_space": "녹슨 청회색 컨테이너가 통로 양쪽에 늘어서 있고, 왼쪽에 큰 창살 창 하나, 오른쪽 가까이에 창살 창 하나와 뒤쪽의 작은 개구부들이 보인다. 바닥에는 노란 선 하나가 길게 이어진다. 경찰은 왼쪽 벽 바로 앞, 페드로는 통로 중앙에 있다. 재료와 노후감은 참고와 비슷하지만, 인물 뒤 통로와 군중을 상당히 많이 노출한다. 참고에서 확인되지 않는 창과 간판이 추가되어 정확히 같은 장소인지는 확인하기 어렵다.",
        "entities": "전경에는 남색 반소매 차림의 젊은 페드로와 동아시아계 성인 남성 경찰이 있다. 페드로의 짙은 머리, 피부색, 얼굴과 체격은 참고에 대체로 부합한다. 경찰은 청록색 제복과 검은 장구 조끼를 착용한다. 그러나 뒤에 경찰 및 주민으로 보이는 인물이 최소 열 명 더 있어 두 사람의 장면이라는 조건을 위반한다. 경찰 소매의 영문 경찰 표기는 읽을 수 있고 가슴에도 문자가 보인다.",
        "hard_violations": [
         "샷 텍스트에 없는 다수의 경찰과 주민을 배경에 추가했다.",
         "경찰 소매에 읽을 수 있는 영문 표기를 노출하여 읽히는 글자 금지 조건을 위반했다."
        ],
        "physics": "페드로의 양손은 경찰 조끼에 실제로 닿고, 앞으로 기울어진 몸통과 뻗은 팔이 미는 힘을 전달한다. 경찰은 벽에 등과 어깨가 가까이 붙어 있으며 뒤로 젖혀진 자세다. 두 사람의 발은 프레임 밖이므로 접지는 확인할 수 없지만, 보이는 신체에 공중 부양이나 불가능한 관절은 없다. 배경 인물들은 통로 바닥에 서 있다."
       },
       {
        "label": "A",
        "direction": "오른쪽 페드로는 왼쪽 경찰의 얼굴을 바라보고 양팔을 경찰 윗가슴으로 뻗는다. 경찰은 페드로 쪽을 보며 몸통이 왼쪽 뒤로 젖혀진다. 양손은 가슴 위 조끼와 옷깃 부근에 닿지만, 펼친 손바닥으로 밀기보다 천을 움켜쥔 모습이다. 힘과 후퇴 방향은 서로 연결되지만 통로의 길이 방향보다는 왼쪽 벽 방향이다.",
        "built_space": "양쪽에 녹슨 청회색 컨테이너 벽이 있고, 왼쪽 전경에는 문 잠금봉과 경첩, 그 뒤에는 철망 구획 하나가 보인다. 통로 끝에는 가로 구조물과 작은 개구부가 있다. 경찰은 왼쪽 컨테이너 문 앞, 페드로는 통로를 가로질러 다리를 벌리고 있다. 재료와 낮의 조명은 참고에 대체로 맞지만 철망과 문 설비는 참고에서 확인되지 않는다. 두 사람의 발과 넓은 바닥까지 담아 지정된 미디엄 숏보다 확실히 넓다.",
        "entities": "보이는 사람은 페드로와 경찰 두 명뿐이다. 페드로는 참고에 가까운 짙은 곱슬머리와 앳된 얼굴, 남색 반소매를 유지한다. 경찰은 동아시아계 젊은 성인 남성으로 검은 제복과 장구 조끼를 착용한다. 이전 사진의 다른 인물이 입었던 녹색 겉옷이나 얼굴 상처를 옮겨오지는 않았다. 조끼에 작은 표식은 있으나 문구가 명확히 읽히지는 않는다.",
        "hard_violations": [],
        "physics": "페드로는 오른쪽으로 뻗은 뒤쪽 발을 바닥에 확실히 붙이고 앞쪽 무릎을 굽혀 체중을 밀어 넣는다. 양손은 경찰의 조끼를 잡고 있어 접촉이 유지된다. 경찰은 한쪽 부츠가 몸 아래 바닥에 남아 있고 다른 쪽 부츠는 앞쪽으로 뻗어 지면을 스치며, 그 주변에 먼지가 인다. 몸통의 후방 기울기는 밀침에 따른 균형 상실로 설명 가능하며, 아무 지지 없이 떠 있는 몸은 아니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.75,
    "B": 1.333
   },
   "adjusted": {
    "A": 1.5,
    "B": 1.083
   },
   "violations": {
    "B": [
     "[gemini-pro] 지문(Shot Text)에 존재하지 않는 다수의 사람들을 배경에 임의로 생성 (발명된 인물)",
     "[gemini-pro] 화면 내 읽을 수 있는 텍스트('POLICE', 이름표) 노출",
     "[gpt-high] 샷 텍스트에 없는 다수의 경찰과 주민을 배경에 추가했다.",
     "[gpt-high] 경찰 소매에 읽을 수 있는 영문 표기를 노출하여 읽히는 글자 금지 조건을 위반했다."
    ],
    "A": [
     "[gemini-pro] 이전 컷의 인물 얼굴을 다른 캐릭터(경찰)에 그대로 사용 (금지된 인물 복사)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1083,
   "A": 1500
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1083,
    "verdict_ko": "미디엄 샷 프레이밍과 주인공의 외형은 잘 맞췄으나, 지문에 없는 다수의 군중을 임의로 추가하고 옷에 텍스트가 노출되어 지침을 위반했습니다.  ★위반: [gemini-pro] 지문(Shot Text)에 존재하지 않는 다수의 사람들을 배경에 임의로 생성 (발명된 인물) / [gemini-pro] 화면 내 읽을 수 있는 텍스트('POLICE', 이름표) 노출 / [gpt-high] 샷 텍스트에 없는 다수의 경찰과 주민을 배경에 추가했다. / [gpt-high] 경찰 소매에 읽을 수 있는 영문 표기를 노출하여 읽히는 글자 금지 조건을 위반했다."
   },
   {
    "label": "A",
    "score": 1500,
    "verdict_ko": "이전 컷에 나온 다른 인물의 얼굴을 사용하지 말라는 강력한 지침을 어기고 경찰에 적용했으며, 요청된 미디엄 샷 대신 전신 샷을 렌더링했습니다.  ★위반: [gemini-pro] 이전 컷의 인물 얼굴을 다른 캐릭터(경찰)에 그대로 사용 (금지된 인물 복사)"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 페드로 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S20sh3_sel.png",
    "asset_id": "9ac1c673-f435-4504-8774-2a11bce01062",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 페드로: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1278830>",
    "asset_id": "b09df655-64d4-4db4-a1b3-2f0bb5d29c95",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-fe04-7439-8ea0-3b2ab922ee62",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S20sh3"
  }
 },
 "S20sh6::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:51:46.059636+00:00",
  "fingerprint": "3e3358a34f2d454aa6b9a4d219a3dfc18ec6ab2674f6a4d30cde4371efa8d6bf",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S20sh6_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S20sh6_sel.png",
  "source_sha256": "d4197c54b25c6287aa034ae875c6836432ae21bb90b588080180eb8e197db6d6",
  "file": "S20sh6_cine.png",
  "staged_sha256": "1730cc5257b988441d1d4bbba0128734c3b5daa2c7093c82d0bd2157003b3b05",
  "latency_ms": 15087
 },
 "S20sh16::signage": {
  "fp": "aaf6824d4f8ea964",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::07e119b2d4378bdf": {
  "subjects": [],
  "subject_text": "인천 난민촌 시장과 상점 골목\n컨테이너 거리 사이로 작은 상점과 상품 진열대가 이어지는 시장 골목. 바나나를 파는 가게와 상자 더미, 더욱 좁아지는 샛길이 있다.",
  "identity": "canonical",
  "scope_id": "L167",
  "scope_role": "location_exterior",
  "scope_sha": "eced2ac782880df0"
 },
 "groupbg::market_capture_alley": {
  "input_fingerprint": "1603bdda72a77d08",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "market_capture_alley",
    "tags": [
     "S20sh16"
    ]
   },
   "context_sig": "502f598537ff5204"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On the ground at a narrow alley off the refugee settlement market, where the fleeing youth is subdued.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n인천 난민촌 시장과 상점 골목: 물건을 파는 매대와 천막들이 복잡하게 얽혀 있는 좁고 혼잡한 시장통. (특징: 조악하게 지어진 상점과 노점상들; 바나나 등 식료품이 진열된 매대; 통로에 쌓여 있는 종이 상자와 물건들)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 더욱 좁은 골목길로 도망가려던 찰나-! 퍽!!\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On the ground at a narrow alley off the refugee settlement market, where the fleeing youth is subdued.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n인천 난민촌 시장과 상점 골목: 물건을 파는 매대와 천막들이 복잡하게 얽혀 있는 좁고 혼잡한 시장통. (특징: 조악하게 지어진 상점과 노점상들; 바나나 등 식료품이 진열된 매대; 통로에 쌓여 있는 종이 상자와 물건들)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 더욱 좁은 골목길로 도망가려던 찰나-! 퍽!!\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_market_capture_alley_13d54d.png",
  "asset_id": "5e0698a7-8a58-4159-968f-d06cd12bcc47",
  "input_asset_ids": [
   "45af8bb5-9f53-4e4a-a819-7a8c7ba17ccd"
  ],
  "origin_tag": "S20sh16",
  "place_text": "On the ground at a narrow alley off the refugee settlement market, where the fleeing youth is subdued.",
  "origin_inputs": {
   "place_text": "On the ground at a narrow alley off the refugee settlement market, where the fleeing youth is subdued.",
   "time_of_day_en": "day",
   "conti_asset_id": "45af8bb5-9f53-4e4a-a819-7a8c7ba17ccd"
  }
 },
 "S20sh16::bgfirst_bg": {
  "input_fingerprint": "74ed9dc23d6bc9bc",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 바닥에 쓰러져 일그러진 현우의 등 위를 무릎으로 짓누른 채 체중을 싣고 고정된 경찰(현우 체포·호송)의 굳은 상체.\n\nLOCATION (lock): On the ground at a narrow alley off the refugee settlement market, where the fleeing youth is subdued.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Ground at the alley restraint point (Supporting 현우's prone body and the kneeling officer); used as Provide a visible margin around the contact points so the restraint does not become an ambiguous overlap of bodies.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain neutral daytime ambient light with controlled contrast and no added injury coloration or stylized lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 바닥에 쓰러져 일그러진 현우의 등 위를 무릎으로 짓누른 채 체중을 싣고 고정된 경찰(현우 체포·호송)의 굳은 상체.\n\nLOCATION (lock): On the ground at a narrow alley off the refugee settlement market, where the fleeing youth is subdued.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Ground at the alley restraint point (Supporting 현우's prone body and the kneeling officer); used as Provide a visible margin around the contact points so the restraint does not become an ambiguous overlap of bodies.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain neutral daytime ambient light with controlled contrast and no added injury coloration or stylized lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S20sh16__bgfirst_bg.png",
  "asset_id": "43c81765-5840-45d5-8f30-f37e76cc5841",
  "input_asset_ids": [
   "45af8bb5-9f53-4e4a-a819-7a8c7ba17ccd",
   "5e0698a7-8a58-4159-968f-d06cd12bcc47"
  ]
 },
 "S20sh16": {
  "input_fingerprint": "ca744d6f889d3043",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바닥에 쓰러져 일그러진 현우의 등 위를 무릎으로 짓누른 채 체중을 싣고 고정된 경찰(현우 체포·호송)의 굳은 상체.\n\nLOCATION (lock): On the ground at a narrow alley off the refugee settlement market, where the fleeing youth is subdued. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Ground at the alley restraint point (Supporting 현우's prone body and the kneeling officer); used as Provide a visible margin around the contact points so the restraint does not become an ambiguous overlap of bodies.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain neutral daytime ambient light with controlled contrast and no added injury coloration or stylized lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Hyunwoo is pinned face-down on the ground beneath a plainclothes officer's knee pressing into his back. His head's direction and the positions of his arms and legs are not specified.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Boxes knocked over during the escape remain scattered through the market alley. 현우: He has been brought down by a rifle-butt strike, with his outer shirt still on. His earlier facial injuries and dog-bite leg injury remain.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바닥에 쓰러져 일그러진 현우의 등 위를 무릎으로 짓누른 채 체중을 싣고 고정된 경찰(현우 체포·호송)의 굳은 상체.\n\nLOCATION (lock): On the ground at a narrow alley off the refugee settlement market, where the fleeing youth is subdued. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Ground at the alley restraint point (Supporting 현우's prone body and the kneeling officer); used as Provide a visible margin around the contact points so the restraint does not become an ambiguous overlap of bodies.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain neutral daytime ambient light with controlled contrast and no added injury coloration or stylized lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Hyunwoo is pinned face-down on the ground beneath a plainclothes officer's knee pressing into his back. His head's direction and the positions of his arms and legs are not specified.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Boxes knocked over during the escape remain scattered through the market alley. 현우: He has been brought down by a rifle-butt strike, with his outer shirt still on. His earlier facial injuries and dog-bite leg injury remain.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바닥에 쓰러져 일그러진 현우의 등 위를 무릎으로 짓누른 채 체중을 싣고 고정된 경찰(현우 체포·호송)의 굳은 상체.\n\nLOCATION (lock): On the ground at a narrow alley off the refugee settlement market, where the fleeing youth is subdued. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Ground at the alley restraint point (Supporting 현우's prone body and the kneeling officer); used as Provide a visible margin around the contact points so the restraint does not become an ambiguous overlap of bodies.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain neutral daytime ambient light with controlled contrast and no added injury coloration or stylized lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Hyunwoo is pinned face-down on the ground beneath a plainclothes officer's knee pressing into his back. His head's direction and the positions of his arms and legs are not specified.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Boxes knocked over during the escape remain scattered through the market alley. 현우: He has been brought down by a rifle-butt strike, with his outer shirt still on. His earlier facial injuries and dog-bite leg injury remain.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S20sh16__bgfirst_bg.png",
     "asset_id": "43c81765-5840-45d5-8f30-f37e76cc5841",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S20sh16.png",
     "asset_id": "45af8bb5-9f53-4e4a-a819-7a8c7ba17ccd",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_market_capture_alley_13d54d.png",
     "asset_id": "5e0698a7-8a58-4159-968f-d06cd12bcc47",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "경찰은 엎드린 현우를 내려다보고 있으며, 현우의 시선은 바닥을 향함.",
    "built_space": "참조 이미지의 골목길 구조(상자, 파이프, 문)를 정확히 반영함. 두 인물이 골목 중앙 바닥에 위치함.",
    "entities": "현우는 참조 얼굴과 일치하고 얼굴 상처가 있으나, 지정된 다리의 개 물림 상처는 보이지 않음. 경찰은 사복 차림을 준수함.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 해부학 (경찰의 오른 무릎 아래로 튀어나와 현우의 옆구리에 얹힌 불가능한 각도의 세 번째 신발/발)"
    ],
    "physics": "현우는 바닥에 누워 지탱됨. 경찰은 왼발과 오른 무릎을 지지대로 체중을 싣고 있으나, 잘려 나오거나 꺾인 여분의 발이 물리적 법칙을 위반함."
   },
   {
    "label": "B",
    "direction": "경찰은 아래를 향해 현우의 등/어깨 쪽을 응시하며, 현우는 바닥에 얼굴을 대고 있음.",
    "built_space": "참조 이미지의 골목 배경을 충실히 구현했으며 공간 내 인물들의 위치는 적절함.",
    "entities": "현우는 참조 얼굴과 일치하며 얼굴 상처와 오른쪽 다리의 개 물림 상처(피 묻은 카키색 바지)가 잘 표현됨. 경찰은 사복 차림임.",
    "hard_violations": [
     "[gemini-pro] 지시된 연출 위치 위반 (경찰이 무릎으로 등을 짓누르지 않고 현우의 옆 바닥에 무릎을 꿇은 채 배치됨)",
     "[gemini-pro] 물리적으로 불가능한 해부학 (현우의 상체와 하체가 이어지는 허리 부분의 심하게 왜곡되고 늘어난 비율)",
     "[gpt-high] 현우의 가슴이 카메라 쪽으로 돌아나온 옆누운 자세이며, 보이는 경찰의 무릎과 정강이는 등 위가 아니라 몸 옆 바닥에 놓여 있다. 엎드린 현우의 등을 무릎으로 누르는 필수 배치와 다르다."
    ],
    "physics": "현우는 바닥에 밀착되어 지지됨. 경찰은 지면에 닿은 왼쪽 무릎과 오른쪽 발, 현우의 몸을 짚은 양손으로 체중을 지탱함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정된 '무릎으로 등을 짓누르는' 연출을 시도했으나, 경찰의 오른 무릎 아래에 세 번째 발이 기형적으로 생성되는 치명적인 해부학적 오류가 발생했고 현우의 다리 상처가 누락되었습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "캐릭터 외형과 상처 디테일은 우수하나, 필수 지침인 '무릎으로 등을 짓누르는' 자세를 완전히 무시하고 옆에 무릎을 꿇은 채 손으로 제압하고 있어 핵심 연출을 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "경찰은 엎드린 현우를 내려다보고 있으며, 현우의 시선은 바닥을 향함.",
        "built_space": "참조 이미지의 골목길 구조(상자, 파이프, 문)를 정확히 반영함. 두 인물이 골목 중앙 바닥에 위치함.",
        "entities": "현우는 참조 얼굴과 일치하고 얼굴 상처가 있으나, 지정된 다리의 개 물림 상처는 보이지 않음. 경찰은 사복 차림을 준수함.",
        "hard_violations": [
         "물리적으로 불가능한 해부학 (경찰의 오른 무릎 아래로 튀어나와 현우의 옆구리에 얹힌 불가능한 각도의 세 번째 신발/발)"
        ],
        "physics": "현우는 바닥에 누워 지탱됨. 경찰은 왼발과 오른 무릎을 지지대로 체중을 싣고 있으나, 잘려 나오거나 꺾인 여분의 발이 물리적 법칙을 위반함."
       },
       {
        "label": "B",
        "direction": "경찰은 아래를 향해 현우의 등/어깨 쪽을 응시하며, 현우는 바닥에 얼굴을 대고 있음.",
        "built_space": "참조 이미지의 골목 배경을 충실히 구현했으며 공간 내 인물들의 위치는 적절함.",
        "entities": "현우는 참조 얼굴과 일치하며 얼굴 상처와 오른쪽 다리의 개 물림 상처(피 묻은 카키색 바지)가 잘 표현됨. 경찰은 사복 차림임.",
        "hard_violations": [
         "지시된 연출 위치 위반 (경찰이 무릎으로 등을 짓누르지 않고 현우의 옆 바닥에 무릎을 꿇은 채 배치됨)",
         "물리적으로 불가능한 해부학 (현우의 상체와 하체가 이어지는 허리 부분의 심하게 왜곡되고 늘어난 비율)"
        ],
        "physics": "현우는 바닥에 밀착되어 지지됨. 경찰은 지면에 닿은 왼쪽 무릎과 오른쪽 발, 현우의 몸을 짚은 양손으로 체중을 지탱함."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정된 '무릎으로 등을 짓누르는' 연출을 시도했으나, 경찰의 오른 무릎 아래에 세 번째 발이 기형적으로 생성되는 치명적인 해부학적 오류가 발생했고 현우의 다리 상처가 누락되었습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "캐릭터 외형과 상처 디테일은 우수하나, 필수 지침인 '무릎으로 등을 짓누르는' 자세를 완전히 무시하고 옆에 무릎을 꿇은 채 손으로 제압하고 있어 핵심 연출을 위반했습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "경찰은 엎드린 현우를 내려다보고 있으며, 현우의 시선은 바닥을 향함.",
        "built_space": "참조 이미지의 골목길 구조(상자, 파이프, 문)를 정확히 반영함. 두 인물이 골목 중앙 바닥에 위치함.",
        "entities": "현우는 참조 얼굴과 일치하고 얼굴 상처가 있으나, 지정된 다리의 개 물림 상처는 보이지 않음. 경찰은 사복 차림을 준수함.",
        "hard_violations": [
         "물리적으로 불가능한 해부학 (경찰의 오른 무릎 아래로 튀어나와 현우의 옆구리에 얹힌 불가능한 각도의 세 번째 신발/발)"
        ],
        "physics": "현우는 바닥에 누워 지탱됨. 경찰은 왼발과 오른 무릎을 지지대로 체중을 싣고 있으나, 잘려 나오거나 꺾인 여분의 발이 물리적 법칙을 위반함."
       },
       {
        "label": "B",
        "direction": "경찰은 아래를 향해 현우의 등/어깨 쪽을 응시하며, 현우는 바닥에 얼굴을 대고 있음.",
        "built_space": "참조 이미지의 골목 배경을 충실히 구현했으며 공간 내 인물들의 위치는 적절함.",
        "entities": "현우는 참조 얼굴과 일치하며 얼굴 상처와 오른쪽 다리의 개 물림 상처(피 묻은 카키색 바지)가 잘 표현됨. 경찰은 사복 차림임.",
        "hard_violations": [
         "지시된 연출 위치 위반 (경찰이 무릎으로 등을 짓누르지 않고 현우의 옆 바닥에 무릎을 꿇은 채 배치됨)",
         "물리적으로 불가능한 해부학 (현우의 상체와 하체가 이어지는 허리 부분의 심하게 왜곡되고 늘어난 비율)"
        ],
        "physics": "현우는 바닥에 밀착되어 지지됨. 경찰은 지면에 닿은 왼쪽 무릎과 오른쪽 발, 현우의 몸을 짚은 양손으로 체중을 지탱함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "등에 닿은 경찰의 무릎과 바닥을 딛는 반대쪽 발이 체중을 실은 제압을 명확히 보여주지만, 상체 중심의 미디엄 숏보다는 프레임이 넓다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "장소와 인물 외형은 대체로 맞지만, 현우가 옆으로 돌아누웠고 경찰의 무릎은 등보다 바닥에 놓여 핵심 제압 자세를 구현하지 못했다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "경찰은 고개를 아래로 숙여 현우의 상체와 붙잡은 팔 쪽을 향한다. 현우의 얼굴은 화면 왼쪽 아래로 돌아가 바닥에 닿아 있고 눈은 감겨 있다. 겨누는 무기나 이동 동작은 없다.",
        "built_space": "왼쪽에는 천막과 금속 지주, 상자와 과일 진열대가 이어지고 오른쪽에는 낡은 벽, 금속 문, 굵은 수직 배관과 가는 배관이 보인다. 파란 통 하나와 붉은 판자 하나도 참조와 대응한다. 두 사람은 오른쪽 벽 가까운 골목 바닥에 있으나 경찰의 몸과 다리가 제압 접점을 가려, 등을 무릎으로 누른다는 공간 관계가 성립하지 않는다.",
        "entities": "주요 인물은 사복 차림의 성인 남성 경찰과 앳된 동아시아계 남성 현우 두 명이다. 현우의 헝클어진 검은 머리와 얼굴은 참조에 대체로 가깝고, 회색 겉셔츠를 입고 있으며 얼굴의 작은 상처와 바짓단 부근의 붉은 부상 흔적이 보인다. 국적이나 정확한 나이는 외형만으로 확정할 수 없다. 넘어지거나 기울어진 상자들이 남아 있고, 읽을 수 있는 문구나 그래픽 오버레이는 보이지 않는다.",
        "hard_violations": [
         "현우의 가슴이 카메라 쪽으로 돌아나온 옆누운 자세이며, 보이는 경찰의 무릎과 정강이는 등 위가 아니라 몸 옆 바닥에 놓여 있다. 엎드린 현우의 등을 무릎으로 누르는 필수 배치와 다르다."
        ],
        "physics": "현우의 머리와 몸 옆면, 보이는 다리는 바닥에 지지된다. 뒤로 올라간 팔 부근에는 경찰의 손이 닿아 있어 완전히 무지지로 떠 있는 자세는 아니다. 경찰도 접힌 다리와 반대쪽 발로 지지되지만, 그 하중이 현우의 등 위 무릎 접촉으로 전달되는 모습은 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "경찰은 몸과 고개를 현우의 등 쪽으로 숙이고 있다. 현우는 얼굴을 화면 오른쪽으로 돌린 채 뺨을 바닥에 대고 있으며 시선은 아래로 떨어져 있다. 무전기는 허리에 매달려 있고 조작하거나 누군가를 겨누는 물체는 없다.",
        "built_space": "왼쪽 천막 지주와 상자 더미, 오른쪽 낡은 벽과 금속 문, 굵고 가는 수직 배관, 파란 통 하나, 붉은 판자 하나가 참조 장소와 대응한다. 오른쪽 문 앞의 목재 통 하나와 경찰 발 근처의 사각 바닥 덮개도 보인다. 현우는 골목 바닥에 길게 엎드려 있고 경찰은 그 오른쪽에서 등을 누른다. 무릎 접촉부와 접지한 발 주변에 바닥 여백이 있어 두 신체의 배치가 명료하다. 다만 두 사람의 거의 전신과 긴 골목까지 보여 상체 중심 미디엄 숏보다 넓다.",
        "entities": "사복 재킷과 밝은 바지를 입은 성인 남성 경찰, 헝클어진 검은 머리의 앳된 동아시아계 남성 현우가 보인다. 현우의 얼굴과 체격은 참조에 대체로 부합하며 어두운 겉셔츠를 입고 얼굴에는 작은 상처가 남아 있다. 다리는 바지로 덮여 개에 물린 상처의 지속 여부는 확인하기 어렵다. 경찰 허리에는 무전기와 장비 주머니가 부착되어 있다. 골목에는 기울어진 상자와 흩어진 과일이 남아 있으며 판독 가능한 글자나 오버레이는 보이지 않는다.",
        "hard_violations": [],
        "physics": "현우의 뺨, 가슴과 배, 보이는 팔과 손, 다리가 바닥에 놓여 있다. 경찰의 한쪽 무릎은 현우의 등 위에 명확히 닿고 반대쪽 발은 바닥을 딛는다. 굽힌 무릎과 접지한 발 사이로 체중을 전달할 수 있는 자세이며, 경찰의 손도 현우의 몸에 닿아 있다. 지지 없이 떠 있는 신체나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "등에 닿은 경찰의 무릎과 바닥을 딛는 반대쪽 발이 체중을 실은 제압을 명확히 보여주지만, 상체 중심의 미디엄 숏보다는 프레임이 넓다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "장소와 인물 외형은 대체로 맞지만, 현우가 옆으로 돌아누웠고 경찰의 무릎은 등보다 바닥에 놓여 핵심 제압 자세를 구현하지 못했다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "경찰은 고개를 아래로 숙여 현우의 상체와 붙잡은 팔 쪽을 향한다. 현우의 얼굴은 화면 왼쪽 아래로 돌아가 바닥에 닿아 있고 눈은 감겨 있다. 겨누는 무기나 이동 동작은 없다.",
        "built_space": "왼쪽에는 천막과 금속 지주, 상자와 과일 진열대가 이어지고 오른쪽에는 낡은 벽, 금속 문, 굵은 수직 배관과 가는 배관이 보인다. 파란 통 하나와 붉은 판자 하나도 참조와 대응한다. 두 사람은 오른쪽 벽 가까운 골목 바닥에 있으나 경찰의 몸과 다리가 제압 접점을 가려, 등을 무릎으로 누른다는 공간 관계가 성립하지 않는다.",
        "entities": "주요 인물은 사복 차림의 성인 남성 경찰과 앳된 동아시아계 남성 현우 두 명이다. 현우의 헝클어진 검은 머리와 얼굴은 참조에 대체로 가깝고, 회색 겉셔츠를 입고 있으며 얼굴의 작은 상처와 바짓단 부근의 붉은 부상 흔적이 보인다. 국적이나 정확한 나이는 외형만으로 확정할 수 없다. 넘어지거나 기울어진 상자들이 남아 있고, 읽을 수 있는 문구나 그래픽 오버레이는 보이지 않는다.",
        "hard_violations": [
         "현우의 가슴이 카메라 쪽으로 돌아나온 옆누운 자세이며, 보이는 경찰의 무릎과 정강이는 등 위가 아니라 몸 옆 바닥에 놓여 있다. 엎드린 현우의 등을 무릎으로 누르는 필수 배치와 다르다."
        ],
        "physics": "현우의 머리와 몸 옆면, 보이는 다리는 바닥에 지지된다. 뒤로 올라간 팔 부근에는 경찰의 손이 닿아 있어 완전히 무지지로 떠 있는 자세는 아니다. 경찰도 접힌 다리와 반대쪽 발로 지지되지만, 그 하중이 현우의 등 위 무릎 접촉으로 전달되는 모습은 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "경찰은 몸과 고개를 현우의 등 쪽으로 숙이고 있다. 현우는 얼굴을 화면 오른쪽으로 돌린 채 뺨을 바닥에 대고 있으며 시선은 아래로 떨어져 있다. 무전기는 허리에 매달려 있고 조작하거나 누군가를 겨누는 물체는 없다.",
        "built_space": "왼쪽 천막 지주와 상자 더미, 오른쪽 낡은 벽과 금속 문, 굵고 가는 수직 배관, 파란 통 하나, 붉은 판자 하나가 참조 장소와 대응한다. 오른쪽 문 앞의 목재 통 하나와 경찰 발 근처의 사각 바닥 덮개도 보인다. 현우는 골목 바닥에 길게 엎드려 있고 경찰은 그 오른쪽에서 등을 누른다. 무릎 접촉부와 접지한 발 주변에 바닥 여백이 있어 두 신체의 배치가 명료하다. 다만 두 사람의 거의 전신과 긴 골목까지 보여 상체 중심 미디엄 숏보다 넓다.",
        "entities": "사복 재킷과 밝은 바지를 입은 성인 남성 경찰, 헝클어진 검은 머리의 앳된 동아시아계 남성 현우가 보인다. 현우의 얼굴과 체격은 참조에 대체로 부합하며 어두운 겉셔츠를 입고 얼굴에는 작은 상처가 남아 있다. 다리는 바지로 덮여 개에 물린 상처의 지속 여부는 확인하기 어렵다. 경찰 허리에는 무전기와 장비 주머니가 부착되어 있다. 골목에는 기울어진 상자와 흩어진 과일이 남아 있으며 판독 가능한 글자나 오버레이는 보이지 않는다.",
        "hard_violations": [],
        "physics": "현우의 뺨, 가슴과 배, 보이는 팔과 손, 다리가 바닥에 놓여 있다. 경찰의 한쪽 무릎은 현우의 등 위에 명확히 닿고 반대쪽 발은 바닥을 딛는다. 굽힌 무릎과 접지한 발 사이로 체중을 전달할 수 있는 자세이며, 경찰의 손도 현우의 몸에 닿아 있다. 지지 없이 떠 있는 신체나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.042
   },
   "adjusted": {
    "A": 1.75,
    "B": 0.792
   },
   "violations": {
    "A": [
     "[gemini-pro] 물리적으로 불가능한 해부학 (경찰의 오른 무릎 아래로 튀어나와 현우의 옆구리에 얹힌 불가능한 각도의 세 번째 신발/발)"
    ],
    "B": [
     "[gemini-pro] 지시된 연출 위치 위반 (경찰이 무릎으로 등을 짓누르지 않고 현우의 옆 바닥에 무릎을 꿇은 채 배치됨)",
     "[gemini-pro] 물리적으로 불가능한 해부학 (현우의 상체와 하체가 이어지는 허리 부분의 심하게 왜곡되고 늘어난 비율)",
     "[gpt-high] 현우의 가슴이 카메라 쪽으로 돌아나온 옆누운 자세이며, 보이는 경찰의 무릎과 정강이는 등 위가 아니라 몸 옆 바닥에 놓여 있다. 엎드린 현우의 등을 무릎으로 누르는 필수 배치와 다르다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 792
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "지정된 '무릎으로 등을 짓누르는' 연출을 시도했으나, 경찰의 오른 무릎 아래에 세 번째 발이 기형적으로 생성되는 치명적인 해부학적 오류가 발생했고 현우의 다리 상처가 누락되었습니다.  ★위반: [gemini-pro] 물리적으로 불가능한 해부학 (경찰의 오른 무릎 아래로 튀어나와 현우의 옆구리에 얹힌 불가능한 각도의 세 번째 신발/발)"
   },
   {
    "label": "B",
    "score": 792,
    "verdict_ko": "캐릭터 외형과 상처 디테일은 우수하나, 필수 지침인 '무릎으로 등을 짓누르는' 자세를 완전히 무시하고 옆에 무릎을 꿇은 채 손으로 제압하고 있어 핵심 연출을 위반했습니다.  ★위반: [gemini-pro] 지시된 연출 위치 위반 (경찰이 무릎으로 등을 짓누르지 않고 현우의 옆 바닥에 무릎을 꿇은 채 배치됨) / [gemini-pro] 물리적으로 불가능한 해부학 (현우의 상체와 하체가 이어지는 허리 부분의 심하게 왜곡되고 늘어난 비율) / [gpt-high] 현우의 가슴이 카메라 쪽으로 돌아나온 옆누운 자세이며, 보이는 경찰의 무릎과 정강이는 등 위가 아니라 몸 옆 바닥에 놓여 있다. 엎드린 현우의 등을 무릎으로 누르는 필수 배치와 다르다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_market_capture_alley_13d54d.png",
    "asset_id": "5e0698a7-8a58-4159-968f-d06cd12bcc47",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-ffc8-7872-8dd3-ce2c443a66f0",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S20sh16__bgfirst_bg.png",
   "bg_asset_id": "43c81765-5840-45d5-8f30-f37e76cc5841",
   "bg_record_key": "S20sh16::bgfirst_bg",
   "chain_winner": true,
   "authority": "groupbg",
   "group_key": "market_capture_alley",
   "groupbg_asset_id": "5e0698a7-8a58-4159-968f-d06cd12bcc47"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S20sh16::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:54:18.895258+00:00",
  "fingerprint": "afaa684f5d915c1bbb94bf815009689f3ad3c31a4746fb1466a57e565bbe5e01",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S20sh16_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S20sh16_sel.png",
  "source_sha256": "91ea26bee321a7679ad0db040c28ab24382f8b9794315f6a08d33944521c1b02",
  "file": "S20sh16_cine.png",
  "staged_sha256": "004cd990a9b509db963ab443e017fbcbbcbdd3ea32d06a1b0cfa06badca85f15",
  "latency_ms": 12126
 },
 "S21sh2::signage": {
  "fp": "6f54fd162b569634",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::885ecb7c61f2a28a": {
  "subjects": [],
  "subject_text": "라울의 컨테이너 내부\n잠자리를 둘 수 있는 소박한 컨테이너 주거 공간. 좁은 직사각형 실내에 출입문과 간단한 침구가 있다.",
  "identity": "canonical",
  "scope_id": "L170",
  "scope_role": "location_interior",
  "scope_sha": "bfeabefdb3bedf56"
 },
 "S21sh2::bgfirst_bg": {
  "input_fingerprint": "0d65fd073a5c8b80",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 거칠게 열린 문 안으로 불쑥 몸을 들이민 라울의 다급한 전신.\n\nLOCATION (lock): Just inside the open doorway of a container home, with daylight entering the room where a child has been sleeping.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Container doorway and door (The door has been abruptly opened and 라울 is entering) — Viewed obliquely from inside, with the opening and part of the open door visible beside his body; used as Frame his full-height entry and establish the threshold without enlarging the door through foreground distortion.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daytime ambient illumination appropriate to the container interior, keeping the doorway and entering figure readable without assigning an unsupported lighting source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 거칠게 열린 문 안으로 불쑥 몸을 들이민 라울의 다급한 전신.\n\nLOCATION (lock): Just inside the open doorway of a container home, with daylight entering the room where a child has been sleeping.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Container doorway and door (The door has been abruptly opened and 라울 is entering) — Viewed obliquely from inside, with the opening and part of the open door visible beside his body; used as Frame his full-height entry and establish the threshold without enlarging the door through foreground distortion.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daytime ambient illumination appropriate to the container interior, keeping the doorway and entering figure readable without assigning an unsupported lighting source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S21sh2__bgfirst_bg.png",
  "asset_id": "c6f5da55-c68d-4a4c-892e-04c73c186cf0",
  "input_asset_ids": [
   "f81f4ae2-7f77-4950-92a2-abc01a2199a4",
   "948d45e7-41c1-4297-974a-8863400e8720"
  ]
 },
 "S21sh2": {
  "input_fingerprint": "f2c3323d1f281ad6",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 거칠게 열린 문 안으로 불쑥 몸을 들이민 라울의 다급한 전신.\n\nLOCATION (lock): Just inside the open doorway of a container home, with daylight entering the room where a child has been sleeping. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Container doorway and door (The door has been abruptly opened and 라울 is entering) — Viewed obliquely from inside, with the opening and part of the open door visible beside his body; used as Frame his full-height entry and establish the threshold without enlarging the door through foreground distortion.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daytime ambient illumination appropriate to the container interior, keeping the doorway and entering figure readable without assigning an unsupported lighting source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The door of Raul's container has been thrown open. 라울: He has just entered the container through the open doorway.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 거칠게 열린 문 안으로 불쑥 몸을 들이민 라울의 다급한 전신.\n\nLOCATION (lock): Just inside the open doorway of a container home, with daylight entering the room where a child has been sleeping. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Container doorway and door (The door has been abruptly opened and 라울 is entering) — Viewed obliquely from inside, with the opening and part of the open door visible beside his body; used as Frame his full-height entry and establish the threshold without enlarging the door through foreground distortion.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daytime ambient illumination appropriate to the container interior, keeping the doorway and entering figure readable without assigning an unsupported lighting source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The door of Raul's container has been thrown open. 라울: He has just entered the container through the open doorway.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 거칠게 열린 문 안으로 불쑥 몸을 들이민 라울의 다급한 전신.\n\nLOCATION (lock): Just inside the open doorway of a container home, with daylight entering the room where a child has been sleeping. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Container doorway and door (The door has been abruptly opened and 라울 is entering) — Viewed obliquely from inside, with the opening and part of the open door visible beside his body; used as Frame his full-height entry and establish the threshold without enlarging the door through foreground distortion.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daytime ambient illumination appropriate to the container interior, keeping the doorway and entering figure readable without assigning an unsupported lighting source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The door of Raul's container has been thrown open. 라울: He has just entered the container through the open doorway.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S21sh2__bgfirst_bg.png",
     "asset_id": "c6f5da55-c68d-4a4c-892e-04c73c186cf0",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S21sh2.png",
     "asset_id": "f81f4ae2-7f77-4950-92a2-abc01a2199a4",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163202>",
     "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L170B01.png",
     "asset_id": "948d45e7-41c1-4297-974a-8863400e8720",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163202>",
     "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "소년은 약간 왼쪽을 바라보며 방 안으로 들어오고 있음.",
    "built_space": "복도의 방향은 유지되었으나 방 내부의 구조가 완전히 반전되어 침대가 왼쪽, 문이 오른쪽에 배치됨.",
    "entities": "라울의 인상착의(연령, 피부색, 의상)가 일치함.",
    "hard_violations": [
     "[gemini-pro] physically impossible staging (고정된 로케이션 레퍼런스의 방 내부 구조를 임의로 좌우 반전시킴)"
    ],
    "physics": "오른발은 복도 바닥을 딛고 왼발은 문턱을 넘고 있어 지지력이 자연스러움."
   },
   {
    "label": "B",
    "direction": "소년의 시선은 정면을 향하며 방 안으로 진입하고 있음.",
    "built_space": "레퍼런스와 동일하게 문간이 왼쪽, 침대가 오른쪽에 위치함. 그러나 전경 왼쪽에 문이 존재함에도 불구하고 문틀에 복도 쪽으로 열리는 또 다른 문이 중복으로 생성되어 있음.",
    "entities": "라울(10세 남성, 꽁지머리, 파란 티셔츠)의 인상착의가 일치함.",
    "hard_violations": [
     "[gemini-pro] duplicated fitting (하나의 문틀에 두 개의 문이 존재함)",
     "[gpt-high] 단일 침실 출입구에 손잡이가 달린 전경 문짝과 그 뒤의 별도 문짝을 중복 생성했다."
    ],
    "physics": "왼발은 방바닥을 딛고 있으며, 오른발은 복도 바닥을 차고 들어오는 자세로 지지력이 정상적임."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "다급하게 들어오는 인물의 동작과 프레이밍은 제시문을 잘 따랐으나, 하나의 문틀에 전경의 문과 복도 쪽으로 열린 문이 중복 생성되는 치명적인 구조 오류가 발생했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정된 로케이션의 방 구조(침대와 문의 위치)를 좌우로 임의 변경하여 물리적 공간 제약을 어기는 심각한 오류가 있습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "소년의 시선은 정면을 향하며 방 안으로 진입하고 있음.",
        "built_space": "레퍼런스와 동일하게 문간이 왼쪽, 침대가 오른쪽에 위치함. 그러나 전경 왼쪽에 문이 존재함에도 불구하고 문틀에 복도 쪽으로 열리는 또 다른 문이 중복으로 생성되어 있음.",
        "entities": "라울(10세 남성, 꽁지머리, 파란 티셔츠)의 인상착의가 일치함.",
        "hard_violations": [
         "duplicated fitting (하나의 문틀에 두 개의 문이 존재함)"
        ],
        "physics": "왼발은 방바닥을 딛고 있으며, 오른발은 복도 바닥을 차고 들어오는 자세로 지지력이 정상적임."
       },
       {
        "label": "A",
        "direction": "소년은 약간 왼쪽을 바라보며 방 안으로 들어오고 있음.",
        "built_space": "복도의 방향은 유지되었으나 방 내부의 구조가 완전히 반전되어 침대가 왼쪽, 문이 오른쪽에 배치됨.",
        "entities": "라울의 인상착의(연령, 피부색, 의상)가 일치함.",
        "hard_violations": [
         "physically impossible staging (고정된 로케이션 레퍼런스의 방 내부 구조를 임의로 좌우 반전시킴)"
        ],
        "physics": "오른발은 복도 바닥을 딛고 왼발은 문턱을 넘고 있어 지지력이 자연스러움."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "다급하게 들어오는 인물의 동작과 프레이밍은 제시문을 잘 따랐으나, 하나의 문틀에 전경의 문과 복도 쪽으로 열린 문이 중복 생성되는 치명적인 구조 오류가 발생했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정된 로케이션의 방 구조(침대와 문의 위치)를 좌우로 임의 변경하여 물리적 공간 제약을 어기는 심각한 오류가 있습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "소년의 시선은 정면을 향하며 방 안으로 진입하고 있음.",
        "built_space": "레퍼런스와 동일하게 문간이 왼쪽, 침대가 오른쪽에 위치함. 그러나 전경 왼쪽에 문이 존재함에도 불구하고 문틀에 복도 쪽으로 열리는 또 다른 문이 중복으로 생성되어 있음.",
        "entities": "라울(10세 남성, 꽁지머리, 파란 티셔츠)의 인상착의가 일치함.",
        "hard_violations": [
         "duplicated fitting (하나의 문틀에 두 개의 문이 존재함)"
        ],
        "physics": "왼발은 방바닥을 딛고 있으며, 오른발은 복도 바닥을 차고 들어오는 자세로 지지력이 정상적임."
       },
       {
        "label": "A",
        "direction": "소년은 약간 왼쪽을 바라보며 방 안으로 들어오고 있음.",
        "built_space": "복도의 방향은 유지되었으나 방 내부의 구조가 완전히 반전되어 침대가 왼쪽, 문이 오른쪽에 배치됨.",
        "entities": "라울의 인상착의(연령, 피부색, 의상)가 일치함.",
        "hard_violations": [
         "physically impossible staging (고정된 로케이션 레퍼런스의 방 내부 구조를 임의로 좌우 반전시킴)"
        ],
        "physics": "오른발은 복도 바닥을 딛고 왼발은 문턱을 넘고 있어 지지력이 자연스러움."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "다급하게 방 안으로 들어오는 전신과 시선은 정확하지만, 같은 출입구에 문짝이 중복되어 공간 구조에 치명적인 오류가 있다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "문 옆으로 진입하는 전신과 지지 동작은 자연스럽지만, 참조 공간을 좌우 반전했고 아직 문턱 바깥에 있어 막 들어온 순간의 구현이 덜 정확하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "라울은 화면 오른쪽 침대가 있는 방 안을 바라보며 그쪽으로 상체를 기울인다. 앞발은 문턱 안쪽에 들어왔고 뒷발은 들려 있어, 밖에서 방 안으로 급히 진입하는 방향이 명확하다. 손에 든 물건이나 조준 대상은 없다.",
        "built_space": "오른쪽 목제 침대 한 개, 체크 베개 한 개, 벽 사진 네 장, 문 옆 스위치 한 개와 뒤쪽 생활공간이 참조와 대체로 같은 위치에 있다. 그러나 왼쪽 전경에 손잡이가 달린 흰 문짝이 있고, 그 바로 뒤 같은 출입구에 경첩과 상단 장치가 보이는 별도 문짝이 또 있다. 참조의 단일 침실문을 이중 문짝으로 만들었다. 실내에서 출입구를 비스듬히 보는 와이드 구도와 인물 전신 크기는 요청에 맞는다.",
        "entities": "보이는 사람은 라울 한 명뿐이다. 어린 얼굴, 갈색 피부, 뒤로 묶은 곱슬머리, 마른 아동 체형과 남색 반팔은 인물 참조에 가깝다. 외형은 지정된 혼혈 남자아이 설정과 모순되지 않는다. 긴 바지와 어두운 신발도 보이지만 참조에는 하의가 없어 일치 여부를 단정할 수 없다. 추가 인물이나 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "단일 침실 출입구에 손잡이가 달린 전경 문짝과 그 뒤의 별도 문짝을 중복 생성했다."
        ],
        "physics": "앞발이 방 바닥에 닿아 체중을 받으며, 굽힌 뒷다리와 서로 다른 위치의 팔이 달려 들어오는 동작을 만든다. 공중에 든 뒷발은 지지발과 연결된 정상적인 보행 단계이며 몸이 떠 있지 않다. 침대와 주변 가구는 바닥에 지지되어 있다."
       },
       {
        "label": "B",
        "direction": "라울은 화면 왼쪽 침대 쪽을 바라보며 상체를 방 안으로 내민다. 다리와 팔의 방향도 복도에서 침실로 들어오는 움직임을 나타낸다. 다만 지지발이 아직 복도 쪽에 있어, 이미 들어온 순간보다는 문턱을 넘기 직전으로 읽힌다. 손에 든 물건은 없다.",
        "built_space": "침대 한 개, 체크 베개 한 개, 사진 네 장, 스위치 한 개와 열린 침실문 한 개가 보인다. 문짝 중복은 없다. 그러나 침대와 사진은 왼쪽, 문짝은 오른쪽이며, 뒤쪽 냉장고와 신발장까지 참조와 좌우가 뒤집혀 있다. 이는 단순한 카메라 각도 변경으로 설명하기 어려운 공간 반전이다. 실내의 비스듬한 와이드 구도에서 전신과 문턱, 열린 문 일부는 모두 보인다.",
        "entities": "라울로 읽히는 아동 한 명만 등장한다. 어린 얼굴, 갈색 피부, 뒤로 묶은 머리와 아동 체형은 참조에 가깝고 지정된 인물 설정과 모순되지 않는다. 반팔은 참조의 짙은 남색보다 밝은 청색이다. 반바지, 흰 양말과 운동화는 보이지만 참조에 해당 부위가 없어 불일치로 단정하지 않는다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "앞쪽 운동화가 복도 바닥에 닿아 몸을 지지하고, 반대쪽 무릎이 굽혀져 발이 뒤로 들려 있다. 상체의 전진과 팔의 벌어짐은 급히 들어오는 한 걸음으로 가능하다. 문은 문틀 쪽에 연결되어 있고, 침대와 가구도 바닥에 놓여 있어 지지 없는 부유는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "다급하게 방 안으로 들어오는 전신과 시선은 정확하지만, 같은 출입구에 문짝이 중복되어 공간 구조에 치명적인 오류가 있다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "문 옆으로 진입하는 전신과 지지 동작은 자연스럽지만, 참조 공간을 좌우 반전했고 아직 문턱 바깥에 있어 막 들어온 순간의 구현이 덜 정확하다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "라울은 화면 오른쪽 침대가 있는 방 안을 바라보며 그쪽으로 상체를 기울인다. 앞발은 문턱 안쪽에 들어왔고 뒷발은 들려 있어, 밖에서 방 안으로 급히 진입하는 방향이 명확하다. 손에 든 물건이나 조준 대상은 없다.",
        "built_space": "오른쪽 목제 침대 한 개, 체크 베개 한 개, 벽 사진 네 장, 문 옆 스위치 한 개와 뒤쪽 생활공간이 참조와 대체로 같은 위치에 있다. 그러나 왼쪽 전경에 손잡이가 달린 흰 문짝이 있고, 그 바로 뒤 같은 출입구에 경첩과 상단 장치가 보이는 별도 문짝이 또 있다. 참조의 단일 침실문을 이중 문짝으로 만들었다. 실내에서 출입구를 비스듬히 보는 와이드 구도와 인물 전신 크기는 요청에 맞는다.",
        "entities": "보이는 사람은 라울 한 명뿐이다. 어린 얼굴, 갈색 피부, 뒤로 묶은 곱슬머리, 마른 아동 체형과 남색 반팔은 인물 참조에 가깝다. 외형은 지정된 혼혈 남자아이 설정과 모순되지 않는다. 긴 바지와 어두운 신발도 보이지만 참조에는 하의가 없어 일치 여부를 단정할 수 없다. 추가 인물이나 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "단일 침실 출입구에 손잡이가 달린 전경 문짝과 그 뒤의 별도 문짝을 중복 생성했다."
        ],
        "physics": "앞발이 방 바닥에 닿아 체중을 받으며, 굽힌 뒷다리와 서로 다른 위치의 팔이 달려 들어오는 동작을 만든다. 공중에 든 뒷발은 지지발과 연결된 정상적인 보행 단계이며 몸이 떠 있지 않다. 침대와 주변 가구는 바닥에 지지되어 있다."
       },
       {
        "label": "A",
        "direction": "라울은 화면 왼쪽 침대 쪽을 바라보며 상체를 방 안으로 내민다. 다리와 팔의 방향도 복도에서 침실로 들어오는 움직임을 나타낸다. 다만 지지발이 아직 복도 쪽에 있어, 이미 들어온 순간보다는 문턱을 넘기 직전으로 읽힌다. 손에 든 물건은 없다.",
        "built_space": "침대 한 개, 체크 베개 한 개, 사진 네 장, 스위치 한 개와 열린 침실문 한 개가 보인다. 문짝 중복은 없다. 그러나 침대와 사진은 왼쪽, 문짝은 오른쪽이며, 뒤쪽 냉장고와 신발장까지 참조와 좌우가 뒤집혀 있다. 이는 단순한 카메라 각도 변경으로 설명하기 어려운 공간 반전이다. 실내의 비스듬한 와이드 구도에서 전신과 문턱, 열린 문 일부는 모두 보인다.",
        "entities": "라울로 읽히는 아동 한 명만 등장한다. 어린 얼굴, 갈색 피부, 뒤로 묶은 머리와 아동 체형은 참조에 가깝고 지정된 인물 설정과 모순되지 않는다. 반팔은 참조의 짙은 남색보다 밝은 청색이다. 반바지, 흰 양말과 운동화는 보이지만 참조에 해당 부위가 없어 불일치로 단정하지 않는다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "앞쪽 운동화가 복도 바닥에 닿아 몸을 지지하고, 반대쪽 무릎이 굽혀져 발이 뒤로 들려 있다. 상체의 전진과 팔의 벌어짐은 급히 들어오는 한 걸음으로 가능하다. 문은 문틀 쪽에 연결되어 있고, 침대와 가구도 바닥에 놓여 있어 지지 없는 부유는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.75,
    "B": 1.5
   },
   "adjusted": {
    "A": 1.5,
    "B": 1.25
   },
   "violations": {
    "B": [
     "[gemini-pro] duplicated fitting (하나의 문틀에 두 개의 문이 존재함)",
     "[gpt-high] 단일 침실 출입구에 손잡이가 달린 전경 문짝과 그 뒤의 별도 문짝을 중복 생성했다."
    ],
    "A": [
     "[gemini-pro] physically impossible staging (고정된 로케이션 레퍼런스의 방 내부 구조를 임의로 좌우 반전시킴)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1250,
   "A": 1500
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1250,
    "verdict_ko": "다급하게 들어오는 인물의 동작과 프레이밍은 제시문을 잘 따랐으나, 하나의 문틀에 전경의 문과 복도 쪽으로 열린 문이 중복 생성되는 치명적인 구조 오류가 발생했습니다.  ★위반: [gemini-pro] duplicated fitting (하나의 문틀에 두 개의 문이 존재함) / [gpt-high] 단일 침실 출입구에 손잡이가 달린 전경 문짝과 그 뒤의 별도 문짝을 중복 생성했다."
   },
   {
    "label": "A",
    "score": 1500,
    "verdict_ko": "지정된 로케이션의 방 구조(침대와 문의 위치)를 좌우로 임의 변경하여 물리적 공간 제약을 어기는 심각한 오류가 있습니다.  ★위반: [gemini-pro] physically impossible staging (고정된 로케이션 레퍼런스의 방 내부 구조를 임의로 좌우 반전시킴)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L170B01.png",
    "asset_id": "948d45e7-41c1-4297-974a-8863400e8720",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163202>",
    "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-04a7-72a7-bcb4-ea453348b033",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S21sh2__bgfirst_bg.png",
   "bg_asset_id": "c6f5da55-c68d-4a4c-892e-04c73c186cf0",
   "bg_record_key": "S21sh2::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S21sh2::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:56:24.583352+00:00",
  "fingerprint": "c8264c334541176bcb10f9931549410d20bebd305302b74338cb728b7048d2d4",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S21sh2_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S21sh2_sel.png",
  "source_sha256": "61d138c011dbe3a39123487a87b48928aae0f9acc55e66d01b9c397372e39e44",
  "file": "S21sh2_cine.png",
  "staged_sha256": "a4b5a7e0f361342aa3056516a3767084d896e387a727c5ae29458b58d4395e68",
  "latency_ms": 63922
 },
 "S21sh5::signage": {
  "fp": "764871cb15ce99ad",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S21sh5": {
  "input_fingerprint": "8867133028ca65ef",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 라울의 손가락 끝이 가리키는 텅 빈 공간을 응시하며 눈동자가 커진 앰버의 놀란 얼굴 클로즈업.\n\nLOCATION (lock): In the sleeping area inside the container home, facing an empty patch of room in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container interior beside 앰버 (Only a narrow, unoccupied portion is visible at the edge of the close framing); used as Retain breathing room on the side of her attention without inventing an object to explain 찰리's absence.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the interior's neutral ambient illumination and soft tonal separation, presenting the empty-space reaction as direct reality without a visual distortion.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The container door remains open after the abrupt entrance. 앰버: She has roused from sleep and is sitting up. Her mask and waist tool pouch remain established belongings without a stated removal.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 라울의 손가락 끝이 가리키는 텅 빈 공간을 응시하며 눈동자가 커진 앰버의 놀란 얼굴 클로즈업.\n\nLOCATION (lock): In the sleeping area inside the container home, facing an empty patch of room in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container interior beside 앰버 (Only a narrow, unoccupied portion is visible at the edge of the close framing); used as Retain breathing room on the side of her attention without inventing an object to explain 찰리's absence.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the interior's neutral ambient illumination and soft tonal separation, presenting the empty-space reaction as direct reality without a visual distortion.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The container door remains open after the abrupt entrance. 앰버: She has roused from sleep and is sitting up. Her mask and waist tool pouch remain established belongings without a stated removal.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 라울의 손가락 끝이 가리키는 텅 빈 공간을 응시하며 눈동자가 커진 앰버의 놀란 얼굴 클로즈업.\n\nLOCATION (lock): In the sleeping area inside the container home, facing an empty patch of room in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container interior beside 앰버 (Only a narrow, unoccupied portion is visible at the edge of the close framing); used as Retain breathing room on the side of her attention without inventing an object to explain 찰리's absence.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the interior's neutral ambient illumination and soft tonal separation, presenting the empty-space reaction as direct reality without a visual distortion.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The container door remains open after the abrupt entrance. 앰버: She has roused from sleep and is sitting up. Her mask and waist tool pouch remain established belongings without a stated removal.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "B",
    "direction": "앰버는 눈을 크게 뜨고 화면 왼쪽 전경의 손이 있는 방향을 바라봅니다. 왼쪽에서 들어온 검지는 오른쪽 위, 앰버의 상체 쪽으로 향합니다. 검지가 가리키는 방향과 앰버의 시선이 같은 빈 공간에 모이는 관계로 보이지 않습니다.",
    "built_space": "중앙 왼쪽에 열린 실내 문 하나, 그 너머 밝은 외부 출입구 하나와 천장등 하나가 보입니다. 연한 세로 패널 벽과 복도는 기준 장소와 유사합니다. 그러나 왼쪽 가장자리의 기존 침대 침구와 별도로 오른쪽 앰버 아래에 목재 끝판과 침구가 있는 두 번째 침대가 보입니다. 배경과 허리까지 크게 포함하여 좁은 배경만 남기는 얼굴 클로즈업이 아닙니다.",
    "entities": "금발의 어린 여자아이 한 명은 둥근 얼굴, 큰 눈, 남색 반팔 등 앰버 기준과 대체로 맞습니다. 한국계 백인 혼혈이라는 설정은 외모만으로 확정할 수 없습니다. 허리에 갈색 공구 주머니와 파란 마스크가 보입니다. 전경에는 라울의 손으로 해석되는 손 하나가 있으나, 제공된 자료로 그 인물의 정확한 손 외형은 확인할 수 없습니다. 읽을 수 있는 글자는 없습니다.",
    "hard_violations": [
     "기준 침실의 왼쪽 침대를 남겨둔 채 앰버 아래 오른쪽에 별도의 목재 침대를 추가하여, 고정된 침실 배치에 없는 침대를 만들었습니다."
    ],
    "physics": "앰버의 하체는 화면 아래로 잘렸지만 몸통은 침대 위에 앉은 자세로 이어지며 침구가 지지면입니다. 전경 손은 왼쪽 프레임 밖으로 이어지는 손목과 팔에 연결되어 있어 공중에 떠 있는 손은 아닙니다. 공구 주머니와 마스크는 허리 부근에 걸려 있습니다. 뚜렷한 해부학적 불가능은 없습니다."
   },
   {
    "label": "A",
    "direction": "앰버의 얼굴과 시선은 화면 오른쪽의 사람이 없는 공간을 향합니다. 눈을 크게 뜨고 입을 조금 벌려 놀란 반응을 보이며 렌즈를 직접 보지 않습니다. 라울의 손은 프레임 밖이므로 손끝과 목표 공간의 정렬 자체는 확인할 수 없지만, 보이는 시선 방향에는 별도의 대상 물체가 없습니다.",
    "built_space": "오른쪽 아래에 침대 하나와 청회색 침구가 보입니다. 뒤에는 가로 보강대가 있는 철제 문짝 형태의 면과 오른쪽 세로 골이 있는 금속 벽이 드러납니다. 기준의 매끈한 연색 실내 패널 및 일반 실내문과 마감이 다릅니다. 열린 출입문은 이 구도에서 확인되지 않습니다. 앰버의 허리와 넓은 침구까지 포함해 얼굴 클로즈업 및 좁은 배경이라는 지시에는 미달합니다.",
    "entities": "금발, 둥근 얼굴, 큰 눈을 가진 어린 여자아이 한 명이며 남색 반팔을 입어 앰버 기준과 대체로 일치합니다. 혼혈 배경은 외모만으로 확정할 수 없습니다. 흰색 방진 마스크가 목에 걸려 있고 갈색 공구 주머니가 허리띠에 부착되어 있어 소지품은 유지됩니다. 추가 인물이나 부유하는 손, 읽을 수 있는 글자는 없습니다.",
    "hard_violations": [],
    "physics": "앰버는 침대 가장자리에 상체를 세워 앉아 있으며, 하체 일부는 잘렸지만 골반 위치가 매트리스와 이어집니다. 마스크는 목 주변 끈으로, 공구 주머니는 허리띠로 지지됩니다. 머리카락과 옷은 중력에 맞게 늘어져 있고 지지 없이 떠 있는 신체나 물체는 없습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": null,
     "normalized": null,
     "ok": false
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "손가락이 빈 공간보다 앰버 쪽을 향하고 침대가 추가된 배치이며, 허리까지 보이는 넓은 구도가 얼굴 클로즈업 지시를 어깁니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "빈 오른쪽 공간을 응시하는 놀란 반응은 더 정확하지만, 얼굴 클로즈업보다 넓고 노출된 철제 벽·문이 기준 침실의 마감과 다릅니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버는 눈을 크게 뜨고 화면 왼쪽 전경의 손이 있는 방향을 바라봅니다. 왼쪽에서 들어온 검지는 오른쪽 위, 앰버의 상체 쪽으로 향합니다. 검지가 가리키는 방향과 앰버의 시선이 같은 빈 공간에 모이는 관계로 보이지 않습니다.",
        "built_space": "중앙 왼쪽에 열린 실내 문 하나, 그 너머 밝은 외부 출입구 하나와 천장등 하나가 보입니다. 연한 세로 패널 벽과 복도는 기준 장소와 유사합니다. 그러나 왼쪽 가장자리의 기존 침대 침구와 별도로 오른쪽 앰버 아래에 목재 끝판과 침구가 있는 두 번째 침대가 보입니다. 배경과 허리까지 크게 포함하여 좁은 배경만 남기는 얼굴 클로즈업이 아닙니다.",
        "entities": "금발의 어린 여자아이 한 명은 둥근 얼굴, 큰 눈, 남색 반팔 등 앰버 기준과 대체로 맞습니다. 한국계 백인 혼혈이라는 설정은 외모만으로 확정할 수 없습니다. 허리에 갈색 공구 주머니와 파란 마스크가 보입니다. 전경에는 라울의 손으로 해석되는 손 하나가 있으나, 제공된 자료로 그 인물의 정확한 손 외형은 확인할 수 없습니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "기준 침실의 왼쪽 침대를 남겨둔 채 앰버 아래 오른쪽에 별도의 목재 침대를 추가하여, 고정된 침실 배치에 없는 침대를 만들었습니다."
        ],
        "physics": "앰버의 하체는 화면 아래로 잘렸지만 몸통은 침대 위에 앉은 자세로 이어지며 침구가 지지면입니다. 전경 손은 왼쪽 프레임 밖으로 이어지는 손목과 팔에 연결되어 있어 공중에 떠 있는 손은 아닙니다. 공구 주머니와 마스크는 허리 부근에 걸려 있습니다. 뚜렷한 해부학적 불가능은 없습니다."
       },
       {
        "label": "B",
        "direction": "앰버의 얼굴과 시선은 화면 오른쪽의 사람이 없는 공간을 향합니다. 눈을 크게 뜨고 입을 조금 벌려 놀란 반응을 보이며 렌즈를 직접 보지 않습니다. 라울의 손은 프레임 밖이므로 손끝과 목표 공간의 정렬 자체는 확인할 수 없지만, 보이는 시선 방향에는 별도의 대상 물체가 없습니다.",
        "built_space": "오른쪽 아래에 침대 하나와 청회색 침구가 보입니다. 뒤에는 가로 보강대가 있는 철제 문짝 형태의 면과 오른쪽 세로 골이 있는 금속 벽이 드러납니다. 기준의 매끈한 연색 실내 패널 및 일반 실내문과 마감이 다릅니다. 열린 출입문은 이 구도에서 확인되지 않습니다. 앰버의 허리와 넓은 침구까지 포함해 얼굴 클로즈업 및 좁은 배경이라는 지시에는 미달합니다.",
        "entities": "금발, 둥근 얼굴, 큰 눈을 가진 어린 여자아이 한 명이며 남색 반팔을 입어 앰버 기준과 대체로 일치합니다. 혼혈 배경은 외모만으로 확정할 수 없습니다. 흰색 방진 마스크가 목에 걸려 있고 갈색 공구 주머니가 허리띠에 부착되어 있어 소지품은 유지됩니다. 추가 인물이나 부유하는 손, 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "앰버는 침대 가장자리에 상체를 세워 앉아 있으며, 하체 일부는 잘렸지만 골반 위치가 매트리스와 이어집니다. 마스크는 목 주변 끈으로, 공구 주머니는 허리띠로 지지됩니다. 머리카락과 옷은 중력에 맞게 늘어져 있고 지지 없이 떠 있는 신체나 물체는 없습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "손가락이 빈 공간보다 앰버 쪽을 향하고 침대가 추가된 배치이며, 허리까지 보이는 넓은 구도가 얼굴 클로즈업 지시를 어깁니다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "빈 오른쪽 공간을 응시하는 놀란 반응은 더 정확하지만, 얼굴 클로즈업보다 넓고 노출된 철제 벽·문이 기준 침실의 마감과 다릅니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "앰버는 눈을 크게 뜨고 화면 왼쪽 전경의 손이 있는 방향을 바라봅니다. 왼쪽에서 들어온 검지는 오른쪽 위, 앰버의 상체 쪽으로 향합니다. 검지가 가리키는 방향과 앰버의 시선이 같은 빈 공간에 모이는 관계로 보이지 않습니다.",
        "built_space": "중앙 왼쪽에 열린 실내 문 하나, 그 너머 밝은 외부 출입구 하나와 천장등 하나가 보입니다. 연한 세로 패널 벽과 복도는 기준 장소와 유사합니다. 그러나 왼쪽 가장자리의 기존 침대 침구와 별도로 오른쪽 앰버 아래에 목재 끝판과 침구가 있는 두 번째 침대가 보입니다. 배경과 허리까지 크게 포함하여 좁은 배경만 남기는 얼굴 클로즈업이 아닙니다.",
        "entities": "금발의 어린 여자아이 한 명은 둥근 얼굴, 큰 눈, 남색 반팔 등 앰버 기준과 대체로 맞습니다. 한국계 백인 혼혈이라는 설정은 외모만으로 확정할 수 없습니다. 허리에 갈색 공구 주머니와 파란 마스크가 보입니다. 전경에는 라울의 손으로 해석되는 손 하나가 있으나, 제공된 자료로 그 인물의 정확한 손 외형은 확인할 수 없습니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "기준 침실의 왼쪽 침대를 남겨둔 채 앰버 아래 오른쪽에 별도의 목재 침대를 추가하여, 고정된 침실 배치에 없는 침대를 만들었습니다."
        ],
        "physics": "앰버의 하체는 화면 아래로 잘렸지만 몸통은 침대 위에 앉은 자세로 이어지며 침구가 지지면입니다. 전경 손은 왼쪽 프레임 밖으로 이어지는 손목과 팔에 연결되어 있어 공중에 떠 있는 손은 아닙니다. 공구 주머니와 마스크는 허리 부근에 걸려 있습니다. 뚜렷한 해부학적 불가능은 없습니다."
       },
       {
        "label": "A",
        "direction": "앰버의 얼굴과 시선은 화면 오른쪽의 사람이 없는 공간을 향합니다. 눈을 크게 뜨고 입을 조금 벌려 놀란 반응을 보이며 렌즈를 직접 보지 않습니다. 라울의 손은 프레임 밖이므로 손끝과 목표 공간의 정렬 자체는 확인할 수 없지만, 보이는 시선 방향에는 별도의 대상 물체가 없습니다.",
        "built_space": "오른쪽 아래에 침대 하나와 청회색 침구가 보입니다. 뒤에는 가로 보강대가 있는 철제 문짝 형태의 면과 오른쪽 세로 골이 있는 금속 벽이 드러납니다. 기준의 매끈한 연색 실내 패널 및 일반 실내문과 마감이 다릅니다. 열린 출입문은 이 구도에서 확인되지 않습니다. 앰버의 허리와 넓은 침구까지 포함해 얼굴 클로즈업 및 좁은 배경이라는 지시에는 미달합니다.",
        "entities": "금발, 둥근 얼굴, 큰 눈을 가진 어린 여자아이 한 명이며 남색 반팔을 입어 앰버 기준과 대체로 일치합니다. 혼혈 배경은 외모만으로 확정할 수 없습니다. 흰색 방진 마스크가 목에 걸려 있고 갈색 공구 주머니가 허리띠에 부착되어 있어 소지품은 유지됩니다. 추가 인물이나 부유하는 손, 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "앰버는 침대 가장자리에 상체를 세워 앉아 있으며, 하체 일부는 잘렸지만 골반 위치가 매트리스와 이어집니다. 마스크는 목 주변 끈으로, 공구 주머니는 허리띠로 지지됩니다. 머리카락과 옷은 중력에 맞게 늘어져 있고 지지 없이 떠 있는 신체나 물체는 없습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gemini-pro"
   ],
   "route": "single_reverse"
  },
  "totals": {
   "B": 3,
   "A": 5
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 3,
    "verdict_ko": "손가락이 빈 공간보다 앰버 쪽을 향하고 침대가 추가된 배치이며, 허리까지 보이는 넓은 구도가 얼굴 클로즈업 지시를 어깁니다."
   },
   {
    "label": "A",
    "score": 5,
    "verdict_ko": "빈 오른쪽 공간을 응시하는 놀란 반응은 더 정확하지만, 얼굴 클로즈업보다 넓고 노출된 철제 벽·문이 기준 침실의 마감과 다릅니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S21sh2_sel.png",
    "asset_id": "c19425fd-7b0d-4fdc-98cc-13dfe4ae3f44",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-080a-704d-b2ee-5ec3f7b60e2d",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S21sh2"
  }
 },
 "S21sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T04:58:13.704133+00:00",
  "fingerprint": "8903217354ac2542f4b2b0a154d9f93e79077d5865e533969931afb57f605eea",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S21sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S21sh5_sel.png",
  "source_sha256": "40c57abbb8058902de0bc857c0c0af591e6f4f3073900f8583904432bbb42c5c",
  "file": "S21sh5_cine.png",
  "staged_sha256": "4efc4eca85d6436192c843a31ee59260d7b016c5b9fd361bb2570af396caf2e0",
  "latency_ms": 14783
 },
 "S22sh3::signage": {
  "fp": "4ca48b7b76a20d29",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S22sh3": {
  "input_fingerprint": "c1a7f60912649371",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바나나를 향해 커다란 금속 손가락을 뻗은 찰리의 손 클로즈업.\n\nLOCATION (lock): At the street-facing banana display of an outdoor market stall in the refugee settlement. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Bananas (Offered for sale and not yet touched by 찰리); used as Provide the clearly visible destination of the fingers while occupying less than a third of the image; Banana stall (Visible in a limited area around the offered fruit) — The camera views the selling area obliquely from the side of 찰리's reaching arm; used as Anchor the hand and fruit within the market rather than isolating them against an undefined background.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daytime ambient light with restrained contrast, allowing the metal hand and the bananas' own color to remain distinct without adding a colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Bananas are displayed for sale at the market stall. Charlie's washed but worn gorilla-shaped body retains its blue-lit eyes and worn Ubik logo, with the old coat and hat still serving as his disguise.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바나나를 향해 커다란 금속 손가락을 뻗은 찰리의 손 클로즈업.\n\nLOCATION (lock): At the street-facing banana display of an outdoor market stall in the refugee settlement. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Bananas (Offered for sale and not yet touched by 찰리); used as Provide the clearly visible destination of the fingers while occupying less than a third of the image; Banana stall (Visible in a limited area around the offered fruit) — The camera views the selling area obliquely from the side of 찰리's reaching arm; used as Anchor the hand and fruit within the market rather than isolating them against an undefined background.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daytime ambient light with restrained contrast, allowing the metal hand and the bananas' own color to remain distinct without adding a colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Bananas are displayed for sale at the market stall. Charlie's washed but worn gorilla-shaped body retains its blue-lit eyes and worn Ubik logo, with the old coat and hat still serving as his disguise.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바나나를 향해 커다란 금속 손가락을 뻗은 찰리의 손 클로즈업.\n\nLOCATION (lock): At the street-facing banana display of an outdoor market stall in the refugee settlement. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Bananas (Offered for sale and not yet touched by 찰리); used as Provide the clearly visible destination of the fingers while occupying less than a third of the image; Banana stall (Visible in a limited area around the offered fruit) — The camera views the selling area obliquely from the side of 찰리's reaching arm; used as Anchor the hand and fruit within the market rather than isolating them against an undefined background.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daytime ambient light with restrained contrast, allowing the metal hand and the bananas' own color to remain distinct without adding a colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Bananas are displayed for sale at the market stall. Charlie's washed but worn gorilla-shaped body retains its blue-lit eyes and worn Ubik logo, with the old coat and hat still serving as his disguise.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S22sh3__bgfirst_bg.png",
     "asset_id": "e6ccf295-a01a-4188-a4ca-5ce36e1609e9",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S22sh3.png",
     "asset_id": "774388e6-e167-4f76-a648-7a8dce68ade1",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_banana_market_stall_8677ef.png",
     "asset_id": "0333daa2-9ee6-4388-bf59-c510fb37f958",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "금속 손가락이 우측 하단의 바나나 무더기를 향함.",
    "built_space": "참조 이미지의 야외 시장. 천막과 가판대 배치가 일치함.",
    "entities": "찰리의 얼굴과 상반신 전체 노출, 팔에 Ubik 로고, 상자에 불필요한 문자 존재. 바나나는 화면 우측에 위치함.",
    "hard_violations": [
     "[gemini-pro] 나무 상자에 프롬프트에 없는 텍스트(문자) 누출",
     "[gpt-high] 전완에 읽을 수 있는 UBik 글자와 선명한 도형 로고가 노출되어, 글자와 로고를 금지한 명시 조건을 위반한다."
    ],
    "physics": "몸통에 연결된 팔이 자연스럽게 손을 지지함."
   },
   {
    "label": "B",
    "direction": "금속 손가락이 우측 바나나에 닿기 직전으로 뻗어 있음.",
    "built_space": "참조 이미지의 시장 배경. 우측에 바나나 가판대와 골목 배치가 일치함.",
    "entities": "코트 소매를 입은 팔과 손만 클로즈업됨. 손등에 Ubik 로고가 있으며, 바나나가 화면의 3분의 1 이하를 차지함.",
    "hard_violations": [
     "[gpt-high] 손등에 읽을 수 있는 Ubik 글자가 노출되어, 읽히는 글자와 로고를 금지한 명시 조건을 위반한다."
    ],
    "physics": "코트 소매에서 뻗어나온 팔이 금속 손을 지지함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "요구된 '손 클로즈업' 프레이밍을 정확히 구현했으며, 바나나와 배경 요소의 배치 규정을 충실히 따랐습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "'손 클로즈업' 지시를 무시하고 상반신 전체를 보여주었으며, 나무 상자에 불필요한 문자가 누출되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "금속 손가락이 우측 하단의 바나나 무더기를 향함.",
        "built_space": "참조 이미지의 야외 시장. 천막과 가판대 배치가 일치함.",
        "entities": "찰리의 얼굴과 상반신 전체 노출, 팔에 Ubik 로고, 상자에 불필요한 문자 존재. 바나나는 화면 우측에 위치함.",
        "hard_violations": [
         "나무 상자에 프롬프트에 없는 텍스트(문자) 누출"
        ],
        "physics": "몸통에 연결된 팔이 자연스럽게 손을 지지함."
       },
       {
        "label": "B",
        "direction": "금속 손가락이 우측 바나나에 닿기 직전으로 뻗어 있음.",
        "built_space": "참조 이미지의 시장 배경. 우측에 바나나 가판대와 골목 배치가 일치함.",
        "entities": "코트 소매를 입은 팔과 손만 클로즈업됨. 손등에 Ubik 로고가 있으며, 바나나가 화면의 3분의 1 이하를 차지함.",
        "hard_violations": [],
        "physics": "코트 소매에서 뻗어나온 팔이 금속 손을 지지함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "요구된 '손 클로즈업' 프레이밍을 정확히 구현했으며, 바나나와 배경 요소의 배치 규정을 충실히 따랐습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "'손 클로즈업' 지시를 무시하고 상반신 전체를 보여주었으며, 나무 상자에 불필요한 문자가 누출되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "금속 손가락이 우측 하단의 바나나 무더기를 향함.",
        "built_space": "참조 이미지의 야외 시장. 천막과 가판대 배치가 일치함.",
        "entities": "찰리의 얼굴과 상반신 전체 노출, 팔에 Ubik 로고, 상자에 불필요한 문자 존재. 바나나는 화면 우측에 위치함.",
        "hard_violations": [
         "나무 상자에 프롬프트에 없는 텍스트(문자) 누출"
        ],
        "physics": "몸통에 연결된 팔이 자연스럽게 손을 지지함."
       },
       {
        "label": "B",
        "direction": "금속 손가락이 우측 바나나에 닿기 직전으로 뻗어 있음.",
        "built_space": "참조 이미지의 시장 배경. 우측에 바나나 가판대와 골목 배치가 일치함.",
        "entities": "코트 소매를 입은 팔과 손만 클로즈업됨. 손등에 Ubik 로고가 있으며, 바나나가 화면의 3분의 1 이하를 차지함.",
        "hard_violations": [],
        "physics": "코트 소매에서 뻗어나온 팔이 금속 손을 지지함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "팔 옆에서 금속 손과 바나나를 잡은 구도가 손 클로즈업 지시에 더 가깝지만, 선명한 Ubik 글자가 금지 조건을 위반하고 접촉 전 간격도 불명확하다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "바나나를 향한 손가락과 접촉 전 간격은 명확하지만, 얼굴과 상체까지 드러낸 구도가 손 중심의 제한된 클로즈업에서 벗어나며 읽히는 로고도 금지 조건을 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽에서 나온 팔의 금속 손가락들이 오른쪽 아래 판매대의 바나나를 향한다. 목표는 머리 위에 매달린 송이가 아니라 손 바로 앞에 놓인 바나나다. 얼굴과 시선은 프레임 밖이다. 손끝이 과일 윤곽과 겹쳐 아직 닿지 않았다는 간격은 명확하지 않다.",
        "built_space": "오른쪽 앞에 바나나용 나무 상자 하나, 뒤쪽에 감귤류 상자와 저울 하나, 오른쪽 위에 매달린 바나나 한 송이가 보인다. 금속 기둥들, 적색·베이지 차양, 비닐막과 중앙 시장 통로는 장소 참고와 부합한다. 왼쪽 아래에는 플라스틱 상자를 담은 선반이 보인다. 카메라는 뻗은 팔의 옆에 있지만 판매대 주변뿐 아니라 통로와 건물 배경도 상당히 넓게 담는다.",
        "entities": "보이는 인물 부분은 찰리의 낡은 갈색 코트와 샌드 베이지 장갑판으로 덮인 육중한 기계 팔·손이다. 검은 관절과 마모된 금속 표면은 참고의 정체성에 부합한다. 얼굴, 모자와 눈은 이 구도에서 확인할 수 없으며 누락으로 보지 않는다. 판매용 바나나는 선명하고 화면의 약 3분의 1 미만을 차지한다. 손등의 Ubik 글자는 명확히 읽힌다.",
        "hard_violations": [
         "손등에 읽을 수 있는 Ubik 글자가 노출되어, 읽히는 글자와 로고를 금지한 명시 조건을 위반한다."
        ],
        "physics": "손은 손목 관절과 전완에 연결되고 팔은 코트 소매 안으로 이어져 지지 관계가 자연스럽다. 바나나는 나무 상자와 다른 과일 위에 놓여 있으며, 위쪽 송이는 줄로 매달려 있다. 근거 없이 떠 있는 물체는 없다. 다만 손끝과 바나나 사이가 겹쳐 보이므로 비접촉 순간인지 확정하기 어렵다."
       },
       {
        "label": "B",
        "direction": "크게 편 금속 검지가 오른쪽 아래 상자에 놓인 바나나의 왼쪽 끝을 향하며, 손끝과 과일 사이에는 눈에 보이는 간격이 있다. 찰리의 얼굴도 아래쪽 손과 과일 방향으로 기울어 있다. 다른 손가락은 아래로 굽혀져 있어 검지의 도달 목표가 분명하다.",
        "built_space": "오른쪽 앞의 바나나 나무 상자 하나, 뒤의 감귤류 진열, 저울 하나, 오른쪽 위의 매달린 바나나 한 송이와 금속 지지 기둥들이 보인다. 차양, 비닐막, 컨테이너형 건물과 중앙 통로는 참고 장소와 잘 대응한다. 다만 손뿐 아니라 찰리의 얼굴·어깨·상체와 시장 원경까지 크게 포함하여 제한된 손 클로즈업보다 인물이 과일에 손을 뻗는 넓은 장면으로 읽힌다.",
        "entities": "찰리의 흰 각진 마스크형 얼굴, 푸른 눈, 샌드 베이지 금속 팔, 낡은 갈색 코트와 모자가 보인다. 손의 장갑판과 관절 형상은 참고보다 두껍고 단순화되어 있으나 동일 계열의 기계 몸체로 읽힌다. 바나나는 판매대에 놓여 있고 화면의 3분의 1 미만을 차지한다. 전완에는 읽을 수 있는 UBik 글자와 도형 로고가 있다.",
        "hard_violations": [
         "전완에 읽을 수 있는 UBik 글자와 선명한 도형 로고가 노출되어, 글자와 로고를 금지한 명시 조건을 위반한다."
        ],
        "physics": "손과 손가락은 기계 관절을 통해 전완과 상완으로 연결된다. 가까운 손이 크게 보이는 원근은 성립하며, 손가락을 펴고 나머지를 굽힌 자세도 가능한 동작이다. 진열 바나나는 상자에 받쳐져 있고 위쪽 송이는 줄에 매달려 있다. 지지 없는 부유 물체는 없으며 바나나에 닿기 전 순간도 명확하다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "팔 옆에서 금속 손과 바나나를 잡은 구도가 손 클로즈업 지시에 더 가깝지만, 선명한 Ubik 글자가 금지 조건을 위반하고 접촉 전 간격도 불명확하다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "바나나를 향한 손가락과 접촉 전 간격은 명확하지만, 얼굴과 상체까지 드러낸 구도가 손 중심의 제한된 클로즈업에서 벗어나며 읽히는 로고도 금지 조건을 위반한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽에서 나온 팔의 금속 손가락들이 오른쪽 아래 판매대의 바나나를 향한다. 목표는 머리 위에 매달린 송이가 아니라 손 바로 앞에 놓인 바나나다. 얼굴과 시선은 프레임 밖이다. 손끝이 과일 윤곽과 겹쳐 아직 닿지 않았다는 간격은 명확하지 않다.",
        "built_space": "오른쪽 앞에 바나나용 나무 상자 하나, 뒤쪽에 감귤류 상자와 저울 하나, 오른쪽 위에 매달린 바나나 한 송이가 보인다. 금속 기둥들, 적색·베이지 차양, 비닐막과 중앙 시장 통로는 장소 참고와 부합한다. 왼쪽 아래에는 플라스틱 상자를 담은 선반이 보인다. 카메라는 뻗은 팔의 옆에 있지만 판매대 주변뿐 아니라 통로와 건물 배경도 상당히 넓게 담는다.",
        "entities": "보이는 인물 부분은 찰리의 낡은 갈색 코트와 샌드 베이지 장갑판으로 덮인 육중한 기계 팔·손이다. 검은 관절과 마모된 금속 표면은 참고의 정체성에 부합한다. 얼굴, 모자와 눈은 이 구도에서 확인할 수 없으며 누락으로 보지 않는다. 판매용 바나나는 선명하고 화면의 약 3분의 1 미만을 차지한다. 손등의 Ubik 글자는 명확히 읽힌다.",
        "hard_violations": [
         "손등에 읽을 수 있는 Ubik 글자가 노출되어, 읽히는 글자와 로고를 금지한 명시 조건을 위반한다."
        ],
        "physics": "손은 손목 관절과 전완에 연결되고 팔은 코트 소매 안으로 이어져 지지 관계가 자연스럽다. 바나나는 나무 상자와 다른 과일 위에 놓여 있으며, 위쪽 송이는 줄로 매달려 있다. 근거 없이 떠 있는 물체는 없다. 다만 손끝과 바나나 사이가 겹쳐 보이므로 비접촉 순간인지 확정하기 어렵다."
       },
       {
        "label": "A",
        "direction": "크게 편 금속 검지가 오른쪽 아래 상자에 놓인 바나나의 왼쪽 끝을 향하며, 손끝과 과일 사이에는 눈에 보이는 간격이 있다. 찰리의 얼굴도 아래쪽 손과 과일 방향으로 기울어 있다. 다른 손가락은 아래로 굽혀져 있어 검지의 도달 목표가 분명하다.",
        "built_space": "오른쪽 앞의 바나나 나무 상자 하나, 뒤의 감귤류 진열, 저울 하나, 오른쪽 위의 매달린 바나나 한 송이와 금속 지지 기둥들이 보인다. 차양, 비닐막, 컨테이너형 건물과 중앙 통로는 참고 장소와 잘 대응한다. 다만 손뿐 아니라 찰리의 얼굴·어깨·상체와 시장 원경까지 크게 포함하여 제한된 손 클로즈업보다 인물이 과일에 손을 뻗는 넓은 장면으로 읽힌다.",
        "entities": "찰리의 흰 각진 마스크형 얼굴, 푸른 눈, 샌드 베이지 금속 팔, 낡은 갈색 코트와 모자가 보인다. 손의 장갑판과 관절 형상은 참고보다 두껍고 단순화되어 있으나 동일 계열의 기계 몸체로 읽힌다. 바나나는 판매대에 놓여 있고 화면의 3분의 1 미만을 차지한다. 전완에는 읽을 수 있는 UBik 글자와 도형 로고가 있다.",
        "hard_violations": [
         "전완에 읽을 수 있는 UBik 글자와 선명한 도형 로고가 노출되어, 글자와 로고를 금지한 명시 조건을 위반한다."
        ],
        "physics": "손과 손가락은 기계 관절을 통해 전완과 상완으로 연결된다. 가까운 손이 크게 보이는 원근은 성립하며, 손가락을 펴고 나머지를 굽힌 자세도 가능한 동작이다. 진열 바나나는 상자에 받쳐져 있고 위쪽 송이는 줄에 매달려 있다. 지지 없는 부유 물체는 없으며 바나나에 닿기 전 순간도 명확하다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.179,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.929,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 나무 상자에 프롬프트에 없는 텍스트(문자) 누출",
     "[gpt-high] 전완에 읽을 수 있는 UBik 글자와 선명한 도형 로고가 노출되어, 글자와 로고를 금지한 명시 조건을 위반한다."
    ],
    "B": [
     "[gpt-high] 손등에 읽을 수 있는 Ubik 글자가 노출되어, 읽히는 글자와 로고를 금지한 명시 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 1750,
   "A": 929
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "요구된 '손 클로즈업' 프레이밍을 정확히 구현했으며, 바나나와 배경 요소의 배치 규정을 충실히 따랐습니다.  ★위반: [gpt-high] 손등에 읽을 수 있는 Ubik 글자가 노출되어, 읽히는 글자와 로고를 금지한 명시 조건을 위반한다."
   },
   {
    "label": "A",
    "score": 929,
    "verdict_ko": "'손 클로즈업' 지시를 무시하고 상반신 전체를 보여주었으며, 나무 상자에 불필요한 문자가 누출되었습니다.  ★위반: [gemini-pro] 나무 상자에 프롬프트에 없는 텍스트(문자) 누출 / [gpt-high] 전완에 읽을 수 있는 UBik 글자와 선명한 도형 로고가 노출되어, 글자와 로고를 금지한 명시 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_banana_market_stall_8677ef.png",
    "asset_id": "0333daa2-9ee6-4388-bf59-c510fb37f958",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-09bd-7566-9239-cad73bc13b3f",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S22sh3__bgfirst_bg.png",
   "bg_asset_id": "e6ccf295-a01a-4188-a4ca-5ce36e1609e9",
   "bg_record_key": "S22sh3::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "banana_market_stall",
   "groupbg_asset_id": "0333daa2-9ee6-4388-bf59-c510fb37f958"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S22sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:07:34.125785+00:00",
  "fingerprint": "c8bd5b4d26af04fcd73cbe4695e488daa706e31f01895d026081d22647ecf139",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S22sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S22sh3_sel.png",
  "source_sha256": "596d9abe42d6dde94a7aa06d8ceb6d6529a1f1fd163468d43b89849b0ec94362",
  "file": "S22sh3_cine.png",
  "staged_sha256": "32bfb00baec16c4a32cafd270547cc8aad9899cf8cc38b8400fc6b30b8be9a40",
  "latency_ms": 9090
 },
 "S22sh7::signage": {
  "fp": "f1377a0da1e411fd",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S22sh7": {
  "input_fingerprint": "44a0606afe0117f7",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 수많은 사람들의 인파 속을 향해 한 발을 앞으로 내디딘 mid-stride 자세의 찰리의 거대한 뒷모습.\n\nLOCATION (lock): In the crowded outdoor market lane of the refugee settlement, beyond the banana stall. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Route into the crowd in the upper-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Market passage through the crowd (Occupied by numerous people with naturally varying gaps and unsynchronized steps) — The visible route recedes ahead of 찰리 toward the upper center; used as Give his forward step a definite destination while allowing the crowd to progressively conceal him.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Continue neutral daytime ambient illumination with controlled contrast, separating 찰리 from the crowd through tonal organization rather than a new light cue.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The banana stall remains stocked, with no banana taken. Charlie moves into the market crowd still wearing the old coat and hat over his washed, worn robot body.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 수많은 사람들의 인파 속을 향해 한 발을 앞으로 내디딘 mid-stride 자세의 찰리의 거대한 뒷모습.\n\nLOCATION (lock): In the crowded outdoor market lane of the refugee settlement, beyond the banana stall. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Route into the crowd in the upper-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Market passage through the crowd (Occupied by numerous people with naturally varying gaps and unsynchronized steps) — The visible route recedes ahead of 찰리 toward the upper center; used as Give his forward step a definite destination while allowing the crowd to progressively conceal him.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Continue neutral daytime ambient illumination with controlled contrast, separating 찰리 from the crowd through tonal organization rather than a new light cue.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The banana stall remains stocked, with no banana taken. Charlie moves into the market crowd still wearing the old coat and hat over his washed, worn robot body.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 수많은 사람들의 인파 속을 향해 한 발을 앞으로 내디딘 mid-stride 자세의 찰리의 거대한 뒷모습.\n\nLOCATION (lock): In the crowded outdoor market lane of the refugee settlement, beyond the banana stall. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Route into the crowd in the upper-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Market passage through the crowd (Occupied by numerous people with naturally varying gaps and unsynchronized steps) — The visible route recedes ahead of 찰리 toward the upper center; used as Give his forward step a definite destination while allowing the crowd to progressively conceal him.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Continue neutral daytime ambient illumination with controlled contrast, separating 찰리 from the crowd through tonal organization rather than a new light cue.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The banana stall remains stocked, with no banana taken. Charlie moves into the market crowd still wearing the old coat and hat over his washed, worn robot body.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "찰리가 카메라를 등진 채 전방의 군중 통로를 향하고 있음.",
    "built_space": "양옆으로 천막과 매대가 늘어선 좁고 혼잡한 시장 골목. 왼쪽 전경에 과일 매대가 공간에 맞게 배치됨.",
    "entities": "찰리는 낡은 코트를 입고 샌드 베이지색 로봇 신체를 지니고 있음. 다수의 다국적 군중이 통로를 채움.",
    "hard_violations": [],
    "physics": "찰리의 양발이 바닥에 거의 평평하게 닿아 있어 프롬프트가 요구한 명확한 보행 중간(mid-stride) 자세로 보기 어려움."
   },
   {
    "label": "B",
    "direction": "찰리가 카메라를 등지고 시장 안쪽 사람들을 향해 나아가고 있음.",
    "built_space": "판잣집과 상점이 늘어선 길목 구조이나, 우측 시야가 완전히 가려짐.",
    "entities": "코트를 입은 찰리의 로봇 외형과 주변의 군중들이 지시된 대로 나타남.",
    "hard_violations": [
     "[gemini-pro] 화면 우측 전경에 바나나 더미가 물리적 원근감과 스케일을 완전히 무시한 채 평면적인 스티커처럼 거대하게 합성됨(콜라주/오버레이 금지 위반).",
     "[gpt-high] 오른쪽 바나나 가판이 명확한 수직 절단 경계를 가진 별도 패널처럼 붙어 있어, 콜라주 없는 하나의 연속된 실사 장면이라는 조건을 위반한다."
    ],
    "physics": "찰리의 오른발이 땅에서 떨어져 있어 걷는 도중의 동작과 지탱이 잘 표현됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "시장 골목의 구도와 찰리의 뒷모습, 의상을 공간에 맞게 자연스럽게 구현했으나 역동적인 걷는 자세가 다소 정적임."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "찰리의 걷는 동작은 명확하나, 우측의 바나나 이미지가 비정상적인 크기로 합성되어 지시사항을 심각하게 위반함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리가 카메라를 등진 채 전방의 군중 통로를 향하고 있음.",
        "built_space": "양옆으로 천막과 매대가 늘어선 좁고 혼잡한 시장 골목. 왼쪽 전경에 과일 매대가 공간에 맞게 배치됨.",
        "entities": "찰리는 낡은 코트를 입고 샌드 베이지색 로봇 신체를 지니고 있음. 다수의 다국적 군중이 통로를 채움.",
        "hard_violations": [],
        "physics": "찰리의 양발이 바닥에 거의 평평하게 닿아 있어 프롬프트가 요구한 명확한 보행 중간(mid-stride) 자세로 보기 어려움."
       },
       {
        "label": "B",
        "direction": "찰리가 카메라를 등지고 시장 안쪽 사람들을 향해 나아가고 있음.",
        "built_space": "판잣집과 상점이 늘어선 길목 구조이나, 우측 시야가 완전히 가려짐.",
        "entities": "코트를 입은 찰리의 로봇 외형과 주변의 군중들이 지시된 대로 나타남.",
        "hard_violations": [
         "화면 우측 전경에 바나나 더미가 물리적 원근감과 스케일을 완전히 무시한 채 평면적인 스티커처럼 거대하게 합성됨(콜라주/오버레이 금지 위반)."
        ],
        "physics": "찰리의 오른발이 땅에서 떨어져 있어 걷는 도중의 동작과 지탱이 잘 표현됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "시장 골목의 구도와 찰리의 뒷모습, 의상을 공간에 맞게 자연스럽게 구현했으나 역동적인 걷는 자세가 다소 정적임."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "찰리의 걷는 동작은 명확하나, 우측의 바나나 이미지가 비정상적인 크기로 합성되어 지시사항을 심각하게 위반함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리가 카메라를 등진 채 전방의 군중 통로를 향하고 있음.",
        "built_space": "양옆으로 천막과 매대가 늘어선 좁고 혼잡한 시장 골목. 왼쪽 전경에 과일 매대가 공간에 맞게 배치됨.",
        "entities": "찰리는 낡은 코트를 입고 샌드 베이지색 로봇 신체를 지니고 있음. 다수의 다국적 군중이 통로를 채움.",
        "hard_violations": [],
        "physics": "찰리의 양발이 바닥에 거의 평평하게 닿아 있어 프롬프트가 요구한 명확한 보행 중간(mid-stride) 자세로 보기 어려움."
       },
       {
        "label": "B",
        "direction": "찰리가 카메라를 등지고 시장 안쪽 사람들을 향해 나아가고 있음.",
        "built_space": "판잣집과 상점이 늘어선 길목 구조이나, 우측 시야가 완전히 가려짐.",
        "entities": "코트를 입은 찰리의 로봇 외형과 주변의 군중들이 지시된 대로 나타남.",
        "hard_violations": [
         "화면 우측 전경에 바나나 더미가 물리적 원근감과 스케일을 완전히 무시한 채 평면적인 스티커처럼 거대하게 합성됨(콜라주/오버레이 금지 위반)."
        ],
        "physics": "찰리의 오른발이 땅에서 떨어져 있어 걷는 도중의 동작과 지탱이 잘 표현됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "인파를 향한 뒷모습과 보행 순간은 맞지만, 오른쪽 바나나 가판이 수직 경계로 붙인 별도 패널처럼 나타나 단일 실사 장면 조건을 위반한다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "거대한 찰리의 뒷모습과 상단 중앙의 인파 속 진로를 일관된 와이드 숏으로 구현했지만, 모자가 없고 바나나 가판의 위치·구성이 이전 장면과 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 머리와 몸통은 카메라 반대편, 화면 상단 중앙의 시장 인파를 향한다. 왼발로 지면을 딛고 오른발 뒤꿈치를 들어 앞으로 이동하는 순간으로 보인다. 주변 사람들은 서로 다른 방향으로 걷거나 옆을 보고 있으며, 찰리가 향하는 통로는 보인다.",
        "built_space": "왼쪽에는 컨테이너형 상점과 여러 천막, 중앙에는 군중이 있는 통로가 있다. 오른쪽에는 매달린 바나나 한 송이, 바나나가 쌓인 진열대, 감귤류 상자와 금속 기둥들이 보인다. 그러나 오른쪽 가판 영역이 화면 약 4분의 3 지점의 곧은 수직 경계에서 시작하며, 중앙 장면의 사람과 바닥을 잘라 별도 화면처럼 연결된다.",
        "entities": "찰리는 한 명이며, 넓은 어깨와 긴 기계 팔, 짧은 다리, 마모된 샌드 베이지 장갑을 갖췄다. 낡은 갈색 외투는 있으나 요구된 모자 대신 기계 머리가 드러난다. 외투 바깥으로 큰 등·어깨 장갑이 노출된다. 얼굴은 뒷모습이라 확인할 수 없다. 군중은 주로 동아시아계로 보이는 성인 남녀이며, 바나나는 가판에 남아 있고 찰리의 손은 비어 있다.",
        "hard_violations": [
         "오른쪽 바나나 가판이 명확한 수직 절단 경계를 가진 별도 패널처럼 붙어 있어, 콜라주 없는 하나의 연속된 실사 장면이라는 조건을 위반한다."
        ],
        "physics": "찰리의 왼발은 지면에 놓이고 오른발은 발끝 쪽을 낮춘 채 뒤꿈치가 들려 있어 보행의 지지가 성립한다. 팔은 어깨 관절에 연결되어 자연스럽게 내려와 있다. 군중도 지면 위에 서거나 걷는다. 바나나는 진열대에 놓이거나 위쪽 끈에 매달려 있으며, 지지 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "찰리의 등과 머리는 상단 중앙으로 길게 이어지는 군중 속 통로를 향한다. 오른발이 왼발보다 앞쪽에 놓여 전진 중인 자세로 읽힌다. 가까운 군중은 카메라 쪽으로 걷거나 옆을 보고 있고, 먼 군중은 통로를 따라 밀집해 있어 찰리의 목적지가 분명하다.",
        "built_space": "양옆에 이층 규모의 임시 상점과 컨테이너 벽, 천막, 지지 기둥들이 있고 중앙 통로가 상단 중앙으로 이어진다. 바나나 진열대는 왼쪽에 있으며 위에는 여러 송이가 매달려 있다. 오른쪽에는 별도의 과일 진열대가 있다. 공간은 하나의 연속된 시장으로 보이지만, 이전 장면의 오른쪽 바나나 가판과 매달린 한 송이, 금속 선반 등의 구체적 배치는 그대로 확인되지 않는다. 선택 가능한 카메라 각도만으로 동일한 가판임을 확신하기는 어렵다.",
        "entities": "찰리 한 명의 육중한 뒷모습이 중앙에 크게 보이며, 긴 팔과 짧은 다리, 샌드 베이지 기계 장갑은 인물 참조와 부합한다. 낡은 갈색 외투는 있으나 모자가 없고, 외투 밑단은 이전 장면보다 훨씬 심하게 찢어져 있다. 얼굴은 올바르게 가려져 있다. 군중에는 여러 연령대의 성인 남녀와 서로 다른 외형의 사람들이 포함되어 난민 시장의 다민족 설정을 뒷받침한다. 양손에 바나나는 없으며 가판은 충분히 채워져 있다.",
        "hard_violations": [],
        "physics": "찰리의 두 발은 앞뒤로 어긋나 지면에 닿아 있고, 다리 관절과 몸통의 연결도 유지된다. 큰 보폭은 아니지만 발을 앞으로 옮기는 보행 단계로 가능한 자세다. 군중은 각자의 발로 서거나 걷고, 외투는 몸에서 아래로 처진다. 진열 과일은 상자와 선반이 받치고 매달린 바나나는 끈이 지지한다. 지지 없는 부유나 불가능한 신체 배치는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "인파를 향한 뒷모습과 보행 순간은 맞지만, 오른쪽 바나나 가판이 수직 경계로 붙인 별도 패널처럼 나타나 단일 실사 장면 조건을 위반한다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "거대한 찰리의 뒷모습과 상단 중앙의 인파 속 진로를 일관된 와이드 숏으로 구현했지만, 모자가 없고 바나나 가판의 위치·구성이 이전 장면과 다르다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 머리와 몸통은 카메라 반대편, 화면 상단 중앙의 시장 인파를 향한다. 왼발로 지면을 딛고 오른발 뒤꿈치를 들어 앞으로 이동하는 순간으로 보인다. 주변 사람들은 서로 다른 방향으로 걷거나 옆을 보고 있으며, 찰리가 향하는 통로는 보인다.",
        "built_space": "왼쪽에는 컨테이너형 상점과 여러 천막, 중앙에는 군중이 있는 통로가 있다. 오른쪽에는 매달린 바나나 한 송이, 바나나가 쌓인 진열대, 감귤류 상자와 금속 기둥들이 보인다. 그러나 오른쪽 가판 영역이 화면 약 4분의 3 지점의 곧은 수직 경계에서 시작하며, 중앙 장면의 사람과 바닥을 잘라 별도 화면처럼 연결된다.",
        "entities": "찰리는 한 명이며, 넓은 어깨와 긴 기계 팔, 짧은 다리, 마모된 샌드 베이지 장갑을 갖췄다. 낡은 갈색 외투는 있으나 요구된 모자 대신 기계 머리가 드러난다. 외투 바깥으로 큰 등·어깨 장갑이 노출된다. 얼굴은 뒷모습이라 확인할 수 없다. 군중은 주로 동아시아계로 보이는 성인 남녀이며, 바나나는 가판에 남아 있고 찰리의 손은 비어 있다.",
        "hard_violations": [
         "오른쪽 바나나 가판이 명확한 수직 절단 경계를 가진 별도 패널처럼 붙어 있어, 콜라주 없는 하나의 연속된 실사 장면이라는 조건을 위반한다."
        ],
        "physics": "찰리의 왼발은 지면에 놓이고 오른발은 발끝 쪽을 낮춘 채 뒤꿈치가 들려 있어 보행의 지지가 성립한다. 팔은 어깨 관절에 연결되어 자연스럽게 내려와 있다. 군중도 지면 위에 서거나 걷는다. 바나나는 진열대에 놓이거나 위쪽 끈에 매달려 있으며, 지지 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "찰리의 등과 머리는 상단 중앙으로 길게 이어지는 군중 속 통로를 향한다. 오른발이 왼발보다 앞쪽에 놓여 전진 중인 자세로 읽힌다. 가까운 군중은 카메라 쪽으로 걷거나 옆을 보고 있고, 먼 군중은 통로를 따라 밀집해 있어 찰리의 목적지가 분명하다.",
        "built_space": "양옆에 이층 규모의 임시 상점과 컨테이너 벽, 천막, 지지 기둥들이 있고 중앙 통로가 상단 중앙으로 이어진다. 바나나 진열대는 왼쪽에 있으며 위에는 여러 송이가 매달려 있다. 오른쪽에는 별도의 과일 진열대가 있다. 공간은 하나의 연속된 시장으로 보이지만, 이전 장면의 오른쪽 바나나 가판과 매달린 한 송이, 금속 선반 등의 구체적 배치는 그대로 확인되지 않는다. 선택 가능한 카메라 각도만으로 동일한 가판임을 확신하기는 어렵다.",
        "entities": "찰리 한 명의 육중한 뒷모습이 중앙에 크게 보이며, 긴 팔과 짧은 다리, 샌드 베이지 기계 장갑은 인물 참조와 부합한다. 낡은 갈색 외투는 있으나 모자가 없고, 외투 밑단은 이전 장면보다 훨씬 심하게 찢어져 있다. 얼굴은 올바르게 가려져 있다. 군중에는 여러 연령대의 성인 남녀와 서로 다른 외형의 사람들이 포함되어 난민 시장의 다민족 설정을 뒷받침한다. 양손에 바나나는 없으며 가판은 충분히 채워져 있다.",
        "hard_violations": [],
        "physics": "찰리의 두 발은 앞뒤로 어긋나 지면에 닿아 있고, 다리 관절과 몸통의 연결도 유지된다. 큰 보폭은 아니지만 발을 앞으로 옮기는 보행 단계로 가능한 자세다. 군중은 각자의 발로 서거나 걷고, 외투는 몸에서 아래로 처진다. 진열 과일은 상자와 선반이 받치고 매달린 바나나는 끈이 지지한다. 지지 없는 부유나 불가능한 신체 배치는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.857
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.607
   },
   "violations": {
    "B": [
     "[gemini-pro] 화면 우측 전경에 바나나 더미가 물리적 원근감과 스케일을 완전히 무시한 채 평면적인 스티커처럼 거대하게 합성됨(콜라주/오버레이 금지 위반).",
     "[gpt-high] 오른쪽 바나나 가판이 명확한 수직 절단 경계를 가진 별도 패널처럼 붙어 있어, 콜라주 없는 하나의 연속된 실사 장면이라는 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 607
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "시장 골목의 구도와 찰리의 뒷모습, 의상을 공간에 맞게 자연스럽게 구현했으나 역동적인 걷는 자세가 다소 정적임."
   },
   {
    "label": "B",
    "score": 607,
    "verdict_ko": "찰리의 걷는 동작은 명확하나, 우측의 바나나 이미지가 비정상적인 크기로 합성되어 지시사항을 심각하게 위반함.  ★위반: [gemini-pro] 화면 우측 전경에 바나나 더미가 물리적 원근감과 스케일을 완전히 무시한 채 평면적인 스티커처럼 거대하게 합성됨(콜라주/오버레이 금지 위반). / [gpt-high] 오른쪽 바나나 가판이 명확한 수직 절단 경계를 가진 별도 패널처럼 붙어 있어, 콜라주 없는 하나의 연속된 실사 장면이라는 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S22sh3_sel.png",
    "asset_id": "8a7b12fa-e6b9-4d4b-91e4-5bfad2b42c69",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-0eab-7bdb-957f-0810308eac52",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S22sh3"
  }
 },
 "S22sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:08:29.740697+00:00",
  "fingerprint": "a3780e3453a9033a1d930922d4b207bcfd558729c20ed4a0169f0c44f12c6c7a",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S22sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S22sh7_sel.png",
  "source_sha256": "6a2bb14140c0b6351b16b50cd32c433f2c82c4d9f00b8c3140228ec4380c4682",
  "file": "S22sh7_cine.png",
  "staged_sha256": "4df981089b01b8a1098b9dbad8fae547a32413ddb4744a3fdbb57b9040e8d57f",
  "latency_ms": 11087
 },
 "S22sh9::signage": {
  "fp": "9dbf3f3e354c2eeb",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S22sh9": {
  "input_fingerprint": "9fd85a5e2dc17d2c",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 인파가 사라진 방향을 향해 검지손가락을 길게 뻗은 앰버의 단호한 상체.\n\nLOCATION (lock): At the market-lane junction near the banana stall, looking along the direction the robot vanished into the crowd. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Passage toward 찰리's route in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Market passage in 찰리's direction (찰리 is no longer visible from this framing) — A limited section of the route extends toward the right edge beyond 앰버's pointing hand; used as Give the gesture directional meaning without placing a new object at its endpoint.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain neutral daytime ambient light and restrained contrast consistent with the preceding market action.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the market passage, adjacent shop structures, and daytime lighting visible in the reference. Exclude the departing robot and do not duplicate the reference's passersby.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The banana stall remains in place, while Charlie has moved into the crowd with his coat-and-hat disguise unchanged. 앰버: She has reached the market search location, retaining her mask and waist tool pouch.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 인파가 사라진 방향을 향해 검지손가락을 길게 뻗은 앰버의 단호한 상체.\n\nLOCATION (lock): At the market-lane junction near the banana stall, looking along the direction the robot vanished into the crowd. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Passage toward 찰리's route in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Market passage in 찰리's direction (찰리 is no longer visible from this framing) — A limited section of the route extends toward the right edge beyond 앰버's pointing hand; used as Give the gesture directional meaning without placing a new object at its endpoint.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain neutral daytime ambient light and restrained contrast consistent with the preceding market action.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the market passage, adjacent shop structures, and daytime lighting visible in the reference. Exclude the departing robot and do not duplicate the reference's passersby.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The banana stall remains in place, while Charlie has moved into the crowd with his coat-and-hat disguise unchanged. 앰버: She has reached the market search location, retaining her mask and waist tool pouch.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 인파가 사라진 방향을 향해 검지손가락을 길게 뻗은 앰버의 단호한 상체.\n\nLOCATION (lock): At the market-lane junction near the banana stall, looking along the direction the robot vanished into the crowd. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Passage toward 찰리's route in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Market passage in 찰리's direction (찰리 is no longer visible from this framing) — A limited section of the route extends toward the right edge beyond 앰버's pointing hand; used as Give the gesture directional meaning without placing a new object at its endpoint.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain neutral daytime ambient light and restrained contrast consistent with the preceding market action.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the market passage, adjacent shop structures, and daytime lighting visible in the reference. Exclude the departing robot and do not duplicate the reference's passersby.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The banana stall remains in place, while Charlie has moved into the crowd with his coat-and-hat disguise unchanged. 앰버: She has reached the market search location, retaining her mask and waist tool pouch.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "B",
    "direction": "앰버의 시선과 뻗은 검지는 화면 오른쪽을 향한다. 손끝의 연장선에는 오른쪽 군중과 가게가 있고, 시장 통로는 그보다 왼쪽에서 화면 안쪽으로 이어진다. 찰리가 사라진 통로를 지목한다는 방향 관계가 분명하지 않다.",
    "built_space": "왼쪽 바나나 판매대 하나, 양옆의 낡은 상점 열 두 줄, 왼쪽 베이지색 차양과 오른쪽 파란색·회색 차양, 중앙 통로가 보인다. 장소의 구조와 재질은 참고 사진에 가깝지만, 이전 사진의 넓은 구도를 거의 유지한다. 앰버가 중앙 통로 상당 부분을 가리고 손 너머 오른쪽에는 제한된 경로 대신 군중과 상점이 보인다.",
    "entities": "금발의 어린 여자아이 한 명이 남색 반팔, 흰 마스크, 허리 공구 주머니를 착용했다. 보이는 눈과 머리, 아동 체격은 앰버 참고와 대체로 부합하며 마스크가 하관을 가린다. 찰리는 보이지 않고 바나나 판매대는 남아 있다. 그러나 이전 사진의 앞쪽 머릿수건 여성, 모자 쓴 남성들을 비롯한 수많은 행인이 그대로 등장하여 앰버만 허용한 인물 조건을 위반한다.",
    "hard_violations": [
     "앰버 외에 다수의 인물을 추가했으며, 명시적으로 제외하라고 한 이전 사진의 행인 얼굴·몸·의상을 그대로 재현했다."
    ],
    "physics": "뻗은 팔은 어깨와 팔꿈치에 자연스럽게 연결되고 검지를 편 손동작도 가능한 자세다. 공구 주머니는 허리띠에 걸려 있고 공구는 주머니 안에 수납되어 있다. 발은 화면 밖이지만 서 있는 몸통에 공중 부양을 시사하는 단서는 없다. 과일은 판매대에 놓이거나 위쪽 구조물에 매달려 있다."
   },
   {
    "label": "A",
    "direction": "앰버는 오른쪽을 보며 팔과 검지를 길게 뻗는다. 손끝 바로 너머에는 오른쪽 줄의 남성이 있고, 검지 연장선은 그 뒤 군중을 가로지른다. 통로 자체는 중앙에서 화면 깊숙이 이어지므로 손끝이 찰리의 진행 경로를 정확히 따른다고 보기 어렵다.",
    "built_space": "왼쪽 바나나 판매대 하나와 양쪽 상점 열 두 줄, 베이지색·파란색 차양, 중앙의 열린 노면이 보이며 참고 장소의 건축과 마모를 유지한다. 앰버를 왼쪽에 두어 A보다 중앙·오른쪽 통로가 잘 드러난다. 다만 손 너머 오른쪽 가장자리로 경로가 이어지기보다 사람들과 상점이 놓여 있고, 배경 범위도 요청한 제한된 통로보다 넓다.",
    "entities": "금발, 큰 눈, 어린 여자아이 체격과 남색 반팔이 앰버 참고에 대체로 부합한다. 흰 마스크와 허리 공구 주머니가 유지되며 찰리는 없다. 바나나 판매대도 남아 있다. 그러나 이전 사진에서 보인 머릿수건 여성과 오른쪽 전경 남성들 등 다수의 행인을 그대로 포함하여 인물 제한을 어긴다.",
    "hard_violations": [
     "앰버 외에 다수의 인물을 추가했으며, 명시적으로 제외하라고 한 이전 사진의 행인 얼굴·몸·의상을 그대로 재현했다."
    ],
    "physics": "상체에서 이어지는 팔과 편 검지의 자세는 실제로 가능하다. 허리띠가 공구 주머니를 지탱하고 도구들은 주머니에 꽂혀 있다. 발은 잘렸지만 몸이 떠 있다고 볼 근거는 없다. 배경 인물들은 노면에 서 있고 바나나는 판매대 또는 매달림 구조로 지지된다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": null,
     "normalized": null,
     "ok": false
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "왼쪽의 앰버와 중앙·오른쪽 통로 배치는 더 가깝지만, 금지된 이전 인파를 그대로 재현했고 검지는 통로의 진행 방향보다 오른쪽 사람들을 가리켜 재촬영이 필요하다."
       },
       {
        "label": "A",
        "score": 1,
        "verdict_ko": "마스크와 단호한 지목 동작은 맞지만, 금지된 이전 인파를 그대로 재현하고 앰버가 통로를 가리며 손끝도 이동 경로가 아닌 오른쪽 군중을 향한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버의 시선과 뻗은 검지는 화면 오른쪽을 향한다. 손끝의 연장선에는 오른쪽 군중과 가게가 있고, 시장 통로는 그보다 왼쪽에서 화면 안쪽으로 이어진다. 찰리가 사라진 통로를 지목한다는 방향 관계가 분명하지 않다.",
        "built_space": "왼쪽 바나나 판매대 하나, 양옆의 낡은 상점 열 두 줄, 왼쪽 베이지색 차양과 오른쪽 파란색·회색 차양, 중앙 통로가 보인다. 장소의 구조와 재질은 참고 사진에 가깝지만, 이전 사진의 넓은 구도를 거의 유지한다. 앰버가 중앙 통로 상당 부분을 가리고 손 너머 오른쪽에는 제한된 경로 대신 군중과 상점이 보인다.",
        "entities": "금발의 어린 여자아이 한 명이 남색 반팔, 흰 마스크, 허리 공구 주머니를 착용했다. 보이는 눈과 머리, 아동 체격은 앰버 참고와 대체로 부합하며 마스크가 하관을 가린다. 찰리는 보이지 않고 바나나 판매대는 남아 있다. 그러나 이전 사진의 앞쪽 머릿수건 여성, 모자 쓴 남성들을 비롯한 수많은 행인이 그대로 등장하여 앰버만 허용한 인물 조건을 위반한다.",
        "hard_violations": [
         "앰버 외에 다수의 인물을 추가했으며, 명시적으로 제외하라고 한 이전 사진의 행인 얼굴·몸·의상을 그대로 재현했다."
        ],
        "physics": "뻗은 팔은 어깨와 팔꿈치에 자연스럽게 연결되고 검지를 편 손동작도 가능한 자세다. 공구 주머니는 허리띠에 걸려 있고 공구는 주머니 안에 수납되어 있다. 발은 화면 밖이지만 서 있는 몸통에 공중 부양을 시사하는 단서는 없다. 과일은 판매대에 놓이거나 위쪽 구조물에 매달려 있다."
       },
       {
        "label": "B",
        "direction": "앰버는 오른쪽을 보며 팔과 검지를 길게 뻗는다. 손끝 바로 너머에는 오른쪽 줄의 남성이 있고, 검지 연장선은 그 뒤 군중을 가로지른다. 통로 자체는 중앙에서 화면 깊숙이 이어지므로 손끝이 찰리의 진행 경로를 정확히 따른다고 보기 어렵다.",
        "built_space": "왼쪽 바나나 판매대 하나와 양쪽 상점 열 두 줄, 베이지색·파란색 차양, 중앙의 열린 노면이 보이며 참고 장소의 건축과 마모를 유지한다. 앰버를 왼쪽에 두어 A보다 중앙·오른쪽 통로가 잘 드러난다. 다만 손 너머 오른쪽 가장자리로 경로가 이어지기보다 사람들과 상점이 놓여 있고, 배경 범위도 요청한 제한된 통로보다 넓다.",
        "entities": "금발, 큰 눈, 어린 여자아이 체격과 남색 반팔이 앰버 참고에 대체로 부합한다. 흰 마스크와 허리 공구 주머니가 유지되며 찰리는 없다. 바나나 판매대도 남아 있다. 그러나 이전 사진에서 보인 머릿수건 여성과 오른쪽 전경 남성들 등 다수의 행인을 그대로 포함하여 인물 제한을 어긴다.",
        "hard_violations": [
         "앰버 외에 다수의 인물을 추가했으며, 명시적으로 제외하라고 한 이전 사진의 행인 얼굴·몸·의상을 그대로 재현했다."
        ],
        "physics": "상체에서 이어지는 팔과 편 검지의 자세는 실제로 가능하다. 허리띠가 공구 주머니를 지탱하고 도구들은 주머니에 꽂혀 있다. 발은 잘렸지만 몸이 떠 있다고 볼 근거는 없다. 배경 인물들은 노면에 서 있고 바나나는 판매대 또는 매달림 구조로 지지된다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "왼쪽의 앰버와 중앙·오른쪽 통로 배치는 더 가깝지만, 금지된 이전 인파를 그대로 재현했고 검지는 통로의 진행 방향보다 오른쪽 사람들을 가리켜 재촬영이 필요하다."
       },
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "마스크와 단호한 지목 동작은 맞지만, 금지된 이전 인파를 그대로 재현하고 앰버가 통로를 가리며 손끝도 이동 경로가 아닌 오른쪽 군중을 향한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "앰버의 시선과 뻗은 검지는 화면 오른쪽을 향한다. 손끝의 연장선에는 오른쪽 군중과 가게가 있고, 시장 통로는 그보다 왼쪽에서 화면 안쪽으로 이어진다. 찰리가 사라진 통로를 지목한다는 방향 관계가 분명하지 않다.",
        "built_space": "왼쪽 바나나 판매대 하나, 양옆의 낡은 상점 열 두 줄, 왼쪽 베이지색 차양과 오른쪽 파란색·회색 차양, 중앙 통로가 보인다. 장소의 구조와 재질은 참고 사진에 가깝지만, 이전 사진의 넓은 구도를 거의 유지한다. 앰버가 중앙 통로 상당 부분을 가리고 손 너머 오른쪽에는 제한된 경로 대신 군중과 상점이 보인다.",
        "entities": "금발의 어린 여자아이 한 명이 남색 반팔, 흰 마스크, 허리 공구 주머니를 착용했다. 보이는 눈과 머리, 아동 체격은 앰버 참고와 대체로 부합하며 마스크가 하관을 가린다. 찰리는 보이지 않고 바나나 판매대는 남아 있다. 그러나 이전 사진의 앞쪽 머릿수건 여성, 모자 쓴 남성들을 비롯한 수많은 행인이 그대로 등장하여 앰버만 허용한 인물 조건을 위반한다.",
        "hard_violations": [
         "앰버 외에 다수의 인물을 추가했으며, 명시적으로 제외하라고 한 이전 사진의 행인 얼굴·몸·의상을 그대로 재현했다."
        ],
        "physics": "뻗은 팔은 어깨와 팔꿈치에 자연스럽게 연결되고 검지를 편 손동작도 가능한 자세다. 공구 주머니는 허리띠에 걸려 있고 공구는 주머니 안에 수납되어 있다. 발은 화면 밖이지만 서 있는 몸통에 공중 부양을 시사하는 단서는 없다. 과일은 판매대에 놓이거나 위쪽 구조물에 매달려 있다."
       },
       {
        "label": "A",
        "direction": "앰버는 오른쪽을 보며 팔과 검지를 길게 뻗는다. 손끝 바로 너머에는 오른쪽 줄의 남성이 있고, 검지 연장선은 그 뒤 군중을 가로지른다. 통로 자체는 중앙에서 화면 깊숙이 이어지므로 손끝이 찰리의 진행 경로를 정확히 따른다고 보기 어렵다.",
        "built_space": "왼쪽 바나나 판매대 하나와 양쪽 상점 열 두 줄, 베이지색·파란색 차양, 중앙의 열린 노면이 보이며 참고 장소의 건축과 마모를 유지한다. 앰버를 왼쪽에 두어 A보다 중앙·오른쪽 통로가 잘 드러난다. 다만 손 너머 오른쪽 가장자리로 경로가 이어지기보다 사람들과 상점이 놓여 있고, 배경 범위도 요청한 제한된 통로보다 넓다.",
        "entities": "금발, 큰 눈, 어린 여자아이 체격과 남색 반팔이 앰버 참고에 대체로 부합한다. 흰 마스크와 허리 공구 주머니가 유지되며 찰리는 없다. 바나나 판매대도 남아 있다. 그러나 이전 사진에서 보인 머릿수건 여성과 오른쪽 전경 남성들 등 다수의 행인을 그대로 포함하여 인물 제한을 어긴다.",
        "hard_violations": [
         "앰버 외에 다수의 인물을 추가했으며, 명시적으로 제외하라고 한 이전 사진의 행인 얼굴·몸·의상을 그대로 재현했다."
        ],
        "physics": "상체에서 이어지는 팔과 편 검지의 자세는 실제로 가능하다. 허리띠가 공구 주머니를 지탱하고 도구들은 주머니에 꽂혀 있다. 발은 잘렸지만 몸이 떠 있다고 볼 근거는 없다. 배경 인물들은 노면에 서 있고 바나나는 판매대 또는 매달림 구조로 지지된다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gemini-pro"
   ],
   "route": "single_reverse"
  },
  "totals": {
   "A": 2,
   "B": 1
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2,
    "verdict_ko": "왼쪽의 앰버와 중앙·오른쪽 통로 배치는 더 가깝지만, 금지된 이전 인파를 그대로 재현했고 검지는 통로의 진행 방향보다 오른쪽 사람들을 가리켜 재촬영이 필요하다."
   },
   {
    "label": "B",
    "score": 1,
    "verdict_ko": "마스크와 단호한 지목 동작은 맞지만, 금지된 이전 인파를 그대로 재현하고 앰버가 통로를 가리며 손끝도 이동 경로가 아닌 오른쪽 군중을 향한다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S22sh7_sel.png",
    "asset_id": "b286f35e-4251-48d8-9a7b-0f76b3edbf3b",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-106c-7775-ad0f-58e5282b935e",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S22sh7"
  }
 },
 "S22sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:09:23.151776+00:00",
  "fingerprint": "dc52c1c3d7fe8f3f1f5b7ccebb4d25e5c52d3e87b4f2d19a1db9147f18526330",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S22sh9_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S22sh9_sel.png",
  "source_sha256": "d23ebaf13e28ba2945fd06a1555f411a31ea4d97ad104e9cf80aeb77dc52e66b",
  "file": "S22sh9_cine.png",
  "staged_sha256": "be63ff1e5a69b3ceac2056ff77d6bc00a1094f83f85ee5db08c51c43b6d80a50",
  "latency_ms": 12282
 },
 "S23sh1::signage": {
  "fp": "50585ca94203c8f1",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::d817c9595eb14cf4": {
  "subjects": [],
  "subject_text": "구치소 면회실\n넓은 실내 한가운데 테이블과 의자가 놓인 면회 공간. 중앙 좌석 주변으로 여백이 넓게 남아 있는 단순한 구성이다.",
  "identity": "canonical",
  "scope_id": "L173",
  "scope_role": "location_interior",
  "scope_sha": "cf46c6ca176ce145"
 },
 "S23sh1::bgfirst_bg": {
  "input_fingerprint": "69263a57c81a657b",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 얼굴 곳곳에 피멍이 든 채 수갑 찬 양손을 테이블 위로 길게 뻗은 현우의 절박한 상체.\n\nLOCATION (lock): At the central table inside a spacious detention-center visiting room, in plain daytime interior light.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 면회실 테이블 (Positioned in the middle of the visiting room, supporting 현우's extended hands) — The near corner and a shallow view across the top face the camera; used as Carries the diagonal from the restrained hands toward the pleading face; 수갑 (Fastened around both wrists) — Seen obliquely around the wrists above the tabletop; used as A small but readable restraint detail within the upper-body composition; 넓은 면회실 (Open room space remains visible beyond the centrally placed table); used as Negative space separates the isolated seated figure from the surrounding room.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient illumination preserves the facial bruising and hand detail with controlled contrast, without specifying a visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 얼굴 곳곳에 피멍이 든 채 수갑 찬 양손을 테이블 위로 길게 뻗은 현우의 절박한 상체.\n\nLOCATION (lock): At the central table inside a spacious detention-center visiting room, in plain daytime interior light.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 면회실 테이블 (Positioned in the middle of the visiting room, supporting 현우's extended hands) — The near corner and a shallow view across the top face the camera; used as Carries the diagonal from the restrained hands toward the pleading face; 수갑 (Fastened around both wrists) — Seen obliquely around the wrists above the tabletop; used as A small but readable restraint detail within the upper-body composition; 넓은 면회실 (Open room space remains visible beyond the centrally placed table); used as Negative space separates the isolated seated figure from the surrounding room.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient illumination preserves the facial bruising and hand detail with controlled contrast, without specifying a visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S23sh1__bgfirst_bg.png",
  "asset_id": "10a8564f-8f3e-46a6-8c49-5dd05488e9d9",
  "input_asset_ids": [
   "69e13ce4-f531-431c-a45d-c3679c68fd0e",
   "0977c2dd-72ef-4442-9b57-c56745b38ac5"
  ]
 },
 "S23sh1": {
  "input_fingerprint": "3fd80190ed57f37f",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 얼굴 곳곳에 피멍이 든 채 수갑 찬 양손을 테이블 위로 길게 뻗은 현우의 절박한 상체.\n\nLOCATION (lock): At the central table inside a spacious detention-center visiting room, in plain daytime interior light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 면회실 테이블 (Positioned in the middle of the visiting room, supporting 현우's extended hands) — The near corner and a shallow view across the top face the camera; used as Carries the diagonal from the restrained hands toward the pleading face; 수갑 (Fastened around both wrists) — Seen obliquely around the wrists above the tabletop; used as A small but readable restraint detail within the upper-body composition; 넓은 면회실 (Open room space remains visible beyond the centrally placed table); used as Negative space separates the isolated seated figure from the surrounding room.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient illumination preserves the facial bruising and hand detail with controlled contrast, without specifying a visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A table stands in the middle of the large visitation room. 현우: He sits at the table with his wrists still handcuffed and visible bruising on his face. His dog-bitten leg remains injured.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 얼굴 곳곳에 피멍이 든 채 수갑 찬 양손을 테이블 위로 길게 뻗은 현우의 절박한 상체.\n\nLOCATION (lock): At the central table inside a spacious detention-center visiting room, in plain daytime interior light. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 면회실 테이블 (Positioned in the middle of the visiting room, supporting 현우's extended hands) — The near corner and a shallow view across the top face the camera; used as Carries the diagonal from the restrained hands toward the pleading face; 수갑 (Fastened around both wrists) — Seen obliquely around the wrists above the tabletop; used as A small but readable restraint detail within the upper-body composition; 넓은 면회실 (Open room space remains visible beyond the centrally placed table); used as Negative space separates the isolated seated figure from the surrounding room.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient illumination preserves the facial bruising and hand detail with controlled contrast, without specifying a visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A table stands in the middle of the large visitation room. 현우: He sits at the table with his wrists still handcuffed and visible bruising on his face. His dog-bitten leg remains injured.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 얼굴 곳곳에 피멍이 든 채 수갑 찬 양손을 테이블 위로 길게 뻗은 현우의 절박한 상체.\n\nLOCATION (lock): At the central table inside a spacious detention-center visiting room, in plain daytime interior light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 면회실 테이블 (Positioned in the middle of the visiting room, supporting 현우's extended hands) — The near corner and a shallow view across the top face the camera; used as Carries the diagonal from the restrained hands toward the pleading face; 수갑 (Fastened around both wrists) — Seen obliquely around the wrists above the tabletop; used as A small but readable restraint detail within the upper-body composition; 넓은 면회실 (Open room space remains visible beyond the centrally placed table); used as Negative space separates the isolated seated figure from the surrounding room.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient illumination preserves the facial bruising and hand detail with controlled contrast, without specifying a visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A table stands in the middle of the large visitation room. 현우: He sits at the table with his wrists still handcuffed and visible bruising on his face. His dog-bitten leg remains injured.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S23sh1__bgfirst_bg.png",
     "asset_id": "10a8564f-8f3e-46a6-8c49-5dd05488e9d9",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S23sh1.png",
     "asset_id": "69e13ce4-f531-431c-a45d-c3679c68fd0e",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L173B01.png",
     "asset_id": "0977c2dd-72ef-4442-9b57-c56745b38ac5",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선이 프레임 우측 전경의 인물을 향함.",
    "built_space": "창문과 좌측 문이 있는 레퍼런스 면회실 구조와 일치함. 중앙 테이블에 현우가 위치함.",
    "entities": "피멍이 든 얼굴과 수갑을 찬 현우. 프레임 우측에 지시문에 없는 정체불명의 인물이 있음.",
    "hard_violations": [
     "[gemini-pro] 지시문에 없는 추가 인물 등장 (프레임 우측 전경)",
     "[gpt-high] 현우 외에는 아무도 보여서는 안 되는 장면에 두 번째 남성의 머리와 상체를 전경으로 추가했다."
    ],
    "physics": "현우의 팔과 손이 테이블 위에서 자연스럽게 지탱됨."
   },
   {
    "label": "B",
    "direction": "현우의 시선이 프레임 밖 좌측 허공을 향함.",
    "built_space": "열린 문, 게시판, 창문 등 레퍼런스의 면회실 공간이 정확히 묘사됨. 테이블이 중앙에 배치됨.",
    "entities": "지시된 외모와 피멍, 수갑을 착용한 현우만 단독으로 프레임에 존재함. 테이블 위에 흐릿한 문서가 있음.",
    "hard_violations": [
     "[gpt-high] 프롬프트와 장소 참조에 없는 서류·클립보드·베이지색 파일을 테이블 전경에 추가했다."
    ],
    "physics": "현우의 하완과 손이 테이블 표면에 안정적으로 맞닿아 지탱됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지시문대로 현우가 단독으로 등장하며, 뻗은 팔과 수갑, 피멍 등 요구된 디테일을 정확한 프레이밍 내에서 사실적으로 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "샷 텍스트에 명시되지 않은 인물이 프레임 우측 전경에 등장하여 치명적인 하드 위반이 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선이 프레임 우측 전경의 인물을 향함.",
        "built_space": "창문과 좌측 문이 있는 레퍼런스 면회실 구조와 일치함. 중앙 테이블에 현우가 위치함.",
        "entities": "피멍이 든 얼굴과 수갑을 찬 현우. 프레임 우측에 지시문에 없는 정체불명의 인물이 있음.",
        "hard_violations": [
         "지시문에 없는 추가 인물 등장 (프레임 우측 전경)"
        ],
        "physics": "현우의 팔과 손이 테이블 위에서 자연스럽게 지탱됨."
       },
       {
        "label": "B",
        "direction": "현우의 시선이 프레임 밖 좌측 허공을 향함.",
        "built_space": "열린 문, 게시판, 창문 등 레퍼런스의 면회실 공간이 정확히 묘사됨. 테이블이 중앙에 배치됨.",
        "entities": "지시된 외모와 피멍, 수갑을 착용한 현우만 단독으로 프레임에 존재함. 테이블 위에 흐릿한 문서가 있음.",
        "hard_violations": [],
        "physics": "현우의 하완과 손이 테이블 표면에 안정적으로 맞닿아 지탱됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지시문대로 현우가 단독으로 등장하며, 뻗은 팔과 수갑, 피멍 등 요구된 디테일을 정확한 프레이밍 내에서 사실적으로 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "샷 텍스트에 명시되지 않은 인물이 프레임 우측 전경에 등장하여 치명적인 하드 위반이 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선이 프레임 우측 전경의 인물을 향함.",
        "built_space": "창문과 좌측 문이 있는 레퍼런스 면회실 구조와 일치함. 중앙 테이블에 현우가 위치함.",
        "entities": "피멍이 든 얼굴과 수갑을 찬 현우. 프레임 우측에 지시문에 없는 정체불명의 인물이 있음.",
        "hard_violations": [
         "지시문에 없는 추가 인물 등장 (프레임 우측 전경)"
        ],
        "physics": "현우의 팔과 손이 테이블 위에서 자연스럽게 지탱됨."
       },
       {
        "label": "B",
        "direction": "현우의 시선이 프레임 밖 좌측 허공을 향함.",
        "built_space": "열린 문, 게시판, 창문 등 레퍼런스의 면회실 공간이 정확히 묘사됨. 테이블이 중앙에 배치됨.",
        "entities": "지시된 외모와 피멍, 수갑을 착용한 현우만 단독으로 프레임에 존재함. 테이블 위에 흐릿한 문서가 있음.",
        "hard_violations": [],
        "physics": "현우의 하완과 손이 테이블 표면에 안정적으로 맞닿아 지탱됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "현우의 상체와 길게 뻗은 수갑 찬 양손을 담은 미디엄 구도는 더 충실하지만, 지시되지 않은 서류·클립보드·파일을 추가해 엄격한 통과 기준에는 미달한다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "수갑 찬 손을 내미는 절박한 동작은 맞지만, 전경에 두 번째 인물을 추가하여 현우만 보여야 하는 인물 제한과 상체 중심 구도를 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 화면 왼쪽 위의 화면 밖 상대를 바라보며 양팔을 테이블 너머 왼쪽 전경으로 길게 뻗는다. 시선의 실제 상대는 보이지 않으며, 손에서 얼굴로 이어지는 대각선은 요구된 구도와 맞는다. 무기나 별도의 지향성 소품은 없다.",
        "built_space": "중앙 테이블 한 개와 현우 뒤의 의자 등받이 한 개가 보인다. 뒤쪽에는 쇠창살 창 두 개, 왼쪽에는 열린 출입구 한 개와 그 너머 창 한 개가 있으며, 회색 하부 도장과 밝은 상부 벽은 장소 참조와 부합한다. 게시판은 한 개만 식별된다. 현우는 테이블 뒤 의자에 앉아 있고 빈 공간도 남아 있으나, 가까운 테이블 모서리는 프레임 밖으로 잘려 요구된 모서리 구도는 약하다. 불가능한 반사는 보이지 않는다.",
        "entities": "검은 헝클어진 머리와 남색 반팔을 입은 앳된 동아시아계 남성 한 명으로, 현우의 얼굴·나이대·의상 참조에 대체로 가깝다. 국적은 외모만으로 확인할 수 없다. 눈가와 뺨, 입 주변에 멍이 있고 양쪽 손목에 연결된 금속 수갑이 보인다. 다리는 구도 밖이므로 부상 상태는 판단하지 않는다. 왼쪽 전경에 서류, 클립보드, 베이지색 파일이 추가되어 있다. 글자 형태는 있으나 확실히 읽히는 문구는 식별하기 어렵다.",
        "hard_violations": [
         "프롬프트와 장소 참조에 없는 서류·클립보드·베이지색 파일을 테이블 전경에 추가했다."
        ],
        "physics": "몸은 뒤에 보이는 의자가 받치고, 길게 뻗은 팔과 양손은 테이블에 접촉하여 지지된다. 수갑은 각 손목을 감싸고 연결 사슬이 두 손목 사이에 놓인다. 추가된 종이류도 테이블 위에 놓여 있다. 지지 없이 뜬 물체나 불가능한 관절 자세는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "현우의 눈과 얼굴은 오른쪽 전경의 추가 남성을 향하며, 양손도 그 사람 쪽으로 내밀어 손바닥을 펼친다. 호소하는 행동의 방향은 명확하지만, 그 표적 인물 자체가 이 장면에 허용되지 않았다. 전경 남성은 현우 쪽으로 몸을 향하고 있으나 눈은 보이지 않는다.",
        "built_space": "중앙 테이블 한 개, 현우가 앉은 의자 일부 한 개, 왼쪽 닫힌 철문 한 개, 뒤쪽 쇠창살 창 두 개, 오른쪽 벽의 세로 배관과 바닥형 설비 한 개가 보인다. 벽의 이중 도장과 넓은 빈 바닥은 장소 참조에 가깝다. 테이블의 가까운 모서리와 상판은 잘 드러나지만, 추가 인물의 어깨가 오른쪽 전경을 크게 차지하여 단독 상체 구도가 어깨너머 구도로 바뀌었다. 불가능한 반사는 없다.",
        "entities": "중앙에는 헝클어진 검은 머리, 어두운 반팔, 양쪽 눈가의 피멍을 가진 젊은 동아시아계 남성이 있고 양손목에는 금속 수갑이 채워져 있다. 참조보다 얼굴이 다소 성숙하고 상의는 남색보다 검정에 가깝다. 오른쪽에는 비슷한 검은 머리와 어두운 상의를 입은 두 번째 남성의 머리와 상체가 보인다. 이 인물은 허용된 단독 현우 장면에 없는 추가 인물이다. 다리 부상은 프레임 밖이며 읽을 수 있는 글씨는 보이지 않는다.",
        "hard_violations": [
         "현우 외에는 아무도 보여서는 안 되는 장면에 두 번째 남성의 머리와 상체를 전경으로 추가했다."
        ],
        "physics": "현우는 등받이가 뒤에 있는 의자에 앉아 있고, 팔뚝은 테이블이 받친다. 펼친 손은 손목과 팔에 연결되어 자연스럽게 들려 있으며 수갑 사슬도 두 손목 사이에 연결된다. 전경 인물의 하체는 화면 밖이므로 지지 상태를 직접 확인할 수 없지만 공중에 떠 있다는 증거는 없다. 명백한 부유나 불가능한 해부학은 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "현우의 상체와 길게 뻗은 수갑 찬 양손을 담은 미디엄 구도는 더 충실하지만, 지시되지 않은 서류·클립보드·파일을 추가해 엄격한 통과 기준에는 미달한다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "수갑 찬 손을 내미는 절박한 동작은 맞지만, 전경에 두 번째 인물을 추가하여 현우만 보여야 하는 인물 제한과 상체 중심 구도를 위반한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 화면 왼쪽 위의 화면 밖 상대를 바라보며 양팔을 테이블 너머 왼쪽 전경으로 길게 뻗는다. 시선의 실제 상대는 보이지 않으며, 손에서 얼굴로 이어지는 대각선은 요구된 구도와 맞는다. 무기나 별도의 지향성 소품은 없다.",
        "built_space": "중앙 테이블 한 개와 현우 뒤의 의자 등받이 한 개가 보인다. 뒤쪽에는 쇠창살 창 두 개, 왼쪽에는 열린 출입구 한 개와 그 너머 창 한 개가 있으며, 회색 하부 도장과 밝은 상부 벽은 장소 참조와 부합한다. 게시판은 한 개만 식별된다. 현우는 테이블 뒤 의자에 앉아 있고 빈 공간도 남아 있으나, 가까운 테이블 모서리는 프레임 밖으로 잘려 요구된 모서리 구도는 약하다. 불가능한 반사는 보이지 않는다.",
        "entities": "검은 헝클어진 머리와 남색 반팔을 입은 앳된 동아시아계 남성 한 명으로, 현우의 얼굴·나이대·의상 참조에 대체로 가깝다. 국적은 외모만으로 확인할 수 없다. 눈가와 뺨, 입 주변에 멍이 있고 양쪽 손목에 연결된 금속 수갑이 보인다. 다리는 구도 밖이므로 부상 상태는 판단하지 않는다. 왼쪽 전경에 서류, 클립보드, 베이지색 파일이 추가되어 있다. 글자 형태는 있으나 확실히 읽히는 문구는 식별하기 어렵다.",
        "hard_violations": [
         "프롬프트와 장소 참조에 없는 서류·클립보드·베이지색 파일을 테이블 전경에 추가했다."
        ],
        "physics": "몸은 뒤에 보이는 의자가 받치고, 길게 뻗은 팔과 양손은 테이블에 접촉하여 지지된다. 수갑은 각 손목을 감싸고 연결 사슬이 두 손목 사이에 놓인다. 추가된 종이류도 테이블 위에 놓여 있다. 지지 없이 뜬 물체나 불가능한 관절 자세는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "현우의 눈과 얼굴은 오른쪽 전경의 추가 남성을 향하며, 양손도 그 사람 쪽으로 내밀어 손바닥을 펼친다. 호소하는 행동의 방향은 명확하지만, 그 표적 인물 자체가 이 장면에 허용되지 않았다. 전경 남성은 현우 쪽으로 몸을 향하고 있으나 눈은 보이지 않는다.",
        "built_space": "중앙 테이블 한 개, 현우가 앉은 의자 일부 한 개, 왼쪽 닫힌 철문 한 개, 뒤쪽 쇠창살 창 두 개, 오른쪽 벽의 세로 배관과 바닥형 설비 한 개가 보인다. 벽의 이중 도장과 넓은 빈 바닥은 장소 참조에 가깝다. 테이블의 가까운 모서리와 상판은 잘 드러나지만, 추가 인물의 어깨가 오른쪽 전경을 크게 차지하여 단독 상체 구도가 어깨너머 구도로 바뀌었다. 불가능한 반사는 없다.",
        "entities": "중앙에는 헝클어진 검은 머리, 어두운 반팔, 양쪽 눈가의 피멍을 가진 젊은 동아시아계 남성이 있고 양손목에는 금속 수갑이 채워져 있다. 참조보다 얼굴이 다소 성숙하고 상의는 남색보다 검정에 가깝다. 오른쪽에는 비슷한 검은 머리와 어두운 상의를 입은 두 번째 남성의 머리와 상체가 보인다. 이 인물은 허용된 단독 현우 장면에 없는 추가 인물이다. 다리 부상은 프레임 밖이며 읽을 수 있는 글씨는 보이지 않는다.",
        "hard_violations": [
         "현우 외에는 아무도 보여서는 안 되는 장면에 두 번째 남성의 머리와 상체를 전경으로 추가했다."
        ],
        "physics": "현우는 등받이가 뒤에 있는 의자에 앉아 있고, 팔뚝은 테이블이 받친다. 펼친 손은 손목과 팔에 연결되어 자연스럽게 들려 있으며 수갑 사슬도 두 손목 사이에 연결된다. 전경 인물의 하체는 화면 밖이므로 지지 상태를 직접 확인할 수 없지만 공중에 떠 있다는 증거는 없다. 명백한 부유나 불가능한 해부학은 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.929,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.679,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 지시문에 없는 추가 인물 등장 (프레임 우측 전경)",
     "[gpt-high] 현우 외에는 아무도 보여서는 안 되는 장면에 두 번째 남성의 머리와 상체를 전경으로 추가했다."
    ],
    "B": [
     "[gpt-high] 프롬프트와 장소 참조에 없는 서류·클립보드·베이지색 파일을 테이블 전경에 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 1750,
   "A": 679
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "지시문대로 현우가 단독으로 등장하며, 뻗은 팔과 수갑, 피멍 등 요구된 디테일을 정확한 프레이밍 내에서 사실적으로 구현했습니다.  ★위반: [gpt-high] 프롬프트와 장소 참조에 없는 서류·클립보드·베이지색 파일을 테이블 전경에 추가했다."
   },
   {
    "label": "A",
    "score": 679,
    "verdict_ko": "샷 텍스트에 명시되지 않은 인물이 프레임 우측 전경에 등장하여 치명적인 하드 위반이 발생했습니다.  ★위반: [gemini-pro] 지시문에 없는 추가 인물 등장 (프레임 우측 전경) / [gpt-high] 현우 외에는 아무도 보여서는 안 되는 장면에 두 번째 남성의 머리와 상체를 전경으로 추가했다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L173B01.png",
    "asset_id": "0977c2dd-72ef-4442-9b57-c56745b38ac5",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-1241-7483-bda4-7a2fe57b0341",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S23sh1__bgfirst_bg.png",
   "bg_asset_id": "10a8564f-8f3e-46a6-8c49-5dd05488e9d9",
   "bg_record_key": "S23sh1::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S23sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:02:54.098355+00:00",
  "fingerprint": "608a13964217c03ad88f839780eab17a5f16d1a5a5c011da0c4241d44589e4a6",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S23sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S23sh1_sel.png",
  "source_sha256": "24b646eab895044a4e4c603d43973a7617733ef6282668bc96d88ded9ac32d49",
  "file": "S23sh1_cine.png",
  "staged_sha256": "04affc85104b2e864cf97ceb66c9848e3a28b4e721688fe7a6e25f8a88115718",
  "latency_ms": 10761
 },
 "S23sh6::signage": {
  "fp": "7dd4d8c93e26b70b",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S23sh6": {
  "input_fingerprint": "de4b30fe6af10230",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 테이블 위에 놓인 찰리의 사진을 불안한 기색으로 내려다보는 현우의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): At the central visiting table inside the spacious detention-center room, with a robot photograph visible in daytime interior light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 찰리 수배 사진 (Lying on the table, partially visible at the lower frame edge) — The image-bearing face is angled upward toward the camera, showing part of 찰리's likeness without introducing additional readable text; used as Provides the concrete cause of 현우's lowered gaze while remaining subordinate to his face; 면회실 테이블 (Supporting the photograph) — Only a narrow oblique strip of the upper surface is visible; used as Maintains the table-side spatial continuity of the approach.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding restrained ambient illumination, allowing the lowered eyes and bruises to remain legible without a new lighting accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same visitation table and the surrounding room surfaces and lighting. Exclude the earlier handcuffs, which have now been removed.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The wanted photograph depicts Charlie; it is the same photograph brought into the visitation room. 현우: His wrists are now free of handcuffs. He remains at the table with facial bruising and the persistent dog-bite leg injury.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 테이블 위에 놓인 찰리의 사진을 불안한 기색으로 내려다보는 현우의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): At the central visiting table inside the spacious detention-center room, with a robot photograph visible in daytime interior light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 찰리 수배 사진 (Lying on the table, partially visible at the lower frame edge) — The image-bearing face is angled upward toward the camera, showing part of 찰리's likeness without introducing additional readable text; used as Provides the concrete cause of 현우's lowered gaze while remaining subordinate to his face; 면회실 테이블 (Supporting the photograph) — Only a narrow oblique strip of the upper surface is visible; used as Maintains the table-side spatial continuity of the approach.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding restrained ambient illumination, allowing the lowered eyes and bruises to remain legible without a new lighting accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same visitation table and the surrounding room surfaces and lighting. Exclude the earlier handcuffs, which have now been removed.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The wanted photograph depicts Charlie; it is the same photograph brought into the visitation room. 현우: His wrists are now free of handcuffs. He remains at the table with facial bruising and the persistent dog-bite leg injury.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 테이블 위에 놓인 찰리의 사진을 불안한 기색으로 내려다보는 현우의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): At the central visiting table inside the spacious detention-center room, with a robot photograph visible in daytime interior light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 찰리 수배 사진 (Lying on the table, partially visible at the lower frame edge) — The image-bearing face is angled upward toward the camera, showing part of 찰리's likeness without introducing additional readable text; used as Provides the concrete cause of 현우's lowered gaze while remaining subordinate to his face; 면회실 테이블 (Supporting the photograph) — Only a narrow oblique strip of the upper surface is visible; used as Maintains the table-side spatial continuity of the approach.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding restrained ambient illumination, allowing the lowered eyes and bruises to remain legible without a new lighting accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same visitation table and the surrounding room surfaces and lighting. Exclude the earlier handcuffs, which have now been removed.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The wanted photograph depicts Charlie; it is the same photograph brought into the visitation room. 현우: His wrists are now free of handcuffs. He remains at the table with facial bruising and the persistent dog-bite leg injury.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선이 테이블 위에 놓인 로봇 사진을 향해 자연스럽게 아래로 향함.",
    "built_space": "이전 샷과 일치하는 면회실의 창문과 벽 배경이 보이며, 앞쪽에 금속 테이블 표면이 위치함.",
    "entities": "현우의 멍든 얼굴, 헝클어진 머리, 의상이 레퍼런스와 일치하며, 테이블에 텍스트가 없는 로봇 사진이 있음.",
    "hard_violations": [],
    "physics": "현우의 상체가 테이블 너머로 안정적으로 위치해 있고, 사진은 테이블 표면 위에 평평하게 놓여 있음."
   },
   {
    "label": "B",
    "direction": "현우의 시선이 테이블 위에 놓인 로봇 사진을 향해 아래로 향함.",
    "built_space": "이전 샷과 일치하는 면회실 배경과 앞쪽의 금속 테이블이 보임.",
    "entities": "현우의 인상과 멍 자국은 레퍼런스와 일치하나, 로봇 사진 안에 'Charlie'라는 텍스트가 선명하게 포함됨.",
    "hard_violations": [
     "[gemini-pro] 이미지 내에 읽을 수 있는 텍스트('Charlie')가 생성됨 (leaked text)",
     "[gpt-high] 사진 속 로봇 가슴에 영문 이름이 선명하게 읽혀, 어디에도 읽을 수 있는 글자를 넣지 말라는 조건을 위반합니다."
    ],
    "physics": "현우가 테이블 쪽에 몸을 기댄 채 위치하며, 사진은 테이블 위에 지지되어 놓여 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지시된 클로즈업 구도와 레퍼런스의 인물 특징을 정확히 반영했으며, 금지된 텍스트 없이 로봇 사진을 적절히 배치하여 프롬프트를 충실히 이행했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "사진 속 로봇에 읽을 수 있는 텍스트('Charlie')가 포함되어 있어, 모든 텍스트 생성을 금지한 최우선 지침을 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선이 테이블 위에 놓인 로봇 사진을 향해 자연스럽게 아래로 향함.",
        "built_space": "이전 샷과 일치하는 면회실의 창문과 벽 배경이 보이며, 앞쪽에 금속 테이블 표면이 위치함.",
        "entities": "현우의 멍든 얼굴, 헝클어진 머리, 의상이 레퍼런스와 일치하며, 테이블에 텍스트가 없는 로봇 사진이 있음.",
        "hard_violations": [],
        "physics": "현우의 상체가 테이블 너머로 안정적으로 위치해 있고, 사진은 테이블 표면 위에 평평하게 놓여 있음."
       },
       {
        "label": "B",
        "direction": "현우의 시선이 테이블 위에 놓인 로봇 사진을 향해 아래로 향함.",
        "built_space": "이전 샷과 일치하는 면회실 배경과 앞쪽의 금속 테이블이 보임.",
        "entities": "현우의 인상과 멍 자국은 레퍼런스와 일치하나, 로봇 사진 안에 'Charlie'라는 텍스트가 선명하게 포함됨.",
        "hard_violations": [
         "이미지 내에 읽을 수 있는 텍스트('Charlie')가 생성됨 (leaked text)"
        ],
        "physics": "현우가 테이블 쪽에 몸을 기댄 채 위치하며, 사진은 테이블 위에 지지되어 놓여 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지시된 클로즈업 구도와 레퍼런스의 인물 특징을 정확히 반영했으며, 금지된 텍스트 없이 로봇 사진을 적절히 배치하여 프롬프트를 충실히 이행했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "사진 속 로봇에 읽을 수 있는 텍스트('Charlie')가 포함되어 있어, 모든 텍스트 생성을 금지한 최우선 지침을 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선이 테이블 위에 놓인 로봇 사진을 향해 자연스럽게 아래로 향함.",
        "built_space": "이전 샷과 일치하는 면회실의 창문과 벽 배경이 보이며, 앞쪽에 금속 테이블 표면이 위치함.",
        "entities": "현우의 멍든 얼굴, 헝클어진 머리, 의상이 레퍼런스와 일치하며, 테이블에 텍스트가 없는 로봇 사진이 있음.",
        "hard_violations": [],
        "physics": "현우의 상체가 테이블 너머로 안정적으로 위치해 있고, 사진은 테이블 표면 위에 평평하게 놓여 있음."
       },
       {
        "label": "B",
        "direction": "현우의 시선이 테이블 위에 놓인 로봇 사진을 향해 아래로 향함.",
        "built_space": "이전 샷과 일치하는 면회실 배경과 앞쪽의 금속 테이블이 보임.",
        "entities": "현우의 인상과 멍 자국은 레퍼런스와 일치하나, 로봇 사진 안에 'Charlie'라는 텍스트가 선명하게 포함됨.",
        "hard_violations": [
         "이미지 내에 읽을 수 있는 텍스트('Charlie')가 생성됨 (leaked text)"
        ],
        "physics": "현우가 테이블 쪽에 몸을 기댄 채 위치하며, 사진은 테이블 위에 지지되어 놓여 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "현우의 얼굴과 멍은 잘 이어지지만 사진에 읽히는 글자가 있어 실격이며, 사진을 내려다보는 시선과 좁고 비스듬한 상판 구도도 불충분합니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "읽히는 글자 없이 현우와 로봇 사진을 구현해 우세하지만, 시선이 사진보다 정면에 가깝고 상판과 사진이 하단을 지나치게 넓게 차지합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 얼굴은 정면을 향하고 눈은 약간 낮아져 있으나, 바로 아래 탁자에 놓인 사진보다 카메라 쪽 전방을 보는 인상이 강합니다. 사진의 로봇은 관객에게 거꾸로 보이므로 현우가 읽는 방향으로 놓였으며, 인쇄면은 위를 향합니다.",
        "built_space": "금속 상판 하나와 그 위 사진 한 장, 뒤쪽의 밝은 창 구획 두 곳과 밝은 상부·회색 하부 벽이 보입니다. 현우는 탁자 건너편에서 몸을 앞으로 숙이고 있습니다. 상판은 요구된 좁고 비스듬한 띠가 아니라 화면 하단 약 4분의 1을 차지하는 수평 영역입니다. 창의 세부 창살은 흐려 확인하기 어렵고, 불가능한 반사나 중복 설비는 보이지 않습니다.",
        "entities": "앳된 동아시아계 남성 한 명의 얼굴, 헝클어진 검은 머리, 남색 티셔츠와 왼쪽 눈가·볼의 멍이 참조 속 현우와 대체로 일치합니다. 한국계 미국인이라는 국적은 외관만으로 확인할 수 없습니다. 사진에는 로봇이 있으며 가슴 부분의 영문 이름이 읽힙니다. 찰리의 별도 외형 참조가 없어 로봇의 정확한 동일성은 검증할 수 없습니다. 손목과 다리는 프레임 밖이므로 수갑 제거와 다리 부상은 확인 대상이 아닙니다.",
        "hard_violations": [
         "사진 속 로봇 가슴에 영문 이름이 선명하게 읽혀, 어디에도 읽을 수 있는 글자를 넣지 말라는 조건을 위반합니다."
        ],
        "physics": "사진은 금속 상판에 평평하게 놓여 지지됩니다. 현우의 머리는 목과 상체에 자연스럽게 연결되어 있으며 탁자 높이까지 몸을 숙인 자세로 볼 수 있습니다. 의자와 하체는 보이지 않지만 공중에 떠 있다고 판단할 근거는 없습니다."
       },
       {
        "label": "B",
        "direction": "현우는 고개와 눈을 조금 낮추었지만 시선은 여전히 카메라 가까운 정면 아래를 향하며, 바로 아래 로봇 사진에 눈을 고정한 하향 시선은 명확하지 않습니다. 사진의 인쇄면은 위를 향하고 로봇의 머리는 관객 쪽에 있어 현우에게 올바른 방향입니다.",
        "built_space": "금속 탁자 하나와 사진 한 장이 전경에 있고, 뒤에는 창 구획 두 곳과 밝은 상부·회색 하부 벽이 있습니다. 현우는 탁자 건너편에 있으며 의자는 잘려 있습니다. 상판이 하단 약 4분의 1을 수평으로 차지해 좁고 비스듬한 상판 조각이라는 지시와 다릅니다. 뒤 벽의 벗겨진 자국은 참조보다 두드러집니다. 중복 설비나 불가능한 반사는 보이지 않습니다.",
        "entities": "한 명의 앳된 동아시아계 남성, 검은 헝클어진 머리, 남색 티셔츠와 왼쪽 얼굴의 멍이 현우의 참조와 대체로 맞습니다. 국적은 시각적으로 판별할 수 없습니다. 하단 사진에는 로봇 얼굴과 몸통 일부가 보이고 읽히는 글자는 없습니다. 찰리의 정확한 외형을 대조할 별도 참조는 없습니다. 손목과 다리는 클로즈업 밖이므로 해당 상태를 보이지 않는다는 이유로 감점하지 않습니다.",
        "hard_violations": [],
        "physics": "사진은 상판에 놓여 있고 들리거나 떠 있는 부분은 없습니다. 현우는 상체를 앞으로 낮춘 자세이며 머리와 목의 연결도 자연스럽습니다. 하체 지지점은 화면 밖이지만, 보이는 신체에서 불가능한 자세나 무지지 부유는 발견되지 않습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "현우의 얼굴과 멍은 잘 이어지지만 사진에 읽히는 글자가 있어 실격이며, 사진을 내려다보는 시선과 좁고 비스듬한 상판 구도도 불충분합니다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "읽히는 글자 없이 현우와 로봇 사진을 구현해 우세하지만, 시선이 사진보다 정면에 가깝고 상판과 사진이 하단을 지나치게 넓게 차지합니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 얼굴은 정면을 향하고 눈은 약간 낮아져 있으나, 바로 아래 탁자에 놓인 사진보다 카메라 쪽 전방을 보는 인상이 강합니다. 사진의 로봇은 관객에게 거꾸로 보이므로 현우가 읽는 방향으로 놓였으며, 인쇄면은 위를 향합니다.",
        "built_space": "금속 상판 하나와 그 위 사진 한 장, 뒤쪽의 밝은 창 구획 두 곳과 밝은 상부·회색 하부 벽이 보입니다. 현우는 탁자 건너편에서 몸을 앞으로 숙이고 있습니다. 상판은 요구된 좁고 비스듬한 띠가 아니라 화면 하단 약 4분의 1을 차지하는 수평 영역입니다. 창의 세부 창살은 흐려 확인하기 어렵고, 불가능한 반사나 중복 설비는 보이지 않습니다.",
        "entities": "앳된 동아시아계 남성 한 명의 얼굴, 헝클어진 검은 머리, 남색 티셔츠와 왼쪽 눈가·볼의 멍이 참조 속 현우와 대체로 일치합니다. 한국계 미국인이라는 국적은 외관만으로 확인할 수 없습니다. 사진에는 로봇이 있으며 가슴 부분의 영문 이름이 읽힙니다. 찰리의 별도 외형 참조가 없어 로봇의 정확한 동일성은 검증할 수 없습니다. 손목과 다리는 프레임 밖이므로 수갑 제거와 다리 부상은 확인 대상이 아닙니다.",
        "hard_violations": [
         "사진 속 로봇 가슴에 영문 이름이 선명하게 읽혀, 어디에도 읽을 수 있는 글자를 넣지 말라는 조건을 위반합니다."
        ],
        "physics": "사진은 금속 상판에 평평하게 놓여 지지됩니다. 현우의 머리는 목과 상체에 자연스럽게 연결되어 있으며 탁자 높이까지 몸을 숙인 자세로 볼 수 있습니다. 의자와 하체는 보이지 않지만 공중에 떠 있다고 판단할 근거는 없습니다."
       },
       {
        "label": "A",
        "direction": "현우는 고개와 눈을 조금 낮추었지만 시선은 여전히 카메라 가까운 정면 아래를 향하며, 바로 아래 로봇 사진에 눈을 고정한 하향 시선은 명확하지 않습니다. 사진의 인쇄면은 위를 향하고 로봇의 머리는 관객 쪽에 있어 현우에게 올바른 방향입니다.",
        "built_space": "금속 탁자 하나와 사진 한 장이 전경에 있고, 뒤에는 창 구획 두 곳과 밝은 상부·회색 하부 벽이 있습니다. 현우는 탁자 건너편에 있으며 의자는 잘려 있습니다. 상판이 하단 약 4분의 1을 수평으로 차지해 좁고 비스듬한 상판 조각이라는 지시와 다릅니다. 뒤 벽의 벗겨진 자국은 참조보다 두드러집니다. 중복 설비나 불가능한 반사는 보이지 않습니다.",
        "entities": "한 명의 앳된 동아시아계 남성, 검은 헝클어진 머리, 남색 티셔츠와 왼쪽 얼굴의 멍이 현우의 참조와 대체로 맞습니다. 국적은 시각적으로 판별할 수 없습니다. 하단 사진에는 로봇 얼굴과 몸통 일부가 보이고 읽히는 글자는 없습니다. 찰리의 정확한 외형을 대조할 별도 참조는 없습니다. 손목과 다리는 클로즈업 밖이므로 해당 상태를 보이지 않는다는 이유로 감점하지 않습니다.",
        "hard_violations": [],
        "physics": "사진은 상판에 놓여 있고 들리거나 떠 있는 부분은 없습니다. 현우는 상체를 앞으로 낮춘 자세이며 머리와 목의 연결도 자연스럽습니다. 하체 지지점은 화면 밖이지만, 보이는 신체에서 불가능한 자세나 무지지 부유는 발견되지 않습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.929
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.679
   },
   "violations": {
    "B": [
     "[gemini-pro] 이미지 내에 읽을 수 있는 텍스트('Charlie')가 생성됨 (leaked text)",
     "[gpt-high] 사진 속 로봇 가슴에 영문 이름이 선명하게 읽혀, 어디에도 읽을 수 있는 글자를 넣지 말라는 조건을 위반합니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 679
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지시된 클로즈업 구도와 레퍼런스의 인물 특징을 정확히 반영했으며, 금지된 텍스트 없이 로봇 사진을 적절히 배치하여 프롬프트를 충실히 이행했습니다."
   },
   {
    "label": "B",
    "score": 679,
    "verdict_ko": "사진 속 로봇에 읽을 수 있는 텍스트('Charlie')가 포함되어 있어, 모든 텍스트 생성을 금지한 최우선 지침을 위반했습니다.  ★위반: [gemini-pro] 이미지 내에 읽을 수 있는 텍스트('Charlie')가 생성됨 (leaked text) / [gpt-high] 사진 속 로봇 가슴에 영문 이름이 선명하게 읽혀, 어디에도 읽을 수 있는 글자를 넣지 말라는 조건을 위반합니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S23sh1_sel.png",
    "asset_id": "717b18c7-54d5-4426-8cc5-e1c0830b025e",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-15a7-74c0-b5bf-2b09dcbd2048",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S23sh1"
  }
 },
 "S23sh6::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:03:51.342819+00:00",
  "fingerprint": "92ccec68c53ce6644a5efcaa2d9133e418e489280603fcb16c1e6771214dbc4b",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S23sh6_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S23sh6_sel.png",
  "source_sha256": "dab0a2d366ccab80f16ab663f72e351799fce2c4c2b98da938c5d35b2db26e4c",
  "file": "S23sh6_cine.png",
  "staged_sha256": "3b3e95fce901b4136cd4cfeb8984aba3ee405c3eb873b100416ba36f75ab1b48",
  "latency_ms": 11651
 },
 "S23sh11::signage": {
  "fp": "5fa6b125f39d0b83",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S23sh11": {
  "input_fingerprint": "afbfa42db7d37089",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 서로를 마주 보며 팽팽한 긴장감 속에 마주 앉아 있는 현우와 윤성찬의 옆모습 풀샷.\n\nLOCATION (lock): Across the central table inside the spacious detention-center visiting room, under ordinary interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 면회실 테이블 (Separating the two seated men) — Its side and a shallow portion of its top are visible between the opposing profiles; used as Defines the negotiation axis and the physical separation; 두 사람의 좌석 (Occupied on opposite sides of the table) — Viewed from the side, with both seated bodies and feet unobscured; used as Makes the different seated weight distributions readable; 넓은 면회실 (Visible around the central seating arrangement); used as Leaves austere breathing room around the intimate negotiation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the same controlled ambient contrast across both figures and the surrounding visiting room.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the visitation table, nearby room surfaces, and consistent daytime interior lighting. Exclude the handcuffs from the earlier arrival state.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The wanted photograph of Charlie remains available in the visitation room. 현우: He remains at the table with his handcuffs removed, his face bruised and his dog-bitten leg still injured. 윤성찬: He remains in his suit during the negotiation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 서로를 마주 보며 팽팽한 긴장감 속에 마주 앉아 있는 현우와 윤성찬의 옆모습 풀샷.\n\nLOCATION (lock): Across the central table inside the spacious detention-center visiting room, under ordinary interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 면회실 테이블 (Separating the two seated men) — Its side and a shallow portion of its top are visible between the opposing profiles; used as Defines the negotiation axis and the physical separation; 두 사람의 좌석 (Occupied on opposite sides of the table) — Viewed from the side, with both seated bodies and feet unobscured; used as Makes the different seated weight distributions readable; 넓은 면회실 (Visible around the central seating arrangement); used as Leaves austere breathing room around the intimate negotiation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the same controlled ambient contrast across both figures and the surrounding visiting room.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the visitation table, nearby room surfaces, and consistent daytime interior lighting. Exclude the handcuffs from the earlier arrival state.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The wanted photograph of Charlie remains available in the visitation room. 현우: He remains at the table with his handcuffs removed, his face bruised and his dog-bitten leg still injured. 윤성찬: He remains in his suit during the negotiation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 서로를 마주 보며 팽팽한 긴장감 속에 마주 앉아 있는 현우와 윤성찬의 옆모습 풀샷.\n\nLOCATION (lock): Across the central table inside the spacious detention-center visiting room, under ordinary interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 면회실 테이블 (Separating the two seated men) — Its side and a shallow portion of its top are visible between the opposing profiles; used as Defines the negotiation axis and the physical separation; 두 사람의 좌석 (Occupied on opposite sides of the table) — Viewed from the side, with both seated bodies and feet unobscured; used as Makes the different seated weight distributions readable; 넓은 면회실 (Visible around the central seating arrangement); used as Leaves austere breathing room around the intimate negotiation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the same controlled ambient contrast across both figures and the surrounding visiting room.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the visitation table, nearby room surfaces, and consistent daytime interior lighting. Exclude the handcuffs from the earlier arrival state.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The wanted photograph of Charlie remains available in the visitation room. 현우: He remains at the table with his handcuffs removed, his face bruised and his dog-bitten leg still injured. 윤성찬: He remains in his suit during the negotiation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "두 인물이 테이블을 사이에 두고 서로의 눈을 정면으로 응시함.",
    "built_space": "넓은 면회실. 중앙 테이블 양옆에 인물들이 위치하며, 배경에 배치된 다른 테이블과 의자들의 구조가 정상적임.",
    "entities": "현우는 얼굴과 다리에 멍이 있으며 수갑을 차지 않은 상태로 참고 이미지와 일치함. 윤성찬은 정장 차림으로 완벽히 일치함. 테이블 위 사진이 이전 샷의 내용과 동일함.",
    "hard_violations": [
     "[gpt-high] 요청과 참조에 없는 검은 펜과 갈색 서류봉투를 중앙 탁자의 소품으로 추가했다."
    ],
    "physics": "두 인물 모두 의자에 자연스럽게 착석하여 체중이 시트에 실려 있으며, 손과 발의 위치가 안정적임."
   },
   {
    "label": "B",
    "direction": "두 인물이 테이블을 사이에 두고 서로를 마주 봄.",
    "built_space": "면회실 중앙에 테이블과 의자가 배치됨. 배경 좌측에 있는 테이블과 의자 다리들이 형체를 알 수 없이 엉켜 있음.",
    "entities": "현우는 우측 손목에 금속 수갑을 찬 상태이며 긴바지를 입어 상처가 가려짐. 윤성찬은 참고 이미지와 일치함.",
    "hard_violations": [
     "[gemini-pro] 배경 좌측의 의자와 테이블 다리가 물리적으로 불가능하게 융합됨"
    ],
    "physics": "인물들은 의자에 앉아 바닥과 시트에 의해 지지받고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "요구된 발끝 노출 프레이밍은 다소 잘렸으나, 수갑 해제 및 다리 상처 노출 지시를 완벽히 구현했고 이전 샷의 사진 소품까지 정확히 유지함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "발이 보이는 프레이밍은 충족했으나, 제거해야 할 수갑이 여전히 남아있고 배경 가구의 물리적 구조가 무너져 실격됨."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 인물이 테이블을 사이에 두고 서로의 눈을 정면으로 응시함.",
        "built_space": "넓은 면회실. 중앙 테이블 양옆에 인물들이 위치하며, 배경에 배치된 다른 테이블과 의자들의 구조가 정상적임.",
        "entities": "현우는 얼굴과 다리에 멍이 있으며 수갑을 차지 않은 상태로 참고 이미지와 일치함. 윤성찬은 정장 차림으로 완벽히 일치함. 테이블 위 사진이 이전 샷의 내용과 동일함.",
        "hard_violations": [],
        "physics": "두 인물 모두 의자에 자연스럽게 착석하여 체중이 시트에 실려 있으며, 손과 발의 위치가 안정적임."
       },
       {
        "label": "B",
        "direction": "두 인물이 테이블을 사이에 두고 서로를 마주 봄.",
        "built_space": "면회실 중앙에 테이블과 의자가 배치됨. 배경 좌측에 있는 테이블과 의자 다리들이 형체를 알 수 없이 엉켜 있음.",
        "entities": "현우는 우측 손목에 금속 수갑을 찬 상태이며 긴바지를 입어 상처가 가려짐. 윤성찬은 참고 이미지와 일치함.",
        "hard_violations": [
         "배경 좌측의 의자와 테이블 다리가 물리적으로 불가능하게 융합됨"
        ],
        "physics": "인물들은 의자에 앉아 바닥과 시트에 의해 지지받고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "요구된 발끝 노출 프레이밍은 다소 잘렸으나, 수갑 해제 및 다리 상처 노출 지시를 완벽히 구현했고 이전 샷의 사진 소품까지 정확히 유지함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "발이 보이는 프레이밍은 충족했으나, 제거해야 할 수갑이 여전히 남아있고 배경 가구의 물리적 구조가 무너져 실격됨."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "두 인물이 테이블을 사이에 두고 서로의 눈을 정면으로 응시함.",
        "built_space": "넓은 면회실. 중앙 테이블 양옆에 인물들이 위치하며, 배경에 배치된 다른 테이블과 의자들의 구조가 정상적임.",
        "entities": "현우는 얼굴과 다리에 멍이 있으며 수갑을 차지 않은 상태로 참고 이미지와 일치함. 윤성찬은 정장 차림으로 완벽히 일치함. 테이블 위 사진이 이전 샷의 내용과 동일함.",
        "hard_violations": [],
        "physics": "두 인물 모두 의자에 자연스럽게 착석하여 체중이 시트에 실려 있으며, 손과 발의 위치가 안정적임."
       },
       {
        "label": "B",
        "direction": "두 인물이 테이블을 사이에 두고 서로를 마주 봄.",
        "built_space": "면회실 중앙에 테이블과 의자가 배치됨. 배경 좌측에 있는 테이블과 의자 다리들이 형체를 알 수 없이 엉켜 있음.",
        "entities": "현우는 우측 손목에 금속 수갑을 찬 상태이며 긴바지를 입어 상처가 가려짐. 윤성찬은 참고 이미지와 일치함.",
        "hard_violations": [
         "배경 좌측의 의자와 테이블 다리가 물리적으로 불가능하게 융합됨"
        ],
        "physics": "인물들은 의자에 앉아 바닥과 시트에 의해 지지받고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "두 사람이 서로를 바라보는 옆모습과 발까지 담은 풀샷을 구현해 우세하지만, 탁자 상판의 노출이 다소 넓고 현우의 얼굴 멍은 약하다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "상호 응시와 얼굴 부상은 잘 보이지만, 발과 하체를 자른 구도가 옆모습 풀샷 지시를 어기며 탁자에 불필요한 펜과 봉투를 추가했다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 현우는 오른쪽 윤성찬의 얼굴을, 오른쪽 윤성찬은 왼쪽 현우의 얼굴을 바라본다. 두 사람의 몸과 무릎도 탁자를 사이에 두고 서로를 향한다. 카메라를 보는 인물은 없다.",
        "built_space": "중앙 금속 탁자 한 개의 양쪽에 의자 한 개씩이 있고 두 사람이 각각 앉아 있다. 뒤에는 빈 탁자 두 개와 각각의 주변 의자들이 보인다. 뒤 벽의 창 두 구획과 오른쪽 벽의 부분 창 한 구획, 회색 하단 도장과 벗겨진 밝은 상단 벽이 보이며 이전 장면의 재질과 낮 조명을 따른다. 두 사람의 머리부터 신발까지 포함하지만 가까운 탁자 다리와 가로대가 하체 일부에 겹친다. 상판은 요청한 얕은 부분보다 비교적 넓게 드러난다.",
        "entities": "등장인물은 젊은 동아시아계 남성 현우와 고령의 동아시아계 남성 윤성찬 두 명뿐이다. 현우의 검은 머리, 남색 반소매 상의와 앳된 얼굴은 참조에 대체로 맞지만 얼굴 멍은 이전 장면보다 희미하다. 긴 바지가 다리를 덮어 개에게 물린 상처는 확인할 수 없다. 윤성찬의 회색 머리, 안경, 주름진 얼굴, 짙은 재킷과 청색 셔츠는 참조에 부합한다. 탁자에는 사진 한 장이 있으나 찰리의 정확한 모습까지 판별하기는 어렵다. 수갑과 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "두 사람의 엉덩이는 각 의자 좌판에 놓이고 등받이는 몸 뒤에 있다. 손은 허벅지와 무릎 위에 얹혀 있으며 신발은 바닥에 닿는다. 의자와 탁자는 바닥에 선 다리로 지지되고 사진은 상판에 놓여 있다. 지지 없이 떠 있는 인물이나 물체는 없다."
       },
       {
        "label": "B",
        "direction": "현우는 오른쪽 윤성찬의 얼굴을 바라보고 윤성찬은 왼쪽 현우의 얼굴을 바라본다. 두 사람 모두 몸을 상대 쪽으로 향하며, 현우는 상체를 조금 앞으로 기울인다. 사진을 내려다보거나 카메라를 응시하는 구도는 아니다.",
        "built_space": "중앙 금속 탁자 한 개와 두 사람이 사용하는 의자 두 개가 보인다. 중앙 탁자 뒤쪽에도 빈 의자 등받이 두 개가 추가로 드러난다. 배경에는 빈 탁자 두 개가 있으며 가운데 뒤 탁자 양옆으로 의자 네 개가 보인다. 뒤 벽에는 왼쪽 끝의 부분 창을 포함한 세 창 구획, 오른쪽 벽에는 부분 창 한 구획이 보인다. 낡은 투톤 벽과 금속 가구, 낮빛은 참조 장소와 유사하지만, 프레임 아래에서 두 사람의 다리와 발이 잘려 요구한 전신 측면 구도가 아니다. 탁자 상판도 상당히 넓게 보인다.",
        "entities": "현우와 윤성찬에 해당하는 두 남성만 등장한다. 현우는 검은 헝클어진 머리와 남색 반소매 옷을 입고 얼굴과 손에 멍이 있으며, 반바지 아래 드러난 다리에 상처가 보인다. 윤성찬은 회색 머리와 안경, 고령의 얼굴, 짙은 재킷과 청색 셔츠로 참조를 따른다. 탁자에는 사진 한 장 외에 검은 펜과 갈색 서류봉투가 추가되어 있다. 수갑은 없고 봉투의 흔적은 읽을 수 있는 글자로 판별되지 않는다.",
        "hard_violations": [
         "요청과 참조에 없는 검은 펜과 갈색 서류봉투를 중앙 탁자의 소품으로 추가했다."
        ],
        "physics": "두 사람 모두 의자 좌판에 앉아 있고 등받이가 몸 뒤에 있어 착석 구조는 자연스럽다. 현우의 양손과 팔은 상판에 지지되고 윤성찬의 보이는 손은 허벅지에 놓인다. 사진과 펜, 봉투는 상판 위에 놓여 있다. 발은 화면 밖이라 바닥 접촉을 확인할 수 없지만, 보이는 몸은 의자로 지지되며 공중에 떠 있는 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "두 사람이 서로를 바라보는 옆모습과 발까지 담은 풀샷을 구현해 우세하지만, 탁자 상판의 노출이 다소 넓고 현우의 얼굴 멍은 약하다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "상호 응시와 얼굴 부상은 잘 보이지만, 발과 하체를 자른 구도가 옆모습 풀샷 지시를 어기며 탁자에 불필요한 펜과 봉투를 추가했다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽 현우는 오른쪽 윤성찬의 얼굴을, 오른쪽 윤성찬은 왼쪽 현우의 얼굴을 바라본다. 두 사람의 몸과 무릎도 탁자를 사이에 두고 서로를 향한다. 카메라를 보는 인물은 없다.",
        "built_space": "중앙 금속 탁자 한 개의 양쪽에 의자 한 개씩이 있고 두 사람이 각각 앉아 있다. 뒤에는 빈 탁자 두 개와 각각의 주변 의자들이 보인다. 뒤 벽의 창 두 구획과 오른쪽 벽의 부분 창 한 구획, 회색 하단 도장과 벗겨진 밝은 상단 벽이 보이며 이전 장면의 재질과 낮 조명을 따른다. 두 사람의 머리부터 신발까지 포함하지만 가까운 탁자 다리와 가로대가 하체 일부에 겹친다. 상판은 요청한 얕은 부분보다 비교적 넓게 드러난다.",
        "entities": "등장인물은 젊은 동아시아계 남성 현우와 고령의 동아시아계 남성 윤성찬 두 명뿐이다. 현우의 검은 머리, 남색 반소매 상의와 앳된 얼굴은 참조에 대체로 맞지만 얼굴 멍은 이전 장면보다 희미하다. 긴 바지가 다리를 덮어 개에게 물린 상처는 확인할 수 없다. 윤성찬의 회색 머리, 안경, 주름진 얼굴, 짙은 재킷과 청색 셔츠는 참조에 부합한다. 탁자에는 사진 한 장이 있으나 찰리의 정확한 모습까지 판별하기는 어렵다. 수갑과 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "두 사람의 엉덩이는 각 의자 좌판에 놓이고 등받이는 몸 뒤에 있다. 손은 허벅지와 무릎 위에 얹혀 있으며 신발은 바닥에 닿는다. 의자와 탁자는 바닥에 선 다리로 지지되고 사진은 상판에 놓여 있다. 지지 없이 떠 있는 인물이나 물체는 없다."
       },
       {
        "label": "A",
        "direction": "현우는 오른쪽 윤성찬의 얼굴을 바라보고 윤성찬은 왼쪽 현우의 얼굴을 바라본다. 두 사람 모두 몸을 상대 쪽으로 향하며, 현우는 상체를 조금 앞으로 기울인다. 사진을 내려다보거나 카메라를 응시하는 구도는 아니다.",
        "built_space": "중앙 금속 탁자 한 개와 두 사람이 사용하는 의자 두 개가 보인다. 중앙 탁자 뒤쪽에도 빈 의자 등받이 두 개가 추가로 드러난다. 배경에는 빈 탁자 두 개가 있으며 가운데 뒤 탁자 양옆으로 의자 네 개가 보인다. 뒤 벽에는 왼쪽 끝의 부분 창을 포함한 세 창 구획, 오른쪽 벽에는 부분 창 한 구획이 보인다. 낡은 투톤 벽과 금속 가구, 낮빛은 참조 장소와 유사하지만, 프레임 아래에서 두 사람의 다리와 발이 잘려 요구한 전신 측면 구도가 아니다. 탁자 상판도 상당히 넓게 보인다.",
        "entities": "현우와 윤성찬에 해당하는 두 남성만 등장한다. 현우는 검은 헝클어진 머리와 남색 반소매 옷을 입고 얼굴과 손에 멍이 있으며, 반바지 아래 드러난 다리에 상처가 보인다. 윤성찬은 회색 머리와 안경, 고령의 얼굴, 짙은 재킷과 청색 셔츠로 참조를 따른다. 탁자에는 사진 한 장 외에 검은 펜과 갈색 서류봉투가 추가되어 있다. 수갑은 없고 봉투의 흔적은 읽을 수 있는 글자로 판별되지 않는다.",
        "hard_violations": [
         "요청과 참조에 없는 검은 펜과 갈색 서류봉투를 중앙 탁자의 소품으로 추가했다."
        ],
        "physics": "두 사람 모두 의자 좌판에 앉아 있고 등받이가 몸 뒤에 있어 착석 구조는 자연스럽다. 현우의 양손과 팔은 상판에 지지되고 윤성찬의 보이는 손은 허벅지에 놓인다. 사진과 펜, 봉투는 상판 위에 놓여 있다. 발은 화면 밖이라 바닥 접촉을 확인할 수 없지만, 보이는 몸은 의자로 지지되며 공중에 떠 있는 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.375,
    "B": 1.429
   },
   "adjusted": {
    "A": 1.125,
    "B": 1.179
   },
   "violations": {
    "B": [
     "[gemini-pro] 배경 좌측의 의자와 테이블 다리가 물리적으로 불가능하게 융합됨"
    ],
    "A": [
     "[gpt-high] 요청과 참조에 없는 검은 펜과 갈색 서류봉투를 중앙 탁자의 소품으로 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1125,
   "B": 1179
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1125,
    "verdict_ko": "요구된 발끝 노출 프레이밍은 다소 잘렸으나, 수갑 해제 및 다리 상처 노출 지시를 완벽히 구현했고 이전 샷의 사진 소품까지 정확히 유지함.  ★위반: [gpt-high] 요청과 참조에 없는 검은 펜과 갈색 서류봉투를 중앙 탁자의 소품으로 추가했다."
   },
   {
    "label": "B",
    "score": 1179,
    "verdict_ko": "발이 보이는 프레이밍은 충족했으나, 제거해야 할 수갑이 여전히 남아있고 배경 가구의 물리적 구조가 무너져 실격됨.  ★위반: [gemini-pro] 배경 좌측의 의자와 테이블 다리가 물리적으로 불가능하게 융합됨"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S23sh6_sel.png",
    "asset_id": "fa401d32-ddd3-4003-b261-f13417ece72c",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 윤성찬: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1453233>",
    "asset_id": "04d34665-3829-49fb-a6d2-25b9fd051d63",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-1760-7b5e-9f25-517bff122732",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S23sh6"
  }
 },
 "S23sh11::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:05:36.030539+00:00",
  "fingerprint": "adfcb6d3db01370132f6c7d31f6c0e1986d138705912efd694e738552b4e68ae",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S23sh11_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S23sh11_sel.png",
  "source_sha256": "a2eda67b115a54f8dfbbd55a6560b6db976d46c23ffde9347c96a1965a81a4fc",
  "file": "S23sh11_cine.png",
  "staged_sha256": "5e033aa7b838c17ccc8da9b38c5ffb502093f7e3842df11f8d4f4dc0cb362ebe",
  "latency_ms": 12865
 },
 "S24sh4::signage": {
  "fp": "3089b4e860d5cdd8",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S24sh4": {
  "input_fingerprint": "fab2a1a49292432c",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 현우의 어깨를 양손으로 꽉 붙잡은 채 눈물을 글썽이며 환하게 웃는 페드로의 얼굴.\n\nLOCATION (lock): Inside the front living area of the refugee family's container home, with daytime light entering from the entrance. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 현우의 집 내부 (The reunion takes place inside the home); used as A soft, narrow background margin establishes the interior without inventing furnishings.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use unobtrusive ambient illumination appropriate to the interior, preserving the tears and smile without introducing an unsupported source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the container's fixed interior surfaces, household fixtures, and daylight as the location reference. Exclude the girl and robot, and do not restore an undisturbed arrangement of belongings after the search.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Container 7-31 retains the damage and disorder from the militia's search. 현우: He is back inside his container, no longer handcuffed. His facial bruises and injured leg persist. 페드로: He has hurried inside the container and is out of breath.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 현우의 어깨를 양손으로 꽉 붙잡은 채 눈물을 글썽이며 환하게 웃는 페드로의 얼굴.\n\nLOCATION (lock): Inside the front living area of the refugee family's container home, with daytime light entering from the entrance. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 현우의 집 내부 (The reunion takes place inside the home); used as A soft, narrow background margin establishes the interior without inventing furnishings.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use unobtrusive ambient illumination appropriate to the interior, preserving the tears and smile without introducing an unsupported source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the container's fixed interior surfaces, household fixtures, and daylight as the location reference. Exclude the girl and robot, and do not restore an undisturbed arrangement of belongings after the search.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Container 7-31 retains the damage and disorder from the militia's search. 현우: He is back inside his container, no longer handcuffed. His facial bruises and injured leg persist. 페드로: He has hurried inside the container and is out of breath.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 현우의 어깨를 양손으로 꽉 붙잡은 채 눈물을 글썽이며 환하게 웃는 페드로의 얼굴.\n\nLOCATION (lock): Inside the front living area of the refugee family's container home, with daytime light entering from the entrance. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 현우의 집 내부 (The reunion takes place inside the home); used as A soft, narrow background margin establishes the interior without inventing furnishings.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use unobtrusive ambient illumination appropriate to the interior, preserving the tears and smile without introducing an unsupported source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the container's fixed interior surfaces, household fixtures, and daylight as the location reference. Exclude the girl and robot, and do not restore an undisturbed arrangement of belongings after the search.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Container 7-31 retains the damage and disorder from the militia's search. 현우: He is back inside his container, no longer handcuffed. His facial bruises and injured leg persist. 페드로: He has hurried inside the container and is out of breath.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "페드로는 바로 앞의 현우를 바라보며 웃고 있음.",
    "built_space": "컨테이너 내부. 뒤쪽으로 열린 문과 가전제품(냉장고) 등 이전 샷에 없던 구조물들이 흐리게 배치됨.",
    "entities": "페드로의 얼굴과 눈물, 웃는 표정은 레퍼런스에 부합함. 현우 역시 멍든 뺨과 머리스타일이 일치함.",
    "hard_violations": [
     "[gpt-high] 가구를 새로 만들지 말라는 배경 지시에도 불구하고, 장소 참고에 근거 없는 흰 냉장고와 목재 상부 벽장을 추가해 실내를 주방 형태로 재구성했다."
    ],
    "physics": "현우의 양손이 페드로의 양어깨를 붙잡고 몸을 지탱하고 있으며, 페드로의 팔은 아래로 향해 있어 어떤 물체도 잡고 있지 않음."
   },
   {
    "label": "B",
    "direction": "페드로는 마주 선 현우의 얼굴을 향해 시선을 두고 있음.",
    "built_space": "컨테이너 내부 공간. 화면 좌측에 이전 샷과 일치하는 금속 앵글 선반과 벽면 스위치가 보이며 뒤쪽 창문으로 빛이 들어옴.",
    "entities": "페드로는 눈물을 글썽이며 웃고 있는 라틴계 외모로 레퍼런스와 일치함. 현우는 뺨에 멍 자국이 있고 헝클어진 머리를 한 뒷모습으로 일치함.",
    "hard_violations": [],
    "physics": "페드로의 양손이 현우의 양어깨를 안정적으로 감싸 쥐고 있으며, 현우의 왼손 역시 페드로의 오른쪽 어깨에 자연스럽게 얹혀 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "페드로가 양손으로 현우의 어깨를 꽉 잡고 눈물 흘리며 웃는 핵심 연출을 정확히 구현했으며, 이전 샷에 등장한 금속 선반 등의 배경 요소도 충실히 유지했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "페드로가 현우의 어깨를 잡아야 한다는 텍스트와 정반대로 현우가 페드로의 어깨를 잡고 있어 연출 지시를 크게 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "페드로는 마주 선 현우의 얼굴을 향해 시선을 두고 있음.",
        "built_space": "컨테이너 내부 공간. 화면 좌측에 이전 샷과 일치하는 금속 앵글 선반과 벽면 스위치가 보이며 뒤쪽 창문으로 빛이 들어옴.",
        "entities": "페드로는 눈물을 글썽이며 웃고 있는 라틴계 외모로 레퍼런스와 일치함. 현우는 뺨에 멍 자국이 있고 헝클어진 머리를 한 뒷모습으로 일치함.",
        "hard_violations": [],
        "physics": "페드로의 양손이 현우의 양어깨를 안정적으로 감싸 쥐고 있으며, 현우의 왼손 역시 페드로의 오른쪽 어깨에 자연스럽게 얹혀 있음."
       },
       {
        "label": "A",
        "direction": "페드로는 바로 앞의 현우를 바라보며 웃고 있음.",
        "built_space": "컨테이너 내부. 뒤쪽으로 열린 문과 가전제품(냉장고) 등 이전 샷에 없던 구조물들이 흐리게 배치됨.",
        "entities": "페드로의 얼굴과 눈물, 웃는 표정은 레퍼런스에 부합함. 현우 역시 멍든 뺨과 머리스타일이 일치함.",
        "hard_violations": [],
        "physics": "현우의 양손이 페드로의 양어깨를 붙잡고 몸을 지탱하고 있으며, 페드로의 팔은 아래로 향해 있어 어떤 물체도 잡고 있지 않음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "페드로가 양손으로 현우의 어깨를 꽉 잡고 눈물 흘리며 웃는 핵심 연출을 정확히 구현했으며, 이전 샷에 등장한 금속 선반 등의 배경 요소도 충실히 유지했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "페드로가 현우의 어깨를 잡아야 한다는 텍스트와 정반대로 현우가 페드로의 어깨를 잡고 있어 연출 지시를 크게 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "페드로는 마주 선 현우의 얼굴을 향해 시선을 두고 있음.",
        "built_space": "컨테이너 내부 공간. 화면 좌측에 이전 샷과 일치하는 금속 앵글 선반과 벽면 스위치가 보이며 뒤쪽 창문으로 빛이 들어옴.",
        "entities": "페드로는 눈물을 글썽이며 웃고 있는 라틴계 외모로 레퍼런스와 일치함. 현우는 뺨에 멍 자국이 있고 헝클어진 머리를 한 뒷모습으로 일치함.",
        "hard_violations": [],
        "physics": "페드로의 양손이 현우의 양어깨를 안정적으로 감싸 쥐고 있으며, 현우의 왼손 역시 페드로의 오른쪽 어깨에 자연스럽게 얹혀 있음."
       },
       {
        "label": "A",
        "direction": "페드로는 바로 앞의 현우를 바라보며 웃고 있음.",
        "built_space": "컨테이너 내부. 뒤쪽으로 열린 문과 가전제품(냉장고) 등 이전 샷에 없던 구조물들이 흐리게 배치됨.",
        "entities": "페드로의 얼굴과 눈물, 웃는 표정은 레퍼런스에 부합함. 현우 역시 멍든 뺨과 머리스타일이 일치함.",
        "hard_violations": [],
        "physics": "현우의 양손이 페드로의 양어깨를 붙잡고 몸을 지탱하고 있으며, 페드로의 팔은 아래로 향해 있어 어떤 물체도 잡고 있지 않음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "페드로가 현우의 어깨에 손을 얹고 눈물 어린 미소를 보이며 기존 실내도 비교적 유지하지만, 얼굴 클로즈업보다 넓고 현우의 맞잡는 동작이 강조된다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "페드로의 얼굴 크기와 표정은 정확하지만, 보이는 양손 동작은 현우가 페드로를 붙잡는 반대 관계이며 배경에 근거 없는 냉장고와 벽장을 추가했다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 페드로는 오른쪽 현우의 눈을 바라보며 웃고, 현우도 페드로를 향한다. 페드로의 보이는 팔 하나는 오른쪽으로 뻗어 현우의 가까운 어깨에 닿는다. 동시에 현우의 두 손이 페드로의 양어깨를 잡고 있어, 지시된 동작보다 서로 어깨를 붙잡는 장면으로 읽힌다.",
        "built_space": "낡은 밝은색 세로 패널 벽, 금속 선반 한 개, 뒤쪽 창 한 개, 왼쪽 배관과 전기함 한 개가 보인다. 선반과 벽의 재질은 이전 장면과 유사하지만 배경이 좁은 여백 이상으로 드러난다. 두 인물은 실내에서 마주 보며 고정 시설과 충돌하지 않는다. 반사는 없다.",
        "entities": "등장인물은 두 젊은 남성뿐이며 소녀와 로봇은 없다. 페드로는 짙은 머리, 참고 인물과 대체로 유사한 얼굴, 남색 상의, 젖은 눈과 뺨의 눈물, 환한 미소를 보인다. 현우는 검은 헝클어진 머리와 남색 상의, 뺨의 상처가 유지된다. 참고 이미지의 정확한 얼굴 일치는 완벽하지 않다. 다리와 수갑 여부는 구도 밖이며 읽을 수 있는 문자는 없다.",
        "hard_violations": [],
        "physics": "보이는 세 손은 어깨와 옷에 실제로 접촉하며, 팔의 진행 방향도 두 사람 사이의 거리에서 가능하다. 페드로의 다른 손은 가려져 양손 접촉을 직접 확인할 수 없지만, 이를 추가 손이나 부유한 손으로 볼 근거는 없다. 하체와 바닥은 잘려 있으며 공중에 떠 있는 몸이나 물체는 없다."
       },
       {
        "label": "B",
        "direction": "페드로의 시선은 오른쪽 현우의 얼굴에 정확히 향하고 현우도 그를 마주 본다. 그러나 화면에 드러난 양손은 현우의 남색 긴소매 팔에서 이어져 페드로의 양어깨를 잡는다. 페드로가 현우의 어깨를 붙잡는 접촉은 보이지 않아, 가시적인 행동의 주체와 대상이 지시와 반대다.",
        "built_space": "왼쪽에 낡은 출입구와 밝은 개구부 한 곳, 뒤쪽에 흰 냉장고 한 대, 목재 상부 벽장 한 묶음, 오른쪽 창 한 곳과 주방 작업대가 보인다. 기존 컨테이너의 좁고 부드러운 배경 여백 대신 새 주방 설비를 구체적으로 구성했다. 인물의 얼굴은 요구된 클로즈업 크기에 가깝고 두 사람의 공간적 위치 자체는 가능하다. 반사는 없다.",
        "entities": "두 젊은 남성만 등장한다. 페드로의 짙은 머리와 남색 상의, 눈물이 맺힌 눈, 치아를 드러낸 밝은 미소가 확인된다. 현우는 검은 머리와 남색 긴소매, 뺨의 멍을 유지한다. 두 얼굴은 참고 인물과 대체로 유사하나 정확히 같지는 않다. 소녀와 로봇, 읽을 수 있는 글자는 없다. 다리 부상은 프레임 밖이라 판단 대상이 아니다.",
        "hard_violations": [
         "가구를 새로 만들지 말라는 배경 지시에도 불구하고, 장소 참고에 근거 없는 흰 냉장고와 목재 상부 벽장을 추가해 실내를 주방 형태로 재구성했다."
        ],
        "physics": "현우의 두 손은 소매에서 자연스럽게 이어지고 페드로의 어깨에 접촉한다. 손과 팔의 자세는 물리적으로 가능하지만 지시된 행동의 주체가 다르다. 페드로의 손은 프레임 밖이라 접촉 여부를 확인할 수 없다. 하체가 잘린 정상적인 상반신 구도이며, 지지 없이 떠 있는 신체나 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "페드로가 현우의 어깨에 손을 얹고 눈물 어린 미소를 보이며 기존 실내도 비교적 유지하지만, 얼굴 클로즈업보다 넓고 현우의 맞잡는 동작이 강조된다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "페드로의 얼굴 크기와 표정은 정확하지만, 보이는 양손 동작은 현우가 페드로를 붙잡는 반대 관계이며 배경에 근거 없는 냉장고와 벽장을 추가했다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽 페드로는 오른쪽 현우의 눈을 바라보며 웃고, 현우도 페드로를 향한다. 페드로의 보이는 팔 하나는 오른쪽으로 뻗어 현우의 가까운 어깨에 닿는다. 동시에 현우의 두 손이 페드로의 양어깨를 잡고 있어, 지시된 동작보다 서로 어깨를 붙잡는 장면으로 읽힌다.",
        "built_space": "낡은 밝은색 세로 패널 벽, 금속 선반 한 개, 뒤쪽 창 한 개, 왼쪽 배관과 전기함 한 개가 보인다. 선반과 벽의 재질은 이전 장면과 유사하지만 배경이 좁은 여백 이상으로 드러난다. 두 인물은 실내에서 마주 보며 고정 시설과 충돌하지 않는다. 반사는 없다.",
        "entities": "등장인물은 두 젊은 남성뿐이며 소녀와 로봇은 없다. 페드로는 짙은 머리, 참고 인물과 대체로 유사한 얼굴, 남색 상의, 젖은 눈과 뺨의 눈물, 환한 미소를 보인다. 현우는 검은 헝클어진 머리와 남색 상의, 뺨의 상처가 유지된다. 참고 이미지의 정확한 얼굴 일치는 완벽하지 않다. 다리와 수갑 여부는 구도 밖이며 읽을 수 있는 문자는 없다.",
        "hard_violations": [],
        "physics": "보이는 세 손은 어깨와 옷에 실제로 접촉하며, 팔의 진행 방향도 두 사람 사이의 거리에서 가능하다. 페드로의 다른 손은 가려져 양손 접촉을 직접 확인할 수 없지만, 이를 추가 손이나 부유한 손으로 볼 근거는 없다. 하체와 바닥은 잘려 있으며 공중에 떠 있는 몸이나 물체는 없다."
       },
       {
        "label": "A",
        "direction": "페드로의 시선은 오른쪽 현우의 얼굴에 정확히 향하고 현우도 그를 마주 본다. 그러나 화면에 드러난 양손은 현우의 남색 긴소매 팔에서 이어져 페드로의 양어깨를 잡는다. 페드로가 현우의 어깨를 붙잡는 접촉은 보이지 않아, 가시적인 행동의 주체와 대상이 지시와 반대다.",
        "built_space": "왼쪽에 낡은 출입구와 밝은 개구부 한 곳, 뒤쪽에 흰 냉장고 한 대, 목재 상부 벽장 한 묶음, 오른쪽 창 한 곳과 주방 작업대가 보인다. 기존 컨테이너의 좁고 부드러운 배경 여백 대신 새 주방 설비를 구체적으로 구성했다. 인물의 얼굴은 요구된 클로즈업 크기에 가깝고 두 사람의 공간적 위치 자체는 가능하다. 반사는 없다.",
        "entities": "두 젊은 남성만 등장한다. 페드로의 짙은 머리와 남색 상의, 눈물이 맺힌 눈, 치아를 드러낸 밝은 미소가 확인된다. 현우는 검은 머리와 남색 긴소매, 뺨의 멍을 유지한다. 두 얼굴은 참고 인물과 대체로 유사하나 정확히 같지는 않다. 소녀와 로봇, 읽을 수 있는 글자는 없다. 다리 부상은 프레임 밖이라 판단 대상이 아니다.",
        "hard_violations": [
         "가구를 새로 만들지 말라는 배경 지시에도 불구하고, 장소 참고에 근거 없는 흰 냉장고와 목재 상부 벽장을 추가해 실내를 주방 형태로 재구성했다."
        ],
        "physics": "현우의 두 손은 소매에서 자연스럽게 이어지고 페드로의 어깨에 접촉한다. 손과 팔의 자세는 물리적으로 가능하지만 지시된 행동의 주체가 다르다. 페드로의 손은 프레임 밖이라 접촉 여부를 확인할 수 없다. 하체가 잘린 정상적인 상반신 구도이며, 지지 없이 떠 있는 신체나 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.857,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.607,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gpt-high] 가구를 새로 만들지 말라는 배경 지시에도 불구하고, 장소 참고에 근거 없는 흰 냉장고와 목재 상부 벽장을 추가해 실내를 주방 형태로 재구성했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 607
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "페드로가 양손으로 현우의 어깨를 꽉 잡고 눈물 흘리며 웃는 핵심 연출을 정확히 구현했으며, 이전 샷에 등장한 금속 선반 등의 배경 요소도 충실히 유지했습니다."
   },
   {
    "label": "A",
    "score": 607,
    "verdict_ko": "페드로가 현우의 어깨를 잡아야 한다는 텍스트와 정반대로 현우가 페드로의 어깨를 잡고 있어 연출 지시를 크게 위반했습니다.  ★위반: [gpt-high] 가구를 새로 만들지 말라는 배경 지시에도 불구하고, 장소 참고에 근거 없는 흰 냉장고와 목재 상부 벽장을 추가해 실내를 주방 형태로 재구성했다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S14sh9_sel.png",
    "asset_id": "a924d04d-33b1-424c-9a87-80041aaffe5e",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 페드로: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1278830>",
    "asset_id": "b09df655-64d4-4db4-a1b3-2f0bb5d29c95",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-1930-7fcf-bf82-e21966da0576",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S14sh9"
  }
 },
 "S24sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:10:36.761661+00:00",
  "fingerprint": "ec388715cc235ff1ab14850709711d21ecc0870c46420ac1265aca5f1f1b2a1a",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S24sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S24sh4_sel.png",
  "source_sha256": "8006ce7ebeb266f4da7d46abd373a70385837f1bd96063e3d92c3fc9be274f70",
  "file": "S24sh4_cine.png",
  "staged_sha256": "a489189622f281f22db3fda39f35e3e502976102325a0cc158878c00340c97ac",
  "latency_ms": 9892
 },
 "S24sh7::signage": {
  "fp": "a8103b87dcc7eeb2",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S24sh7": {
  "input_fingerprint": "e950c06f3eb6ca7e",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 페드로의 말에 충격을 받은 듯 눈이 커진 채 얼어붙은 현우의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): In the container home's shared interior near the entrance, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 현우의 집 내부 (Visible only as a narrow background around 현우); used as Maintains the location while leaving the reaction unobstructed.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the reunion's restrained ambient lighting and controlled contrast, keeping the widened eyes clearly readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the container's fixed surfaces, household fixtures, and daytime light. Exclude the girl and robot from the reference, and do not undo the disruption caused by the search.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The searched container remains damaged and disordered; its number is 7-31. 현우: He remains inside the container with free wrists, facial bruising and an unhealed leg injury.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 페드로의 말에 충격을 받은 듯 눈이 커진 채 얼어붙은 현우의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): In the container home's shared interior near the entrance, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 현우의 집 내부 (Visible only as a narrow background around 현우); used as Maintains the location while leaving the reaction unobstructed.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the reunion's restrained ambient lighting and controlled contrast, keeping the widened eyes clearly readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the container's fixed surfaces, household fixtures, and daytime light. Exclude the girl and robot from the reference, and do not undo the disruption caused by the search.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The searched container remains damaged and disordered; its number is 7-31. 현우: He remains inside the container with free wrists, facial bruising and an unhealed leg injury.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 페드로의 말에 충격을 받은 듯 눈이 커진 채 얼어붙은 현우의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): In the container home's shared interior near the entrance, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 현우의 집 내부 (Visible only as a narrow background around 현우); used as Maintains the location while leaving the reaction unobstructed.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the reunion's restrained ambient lighting and controlled contrast, keeping the widened eyes clearly readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the container's fixed surfaces, household fixtures, and daytime light. Exclude the girl and robot from the reference, and do not undo the disruption caused by the search.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The searched container remains damaged and disordered; its number is 7-31. 현우: He remains inside the container with free wrists, facial bruising and an unhealed leg injury.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "인물의 시선이 정면 카메라 렌즈를 똑바로 향하고 있습니다. 대화 상대방(페드로)에 반응해야 하는 상황임에도 렌즈를 쳐다보고 있어 다소 부자연스럽습니다.",
    "built_space": "컨테이너 내부 배경이 레퍼런스와 일치합니다. 인물 왼편의 스위치와 뒤쪽의 금속 선반이 올바르게 배치되어 있습니다.",
    "entities": "현우의 특징인 헝클어진 검은 머리, 파란색 티셔츠, 얼굴의 타박상이 캐릭터 레퍼런스 및 텍스트와 정확히 일치합니다.",
    "hard_violations": [],
    "physics": "인물이 올곧게 서 있는 자세를 유지하고 있으며, 옷의 주름이나 신체 구조에 물리적인 오류가 없습니다."
   },
   {
    "label": "B",
    "direction": "인물의 시선이 카메라 왼쪽 화면 밖을 향하고 있습니다. 이는 오프스크린에 있는 페드로를 바라보며 충격받은 상황을 연출하기에 매우 적절합니다.",
    "built_space": "레퍼런스에 등장했던 왼쪽 벽의 스위치, 금속 선반, 그리고 오른쪽의 창문 채광 등 컨테이너 내부 구조가 정확하게 구현되었습니다.",
    "entities": "현우의 앳된 얼굴, 헝클어진 머리, 얼굴 뺨의 멍 자국 등이 지정된 캐릭터 설정과 잘 맞습니다.",
    "hard_violations": [],
    "physics": "자연스럽게 어깨와 목의 무게 중심이 잡혀 있으며, 허공에 떠 있거나 지탱되지 않은 사물 없이 물리적으로 안정적입니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 10,
        "verdict_ko": "프레임 밖을 향한 자연스러운 시선 처리와 충격받아 눈이 커진 굳은 표정 연기가 프롬프트의 영화적 리액션 상황을 완벽하게 묘사합니다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "표정과 배경의 디테일은 훌륭하나, 카메라 렌즈를 정면으로 응시하고 있어 자연스러운 컷보다는 프로필 사진처럼 보인다는 점이 아쉽습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "인물의 시선이 정면 카메라 렌즈를 똑바로 향하고 있습니다. 대화 상대방(페드로)에 반응해야 하는 상황임에도 렌즈를 쳐다보고 있어 다소 부자연스럽습니다.",
        "built_space": "컨테이너 내부 배경이 레퍼런스와 일치합니다. 인물 왼편의 스위치와 뒤쪽의 금속 선반이 올바르게 배치되어 있습니다.",
        "entities": "현우의 특징인 헝클어진 검은 머리, 파란색 티셔츠, 얼굴의 타박상이 캐릭터 레퍼런스 및 텍스트와 정확히 일치합니다.",
        "hard_violations": [],
        "physics": "인물이 올곧게 서 있는 자세를 유지하고 있으며, 옷의 주름이나 신체 구조에 물리적인 오류가 없습니다."
       },
       {
        "label": "B",
        "direction": "인물의 시선이 카메라 왼쪽 화면 밖을 향하고 있습니다. 이는 오프스크린에 있는 페드로를 바라보며 충격받은 상황을 연출하기에 매우 적절합니다.",
        "built_space": "레퍼런스에 등장했던 왼쪽 벽의 스위치, 금속 선반, 그리고 오른쪽의 창문 채광 등 컨테이너 내부 구조가 정확하게 구현되었습니다.",
        "entities": "현우의 앳된 얼굴, 헝클어진 머리, 얼굴 뺨의 멍 자국 등이 지정된 캐릭터 설정과 잘 맞습니다.",
        "hard_violations": [],
        "physics": "자연스럽게 어깨와 목의 무게 중심이 잡혀 있으며, 허공에 떠 있거나 지탱되지 않은 사물 없이 물리적으로 안정적입니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 10,
        "verdict_ko": "프레임 밖을 향한 자연스러운 시선 처리와 충격받아 눈이 커진 굳은 표정 연기가 프롬프트의 영화적 리액션 상황을 완벽하게 묘사합니다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "표정과 배경의 디테일은 훌륭하나, 카메라 렌즈를 정면으로 응시하고 있어 자연스러운 컷보다는 프로필 사진처럼 보인다는 점이 아쉽습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "인물의 시선이 정면 카메라 렌즈를 똑바로 향하고 있습니다. 대화 상대방(페드로)에 반응해야 하는 상황임에도 렌즈를 쳐다보고 있어 다소 부자연스럽습니다.",
        "built_space": "컨테이너 내부 배경이 레퍼런스와 일치합니다. 인물 왼편의 스위치와 뒤쪽의 금속 선반이 올바르게 배치되어 있습니다.",
        "entities": "현우의 특징인 헝클어진 검은 머리, 파란색 티셔츠, 얼굴의 타박상이 캐릭터 레퍼런스 및 텍스트와 정확히 일치합니다.",
        "hard_violations": [],
        "physics": "인물이 올곧게 서 있는 자세를 유지하고 있으며, 옷의 주름이나 신체 구조에 물리적인 오류가 없습니다."
       },
       {
        "label": "B",
        "direction": "인물의 시선이 카메라 왼쪽 화면 밖을 향하고 있습니다. 이는 오프스크린에 있는 페드로를 바라보며 충격받은 상황을 연출하기에 매우 적절합니다.",
        "built_space": "레퍼런스에 등장했던 왼쪽 벽의 스위치, 금속 선반, 그리고 오른쪽의 창문 채광 등 컨테이너 내부 구조가 정확하게 구현되었습니다.",
        "entities": "현우의 앳된 얼굴, 헝클어진 머리, 얼굴 뺨의 멍 자국 등이 지정된 캐릭터 설정과 잘 맞습니다.",
        "hard_violations": [],
        "physics": "자연스럽게 어깨와 목의 무게 중심이 잡혀 있으며, 허공에 떠 있거나 지탱되지 않은 사물 없이 물리적으로 안정적입니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "화면 밖 상대에게 향한 시선과 커진 눈, 굳은 얼굴의 클로즈업이 페드로의 말을 듣고 얼어붙은 순간을 더 충실하게 구현한다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "인물과 장소, 놀란 표정은 부합하지만 정면 렌즈 응시와 대칭적인 자세가 대화 중 반응보다 인물 참고 사진의 초상 구도에 가깝다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴과 두 눈이 화면 왼쪽의 카메라 밖 상대를 향한다. 대상 자체는 보이지 않지만 페드로의 말을 듣는 시선으로 자연스럽게 연결된다. 눈꺼풀을 들어 눈이 커졌고 입이 조금 벌어져 있으며, 렌즈를 정면으로 응시하지 않는다. 무기나 방향을 확인할 휴대 물체는 없다.",
        "built_space": "뒤쪽 왼편에 금속 선반 한 개와 전선관에 연결된 흰 전기함 한 개, 오른편에 밝은 창 한 개, 맨 왼쪽에 출입구로 보이는 밝은 틈이 있다. 낡은 밝은색 세로 벽면과 생활용품이 놓인 선반은 이전 장면의 재질과 설비를 이어 간다. 현우는 설비 앞쪽 실내에 있으며 얼굴 중심의 클로즈업이다. 중복 설비나 불가능한 반사는 보이지 않는다. 바닥의 수색 흔적은 구도 밖이다.",
        "entities": "현우로 보이는 앳된 동아시아계 남성 한 명만 있다. 한국계 미국인이라는 국적 배경은 외모만으로 확인할 수 없지만, 얼굴과 헝클어진 검은 머리, 남색 둥근 목 티셔츠는 인물 참고와 부합한다. 뺨에 붉은 상처와 멍이 보인다. 다른 사람이나 로봇은 없고 읽을 수 있는 글자도 없다. 손목과 다리는 프레임 밖이므로 결박 여부와 다리 부상은 확인할 수 없다.",
        "hard_violations": [],
        "physics": "머리가 목과 어깨에 정상적으로 연결되어 있고 상체는 화면 아래로 이어진다. 발과 바닥 접촉은 클로즈업 밖이지만 공중에 떠 있다는 징후는 없다. 정지한 얼굴과 약간 돌아간 상체는 충격으로 동작을 멈춘 자세로 가능하다. 선반 위 물건들은 선반에 놓여 있으며 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "얼굴이 거의 완전한 정면이고 두 눈도 렌즈 쪽을 응시한다. 페드로가 카메라 가까이에 있다고 해석할 여지는 있지만, 화면에서 확인되는 시선 대상은 렌즈에 가깝다. 눈은 커져 있고 입은 조금 벌어져 있으나 대화 상대를 향한 반응보다 정면 초상처럼 읽힌다. 무기나 휴대 물체는 없다.",
        "built_space": "인물 뒤에 금속 선반 한 개, 왼편에 전선관과 흰 전기함 한 개, 오른쪽 가장자리에 밝은 창 한 개, 왼쪽 가장자리에 출입구로 보이는 밝은 틈이 있다. 낡은 밝은 벽과 낮의 확산광은 이전 장면과 대체로 일치한다. 현우는 선반 앞에 있으며 얼굴과 목, 어깨 일부만 잡힌다. 설비 중복이나 불가능한 반사는 없다. 바닥과 실내 전체의 어질러진 상태는 보이지 않는다.",
        "entities": "앳된 동아시아계 남성 한 명으로, 참고 속 현우의 얼굴과 검은 머리, 남색 티셔츠를 잘 유지한다. 한국계 미국인이라는 배경 자체는 시각적으로 확정할 수 없다. 화면 오른쪽 뺨의 붉은 멍과 찰과상이 뚜렷하다. 추가 인물, 로봇, 읽을 수 있는 문구는 없다. 손목과 다리의 상태는 올바른 클로즈업 범위 밖이라 평가할 수 없다.",
        "hard_violations": [],
        "physics": "머리와 목, 양어깨의 연결은 자연스럽고 상체는 화면 아래로 이어진다. 하체 지지는 보이지 않지만 부유하거나 넘어지는 모습은 아니다. 정면으로 굳은 자세는 신체적으로 가능하며, 다만 대칭성이 강해 멈춘 동작보다는 준비된 초상 자세에 가깝다. 배경 물건은 선반에 지지되어 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "화면 밖 상대에게 향한 시선과 커진 눈, 굳은 얼굴의 클로즈업이 페드로의 말을 듣고 얼어붙은 순간을 더 충실하게 구현한다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "인물과 장소, 놀란 표정은 부합하지만 정면 렌즈 응시와 대칭적인 자세가 대화 중 반응보다 인물 참고 사진의 초상 구도에 가깝다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴과 두 눈이 화면 왼쪽의 카메라 밖 상대를 향한다. 대상 자체는 보이지 않지만 페드로의 말을 듣는 시선으로 자연스럽게 연결된다. 눈꺼풀을 들어 눈이 커졌고 입이 조금 벌어져 있으며, 렌즈를 정면으로 응시하지 않는다. 무기나 방향을 확인할 휴대 물체는 없다.",
        "built_space": "뒤쪽 왼편에 금속 선반 한 개와 전선관에 연결된 흰 전기함 한 개, 오른편에 밝은 창 한 개, 맨 왼쪽에 출입구로 보이는 밝은 틈이 있다. 낡은 밝은색 세로 벽면과 생활용품이 놓인 선반은 이전 장면의 재질과 설비를 이어 간다. 현우는 설비 앞쪽 실내에 있으며 얼굴 중심의 클로즈업이다. 중복 설비나 불가능한 반사는 보이지 않는다. 바닥의 수색 흔적은 구도 밖이다.",
        "entities": "현우로 보이는 앳된 동아시아계 남성 한 명만 있다. 한국계 미국인이라는 국적 배경은 외모만으로 확인할 수 없지만, 얼굴과 헝클어진 검은 머리, 남색 둥근 목 티셔츠는 인물 참고와 부합한다. 뺨에 붉은 상처와 멍이 보인다. 다른 사람이나 로봇은 없고 읽을 수 있는 글자도 없다. 손목과 다리는 프레임 밖이므로 결박 여부와 다리 부상은 확인할 수 없다.",
        "hard_violations": [],
        "physics": "머리가 목과 어깨에 정상적으로 연결되어 있고 상체는 화면 아래로 이어진다. 발과 바닥 접촉은 클로즈업 밖이지만 공중에 떠 있다는 징후는 없다. 정지한 얼굴과 약간 돌아간 상체는 충격으로 동작을 멈춘 자세로 가능하다. 선반 위 물건들은 선반에 놓여 있으며 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "얼굴이 거의 완전한 정면이고 두 눈도 렌즈 쪽을 응시한다. 페드로가 카메라 가까이에 있다고 해석할 여지는 있지만, 화면에서 확인되는 시선 대상은 렌즈에 가깝다. 눈은 커져 있고 입은 조금 벌어져 있으나 대화 상대를 향한 반응보다 정면 초상처럼 읽힌다. 무기나 휴대 물체는 없다.",
        "built_space": "인물 뒤에 금속 선반 한 개, 왼편에 전선관과 흰 전기함 한 개, 오른쪽 가장자리에 밝은 창 한 개, 왼쪽 가장자리에 출입구로 보이는 밝은 틈이 있다. 낡은 밝은 벽과 낮의 확산광은 이전 장면과 대체로 일치한다. 현우는 선반 앞에 있으며 얼굴과 목, 어깨 일부만 잡힌다. 설비 중복이나 불가능한 반사는 없다. 바닥과 실내 전체의 어질러진 상태는 보이지 않는다.",
        "entities": "앳된 동아시아계 남성 한 명으로, 참고 속 현우의 얼굴과 검은 머리, 남색 티셔츠를 잘 유지한다. 한국계 미국인이라는 배경 자체는 시각적으로 확정할 수 없다. 화면 오른쪽 뺨의 붉은 멍과 찰과상이 뚜렷하다. 추가 인물, 로봇, 읽을 수 있는 문구는 없다. 손목과 다리의 상태는 올바른 클로즈업 범위 밖이라 평가할 수 없다.",
        "hard_violations": [],
        "physics": "머리와 목, 양어깨의 연결은 자연스럽고 상체는 화면 아래로 이어진다. 하체 지지는 보이지 않지만 부유하거나 넘어지는 모습은 아니다. 정면으로 굳은 자세는 신체적으로 가능하며, 다만 대칭성이 강해 멈춘 동작보다는 준비된 초상 자세에 가깝다. 배경 물건은 선반에 지지되어 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.478,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.478,
    "B": 2.0
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1478
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "프레임 밖을 향한 자연스러운 시선 처리와 충격받아 눈이 커진 굳은 표정 연기가 프롬프트의 영화적 리액션 상황을 완벽하게 묘사합니다."
   },
   {
    "label": "A",
    "score": 1478,
    "verdict_ko": "표정과 배경의 디테일은 훌륭하나, 카메라 렌즈를 정면으로 응시하고 있어 자연스러운 컷보다는 프로필 사진처럼 보인다는 점이 아쉽습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S24sh4_sel.png",
    "asset_id": "27614b90-8a30-41cd-b152-71bbcf67885e",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-1b04-7849-82dc-35669831a351",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S24sh4"
  }
 },
 "S24sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:11:35.024532+00:00",
  "fingerprint": "847def52886133ab343696d164deb9ac771a9521b848118f9f4479fb939a7905",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S24sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S24sh7_sel.png",
  "source_sha256": "cd4b16fbf3345a549d274bcac7f6721a6c58de4ddad8d59fd75652e59e9c12db",
  "file": "S24sh7_cine.png",
  "staged_sha256": "89238e0143698795009f3706cfd5362cc3f0afd42031d548cd0fc325c09dfc7d",
  "latency_ms": 9605
 },
 "S24sh10::signage": {
  "fp": "4608fb3a9d635532",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S24sh10": {
  "input_fingerprint": "ce51422058014a26",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 컨테이너 밖을 향해 뒷발로 지면을 강하게 밀어내며 질주하는 mid-stride 자세의 현우 전신 풀샷.\n\nLOCATION (lock): At the front doorway of the container home, opening onto the refugee settlement lane as the youth rushes outside. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Home entrance behind the left-to-right running figure in the middle-left of the frame, background; Clear continuation of the exit route in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: 컨테이너 집 출입구 (Open for 현우's exit) — Viewed obliquely behind him on the left, with the passage into the home visible; used as Makes the direction away from the home unambiguous without overwhelming the figure; 집 밖 지면 (Visible beneath the running figure); used as Shows the rear foot's contact and provides space for the next stride.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use ambient daytime illumination appropriate to the exterior with restrained tonal contrast and no invented directional source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The damage and disorder inside container 7-31 remain uncleared. 현우: He is leaving the container at a run, with his wrists free. His facial bruising and dog-bitten leg remain unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 컨테이너 밖을 향해 뒷발로 지면을 강하게 밀어내며 질주하는 mid-stride 자세의 현우 전신 풀샷.\n\nLOCATION (lock): At the front doorway of the container home, opening onto the refugee settlement lane as the youth rushes outside. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Home entrance behind the left-to-right running figure in the middle-left of the frame, background; Clear continuation of the exit route in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: 컨테이너 집 출입구 (Open for 현우's exit) — Viewed obliquely behind him on the left, with the passage into the home visible; used as Makes the direction away from the home unambiguous without overwhelming the figure; 집 밖 지면 (Visible beneath the running figure); used as Shows the rear foot's contact and provides space for the next stride.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use ambient daytime illumination appropriate to the exterior with restrained tonal contrast and no invented directional source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The damage and disorder inside container 7-31 remain uncleared. 현우: He is leaving the container at a run, with his wrists free. His facial bruising and dog-bitten leg remain unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 컨테이너 밖을 향해 뒷발로 지면을 강하게 밀어내며 질주하는 mid-stride 자세의 현우 전신 풀샷.\n\nLOCATION (lock): At the front doorway of the container home, opening onto the refugee settlement lane as the youth rushes outside. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Home entrance behind the left-to-right running figure in the middle-left of the frame, background; Clear continuation of the exit route in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: 컨테이너 집 출입구 (Open for 현우's exit) — Viewed obliquely behind him on the left, with the passage into the home visible; used as Makes the direction away from the home unambiguous without overwhelming the figure; 집 밖 지면 (Visible beneath the running figure); used as Shows the rear foot's contact and provides space for the next stride.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use ambient daytime illumination appropriate to the exterior with restrained tonal contrast and no invented directional source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The damage and disorder inside container 7-31 remain uncleared. 현우: He is leaving the container at a run, with his wrists free. His facial bruising and dog-bitten leg remain unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "화면 오른쪽을 향해 질주하며 시선도 같은 방향을 향함.",
    "built_space": "왼쪽에 열린 컨테이너 문이 있고 오른쪽으로 난민촌 통로가 길게 이어짐.",
    "entities": "현우. 기준 이미지와 동일한 남색 반팔 티셔츠와 바지를 착용했으며 얼굴에 타박상이 보임.",
    "hard_violations": [],
    "physics": "앞발(오른발)이 지면에 닿아 몸을 지지하며 먼지를 일으키고, 뒷발은 공중에 떠 있음."
   },
   {
    "label": "B",
    "direction": "화면 오른쪽을 향해 질주하며 시선도 앞으로 향함.",
    "built_space": "왼쪽의 컨테이너 출입구와 오른쪽의 통로가 정상적으로 배치됨.",
    "entities": "현우. 남색 티셔츠 위에 기준 이미지에 없는 회색 셔츠를 덧입고 있음.",
    "hard_violations": [
     "[gemini-pro] 이전 장면과 기준 이미지에 없는 회색 겉옷(invented object) 착용",
     "[gemini-pro] 지면에 닿은 발 없이 몸이 공중에 완전히 떠 있음 (unsupported body)"
    ],
    "physics": "두 발이 모두 허공에 떠 있어 몸을 물리적으로 지지하는 곳이 없음. 뒷발 아래에 먼지가 일지만 지면과 닿아있지 않음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "의상과 배경, 인물의 외형은 기준과 잘 일치하나, 프롬프트가 요구한 '뒷발 접촉' 대신 앞발이 지면에 닿아 동작 연출에서 감점이 있습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "이전 장면에 없는 겉옷을 임의로 추가하는 심각한 의상 오류가 발생했으며, 두 발이 모두 공중에 떠 있어 동작 지시를 어겼습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "화면 오른쪽을 향해 질주하며 시선도 같은 방향을 향함.",
        "built_space": "왼쪽에 열린 컨테이너 문이 있고 오른쪽으로 난민촌 통로가 길게 이어짐.",
        "entities": "현우. 기준 이미지와 동일한 남색 반팔 티셔츠와 바지를 착용했으며 얼굴에 타박상이 보임.",
        "hard_violations": [],
        "physics": "앞발(오른발)이 지면에 닿아 몸을 지지하며 먼지를 일으키고, 뒷발은 공중에 떠 있음."
       },
       {
        "label": "B",
        "direction": "화면 오른쪽을 향해 질주하며 시선도 앞으로 향함.",
        "built_space": "왼쪽의 컨테이너 출입구와 오른쪽의 통로가 정상적으로 배치됨.",
        "entities": "현우. 남색 티셔츠 위에 기준 이미지에 없는 회색 셔츠를 덧입고 있음.",
        "hard_violations": [
         "이전 장면과 기준 이미지에 없는 회색 겉옷(invented object) 착용",
         "지면에 닿은 발 없이 몸이 공중에 완전히 떠 있음 (unsupported body)"
        ],
        "physics": "두 발이 모두 허공에 떠 있어 몸을 물리적으로 지지하는 곳이 없음. 뒷발 아래에 먼지가 일지만 지면과 닿아있지 않음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "의상과 배경, 인물의 외형은 기준과 잘 일치하나, 프롬프트가 요구한 '뒷발 접촉' 대신 앞발이 지면에 닿아 동작 연출에서 감점이 있습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "이전 장면에 없는 겉옷을 임의로 추가하는 심각한 의상 오류가 발생했으며, 두 발이 모두 공중에 떠 있어 동작 지시를 어겼습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "화면 오른쪽을 향해 질주하며 시선도 같은 방향을 향함.",
        "built_space": "왼쪽에 열린 컨테이너 문이 있고 오른쪽으로 난민촌 통로가 길게 이어짐.",
        "entities": "현우. 기준 이미지와 동일한 남색 반팔 티셔츠와 바지를 착용했으며 얼굴에 타박상이 보임.",
        "hard_violations": [],
        "physics": "앞발(오른발)이 지면에 닿아 몸을 지지하며 먼지를 일으키고, 뒷발은 공중에 떠 있음."
       },
       {
        "label": "B",
        "direction": "화면 오른쪽을 향해 질주하며 시선도 앞으로 향함.",
        "built_space": "왼쪽의 컨테이너 출입구와 오른쪽의 통로가 정상적으로 배치됨.",
        "entities": "현우. 남색 티셔츠 위에 기준 이미지에 없는 회색 셔츠를 덧입고 있음.",
        "hard_violations": [
         "이전 장면과 기준 이미지에 없는 회색 겉옷(invented object) 착용",
         "지면에 닿은 발 없이 몸이 공중에 완전히 떠 있음 (unsupported body)"
        ],
        "physics": "두 발이 모두 허공에 떠 있어 몸을 물리적으로 지지하는 곳이 없음. 뒷발 아래에 먼지가 일지만 지면과 닿아있지 않음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "출입구를 등지고 왼쪽에서 오른쪽으로 질주하는 전신 구도는 더 정확하지만, 뒷발이 이미 떠 있어 지면을 밀어내는 지정 순간과 다르고 참고에 없는 겉옷을 입었다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "현우의 얼굴·반팔 차림과 발의 지면 접촉은 잘 보이지만, 오른쪽 탈출로가 아닌 화면 앞쪽 왼편으로 달려 지정된 진행 방향을 어겼다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴과 시선은 화면 오른쪽의 빈 골목을 향한다. 몸통의 전경사와 앞으로 뻗은 다리도 오른쪽 이동을 나타내며, 왼쪽 뒤의 열린 집에서 멀어지는 방향이 명확하다. 겨누는 물건이나 휴대 도구는 없다.",
        "built_space": "왼쪽에 열린 출입구 하나와 바깥으로 젖힌 문짝 하나가 있으며, 문 안쪽의 어수선한 물건과 바닥이 보인다. 현우는 문턱 바깥에 있고 오른쪽에는 다음 보폭을 위한 지면과 컨테이너 사이 통로가 이어진다. 녹슨 청회색 골강판과 먼 골목의 가로 구조물은 장소 참고와 유사하다. 참고의 철망 구간은 뚜렷하게 확인되지 않는다. 전신 와이드숏이지만 인물은 지정된 중간 왼쪽보다 중앙에 가깝다.",
        "entities": "젊은 동아시아계 남성 한 명만 보이며, 헝클어진 검은 머리와 앳된 얼굴은 현우의 참고와 대체로 부합한다. 국적은 외관으로 확인할 수 없다. 볼의 멍, 남색 티셔츠, 자유로운 양손목이 보인다. 다만 참고에 없는 회색 긴팔 겉옷을 추가했다. 긴 바지와 부츠가 다리를 덮어 개에게 물린 상처의 유지 여부는 판별할 수 없다. 읽을 수 있는 글자나 다른 사람은 없다.",
        "hard_violations": [],
        "physics": "두 발 모두 지면에서 떨어져 있다. 뒤쪽 문턱 부근의 먼지, 뒤로 접힌 다리, 앞으로 내민 착지 다리는 막 박차고 나온 달리기의 공중 국면으로 해석할 수 있으며 앞발 아래에 착지할 지면도 있다. 따라서 근거 없는 공중 부유는 아니다. 그러나 실제 뒷발의 지면 접촉은 보이지 않아, 뒷발로 땅을 강하게 밀어내는 바로 그 순간이라는 요구는 충족하지 못한다."
       },
       {
        "label": "B",
        "direction": "얼굴과 시선은 화면 왼쪽 앞을 향하고 몸도 카메라 쪽으로 비스듬히 기울어 있다. 출입구 밖으로 나오는 동작은 읽히지만, 화면 오른쪽의 빈 통로를 향한 왼쪽에서 오른쪽 질주는 아니다. 손에는 겨누거나 사용하는 물건이 없다.",
        "built_space": "왼쪽 출입구 하나와 열린 문짝 하나, 안쪽의 쌓인 물건이 보인다. 현우는 문턱 밖 골목 바닥에 있으며 오른쪽 통로는 비어 있다. 녹슨 컨테이너 벽과 먼 가로 구조물은 참고 장소와 잘 이어지지만 철망 구간은 확인하기 어렵다. 전신과 지면을 포함한 와이드숏이나 인물은 중앙에 있고, 실제 동작 축은 오른쪽 탈출 공간을 활용하지 않는다.",
        "entities": "젊은 동아시아계 남성 한 명으로, 검은 헝클어진 머리와 얼굴 생김새가 현우 참고에 가깝다. 국적 자체는 시각적으로 검증할 수 없다. 남색 반팔 티셔츠는 참고 의상과 일치하며 얼굴의 멍과 구속되지 않은 손목도 보인다. 긴 바지와 부츠 때문에 개에게 물린 다리 상처는 확인할 수 없다. 다른 사람이나 읽을 수 있는 문자는 없다.",
        "hard_violations": [],
        "physics": "화면 아래쪽으로 내민 발의 밑창이 흙바닥에 닿아 몸을 지지하고 접촉 부근에 먼지가 일어난다. 반대쪽 다리는 뒤에서 접혀 들려 있다. 기울어진 몸과 팔 동작은 방향을 틀며 달리는 자세로 물리적으로 가능하다. 다만 접지한 것은 앞쪽으로 내민 발이고 뒷발은 떠 있어, 요구된 뒷발의 강한 지면 밀기를 보여주지는 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "출입구를 등지고 왼쪽에서 오른쪽으로 질주하는 전신 구도는 더 정확하지만, 뒷발이 이미 떠 있어 지면을 밀어내는 지정 순간과 다르고 참고에 없는 겉옷을 입었다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "현우의 얼굴·반팔 차림과 발의 지면 접촉은 잘 보이지만, 오른쪽 탈출로가 아닌 화면 앞쪽 왼편으로 달려 지정된 진행 방향을 어겼다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴과 시선은 화면 오른쪽의 빈 골목을 향한다. 몸통의 전경사와 앞으로 뻗은 다리도 오른쪽 이동을 나타내며, 왼쪽 뒤의 열린 집에서 멀어지는 방향이 명확하다. 겨누는 물건이나 휴대 도구는 없다.",
        "built_space": "왼쪽에 열린 출입구 하나와 바깥으로 젖힌 문짝 하나가 있으며, 문 안쪽의 어수선한 물건과 바닥이 보인다. 현우는 문턱 바깥에 있고 오른쪽에는 다음 보폭을 위한 지면과 컨테이너 사이 통로가 이어진다. 녹슨 청회색 골강판과 먼 골목의 가로 구조물은 장소 참고와 유사하다. 참고의 철망 구간은 뚜렷하게 확인되지 않는다. 전신 와이드숏이지만 인물은 지정된 중간 왼쪽보다 중앙에 가깝다.",
        "entities": "젊은 동아시아계 남성 한 명만 보이며, 헝클어진 검은 머리와 앳된 얼굴은 현우의 참고와 대체로 부합한다. 국적은 외관으로 확인할 수 없다. 볼의 멍, 남색 티셔츠, 자유로운 양손목이 보인다. 다만 참고에 없는 회색 긴팔 겉옷을 추가했다. 긴 바지와 부츠가 다리를 덮어 개에게 물린 상처의 유지 여부는 판별할 수 없다. 읽을 수 있는 글자나 다른 사람은 없다.",
        "hard_violations": [],
        "physics": "두 발 모두 지면에서 떨어져 있다. 뒤쪽 문턱 부근의 먼지, 뒤로 접힌 다리, 앞으로 내민 착지 다리는 막 박차고 나온 달리기의 공중 국면으로 해석할 수 있으며 앞발 아래에 착지할 지면도 있다. 따라서 근거 없는 공중 부유는 아니다. 그러나 실제 뒷발의 지면 접촉은 보이지 않아, 뒷발로 땅을 강하게 밀어내는 바로 그 순간이라는 요구는 충족하지 못한다."
       },
       {
        "label": "A",
        "direction": "얼굴과 시선은 화면 왼쪽 앞을 향하고 몸도 카메라 쪽으로 비스듬히 기울어 있다. 출입구 밖으로 나오는 동작은 읽히지만, 화면 오른쪽의 빈 통로를 향한 왼쪽에서 오른쪽 질주는 아니다. 손에는 겨누거나 사용하는 물건이 없다.",
        "built_space": "왼쪽 출입구 하나와 열린 문짝 하나, 안쪽의 쌓인 물건이 보인다. 현우는 문턱 밖 골목 바닥에 있으며 오른쪽 통로는 비어 있다. 녹슨 컨테이너 벽과 먼 가로 구조물은 참고 장소와 잘 이어지지만 철망 구간은 확인하기 어렵다. 전신과 지면을 포함한 와이드숏이나 인물은 중앙에 있고, 실제 동작 축은 오른쪽 탈출 공간을 활용하지 않는다.",
        "entities": "젊은 동아시아계 남성 한 명으로, 검은 헝클어진 머리와 얼굴 생김새가 현우 참고에 가깝다. 국적 자체는 시각적으로 검증할 수 없다. 남색 반팔 티셔츠는 참고 의상과 일치하며 얼굴의 멍과 구속되지 않은 손목도 보인다. 긴 바지와 부츠 때문에 개에게 물린 다리 상처는 확인할 수 없다. 다른 사람이나 읽을 수 있는 문자는 없다.",
        "hard_violations": [],
        "physics": "화면 아래쪽으로 내민 발의 밑창이 흙바닥에 닿아 몸을 지지하고 접촉 부근에 먼지가 일어난다. 반대쪽 다리는 뒤에서 접혀 들려 있다. 기울어진 몸과 팔 동작은 방향을 틀며 달리는 자세로 물리적으로 가능하다. 다만 접지한 것은 앞쪽으로 내민 발이고 뒷발은 떠 있어, 요구된 뒷발의 강한 지면 밀기를 보여주지는 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.714,
    "B": 1.5
   },
   "adjusted": {
    "A": 1.714,
    "B": 1.25
   },
   "violations": {
    "B": [
     "[gemini-pro] 이전 장면과 기준 이미지에 없는 회색 겉옷(invented object) 착용",
     "[gemini-pro] 지면에 닿은 발 없이 몸이 공중에 완전히 떠 있음 (unsupported body)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1714,
   "B": 1250
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1714,
    "verdict_ko": "의상과 배경, 인물의 외형은 기준과 잘 일치하나, 프롬프트가 요구한 '뒷발 접촉' 대신 앞발이 지면에 닿아 동작 연출에서 감점이 있습니다."
   },
   {
    "label": "B",
    "score": 1250,
    "verdict_ko": "이전 장면에 없는 겉옷을 임의로 추가하는 심각한 의상 오류가 발생했으며, 두 발이 모두 공중에 떠 있어 동작 지시를 어겼습니다.  ★위반: [gemini-pro] 이전 장면과 기준 이미지에 없는 회색 겉옷(invented object) 착용 / [gemini-pro] 지면에 닿은 발 없이 몸이 공중에 완전히 떠 있음 (unsupported body)"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S20sh6_sel.png",
    "asset_id": "657afd08-e32e-45ee-b32c-f6cca3ffd1ad",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-1cb8-7c32-9e27-291f4e8a344c",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S20sh6"
  }
 },
 "S24sh10::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:09:15.071335+00:00",
  "fingerprint": "3f337bd0b309135502bc4e24cdd2f8a8d85f0a0e7f2c5e4b7835fcae3b350b75",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S24sh10_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S24sh10_sel.png",
  "source_sha256": "a60b21c93267a6e6cdf734ccc33872ef7c03625f15a0215c2b920e8433ed0fb5",
  "file": "S24sh10_cine.png",
  "staged_sha256": "034b8f1c9a2a16bb258448dd5ffa535ec69a6bde34f95c6689b8799225b00dec",
  "latency_ms": 11270
 },
 "S25sh12::signage": {
  "fp": "edc9b51253de3050",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S25sh12": {
  "input_fingerprint": "6d1e9d9c5f00629d",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 찰리의 가슴에서 뿜어진 빛과 함께 주변 가로등들이 일제히 탁 켜져 빛을 발하는 눈부신 순간의 광경.\n\nLOCATION (lock): In the refugee settlement's open wedding-reception lot, beneath makeshift canopies and newly illuminated streetlights. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Streetlights behind and left of the unobstructed chest in the upper-left of the frame, background; Streetlights behind and right of the unobstructed chest in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: 주변 가로등 (Arranged around the reception area and continuing into the settlement) — Seen from above and obliquely, distributed behind and to both sides of 찰리; used as Their physical distribution gives the expanding event depth and scale; 피로연장 공터 (An open clearing used for the reception); used as Provides visible separation between 찰리 and the surrounding streetlights; 피로연장 천막 (Set up in the clearing) — An oblique upper and side view appears along a background edge; used as Anchors the spectacle to the modest reception setting without blocking the chest; 난민촌 (Extending beyond the reception clearing); used as Supplies the wider environmental scale revealed by the retreat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: At dusk, the ring-shaped chest emission and rapidly relighting streetlights create a dazzling expansion of illumination while controlled exposure preserves 찰리's outline.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Streetlights around the tented reception are relighting after the blackout; the repaired jukebox remains at the party, and some lamps burst from excessive brightness. Charlie's chest ring emits an intensifying light, with his coat-and-hat disguise otherwise unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 찰리의 가슴에서 뿜어진 빛과 함께 주변 가로등들이 일제히 탁 켜져 빛을 발하는 눈부신 순간의 광경.\n\nLOCATION (lock): In the refugee settlement's open wedding-reception lot, beneath makeshift canopies and newly illuminated streetlights. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Streetlights behind and left of the unobstructed chest in the upper-left of the frame, background; Streetlights behind and right of the unobstructed chest in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: 주변 가로등 (Arranged around the reception area and continuing into the settlement) — Seen from above and obliquely, distributed behind and to both sides of 찰리; used as Their physical distribution gives the expanding event depth and scale; 피로연장 공터 (An open clearing used for the reception); used as Provides visible separation between 찰리 and the surrounding streetlights; 피로연장 천막 (Set up in the clearing) — An oblique upper and side view appears along a background edge; used as Anchors the spectacle to the modest reception setting without blocking the chest; 난민촌 (Extending beyond the reception clearing); used as Supplies the wider environmental scale revealed by the retreat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: At dusk, the ring-shaped chest emission and rapidly relighting streetlights create a dazzling expansion of illumination while controlled exposure preserves 찰리's outline.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Streetlights around the tented reception are relighting after the blackout; the repaired jukebox remains at the party, and some lamps burst from excessive brightness. Charlie's chest ring emits an intensifying light, with his coat-and-hat disguise otherwise unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 찰리의 가슴에서 뿜어진 빛과 함께 주변 가로등들이 일제히 탁 켜져 빛을 발하는 눈부신 순간의 광경.\n\nLOCATION (lock): In the refugee settlement's open wedding-reception lot, beneath makeshift canopies and newly illuminated streetlights. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Streetlights behind and left of the unobstructed chest in the upper-left of the frame, background; Streetlights behind and right of the unobstructed chest in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: 주변 가로등 (Arranged around the reception area and continuing into the settlement) — Seen from above and obliquely, distributed behind and to both sides of 찰리; used as Their physical distribution gives the expanding event depth and scale; 피로연장 공터 (An open clearing used for the reception); used as Provides visible separation between 찰리 and the surrounding streetlights; 피로연장 천막 (Set up in the clearing) — An oblique upper and side view appears along a background edge; used as Anchors the spectacle to the modest reception setting without blocking the chest; 난민촌 (Extending beyond the reception clearing); used as Supplies the wider environmental scale revealed by the retreat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: At dusk, the ring-shaped chest emission and rapidly relighting streetlights create a dazzling expansion of illumination while controlled exposure preserves 찰리's outline.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Streetlights around the tented reception are relighting after the blackout; the repaired jukebox remains at the party, and some lamps burst from excessive brightness. Charlie's chest ring emits an intensifying light, with his coat-and-hat disguise otherwise unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S25sh12__bgfirst_bg.png",
     "asset_id": "44dbc893-09f5-4451-985f-95d270aac79f",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S25sh12.png",
     "asset_id": "b30b0e2b-b930-4260-9eaf-f19f33633ca4",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_reception_clearing_3488b1.png",
     "asset_id": "7f28b2fd-054b-41f4-b119-0c0d0c391ef1",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "가슴에서 빛이 뿜어져 나오나, 좌측 바닥을 향하는 인위적 화살표가 함께 그려져 있음.",
    "built_space": "위치 참조와 일치하는 피로연장 배경 및 전경 좌측에 쥬크박스가 배치됨.",
    "entities": "찰리의 외형은 참조와 일치하나 지시된 코트와 모자가 없음.",
    "hard_violations": [
     "[gemini-pro] 화면을 가로지르는 인위적인 화살표 그래픽 노출 (leaked marker)",
     "[gpt-high] 가슴에서 왼쪽 아래로 이어지는 흰 선과 화살촉이 화면 위에 붙인 방향 지시 도식으로 보여, 금지된 마커·다이어그램·오버레이에 해당한다."
    ],
    "physics": "찰리가 바닥에 단단히 서 있음."
   },
   {
    "label": "B",
    "direction": "찰리의 가슴 링에서 정면으로 강한 빛이 뿜어져 나옴.",
    "built_space": "피로연장 공터, 천막, 켜진 가로등들이 참조 이미지와 유사하게 배치됨.",
    "entities": "찰리는 참조된 로봇 외형에 프롬프트가 지시한 코트와 모자를 착용함.",
    "hard_violations": [],
    "physics": "찰리가 두 발로 바닥을 안정적으로 딛고 서 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "요구된 코트와 모자 변장을 정확히 반영하고, 가슴의 링 발광 및 가로등 배경을 충실히 구현함."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "화면에 인위적인 화살표 그래픽이 포함된 하드 위반이 있으며, 코트와 모자 착용 지시도 누락됨."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 가슴 링에서 정면으로 강한 빛이 뿜어져 나옴.",
        "built_space": "피로연장 공터, 천막, 켜진 가로등들이 참조 이미지와 유사하게 배치됨.",
        "entities": "찰리는 참조된 로봇 외형에 프롬프트가 지시한 코트와 모자를 착용함.",
        "hard_violations": [],
        "physics": "찰리가 두 발로 바닥을 안정적으로 딛고 서 있음."
       },
       {
        "label": "A",
        "direction": "가슴에서 빛이 뿜어져 나오나, 좌측 바닥을 향하는 인위적 화살표가 함께 그려져 있음.",
        "built_space": "위치 참조와 일치하는 피로연장 배경 및 전경 좌측에 쥬크박스가 배치됨.",
        "entities": "찰리의 외형은 참조와 일치하나 지시된 코트와 모자가 없음.",
        "hard_violations": [
         "화면을 가로지르는 인위적인 화살표 그래픽 노출 (leaked marker)"
        ],
        "physics": "찰리가 바닥에 단단히 서 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "요구된 코트와 모자 변장을 정확히 반영하고, 가슴의 링 발광 및 가로등 배경을 충실히 구현함."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "화면에 인위적인 화살표 그래픽이 포함된 하드 위반이 있으며, 코트와 모자 착용 지시도 누락됨."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 가슴 링에서 정면으로 강한 빛이 뿜어져 나옴.",
        "built_space": "피로연장 공터, 천막, 켜진 가로등들이 참조 이미지와 유사하게 배치됨.",
        "entities": "찰리는 참조된 로봇 외형에 프롬프트가 지시한 코트와 모자를 착용함.",
        "hard_violations": [],
        "physics": "찰리가 두 발로 바닥을 안정적으로 딛고 서 있음."
       },
       {
        "label": "A",
        "direction": "가슴에서 빛이 뿜어져 나오나, 좌측 바닥을 향하는 인위적 화살표가 함께 그려져 있음.",
        "built_space": "위치 참조와 일치하는 피로연장 배경 및 전경 좌측에 쥬크박스가 배치됨.",
        "entities": "찰리의 외형은 참조와 일치하나 지시된 코트와 모자가 없음.",
        "hard_violations": [
         "화면을 가로지르는 인위적인 화살표 그래픽 노출 (leaked marker)"
        ],
        "physics": "찰리가 바닥에 단단히 서 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "가려지지 않은 가슴의 고리형 발광, 코트와 모자, 양쪽 뒤편 가로등과 일몰의 피로연장을 충실히 구현하지만 찰리의 화면 점유율이 커 공간적 확장감은 다소 약하다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "넓은 공터와 가로등의 동시 점등은 드러나지만 가슴에서 주크박스 쪽으로 향하는 화살표 도식이 금지된 오버레이이며, 코트·모자와 식별 가능한 발광 고리도 빠졌다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 얼굴과 몸통은 대체로 카메라를 향하며 특정 물체를 겨냥하지 않는다. 가슴의 빛은 원형 중심에서 방사상으로 퍼진다. 가로등은 가슴보다 위쪽의 좌우 뒤편에 분포하며 공터와 정착지 쪽을 비춘다. 손에 든 물건이나 무기는 없다.",
        "built_space": "중앙 공터를 두고 왼쪽에 큰 연회 천막군, 오른쪽 뒤에 낮은 천막, 오른쪽 전경에 잘린 천막이 있다. 천막 아래에는 원형 테이블과 의자가 있고 뒤로 컨테이너 주거지, 전신주, 산과 수면이 이어져 참조 장소의 구조와 재료가 잘 유지된다. 가까운 좌우 가장자리의 높은 조명 기둥은 각각 하나이며, 그 사이와 먼 정착지에도 여러 가로등이 이어진다. 주크박스는 하나로 오른쪽 중경에 있어 참조의 왼쪽 전경 위치와 다르다. 찰리는 천막 밖 공터에 서 있고 가슴을 가리는 시설은 없다.",
        "entities": "등장 인물은 찰리 하나뿐이다. 샌드 베이지 장갑, 넓은 어깨, 육중하고 긴 팔, 짧은 다리, 흰 기계식 마스크 얼굴이 보인다. 얼굴의 점무늬는 참조의 매끈한 판금 얼굴과 차이가 있다. 모자와 열린 코트를 착용했고 가슴에는 내부가 구분되는 밝은 발광 고리가 있다. 주크박스, 연회용 천막과 가구, 정착지와 켜진 가로등이 보인다. 터지는 전구나 파편은 명확하지 않다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "찰리의 양발이 흙바닥에 닿아 몸을 지탱하며 무릎과 팔의 관절 연결도 자연스럽다. 코트는 어깨에서 걸려 아래로 늘어지고 모자는 머리에 얹혀 있다. 천막은 기둥과 줄로 지지되고 조명은 기둥 또는 전선에 부착되어 있다. 주크박스와 가구는 지면에 놓여 있다. 가슴 발광이 주변 옷과 바닥을 밝히며 근거 없이 떠 있는 몸이나 물체는 없다."
       },
       {
        "label": "B",
        "direction": "찰리는 화면 왼쪽 앞쪽으로 얼굴과 몸통을 조금 돌리고 있다. 가슴에서 시작하는 여러 흰 직선이 왼쪽 아래 주크박스 옆 지면으로 향하며 끝에 뚜렷한 화살촉이 있다. 이는 주변 가로등으로 확장되는 빛보다 특정 지점을 지시하는 도식으로 읽힌다. 가로등 자체는 가슴의 좌우 뒤편 상단에서 공터를 향해 빛난다.",
        "built_space": "왼쪽 큰 연회 천막군, 오른쪽 뒤편 낮은 천막과 오른쪽 전경 천막, 중앙 흙 공터, 컨테이너 정착지와 산이 참조 장소와 대응한다. 가까운 좌우 끝의 높은 조명 기둥은 각각 하나이고 여러 중·원경 가로등이 양쪽으로 이어진다. 주크박스는 하나로 참조처럼 왼쪽 전경에 놓여 있다. 찰리는 공터 중앙보다 오른쪽에 서 있으며 주변과의 빈 공간이 충분하고 천막의 상면도 보인다.",
        "entities": "찰리 한 명만 있으며 베이지색 장갑, 긴 팔과 짧은 다리, 흰 마스크형 얼굴은 참조 정체성과 잘 맞는다. 그러나 유지되어야 할 코트와 모자가 모두 없다. 가슴은 흰 빛 덩어리로 과노출되어 고리의 형태를 식별하기 어렵다. 주크박스, 천막, 연회 가구, 정착지와 밝게 켜진 가로등은 존재한다. 전구 파열은 명확하지 않다. 읽을 수 있는 글자는 없지만 큰 화살표 표시가 추가되어 있다.",
        "hard_violations": [
         "가슴에서 왼쪽 아래로 이어지는 흰 선과 화살촉이 화면 위에 붙인 방향 지시 도식으로 보여, 금지된 마커·다이어그램·오버레이에 해당한다."
        ],
        "physics": "찰리는 벌린 양발을 땅에 딛고 약간 구부린 다리로 체중을 지탱한다. 팔과 손은 몸에 정상적으로 연결되어 아래로 내려와 있다. 주크박스와 가구는 바닥에 놓이고 천막과 가로등에는 기둥 지지가 보인다. 떠 있는 인체나 지지 없는 소품은 없다. 다만 가슴 빛의 선명한 화살촉은 실제 조명이 공간에서 만드는 광학 현상으로 읽히지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "가려지지 않은 가슴의 고리형 발광, 코트와 모자, 양쪽 뒤편 가로등과 일몰의 피로연장을 충실히 구현하지만 찰리의 화면 점유율이 커 공간적 확장감은 다소 약하다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "넓은 공터와 가로등의 동시 점등은 드러나지만 가슴에서 주크박스 쪽으로 향하는 화살표 도식이 금지된 오버레이이며, 코트·모자와 식별 가능한 발광 고리도 빠졌다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 얼굴과 몸통은 대체로 카메라를 향하며 특정 물체를 겨냥하지 않는다. 가슴의 빛은 원형 중심에서 방사상으로 퍼진다. 가로등은 가슴보다 위쪽의 좌우 뒤편에 분포하며 공터와 정착지 쪽을 비춘다. 손에 든 물건이나 무기는 없다.",
        "built_space": "중앙 공터를 두고 왼쪽에 큰 연회 천막군, 오른쪽 뒤에 낮은 천막, 오른쪽 전경에 잘린 천막이 있다. 천막 아래에는 원형 테이블과 의자가 있고 뒤로 컨테이너 주거지, 전신주, 산과 수면이 이어져 참조 장소의 구조와 재료가 잘 유지된다. 가까운 좌우 가장자리의 높은 조명 기둥은 각각 하나이며, 그 사이와 먼 정착지에도 여러 가로등이 이어진다. 주크박스는 하나로 오른쪽 중경에 있어 참조의 왼쪽 전경 위치와 다르다. 찰리는 천막 밖 공터에 서 있고 가슴을 가리는 시설은 없다.",
        "entities": "등장 인물은 찰리 하나뿐이다. 샌드 베이지 장갑, 넓은 어깨, 육중하고 긴 팔, 짧은 다리, 흰 기계식 마스크 얼굴이 보인다. 얼굴의 점무늬는 참조의 매끈한 판금 얼굴과 차이가 있다. 모자와 열린 코트를 착용했고 가슴에는 내부가 구분되는 밝은 발광 고리가 있다. 주크박스, 연회용 천막과 가구, 정착지와 켜진 가로등이 보인다. 터지는 전구나 파편은 명확하지 않다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "찰리의 양발이 흙바닥에 닿아 몸을 지탱하며 무릎과 팔의 관절 연결도 자연스럽다. 코트는 어깨에서 걸려 아래로 늘어지고 모자는 머리에 얹혀 있다. 천막은 기둥과 줄로 지지되고 조명은 기둥 또는 전선에 부착되어 있다. 주크박스와 가구는 지면에 놓여 있다. 가슴 발광이 주변 옷과 바닥을 밝히며 근거 없이 떠 있는 몸이나 물체는 없다."
       },
       {
        "label": "A",
        "direction": "찰리는 화면 왼쪽 앞쪽으로 얼굴과 몸통을 조금 돌리고 있다. 가슴에서 시작하는 여러 흰 직선이 왼쪽 아래 주크박스 옆 지면으로 향하며 끝에 뚜렷한 화살촉이 있다. 이는 주변 가로등으로 확장되는 빛보다 특정 지점을 지시하는 도식으로 읽힌다. 가로등 자체는 가슴의 좌우 뒤편 상단에서 공터를 향해 빛난다.",
        "built_space": "왼쪽 큰 연회 천막군, 오른쪽 뒤편 낮은 천막과 오른쪽 전경 천막, 중앙 흙 공터, 컨테이너 정착지와 산이 참조 장소와 대응한다. 가까운 좌우 끝의 높은 조명 기둥은 각각 하나이고 여러 중·원경 가로등이 양쪽으로 이어진다. 주크박스는 하나로 참조처럼 왼쪽 전경에 놓여 있다. 찰리는 공터 중앙보다 오른쪽에 서 있으며 주변과의 빈 공간이 충분하고 천막의 상면도 보인다.",
        "entities": "찰리 한 명만 있으며 베이지색 장갑, 긴 팔과 짧은 다리, 흰 마스크형 얼굴은 참조 정체성과 잘 맞는다. 그러나 유지되어야 할 코트와 모자가 모두 없다. 가슴은 흰 빛 덩어리로 과노출되어 고리의 형태를 식별하기 어렵다. 주크박스, 천막, 연회 가구, 정착지와 밝게 켜진 가로등은 존재한다. 전구 파열은 명확하지 않다. 읽을 수 있는 글자는 없지만 큰 화살표 표시가 추가되어 있다.",
        "hard_violations": [
         "가슴에서 왼쪽 아래로 이어지는 흰 선과 화살촉이 화면 위에 붙인 방향 지시 도식으로 보여, 금지된 마커·다이어그램·오버레이에 해당한다."
        ],
        "physics": "찰리는 벌린 양발을 땅에 딛고 약간 구부린 다리로 체중을 지탱한다. 팔과 손은 몸에 정상적으로 연결되어 아래로 내려와 있다. 주크박스와 가구는 바닥에 놓이고 천막과 가로등에는 기둥 지지가 보인다. 떠 있는 인체나 지지 없는 소품은 없다. 다만 가슴 빛의 선명한 화살촉은 실제 조명이 공간에서 만드는 광학 현상으로 읽히지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.679,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.429,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 화면을 가로지르는 인위적인 화살표 그래픽 노출 (leaked marker)",
     "[gpt-high] 가슴에서 왼쪽 아래로 이어지는 흰 선과 화살촉이 화면 위에 붙인 방향 지시 도식으로 보여, 금지된 마커·다이어그램·오버레이에 해당한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 429
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "요구된 코트와 모자 변장을 정확히 반영하고, 가슴의 링 발광 및 가로등 배경을 충실히 구현함."
   },
   {
    "label": "A",
    "score": 429,
    "verdict_ko": "화면에 인위적인 화살표 그래픽이 포함된 하드 위반이 있으며, 코트와 모자 착용 지시도 누락됨.  ★위반: [gemini-pro] 화면을 가로지르는 인위적인 화살표 그래픽 노출 (leaked marker) / [gpt-high] 가슴에서 왼쪽 아래로 이어지는 흰 선과 화살촉이 화면 위에 붙인 방향 지시 도식으로 보여, 금지된 마커·다이어그램·오버레이에 해당한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_reception_clearing_3488b1.png",
    "asset_id": "7f28b2fd-054b-41f4-b119-0c0d0c391ef1",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-1e88-76f0-9a13-323b7a8a6a69",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S25sh12__bgfirst_bg.png",
   "bg_asset_id": "44dbc893-09f5-4451-985f-95d270aac79f",
   "bg_record_key": "S25sh12::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "reception_clearing",
   "groupbg_asset_id": "7f28b2fd-054b-41f4-b119-0c0d0c391ef1"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S25sh12::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:13:31.704974+00:00",
  "fingerprint": "b4143b1440b2b39abb2417211f5277da3302c36577d7174636917439de068971",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S25sh12_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S25sh12_sel.png",
  "source_sha256": "9cb9736af602521feaec132664873ecf03c5706ba5c0ebeb613d49eaadf80b94",
  "file": "S25sh12_cine.png",
  "staged_sha256": "9488d8676e49a12af2a921374ccb6d1d182bb40ea12a9a9c12536f753257022d",
  "latency_ms": 11351
 },
 "S25sh19::signage": {
  "fp": "2168f4ffa565587f",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S25sh19": {
  "input_fingerprint": "e83928c59b3d7c22",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 트럭에서 내린 채 험악한 표정으로 앞을 노려보는 박철진과 그 뒤의 민병대원들 전신.\n\nLOCATION (lock): At the vehicle-access edge of the outdoor reception lot, beside a newly arrived militia truck. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 민병대 트럭 (Stopped after arriving, with the men now dismounted) — Its front-side quarter is visible behind the figures, occupying no more than the rear-right third; used as Establishes the source of the intrusion and supplies a grounded scale reference; 피로연장 입구와 공터 (The arrival interrupts the gathering in the clearing); used as Open space toward the left edge receives the men's outward-directed attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The reception remains brightly illuminated after the restoration of power, with controlled contrast preserving the severity of 박철진's expression.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the reception clearing, temporary tent structures, and streetlights that remain illuminated. Exclude the transient radiance emitted by the robot and any bulbs caught in the act of bursting.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The reception and refugee settlement remain brightly illuminated, with some lamps burst and the jukebox playing again. A militia truck has stopped at the tented clearing, and Charlie retains his old coat and hat. 박철진: He has disembarked from the militia truck at the reception.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 트럭에서 내린 채 험악한 표정으로 앞을 노려보는 박철진과 그 뒤의 민병대원들 전신.\n\nLOCATION (lock): At the vehicle-access edge of the outdoor reception lot, beside a newly arrived militia truck. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 민병대 트럭 (Stopped after arriving, with the men now dismounted) — Its front-side quarter is visible behind the figures, occupying no more than the rear-right third; used as Establishes the source of the intrusion and supplies a grounded scale reference; 피로연장 입구와 공터 (The arrival interrupts the gathering in the clearing); used as Open space toward the left edge receives the men's outward-directed attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The reception remains brightly illuminated after the restoration of power, with controlled contrast preserving the severity of 박철진's expression.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the reception clearing, temporary tent structures, and streetlights that remain illuminated. Exclude the transient radiance emitted by the robot and any bulbs caught in the act of bursting.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The reception and refugee settlement remain brightly illuminated, with some lamps burst and the jukebox playing again. A militia truck has stopped at the tented clearing, and Charlie retains his old coat and hat. 박철진: He has disembarked from the militia truck at the reception.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 트럭에서 내린 채 험악한 표정으로 앞을 노려보는 박철진과 그 뒤의 민병대원들 전신.\n\nLOCATION (lock): At the vehicle-access edge of the outdoor reception lot, beside a newly arrived militia truck. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 민병대 트럭 (Stopped after arriving, with the men now dismounted) — Its front-side quarter is visible behind the figures, occupying no more than the rear-right third; used as Establishes the source of the intrusion and supplies a grounded scale reference; 피로연장 입구와 공터 (The arrival interrupts the gathering in the clearing); used as Open space toward the left edge receives the men's outward-directed attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The reception remains brightly illuminated after the restoration of power, with controlled contrast preserving the severity of 박철진's expression.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the reception clearing, temporary tent structures, and streetlights that remain illuminated. Exclude the transient radiance emitted by the robot and any bulbs caught in the act of bursting.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The reception and refugee settlement remain brightly illuminated, with some lamps burst and the jukebox playing again. A militia truck has stopped at the tented clearing, and Charlie retains his old coat and hat. 박철진: He has disembarked from the militia truck at the reception.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "박철진과 그의 뒤에 선 민병대원들 모두 정면(카메라 방향)을 향해 험악한 시선을 던지고 있습니다.",
    "built_space": "화면 좌측에는 레퍼런스와 일치하는 피로연장의 텐트, 줄조명, 테이블들이 배치되어 있고, 우측에는 지시된 대로 민병대 트럭이 공간을 차지하고 있습니다.",
    "entities": "박철진은 제공된 레퍼런스 인물의 외모와 일치하며, 뒤에 선 다인종 민병대원들과 우측의 트럭 모두 프롬프트의 요구사항을 반영하고 있습니다.",
    "hard_violations": [
     "[gemini-pro] 읽을 수 있는 글씨 금지(No readable writing) 지시를 위반하고 트럭 문짝에 명확하게 읽히는 텍스트가 묘사됨.",
     "[gpt-high] 트럭 문과 차체에 읽을 수 있는 한글 및 숫자 표기가 노출되어, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 명시적 금지 조건을 위반한다."
    ],
    "physics": "인물들의 발이 지면에 닿아 무게중심이 잡혀 있으며, 우측의 트럭 역시 지면에 안정적으로 고정되어 있습니다."
   },
   {
    "label": "B",
    "direction": "중앙의 박철진과 주변의 인물들 모두 정면을 노려보고 있습니다.",
    "built_space": "좌측의 피로연장 시설(조명, 천막, 원형 테이블)과 우측 트럭의 배치가 주어진 샷 텍스트의 공간적 조건을 잘 따르고 있습니다.",
    "entities": "중앙의 박철진은 레퍼런스와 일치하지만, 프레임에 등장하는 나머지 민병대원 5명의 얼굴이 모두 박철진과 완전히 똑같이 묘사되었습니다.",
    "hard_violations": [
     "[gemini-pro] 인물 복제(Duplicated bodies): 배경에 서 있는 민병대원 전원이 특정 레퍼런스 인물과 동일한 얼굴로 복제되는 치명적인 오류가 발생함."
    ],
    "physics": "모든 인물이 땅을 디디고 서 있으며 트럭 바퀴와 구조물 역시 지면에 정상적으로 닿아 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "구도와 인물 표현은 전반적으로 우수하나, 엄격히 금지된 명확한 텍스트가 트럭 측면에 생성되어 큰 감점 요인이 되었습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "모든 민병대원이 메인 캐릭터와 완벽히 동일한 얼굴로 복제되어 실사 영화 스틸컷으로서의 사실성을 상실했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진과 그의 뒤에 선 민병대원들 모두 정면(카메라 방향)을 향해 험악한 시선을 던지고 있습니다.",
        "built_space": "화면 좌측에는 레퍼런스와 일치하는 피로연장의 텐트, 줄조명, 테이블들이 배치되어 있고, 우측에는 지시된 대로 민병대 트럭이 공간을 차지하고 있습니다.",
        "entities": "박철진은 제공된 레퍼런스 인물의 외모와 일치하며, 뒤에 선 다인종 민병대원들과 우측의 트럭 모두 프롬프트의 요구사항을 반영하고 있습니다.",
        "hard_violations": [
         "읽을 수 있는 글씨 금지(No readable writing) 지시를 위반하고 트럭 문짝에 명확하게 읽히는 텍스트가 묘사됨."
        ],
        "physics": "인물들의 발이 지면에 닿아 무게중심이 잡혀 있으며, 우측의 트럭 역시 지면에 안정적으로 고정되어 있습니다."
       },
       {
        "label": "B",
        "direction": "중앙의 박철진과 주변의 인물들 모두 정면을 노려보고 있습니다.",
        "built_space": "좌측의 피로연장 시설(조명, 천막, 원형 테이블)과 우측 트럭의 배치가 주어진 샷 텍스트의 공간적 조건을 잘 따르고 있습니다.",
        "entities": "중앙의 박철진은 레퍼런스와 일치하지만, 프레임에 등장하는 나머지 민병대원 5명의 얼굴이 모두 박철진과 완전히 똑같이 묘사되었습니다.",
        "hard_violations": [
         "인물 복제(Duplicated bodies): 배경에 서 있는 민병대원 전원이 특정 레퍼런스 인물과 동일한 얼굴로 복제되는 치명적인 오류가 발생함."
        ],
        "physics": "모든 인물이 땅을 디디고 서 있으며 트럭 바퀴와 구조물 역시 지면에 정상적으로 닿아 있습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "구도와 인물 표현은 전반적으로 우수하나, 엄격히 금지된 명확한 텍스트가 트럭 측면에 생성되어 큰 감점 요인이 되었습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "모든 민병대원이 메인 캐릭터와 완벽히 동일한 얼굴로 복제되어 실사 영화 스틸컷으로서의 사실성을 상실했습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "박철진과 그의 뒤에 선 민병대원들 모두 정면(카메라 방향)을 향해 험악한 시선을 던지고 있습니다.",
        "built_space": "화면 좌측에는 레퍼런스와 일치하는 피로연장의 텐트, 줄조명, 테이블들이 배치되어 있고, 우측에는 지시된 대로 민병대 트럭이 공간을 차지하고 있습니다.",
        "entities": "박철진은 제공된 레퍼런스 인물의 외모와 일치하며, 뒤에 선 다인종 민병대원들과 우측의 트럭 모두 프롬프트의 요구사항을 반영하고 있습니다.",
        "hard_violations": [
         "읽을 수 있는 글씨 금지(No readable writing) 지시를 위반하고 트럭 문짝에 명확하게 읽히는 텍스트가 묘사됨."
        ],
        "physics": "인물들의 발이 지면에 닿아 무게중심이 잡혀 있으며, 우측의 트럭 역시 지면에 안정적으로 고정되어 있습니다."
       },
       {
        "label": "B",
        "direction": "중앙의 박철진과 주변의 인물들 모두 정면을 노려보고 있습니다.",
        "built_space": "좌측의 피로연장 시설(조명, 천막, 원형 테이블)과 우측 트럭의 배치가 주어진 샷 텍스트의 공간적 조건을 잘 따르고 있습니다.",
        "entities": "중앙의 박철진은 레퍼런스와 일치하지만, 프레임에 등장하는 나머지 민병대원 5명의 얼굴이 모두 박철진과 완전히 똑같이 묘사되었습니다.",
        "hard_violations": [
         "인물 복제(Duplicated bodies): 배경에 서 있는 민병대원 전원이 특정 레퍼런스 인물과 동일한 얼굴로 복제되는 치명적인 오류가 발생함."
        ],
        "physics": "모든 인물이 땅을 디디고 서 있으며 트럭 바퀴와 구조물 역시 지면에 정상적으로 닿아 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "전신의 박철진과 뒤따르는 한국인 민병대원, 일몰의 피로연장은 살렸지만, 시선이 왼쪽 공터가 아닌 카메라를 향하고 트럭이 지나치게 크며 참고의 정장도 바뀌었다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "트럭의 읽을 수 있는 문자가 금지 조건을 위반하며, 왼쪽 공터를 향하지 않는 시선·과도한 트럭 비중·참고와 다른 복장에 한국인 현지인 설정 불일치까지 더해진다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진은 정면의 카메라 쪽을 험악하게 바라보고, 뒤의 다섯 대원도 대부분 같은 정면을 본다. 지정된 왼쪽 가장자리의 빈 공터를 주시하는 배치는 아니다. 뚜렷한 총구나 겨누는 물체는 보이지 않는다. 트럭 앞부분은 화면 왼쪽 앞을 향한다.",
        "built_space": "트럭 한 대가 인물 뒤 오른쪽에 있지만 화면 폭의 약 40%와 상당한 높이를 차지하여 후방 오른쪽 3분의 1 제한을 넘는다. 왼쪽에는 연결된 밝은 천막 지붕 두 구획, 여러 원탁과 접이식 의자, 켜진 전구 줄과 투광등 기둥들이 있다. 드럼통은 왼쪽 전경 두 개와 천막 주변 세 개가 뚜렷하다. 흙·자갈·물웅덩이 공터와 뒤편 컨테이너 정착지는 참고 장소와 잘 이어진다. 박철진은 앞 중앙, 대원 다섯 명은 그 뒤 양옆에 있고 일부 대원의 하체는 다른 인물에 가려진다. 주크박스가 있던 오른쪽 배경은 트럭에 가려 확인할 수 없다.",
        "entities": "박철진 한 명과 민병대원 다섯 명이 보이며, 모두 한국인 설정과 양립하는 동아시아계 성인 남성 외형이다. 박철진은 중년의 얼굴과 짧고 뒤로 넘긴 검은 머리로 참고와 대체로 닮았다. 그러나 참고의 남색 정장·흰 셔츠·줄무늬 넥타이 대신 올리브색 야전복과 군화를 착용한다. 트럭, 임시 천막, 피로연 가구, 점등된 조명과 일몰은 구현되어 있다. 이전 장면의 로봇과 방사광은 없고 읽을 수 있는 문자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "박철진과 발이 드러난 대원들은 군화로 지면을 딛고 있다. 가운데 뒤쪽 대원은 하체가 가려졌지만 공중에 뜬 몸으로 보이지 않는다. 트럭은 타이어로 땅을 지지하고 천막은 기둥과 줄에 매달려 있으며 테이블과 의자는 다리로 서 있다. 하차를 마치고 서 있는 상황은 물리적으로 가능하다. 다만 박철진은 팔을 내리고 정면으로 선 다소 경직된 자세다."
       },
       {
        "label": "B",
        "direction": "박철진과 뒤의 여섯 대원은 대체로 카메라 또는 카메라 바로 옆 전방을 바라본다. 왼쪽 빈 공터로 향하는 집단의 주의는 드러나지 않는다. 확인되는 소총 네 정의 총구는 대원들의 몸 앞에서 아래쪽 땅을 향하며 특정 인물을 겨누지 않는다. 트럭은 화면 왼쪽 앞 방향으로 정차해 있다.",
        "built_space": "트럭 한 대의 전면과 긴 적재함 측면이 오른쪽 배경을 차지하여 지정된 후방 오른쪽 3분의 1보다 넓다. 박철진은 앞 중앙, 대원 여섯 명은 뒤에 모여 있으며 두 명은 다른 인물에 상당 부분 가려진다. 왼쪽의 연결된 천막 지붕 두 구획과 원탁·접이식 의자, 전구 줄, 점등된 투광등, 컨테이너와 산 능선은 참고 장소를 재현한다. 주크박스가 있던 영역은 차량과 인물에 가려져 있다. 건축물 내부나 거울 반사는 없다.",
        "entities": "박철진 한 명과 민병대원 여섯 명이 보인다. 박철진의 중년 동아시아계 남성 외형과 검은 머리는 참고에 가깝지만 정장이 야전복으로 바뀌었다. 대원 중에는 흑인으로 보이는 남성 두 명과 백인으로 보이는 남성이 있어, 별도 예외가 없는 현지인은 한국인이라는 설정에 부합하지 않는다. 민병대 트럭과 소총 네 정, 밝은 천막 피로연장, 일몰이 보인다. 트럭 문과 차체에는 판독 가능한 한글 및 숫자 표기가 있다. 이전 장면의 로봇은 없다.",
        "hard_violations": [
         "트럭 문과 차체에 읽을 수 있는 한글 및 숫자 표기가 노출되어, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 명시적 금지 조건을 위반한다."
        ],
        "physics": "박철진과 하체가 보이는 대원들의 발은 지면에 닿아 있다. 뒤쪽 두 대원은 몸이 가려졌을 뿐 지지 없이 떠 있는 것으로 보이지 않는다. 소총은 손과 몸 앞의 멜빵으로 지지되며 아래로 내려 든 자세가 가능하다. 트럭의 바퀴들은 땅에 놓여 있고 천막과 가구도 정상적인 지지 구조를 갖는다. 정차 후 하차한 인물들의 정지 상태로 성립한다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "전신의 박철진과 뒤따르는 한국인 민병대원, 일몰의 피로연장은 살렸지만, 시선이 왼쪽 공터가 아닌 카메라를 향하고 트럭이 지나치게 크며 참고의 정장도 바뀌었다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "트럭의 읽을 수 있는 문자가 금지 조건을 위반하며, 왼쪽 공터를 향하지 않는 시선·과도한 트럭 비중·참고와 다른 복장에 한국인 현지인 설정 불일치까지 더해진다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "박철진은 정면의 카메라 쪽을 험악하게 바라보고, 뒤의 다섯 대원도 대부분 같은 정면을 본다. 지정된 왼쪽 가장자리의 빈 공터를 주시하는 배치는 아니다. 뚜렷한 총구나 겨누는 물체는 보이지 않는다. 트럭 앞부분은 화면 왼쪽 앞을 향한다.",
        "built_space": "트럭 한 대가 인물 뒤 오른쪽에 있지만 화면 폭의 약 40%와 상당한 높이를 차지하여 후방 오른쪽 3분의 1 제한을 넘는다. 왼쪽에는 연결된 밝은 천막 지붕 두 구획, 여러 원탁과 접이식 의자, 켜진 전구 줄과 투광등 기둥들이 있다. 드럼통은 왼쪽 전경 두 개와 천막 주변 세 개가 뚜렷하다. 흙·자갈·물웅덩이 공터와 뒤편 컨테이너 정착지는 참고 장소와 잘 이어진다. 박철진은 앞 중앙, 대원 다섯 명은 그 뒤 양옆에 있고 일부 대원의 하체는 다른 인물에 가려진다. 주크박스가 있던 오른쪽 배경은 트럭에 가려 확인할 수 없다.",
        "entities": "박철진 한 명과 민병대원 다섯 명이 보이며, 모두 한국인 설정과 양립하는 동아시아계 성인 남성 외형이다. 박철진은 중년의 얼굴과 짧고 뒤로 넘긴 검은 머리로 참고와 대체로 닮았다. 그러나 참고의 남색 정장·흰 셔츠·줄무늬 넥타이 대신 올리브색 야전복과 군화를 착용한다. 트럭, 임시 천막, 피로연 가구, 점등된 조명과 일몰은 구현되어 있다. 이전 장면의 로봇과 방사광은 없고 읽을 수 있는 문자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "박철진과 발이 드러난 대원들은 군화로 지면을 딛고 있다. 가운데 뒤쪽 대원은 하체가 가려졌지만 공중에 뜬 몸으로 보이지 않는다. 트럭은 타이어로 땅을 지지하고 천막은 기둥과 줄에 매달려 있으며 테이블과 의자는 다리로 서 있다. 하차를 마치고 서 있는 상황은 물리적으로 가능하다. 다만 박철진은 팔을 내리고 정면으로 선 다소 경직된 자세다."
       },
       {
        "label": "A",
        "direction": "박철진과 뒤의 여섯 대원은 대체로 카메라 또는 카메라 바로 옆 전방을 바라본다. 왼쪽 빈 공터로 향하는 집단의 주의는 드러나지 않는다. 확인되는 소총 네 정의 총구는 대원들의 몸 앞에서 아래쪽 땅을 향하며 특정 인물을 겨누지 않는다. 트럭은 화면 왼쪽 앞 방향으로 정차해 있다.",
        "built_space": "트럭 한 대의 전면과 긴 적재함 측면이 오른쪽 배경을 차지하여 지정된 후방 오른쪽 3분의 1보다 넓다. 박철진은 앞 중앙, 대원 여섯 명은 뒤에 모여 있으며 두 명은 다른 인물에 상당 부분 가려진다. 왼쪽의 연결된 천막 지붕 두 구획과 원탁·접이식 의자, 전구 줄, 점등된 투광등, 컨테이너와 산 능선은 참고 장소를 재현한다. 주크박스가 있던 영역은 차량과 인물에 가려져 있다. 건축물 내부나 거울 반사는 없다.",
        "entities": "박철진 한 명과 민병대원 여섯 명이 보인다. 박철진의 중년 동아시아계 남성 외형과 검은 머리는 참고에 가깝지만 정장이 야전복으로 바뀌었다. 대원 중에는 흑인으로 보이는 남성 두 명과 백인으로 보이는 남성이 있어, 별도 예외가 없는 현지인은 한국인이라는 설정에 부합하지 않는다. 민병대 트럭과 소총 네 정, 밝은 천막 피로연장, 일몰이 보인다. 트럭 문과 차체에는 판독 가능한 한글 및 숫자 표기가 있다. 이전 장면의 로봇은 없다.",
        "hard_violations": [
         "트럭 문과 차체에 읽을 수 있는 한글 및 숫자 표기가 노출되어, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 명시적 금지 조건을 위반한다."
        ],
        "physics": "박철진과 하체가 보이는 대원들의 발은 지면에 닿아 있다. 뒤쪽 두 대원은 몸이 가려졌을 뿐 지지 없이 떠 있는 것으로 보이지 않는다. 소총은 손과 몸 앞의 멜빵으로 지지되며 아래로 내려 든 자세가 가능하다. 트럭의 바퀴들은 땅에 놓여 있고 천막과 가구도 정상적인 지지 구조를 갖는다. 정차 후 하차한 인물들의 정지 상태로 성립한다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.4,
    "B": 1.75
   },
   "adjusted": {
    "A": 1.15,
    "B": 1.5
   },
   "violations": {
    "A": [
     "[gemini-pro] 읽을 수 있는 글씨 금지(No readable writing) 지시를 위반하고 트럭 문짝에 명확하게 읽히는 텍스트가 묘사됨.",
     "[gpt-high] 트럭 문과 차체에 읽을 수 있는 한글 및 숫자 표기가 노출되어, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 명시적 금지 조건을 위반한다."
    ],
    "B": [
     "[gemini-pro] 인물 복제(Duplicated bodies): 배경에 서 있는 민병대원 전원이 특정 레퍼런스 인물과 동일한 얼굴로 복제되는 치명적인 오류가 발생함."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1150,
   "B": 1500
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1150,
    "verdict_ko": "구도와 인물 표현은 전반적으로 우수하나, 엄격히 금지된 명확한 텍스트가 트럭 측면에 생성되어 큰 감점 요인이 되었습니다.  ★위반: [gemini-pro] 읽을 수 있는 글씨 금지(No readable writing) 지시를 위반하고 트럭 문짝에 명확하게 읽히는 텍스트가 묘사됨. / [gpt-high] 트럭 문과 차체에 읽을 수 있는 한글 및 숫자 표기가 노출되어, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 명시적 금지 조건을 위반한다."
   },
   {
    "label": "B",
    "score": 1500,
    "verdict_ko": "모든 민병대원이 메인 캐릭터와 완벽히 동일한 얼굴로 복제되어 실사 영화 스틸컷으로서의 사실성을 상실했습니다.  ★위반: [gemini-pro] 인물 복제(Duplicated bodies): 배경에 서 있는 민병대원 전원이 특정 레퍼런스 인물과 동일한 얼굴로 복제되는 치명적인 오류가 발생함."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S25sh12_sel.png",
    "asset_id": "f62eb55b-5498-4c20-8395-21b23b6ebcd0",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1401722>",
    "asset_id": "fee7383c-fb61-4b3a-ba7c-79f2555de00b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-2370-764e-9dac-ad00021c269c",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S25sh12"
  }
 },
 "S25sh19::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:14:59.549632+00:00",
  "fingerprint": "39d7c4731da770ce51d418af7e809dfe610f466be6f8d41bccd83ce56a658ead",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S25sh19_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S25sh19_sel.png",
  "source_sha256": "868fcfc7176f20f6172b8b850a5b084873c5bc459c29c3acf5b7ed4b45613bbb",
  "file": "S25sh19_cine.png",
  "staged_sha256": "832ef6e560dc754ae1a6789c70a109ab9570a931adef220c6f280e0786b6c353",
  "latency_ms": 9589
 },
 "S25sh22::signage": {
  "fp": "5063e84ef22f5e3f",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::6fefa20797430b9e": {
  "subjects": [],
  "subject_text": "인천 난민촌 피로연장\n컨테이너 주거지의 빈 공터에 천막을 친 간이 행사 공간. 주변에 전등과 가로등이 있으며 한쪽에 주크박스가 놓여 있다.",
  "identity": "canonical",
  "scope_id": "L174",
  "scope_role": "location_exterior",
  "scope_sha": "d226e83a7d00860b"
 },
 "S25sh22::bgfirst_bg": {
  "input_fingerprint": "c7f97fb556fb8097",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 좁고 어두운 골목길 벽에 바짝 기대어 선 채 밖의 눈치를 살피는 현우 일행의 전신.\n\nLOCATION (lock): Against a wall in a narrow, dark alley of the refugee settlement, away from the reception lot.\n\nTIME OF DAY (lock): sunset.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Narrow exposed opening beyond the concealing corner in the upper-right of the frame, background; Inner wall shielding all five figures in the middle-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: 골목 벽 (Concealing the group from the militia beyond the corner) — The camera looks along its inner face at a shallow angle toward the corner; used as Forms the shared hiding boundary and a receding line through the five figures; 골목 모퉁이와 바깥으로 열린 틈 (Only a narrow view beyond the hiding place is available) — The corner interrupts the view at upper right, with the exposed passage visible beyond its edge; used as Makes the danger direction legible without revealing the group's bodies to it; 좁은 골목 바닥 (Visible under the group's feet); used as Preserves full-body readability and shows their uneven, compressed spacing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the nighttime alley dark but readable with restrained ambient separation, without carrying a specific reception light source into this location.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 좁고 어두운 골목길 벽에 바짝 기대어 선 채 밖의 눈치를 살피는 현우 일행의 전신.\n\nLOCATION (lock): Against a wall in a narrow, dark alley of the refugee settlement, away from the reception lot.\n\nTIME OF DAY (lock): sunset.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Narrow exposed opening beyond the concealing corner in the upper-right of the frame, background; Inner wall shielding all five figures in the middle-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: 골목 벽 (Concealing the group from the militia beyond the corner) — The camera looks along its inner face at a shallow angle toward the corner; used as Forms the shared hiding boundary and a receding line through the five figures; 골목 모퉁이와 바깥으로 열린 틈 (Only a narrow view beyond the hiding place is available) — The corner interrupts the view at upper right, with the exposed passage visible beyond its edge; used as Makes the danger direction legible without revealing the group's bodies to it; 좁은 골목 바닥 (Visible under the group's feet); used as Preserves full-body readability and shows their uneven, compressed spacing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the nighttime alley dark but readable with restrained ambient separation, without carrying a specific reception light source into this location.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S25sh22__bgfirst_bg.png",
  "asset_id": "4ea5365d-bc0a-4f55-9bd5-c4357af778b0",
  "input_asset_ids": [
   "ed35ed23-b843-4470-bb40-db03b4117fa0",
   "eac29610-01b6-43a7-9e6e-9e2387ddd674"
  ]
 },
 "S25sh22": {
  "input_fingerprint": "54f343ad7c744ffc",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 좁고 어두운 골목길 벽에 바짝 기대어 선 채 밖의 눈치를 살피는 현우 일행의 전신.\n\nLOCATION (lock): Against a wall in a narrow, dark alley of the refugee settlement, away from the reception lot. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Narrow exposed opening beyond the concealing corner in the upper-right of the frame, background; Inner wall shielding all five figures in the middle-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: 골목 벽 (Concealing the group from the militia beyond the corner) — The camera looks along its inner face at a shallow angle toward the corner; used as Forms the shared hiding boundary and a receding line through the five figures; 골목 모퉁이와 바깥으로 열린 틈 (Only a narrow view beyond the hiding place is available) — The corner interrupts the view at upper right, with the exposed passage visible beyond its edge; used as Makes the danger direction legible without revealing the group's bodies to it; 좁은 골목 바닥 (Visible under the group's feet); used as Preserves full-body readability and shows their uneven, compressed spacing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the nighttime alley dark but readable with restrained ambient separation, without carrying a specific reception light source into this location.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The settlement's restored lighting remains established, although this alley provides concealment behind a wall. At the reception, the tents, parked militia truck, jukebox and burst lamps remain; Charlie still wears his old coat and hat. 현우: He has taken cover behind the alley wall, still wearing his outer shirt. His facial bruises and leg injury persist. 앰버: She is hiding behind the alley wall with her mask and waist tool pouch retained. 라울: He is hiding behind the alley wall after running. 페드로: He is hiding behind the alley wall after running.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 좁고 어두운 골목길 벽에 바짝 기대어 선 채 밖의 눈치를 살피는 현우 일행의 전신.\n\nLOCATION (lock): Against a wall in a narrow, dark alley of the refugee settlement, away from the reception lot. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Narrow exposed opening beyond the concealing corner in the upper-right of the frame, background; Inner wall shielding all five figures in the middle-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: 골목 벽 (Concealing the group from the militia beyond the corner) — The camera looks along its inner face at a shallow angle toward the corner; used as Forms the shared hiding boundary and a receding line through the five figures; 골목 모퉁이와 바깥으로 열린 틈 (Only a narrow view beyond the hiding place is available) — The corner interrupts the view at upper right, with the exposed passage visible beyond its edge; used as Makes the danger direction legible without revealing the group's bodies to it; 좁은 골목 바닥 (Visible under the group's feet); used as Preserves full-body readability and shows their uneven, compressed spacing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the nighttime alley dark but readable with restrained ambient separation, without carrying a specific reception light source into this location.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The settlement's restored lighting remains established, although this alley provides concealment behind a wall. At the reception, the tents, parked militia truck, jukebox and burst lamps remain; Charlie still wears his old coat and hat. 현우: He has taken cover behind the alley wall, still wearing his outer shirt. His facial bruises and leg injury persist. 앰버: She is hiding behind the alley wall with her mask and waist tool pouch retained. 라울: He is hiding behind the alley wall after running. 페드로: He is hiding behind the alley wall after running.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 좁고 어두운 골목길 벽에 바짝 기대어 선 채 밖의 눈치를 살피는 현우 일행의 전신.\n\nLOCATION (lock): Against a wall in a narrow, dark alley of the refugee settlement, away from the reception lot. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Narrow exposed opening beyond the concealing corner in the upper-right of the frame, background; Inner wall shielding all five figures in the middle-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: 골목 벽 (Concealing the group from the militia beyond the corner) — The camera looks along its inner face at a shallow angle toward the corner; used as Forms the shared hiding boundary and a receding line through the five figures; 골목 모퉁이와 바깥으로 열린 틈 (Only a narrow view beyond the hiding place is available) — The corner interrupts the view at upper right, with the exposed passage visible beyond its edge; used as Makes the danger direction legible without revealing the group's bodies to it; 좁은 골목 바닥 (Visible under the group's feet); used as Preserves full-body readability and shows their uneven, compressed spacing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the nighttime alley dark but readable with restrained ambient separation, without carrying a specific reception light source into this location.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The settlement's restored lighting remains established, although this alley provides concealment behind a wall. At the reception, the tents, parked militia truck, jukebox and burst lamps remain; Charlie still wears his old coat and hat. 현우: He has taken cover behind the alley wall, still wearing his outer shirt. His facial bruises and leg injury persist. 앰버: She is hiding behind the alley wall with her mask and waist tool pouch retained. 라울: He is hiding behind the alley wall after running. 페드로: He is hiding behind the alley wall after running.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S25sh22__bgfirst_bg.png",
     "asset_id": "4ea5365d-bc0a-4f55-9bd5-c4357af778b0",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S25sh22.png",
     "asset_id": "ed35ed23-b843-4470-bb40-db03b4117fa0",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163202>",
     "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 페드로: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1278830>",
     "asset_id": "b09df655-64d4-4db4-a1b3-2f0bb5d29c95",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_outdoor_reception_sel.png",
     "asset_id": "eac29610-01b6-43a7-9e6e-9e2387ddd674",
     "role": "location_seed_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163202>",
     "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 페드로: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1278830>",
     "asset_id": "b09df655-64d4-4db4-a1b3-2f0bb5d29c95",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "인물들의 시선이 우측 상단 모퉁이 너머의 노출된 공간을 향함.",
    "built_space": "중앙의 내벽, 우측 상단의 노출된 공간, 발밑 바닥이 프레임 구도 지침과 정확히 일치함.",
    "entities": "현우(얼굴 상처, 다리 부상, 겉옷), 앰버(마스크, 공구 주머니), 찰리(코트, 모자), 페드로, 라울 5명이 모두 일치하게 묘사됨.",
    "hard_violations": [
     "[gemini-pro] 지정된 목록에 없는 인물(배경의 밀리샤 2명) 추가",
     "[gpt-high] 통로 끝에 지정된 다섯 인물 이외의 사람 두 명을 추가했다."
    ],
    "physics": "모든 인물이 발을 바닥에 딛고 벽에 기대거나 몸을 낮춘 안정적인 자세를 취함."
   },
   {
    "label": "B",
    "direction": "인물들의 시선이 우측 모퉁이 바깥쪽을 향함.",
    "built_space": "좌측 내벽과 우측 모퉁이, 바깥 골목이 묘사됨.",
    "entities": "현우, 앰버, 라울, 페드로가 보이나 찰리가 누락됨. 미지정 노년 여성과 군인들이 등장함.",
    "hard_violations": [
     "[gemini-pro] 지정된 목록에 없는 인물(우측 전경의 노년 여성 및 배경의 군인들) 추가",
     "[gpt-high] 지정 인물 찰리를 누락하고 명단에 없는 노년 여성을 전경에 추가했다.",
     "[gpt-high] 오른쪽 통로에 허용되지 않은 무장 인물들을 추가했다.",
     "[gpt-high] 다섯 명 모두를 가리는 벽 안쪽이 아니라 모퉁이 바깥 면에 전경 여성 한 명을 배치했다."
    ],
    "physics": "인물들이 바닥을 딛고 서서 벽에 기대어 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 5명의 인물과 소품 상태를 정확히 구현하고 구도 지침을 잘 따랐으나, 배경에 미지정 인물(밀리샤)이 묘사된 점이 유일한 흠결임."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "주요 인물인 찰리가 완전히 누락되었으며, 전경에 프롬프트에 없는 노년 여성이 크게 추가되어 인물 제약을 심각하게 위반함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "인물들의 시선이 우측 상단 모퉁이 너머의 노출된 공간을 향함.",
        "built_space": "중앙의 내벽, 우측 상단의 노출된 공간, 발밑 바닥이 프레임 구도 지침과 정확히 일치함.",
        "entities": "현우(얼굴 상처, 다리 부상, 겉옷), 앰버(마스크, 공구 주머니), 찰리(코트, 모자), 페드로, 라울 5명이 모두 일치하게 묘사됨.",
        "hard_violations": [
         "지정된 목록에 없는 인물(배경의 밀리샤 2명) 추가"
        ],
        "physics": "모든 인물이 발을 바닥에 딛고 벽에 기대거나 몸을 낮춘 안정적인 자세를 취함."
       },
       {
        "label": "B",
        "direction": "인물들의 시선이 우측 모퉁이 바깥쪽을 향함.",
        "built_space": "좌측 내벽과 우측 모퉁이, 바깥 골목이 묘사됨.",
        "entities": "현우, 앰버, 라울, 페드로가 보이나 찰리가 누락됨. 미지정 노년 여성과 군인들이 등장함.",
        "hard_violations": [
         "지정된 목록에 없는 인물(우측 전경의 노년 여성 및 배경의 군인들) 추가"
        ],
        "physics": "인물들이 바닥을 딛고 서서 벽에 기대어 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 5명의 인물과 소품 상태를 정확히 구현하고 구도 지침을 잘 따랐으나, 배경에 미지정 인물(밀리샤)이 묘사된 점이 유일한 흠결임."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "주요 인물인 찰리가 완전히 누락되었으며, 전경에 프롬프트에 없는 노년 여성이 크게 추가되어 인물 제약을 심각하게 위반함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "인물들의 시선이 우측 상단 모퉁이 너머의 노출된 공간을 향함.",
        "built_space": "중앙의 내벽, 우측 상단의 노출된 공간, 발밑 바닥이 프레임 구도 지침과 정확히 일치함.",
        "entities": "현우(얼굴 상처, 다리 부상, 겉옷), 앰버(마스크, 공구 주머니), 찰리(코트, 모자), 페드로, 라울 5명이 모두 일치하게 묘사됨.",
        "hard_violations": [
         "지정된 목록에 없는 인물(배경의 밀리샤 2명) 추가"
        ],
        "physics": "모든 인물이 발을 바닥에 딛고 벽에 기대거나 몸을 낮춘 안정적인 자세를 취함."
       },
       {
        "label": "B",
        "direction": "인물들의 시선이 우측 모퉁이 바깥쪽을 향함.",
        "built_space": "좌측 내벽과 우측 모퉁이, 바깥 골목이 묘사됨.",
        "entities": "현우, 앰버, 라울, 페드로가 보이나 찰리가 누락됨. 미지정 노년 여성과 군인들이 등장함.",
        "hard_violations": [
         "지정된 목록에 없는 인물(우측 전경의 노년 여성 및 배경의 군인들) 추가"
        ],
        "physics": "인물들이 바닥을 딛고 서서 벽에 기대어 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "찰리 대신 노년 여성을 넣고 배경 인물까지 추가했으며, 현우 등의 발을 잘라 전신 구도와 다섯 명의 공동 은폐 조건을 충족하지 못한다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "다섯 명의 전신, 벽을 따라 숨은 배치와 오른쪽 모퉁이 경계 행동은 훨씬 충실하지만, 금지된 배경 인물 두 명과 앰버의 성숙한 외형 때문에 완전한 적합작은 아니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 오른쪽 모퉁이 밖을 향해 얼굴과 시선을 내밀고, 페드로·라울·앰버도 대체로 같은 오른쪽을 살핀다. 오른쪽의 노년 여성은 바깥이 아니라 현우 쪽을 본다. 배경 무장 인물들은 골목의 카메라 쪽을 향하고 있으나 총구의 정확한 조준 대상은 판별하기 어렵다.",
        "built_space": "큰 콘크리트 벽 하나가 왼쪽에서 중앙 오른쪽 모퉁이까지 이어지고, 모퉁이에는 수직 배관 하나가 보인다. 그 너머 오른쪽에 좁은 통로와 작은 조명들이 있다. 네 명은 벽 안쪽에 붙어 있지만 노년 여성은 모퉁이 반대 면에 있어 같은 벽으로 함께 은폐되지 않는다. 참조의 낡은 회벽·배관 분위기는 일부 닮았으나 상부 판재 외벽과 긴 콘크리트 벽의 정확한 장소 대응은 확인되지 않는다. 인물들이 크게 잡혀 현우와 여성의 하체, 앰버의 발 일부가 잘린다.",
        "entities": "전경에는 다섯 명이 있지만 지정된 다섯 인물이 아니다. 현우는 앳된 동아시아계 남성, 검은 머리, 겉셔츠와 뺨의 멍이 보이며 다리 부상은 확인되지 않는다. 앰버는 금발 여자아이로 마스크와 허리 공구 주머니를 갖췄다. 라울은 머리를 묶은 어린 남자아이이며 페드로는 짙은 머리의 젊은 남성으로 대체로 해당 역할에 맞는다. 찰리의 장갑 몸체·흰 기계 얼굴·코트·모자는 없고 그 자리에 머릿수건을 쓴 노년 여성이 있다. 오른쪽 배경에는 지정 명단에 없는 무장 인물이 최소 세 명 보인다. 읽을 수 있는 문구는 확인되지 않는다.",
        "hard_violations": [
         "지정 인물 찰리를 누락하고 명단에 없는 노년 여성을 전경에 추가했다.",
         "오른쪽 통로에 허용되지 않은 무장 인물들을 추가했다.",
         "다섯 명 모두를 가리는 벽 안쪽이 아니라 모퉁이 바깥 면에 전경 여성 한 명을 배치했다."
        ],
        "physics": "페드로와 라울은 바닥에 닿은 신발로 체중을 지탱하며 몸을 기울인다. 현우는 벽에 손을 대고 다리를 벌려 몸을 지탱하는 자세이고, 잘린 발 때문에 접지 전체는 확인할 수 없다. 여성도 벽에 손을 대고 서 있으며 발은 화면 밖이다. 앰버의 주머니는 허리에 부착되어 있다. 공중에 떠 있거나 지지 없이 매달린 몸과 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "맨 앞 현우는 오른쪽 모퉁이 밖의 통로를 향해 고개를 돌리고 몸은 벽 안쪽에 남겨 둔다. 앰버와 찰리, 뒤의 페드로와 라울도 오른쪽 선두와 모퉁이 방향을 살핀다. 멀리 있는 두 사람은 대체로 카메라 쪽 통로를 향해 서 있다. 뚜렷하게 조준하는 무기나 손에 든 방향성 소품은 없다.",
        "built_space": "긴 회벽 하나가 다섯 명 뒤에서 오른쪽 모퉁이로 이어지고, 그 끝 너머에 좁은 외부 통로가 보인다. 왼쪽에는 벽돌 건물, 수직 배관과 사각 설비함이 각각 하나씩 보이며, 오른쪽 통로에는 가까운 가로등 하나와 더 먼 작은 등 하나가 있다. 다섯 명 모두 같은 벽의 안쪽에 붙고 바닥과 발까지 드러나 전신 와이드 구도가 성립한다. 참조의 벽돌·낡은 회벽·전선·가로등이라는 재료와 설비 계열은 유지하지만 긴 은폐벽과 포장 바닥의 정확한 구조적 일치는 확인되지 않는다. 노을과 어두운 은폐 공간의 대비가 보인다.",
        "entities": "현우·앰버·찰리·페드로·라울로 읽히는 다섯 인물이 있다. 현우는 검은 머리의 젊은 동아시아계 남성이며 겉셔츠, 얼굴 상처와 무릎 부위의 피가 보인다. 앰버는 금발과 마스크, 허리 주머니를 유지하지만 키와 체형이 참조의 10세 여자아이보다 상당히 성숙해 보인다. 찰리는 베이지 장갑, 흰 기계 얼굴, 낡은 코트와 모자를 갖췄으나 참조보다 어깨와 팔의 육중함이 줄었다. 라울은 어두운 피부와 뒤로 묶은 머리의 어린 남자아이로 보인다. 페드로는 짙은 머리의 젊은 남성이지만 머리를 뒤로 묶어 참조와 다르다. 통로 끝에는 추가 인물 두 명과 작은 트럭이 보인다. 읽을 수 있는 문구는 없다.",
        "hard_violations": [
         "통로 끝에 지정된 다섯 인물 이외의 사람 두 명을 추가했다."
        ],
        "physics": "다섯 명 모두 신발 또는 기계 발이 바닥에 닿아 있다. 라울과 페드로는 무릎을 굽히고 손을 허벅지에 얹어 달린 뒤 몸을 낮춘 자세를 지탱한다. 현우는 두 발을 벌리고 벽 가까이 몸을 기울이며, 앰버도 두 발로 서서 상체를 낮춘다. 찰리의 무게는 넓은 두 기계 발이 받치고 코트는 몸에서 자연스럽게 늘어진다. 공구 주머니는 허리띠에 고정되어 있으며 지지 없는 부유나 불가능한 관절 자세는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "찰리 대신 노년 여성을 넣고 배경 인물까지 추가했으며, 현우 등의 발을 잘라 전신 구도와 다섯 명의 공동 은폐 조건을 충족하지 못한다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "다섯 명의 전신, 벽을 따라 숨은 배치와 오른쪽 모퉁이 경계 행동은 훨씬 충실하지만, 금지된 배경 인물 두 명과 앰버의 성숙한 외형 때문에 완전한 적합작은 아니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 오른쪽 모퉁이 밖을 향해 얼굴과 시선을 내밀고, 페드로·라울·앰버도 대체로 같은 오른쪽을 살핀다. 오른쪽의 노년 여성은 바깥이 아니라 현우 쪽을 본다. 배경 무장 인물들은 골목의 카메라 쪽을 향하고 있으나 총구의 정확한 조준 대상은 판별하기 어렵다.",
        "built_space": "큰 콘크리트 벽 하나가 왼쪽에서 중앙 오른쪽 모퉁이까지 이어지고, 모퉁이에는 수직 배관 하나가 보인다. 그 너머 오른쪽에 좁은 통로와 작은 조명들이 있다. 네 명은 벽 안쪽에 붙어 있지만 노년 여성은 모퉁이 반대 면에 있어 같은 벽으로 함께 은폐되지 않는다. 참조의 낡은 회벽·배관 분위기는 일부 닮았으나 상부 판재 외벽과 긴 콘크리트 벽의 정확한 장소 대응은 확인되지 않는다. 인물들이 크게 잡혀 현우와 여성의 하체, 앰버의 발 일부가 잘린다.",
        "entities": "전경에는 다섯 명이 있지만 지정된 다섯 인물이 아니다. 현우는 앳된 동아시아계 남성, 검은 머리, 겉셔츠와 뺨의 멍이 보이며 다리 부상은 확인되지 않는다. 앰버는 금발 여자아이로 마스크와 허리 공구 주머니를 갖췄다. 라울은 머리를 묶은 어린 남자아이이며 페드로는 짙은 머리의 젊은 남성으로 대체로 해당 역할에 맞는다. 찰리의 장갑 몸체·흰 기계 얼굴·코트·모자는 없고 그 자리에 머릿수건을 쓴 노년 여성이 있다. 오른쪽 배경에는 지정 명단에 없는 무장 인물이 최소 세 명 보인다. 읽을 수 있는 문구는 확인되지 않는다.",
        "hard_violations": [
         "지정 인물 찰리를 누락하고 명단에 없는 노년 여성을 전경에 추가했다.",
         "오른쪽 통로에 허용되지 않은 무장 인물들을 추가했다.",
         "다섯 명 모두를 가리는 벽 안쪽이 아니라 모퉁이 바깥 면에 전경 여성 한 명을 배치했다."
        ],
        "physics": "페드로와 라울은 바닥에 닿은 신발로 체중을 지탱하며 몸을 기울인다. 현우는 벽에 손을 대고 다리를 벌려 몸을 지탱하는 자세이고, 잘린 발 때문에 접지 전체는 확인할 수 없다. 여성도 벽에 손을 대고 서 있으며 발은 화면 밖이다. 앰버의 주머니는 허리에 부착되어 있다. 공중에 떠 있거나 지지 없이 매달린 몸과 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "맨 앞 현우는 오른쪽 모퉁이 밖의 통로를 향해 고개를 돌리고 몸은 벽 안쪽에 남겨 둔다. 앰버와 찰리, 뒤의 페드로와 라울도 오른쪽 선두와 모퉁이 방향을 살핀다. 멀리 있는 두 사람은 대체로 카메라 쪽 통로를 향해 서 있다. 뚜렷하게 조준하는 무기나 손에 든 방향성 소품은 없다.",
        "built_space": "긴 회벽 하나가 다섯 명 뒤에서 오른쪽 모퉁이로 이어지고, 그 끝 너머에 좁은 외부 통로가 보인다. 왼쪽에는 벽돌 건물, 수직 배관과 사각 설비함이 각각 하나씩 보이며, 오른쪽 통로에는 가까운 가로등 하나와 더 먼 작은 등 하나가 있다. 다섯 명 모두 같은 벽의 안쪽에 붙고 바닥과 발까지 드러나 전신 와이드 구도가 성립한다. 참조의 벽돌·낡은 회벽·전선·가로등이라는 재료와 설비 계열은 유지하지만 긴 은폐벽과 포장 바닥의 정확한 구조적 일치는 확인되지 않는다. 노을과 어두운 은폐 공간의 대비가 보인다.",
        "entities": "현우·앰버·찰리·페드로·라울로 읽히는 다섯 인물이 있다. 현우는 검은 머리의 젊은 동아시아계 남성이며 겉셔츠, 얼굴 상처와 무릎 부위의 피가 보인다. 앰버는 금발과 마스크, 허리 주머니를 유지하지만 키와 체형이 참조의 10세 여자아이보다 상당히 성숙해 보인다. 찰리는 베이지 장갑, 흰 기계 얼굴, 낡은 코트와 모자를 갖췄으나 참조보다 어깨와 팔의 육중함이 줄었다. 라울은 어두운 피부와 뒤로 묶은 머리의 어린 남자아이로 보인다. 페드로는 짙은 머리의 젊은 남성이지만 머리를 뒤로 묶어 참조와 다르다. 통로 끝에는 추가 인물 두 명과 작은 트럭이 보인다. 읽을 수 있는 문구는 없다.",
        "hard_violations": [
         "통로 끝에 지정된 다섯 인물 이외의 사람 두 명을 추가했다."
        ],
        "physics": "다섯 명 모두 신발 또는 기계 발이 바닥에 닿아 있다. 라울과 페드로는 무릎을 굽히고 손을 허벅지에 얹어 달린 뒤 몸을 낮춘 자세를 지탱한다. 현우는 두 발을 벌리고 벽 가까이 몸을 기울이며, 앰버도 두 발로 서서 상체를 낮춘다. 찰리의 무게는 넓은 두 기계 발이 받치고 코트는 몸에서 자연스럽게 늘어진다. 공구 주머니는 허리띠에 고정되어 있으며 지지 없는 부유나 불가능한 관절 자세는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.829
   },
   "adjusted": {
    "A": 1.75,
    "B": 0.579
   },
   "violations": {
    "A": [
     "[gemini-pro] 지정된 목록에 없는 인물(배경의 밀리샤 2명) 추가",
     "[gpt-high] 통로 끝에 지정된 다섯 인물 이외의 사람 두 명을 추가했다."
    ],
    "B": [
     "[gemini-pro] 지정된 목록에 없는 인물(우측 전경의 노년 여성 및 배경의 군인들) 추가",
     "[gpt-high] 지정 인물 찰리를 누락하고 명단에 없는 노년 여성을 전경에 추가했다.",
     "[gpt-high] 오른쪽 통로에 허용되지 않은 무장 인물들을 추가했다.",
     "[gpt-high] 다섯 명 모두를 가리는 벽 안쪽이 아니라 모퉁이 바깥 면에 전경 여성 한 명을 배치했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 579
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "지정된 5명의 인물과 소품 상태를 정확히 구현하고 구도 지침을 잘 따랐으나, 배경에 미지정 인물(밀리샤)이 묘사된 점이 유일한 흠결임.  ★위반: [gemini-pro] 지정된 목록에 없는 인물(배경의 밀리샤 2명) 추가 / [gpt-high] 통로 끝에 지정된 다섯 인물 이외의 사람 두 명을 추가했다."
   },
   {
    "label": "B",
    "score": 579,
    "verdict_ko": "주요 인물인 찰리가 완전히 누락되었으며, 전경에 프롬프트에 없는 노년 여성이 크게 추가되어 인물 제약을 심각하게 위반함.  ★위반: [gemini-pro] 지정된 목록에 없는 인물(우측 전경의 노년 여성 및 배경의 군인들) 추가 / [gpt-high] 지정 인물 찰리를 누락하고 명단에 없는 노년 여성을 전경에 추가했다. / [gpt-high] 오른쪽 통로에 허용되지 않은 무장 인물들을 추가했다. / [gpt-high] 다섯 명 모두를 가리는 벽 안쪽이 아니라 모퉁이 바깥 면에 전경 여성 한 명을 배치했다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_outdoor_reception_sel.png",
    "asset_id": "eac29610-01b6-43a7-9e6e-9e2387ddd674",
    "role": "location_seed_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163202>",
    "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 페드로: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1278830>",
    "asset_id": "b09df655-64d4-4db4-a1b3-2f0bb5d29c95",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-2533-72bd-abff-de3543beae67",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S25sh22__bgfirst_bg.png",
   "bg_asset_id": "4ea5365d-bc0a-4f55-9bd5-c4357af778b0",
   "bg_record_key": "S25sh22::bgfirst_bg",
   "chain_winner": true,
   "authority": "seed_bg"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  },
  "lane_policy": "ab_select_ready"
 },
 "S25sh22::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:14:14.167795+00:00",
  "fingerprint": "a7c1f6843535118e391b31a1df02eef28c79878298bde766a36199e7c665ea7a",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S25sh22_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S25sh22_sel.png",
  "source_sha256": "1cc0608956ee036ce48b93f3927d395fc9b943cfa88433804a70541135abd23f",
  "file": "S25sh22_cine.png",
  "staged_sha256": "d6b0faee0d49ea711edc2be70a4894ad23f8e5189df5235dd533c00764989d7d",
  "latency_ms": 139769
 },
 "S26sh7::signage": {
  "fp": "9f818f7c8ea421b3",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S26sh7": {
  "input_fingerprint": "ac723983b41c8017",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 무리 속에 섞인 미연을 향해 서늘한 눈빛을 번뜩이는 박철진의 비열한 상체.\n\nLOCATION (lock): In the open reception lot at night, beside the gathered guests being held under militia guard. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 피로연장 중앙 공터 (The guests have been gathered together under threat); used as Leaves a readable depth interval between the interrogation and 미연 at the front of the gathered guests.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the reception's established nighttime illumination with controlled contrast that keeps both 박철진's glance and 미연's alarm readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The reception remains brightly lit, with tents, the jukebox and burst lamps still in place. The rubbish-covered sewer entrance has had its manhole cover pushed aside, and Charlie's coat-and-hat disguise remains unchanged below ground. 박철진: He remains in the reception clearing conducting the interrogation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 무리 속에 섞인 미연을 향해 서늘한 눈빛을 번뜩이는 박철진의 비열한 상체.\n\nLOCATION (lock): In the open reception lot at night, beside the gathered guests being held under militia guard. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 피로연장 중앙 공터 (The guests have been gathered together under threat); used as Leaves a readable depth interval between the interrogation and 미연 at the front of the gathered guests.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the reception's established nighttime illumination with controlled contrast that keeps both 박철진's glance and 미연's alarm readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The reception remains brightly lit, with tents, the jukebox and burst lamps still in place. The rubbish-covered sewer entrance has had its manhole cover pushed aside, and Charlie's coat-and-hat disguise remains unchanged below ground. 박철진: He remains in the reception clearing conducting the interrogation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 무리 속에 섞인 미연을 향해 서늘한 눈빛을 번뜩이는 박철진의 비열한 상체.\n\nLOCATION (lock): In the open reception lot at night, beside the gathered guests being held under militia guard. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 피로연장 중앙 공터 (The guests have been gathered together under threat); used as Leaves a readable depth interval between the interrogation and 미연 at the front of the gathered guests.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the reception's established nighttime illumination with controlled contrast that keeps both 박철진's glance and 미연's alarm readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The reception remains brightly lit, with tents, the jukebox and burst lamps still in place. The rubbish-covered sewer entrance has had its manhole cover pushed aside, and Charlie's coat-and-hat disguise remains unchanged below ground. 박철진: He remains in the reception clearing conducting the interrogation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "박철진이 화면 왼쪽의 무리를 향해 시선을 던지고 있으나, 무리 맨 앞의 미연을 정확히 응시하지 않고 시선이 약간 빗겨나 있습니다.",
    "built_space": "이전 샷의 피로연장 공터, 천막, 조명, 트럭 등의 요소가 배치되어 있으나 인물 배치와 겹치며 다소 평면적으로 보입니다.",
    "entities": "박철진의 외모는 일치하나, 프롬프트에서 금지한 이전 샷의 배경 인물(트럭 우측 민병대원들)이 복장과 얼굴까지 똑같이 복제되어 등장했습니다.",
    "hard_violations": [
     "[gemini-pro] 이전 샷 레퍼런스에 있던 경비병 인물들이 얼굴과 복장이 복제되어 재등장함 (금지된 인물 포함 및 복제)",
     "[gpt-high] 이전 사진에서 박철진 외 인물은 재등장시키지 말라는 명시적 제한에도, 트럭 앞 경비 두 명의 얼굴과 군복·전술조끼 조합을 이어서 배치했다."
    ],
    "physics": "인물들은 지면에 서 있으며 특별히 물리적으로 불가능하거나 떠 있는 요소는 발견되지 않습니다."
   },
   {
    "label": "B",
    "direction": "박철진의 시선이 화면 오른쪽 무리 맨 앞에 선 미연의 얼굴을 정확하게 향하고 있으며, 미연 또한 그를 마주보며 시선이 교차합니다.",
    "built_space": "야외 피로연장 공터, 천막, 줄조명, 드럼통 등 이전 샷의 구조물들이 올바른 심도와 비례를 유지하며 배치되어 있습니다.",
    "entities": "박철진의 외형과 복장이 레퍼런스와 정확히 일치하며, 다국적/다인종으로 구성된 무리와 한국인 미연의 묘사가 프롬프트 조건에 부합합니다. 이전 샷의 인물은 등장하지 않습니다.",
    "hard_violations": [],
    "physics": "모든 인물들이 지면에 체중을 싣고 서 있으며, 부자연스럽게 떠 있거나 지지되지 않은 요소는 없습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "레퍼런스 이미지의 배경 인물(민병대원들)을 얼굴과 포즈까지 그대로 복제하여 등장시킨 치명적인 위반이 있으며, 시선 처리도 다소 부정확합니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "박철진이 미연을 노려보는 서늘한 시선과 미연의 놀란 표정이 정확한 심도로 잘 연출되었으며, 요구된 다국적 무리와 배경 세팅을 충실히 반영했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진이 화면 왼쪽의 무리를 향해 시선을 던지고 있으나, 무리 맨 앞의 미연을 정확히 응시하지 않고 시선이 약간 빗겨나 있습니다.",
        "built_space": "이전 샷의 피로연장 공터, 천막, 조명, 트럭 등의 요소가 배치되어 있으나 인물 배치와 겹치며 다소 평면적으로 보입니다.",
        "entities": "박철진의 외모는 일치하나, 프롬프트에서 금지한 이전 샷의 배경 인물(트럭 우측 민병대원들)이 복장과 얼굴까지 똑같이 복제되어 등장했습니다.",
        "hard_violations": [
         "이전 샷 레퍼런스에 있던 경비병 인물들이 얼굴과 복장이 복제되어 재등장함 (금지된 인물 포함 및 복제)"
        ],
        "physics": "인물들은 지면에 서 있으며 특별히 물리적으로 불가능하거나 떠 있는 요소는 발견되지 않습니다."
       },
       {
        "label": "B",
        "direction": "박철진의 시선이 화면 오른쪽 무리 맨 앞에 선 미연의 얼굴을 정확하게 향하고 있으며, 미연 또한 그를 마주보며 시선이 교차합니다.",
        "built_space": "야외 피로연장 공터, 천막, 줄조명, 드럼통 등 이전 샷의 구조물들이 올바른 심도와 비례를 유지하며 배치되어 있습니다.",
        "entities": "박철진의 외형과 복장이 레퍼런스와 정확히 일치하며, 다국적/다인종으로 구성된 무리와 한국인 미연의 묘사가 프롬프트 조건에 부합합니다. 이전 샷의 인물은 등장하지 않습니다.",
        "hard_violations": [],
        "physics": "모든 인물들이 지면에 체중을 싣고 서 있으며, 부자연스럽게 떠 있거나 지지되지 않은 요소는 없습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "레퍼런스 이미지의 배경 인물(민병대원들)을 얼굴과 포즈까지 그대로 복제하여 등장시킨 치명적인 위반이 있으며, 시선 처리도 다소 부정확합니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "박철진이 미연을 노려보는 서늘한 시선과 미연의 놀란 표정이 정확한 심도로 잘 연출되었으며, 요구된 다국적 무리와 배경 세팅을 충실히 반영했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "박철진이 화면 왼쪽의 무리를 향해 시선을 던지고 있으나, 무리 맨 앞의 미연을 정확히 응시하지 않고 시선이 약간 빗겨나 있습니다.",
        "built_space": "이전 샷의 피로연장 공터, 천막, 조명, 트럭 등의 요소가 배치되어 있으나 인물 배치와 겹치며 다소 평면적으로 보입니다.",
        "entities": "박철진의 외모는 일치하나, 프롬프트에서 금지한 이전 샷의 배경 인물(트럭 우측 민병대원들)이 복장과 얼굴까지 똑같이 복제되어 등장했습니다.",
        "hard_violations": [
         "이전 샷 레퍼런스에 있던 경비병 인물들이 얼굴과 복장이 복제되어 재등장함 (금지된 인물 포함 및 복제)"
        ],
        "physics": "인물들은 지면에 서 있으며 특별히 물리적으로 불가능하거나 떠 있는 요소는 발견되지 않습니다."
       },
       {
        "label": "B",
        "direction": "박철진의 시선이 화면 오른쪽 무리 맨 앞에 선 미연의 얼굴을 정확하게 향하고 있으며, 미연 또한 그를 마주보며 시선이 교차합니다.",
        "built_space": "야외 피로연장 공터, 천막, 줄조명, 드럼통 등 이전 샷의 구조물들이 올바른 심도와 비례를 유지하며 배치되어 있습니다.",
        "entities": "박철진의 외형과 복장이 레퍼런스와 정확히 일치하며, 다국적/다인종으로 구성된 무리와 한국인 미연의 묘사가 프롬프트 조건에 부합합니다. 이전 샷의 인물은 등장하지 않습니다.",
        "hard_violations": [],
        "physics": "모든 인물들이 지면에 체중을 싣고 서 있으며, 부자연스럽게 떠 있거나 지지되지 않은 요소는 없습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "박철진의 상체 중심 미디엄 숏과 미연의 불안은 잘 보이지만, 옆눈질이 뒤쪽 미연에게 실제로 닿는 시선 연결은 불명확하다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "공터의 깊이는 읽히지만 박철진은 뒤의 미연이 아닌 화면 왼쪽 앞을 바라보며, 이전 사진의 제외 대상 경비 인물도 다시 등장한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진은 얼굴을 화면 오른쪽으로 돌리고 눈도 오른쪽을 본다. 미연으로 읽히는 여성은 그의 오른쪽 뒤에 서서 박철진 쪽을 불안하게 바라본다. 그러나 박철진의 눈은 여성의 얼굴보다 카메라에 가까운 오른쪽 공간을 향하는 것으로 보여, 뒤쪽 여성에게 정확히 꽂히는 시선으로 확정하기 어렵다. 군중은 대체로 박철진 또는 카메라 쪽을 본다. 보이는 조준 무기나 이동 동작은 없다.",
        "built_space": "왼쪽에 두 봉우리의 천막 구역, 그 아래 여러 원형 식탁과 접이식 의자, 전구 줄과 기둥형 조명이 있다. 왼쪽 아래에는 드럼통 두 개, 그 뒤에는 드럼통과 사각 설비가 보인다. 오른쪽 끝에는 천막 일부가 있다. 흙과 웅덩이로 된 공터의 재질은 이전 장소와 잘 맞는다. 박철진은 전경 중앙, 여성과 군중은 오른쪽 뒤에 있지만 여성과 박철진 사이의 깊이 간격은 좁다. 인물과 시설이 불가능하게 겹치거나 중복된 고정 설비는 보이지 않는다.",
        "entities": "전경 남성은 중년 한국인 남성으로 제시된 박철진에 부합하는 외모이며, 짧게 뒤로 넘긴 검은 머리와 올리브색 다중 주머니 야전 재킷이 참조와 대체로 일치한다. 오른쪽 앞줄에는 검은 재킷을 입은 성인 여성이 있어 미연의 역할이 명확하고, 뒤에는 서로 다른 외모의 성인 하객들이 모여 있다. 이 여성의 별도 신원 참조는 없어 정확한 얼굴 일치 여부는 판단할 수 없다. 뚜렷하게 식별되는 민병대 경비는 없다. 주크박스, 하수구와 지하 변장은 보이지 않으나 이 상체 중심 구도에서 반드시 보여야 할 대상은 아니다. 읽을 수 있는 글자나 초자연적으로 변형된 눈은 없다.",
        "hard_violations": [],
        "physics": "박철진의 몸통과 팔은 자연스럽게 아래로 이어지며 하체는 프레임 밖이다. 여성도 팔을 내리고 정상적으로 서 있는 자세이고, 뒤 인물들의 보이는 다리는 공터 바닥으로 이어진다. 발이 잘린 인물에게 부유를 시사하는 징후는 없다. 식탁과 의자는 다리로, 천막은 기둥과 줄로 지지된다. 공중에 뜬 물건이나 지지 없는 신체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "박철진은 화면 왼쪽의 카메라 가까운 방향을 응시한다. 미연으로 추정할 수 있는 앞줄 여성들은 그의 왼쪽 뒤, 군중 속에 있으므로 이 시선은 그들에게 향하지 않는다. 앞줄 여성과 하객들은 전경 또는 카메라 쪽을 바라본다. 따라서 박철진이 무리 속 미연을 노려본다는 핵심 행동이 성립하지 않는다. 명확히 조준하는 총구는 보이지 않는다.",
        "built_space": "왼쪽에 두 봉우리의 천막과 전구 줄, 그 아래 식탁과 의자들이 있고, 왼쪽 아래 드럼통 두 개와 그 뒤 사각 설비가 보인다. 오른쪽에는 덮개 달린 군용 트럭 한 대와 맨 오른쪽 아래 꽃 장식 식탁 일부가 남아 있어 이전 장소의 배치를 강하게 이어간다. 군중은 공터 중경에 밀집하고 박철진은 오른쪽 전경에 크게 배치되어 깊이 간격이 읽힌다. 다만 박철진의 허리 부근까지 보여 주는 미디엄 숏보다는 가슴 중심으로 타이트한 구도다.",
        "entities": "박철진은 짧은 검은 머리의 중년 한국인 남성 외모와 올리브색 야전 재킷을 유지한다. 중경 앞줄에는 갈색 재킷 여성과 회색 후드 위 어두운 재킷을 입은 여성이 함께 있어 어느 쪽이 미연인지 덜 명확하다. 하객 무리는 구체적인 사람들로 표현되어 있다. 트럭 앞 맨 오른쪽의 전술조끼 경비와 박철진 오른쪽 어깨 뒤의 경비는 이전 사진의 동일 위치 경비들의 얼굴·복장 조합을 다시 사용한 것으로 보인다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "이전 사진에서 박철진 외 인물은 재등장시키지 말라는 명시적 제한에도, 트럭 앞 경비 두 명의 얼굴과 군복·전술조끼 조합을 이어서 배치했다."
        ],
        "physics": "중경 하객들의 발과 신발은 흙바닥에 닿아 있고 각자의 몸을 지지한다. 전경 박철진의 하체는 잘렸지만 몸통 자세에 부유나 비정상적인 지지 관계는 없다. 천막은 기둥으로 지지되고 드럼통과 식탁은 바닥에 놓여 있다. 트럭과 인물 사이에도 명백한 관통이나 불가능한 접촉은 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "박철진의 상체 중심 미디엄 숏과 미연의 불안은 잘 보이지만, 옆눈질이 뒤쪽 미연에게 실제로 닿는 시선 연결은 불명확하다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "공터의 깊이는 읽히지만 박철진은 뒤의 미연이 아닌 화면 왼쪽 앞을 바라보며, 이전 사진의 제외 대상 경비 인물도 다시 등장한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "박철진은 얼굴을 화면 오른쪽으로 돌리고 눈도 오른쪽을 본다. 미연으로 읽히는 여성은 그의 오른쪽 뒤에 서서 박철진 쪽을 불안하게 바라본다. 그러나 박철진의 눈은 여성의 얼굴보다 카메라에 가까운 오른쪽 공간을 향하는 것으로 보여, 뒤쪽 여성에게 정확히 꽂히는 시선으로 확정하기 어렵다. 군중은 대체로 박철진 또는 카메라 쪽을 본다. 보이는 조준 무기나 이동 동작은 없다.",
        "built_space": "왼쪽에 두 봉우리의 천막 구역, 그 아래 여러 원형 식탁과 접이식 의자, 전구 줄과 기둥형 조명이 있다. 왼쪽 아래에는 드럼통 두 개, 그 뒤에는 드럼통과 사각 설비가 보인다. 오른쪽 끝에는 천막 일부가 있다. 흙과 웅덩이로 된 공터의 재질은 이전 장소와 잘 맞는다. 박철진은 전경 중앙, 여성과 군중은 오른쪽 뒤에 있지만 여성과 박철진 사이의 깊이 간격은 좁다. 인물과 시설이 불가능하게 겹치거나 중복된 고정 설비는 보이지 않는다.",
        "entities": "전경 남성은 중년 한국인 남성으로 제시된 박철진에 부합하는 외모이며, 짧게 뒤로 넘긴 검은 머리와 올리브색 다중 주머니 야전 재킷이 참조와 대체로 일치한다. 오른쪽 앞줄에는 검은 재킷을 입은 성인 여성이 있어 미연의 역할이 명확하고, 뒤에는 서로 다른 외모의 성인 하객들이 모여 있다. 이 여성의 별도 신원 참조는 없어 정확한 얼굴 일치 여부는 판단할 수 없다. 뚜렷하게 식별되는 민병대 경비는 없다. 주크박스, 하수구와 지하 변장은 보이지 않으나 이 상체 중심 구도에서 반드시 보여야 할 대상은 아니다. 읽을 수 있는 글자나 초자연적으로 변형된 눈은 없다.",
        "hard_violations": [],
        "physics": "박철진의 몸통과 팔은 자연스럽게 아래로 이어지며 하체는 프레임 밖이다. 여성도 팔을 내리고 정상적으로 서 있는 자세이고, 뒤 인물들의 보이는 다리는 공터 바닥으로 이어진다. 발이 잘린 인물에게 부유를 시사하는 징후는 없다. 식탁과 의자는 다리로, 천막은 기둥과 줄로 지지된다. 공중에 뜬 물건이나 지지 없는 신체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "박철진은 화면 왼쪽의 카메라 가까운 방향을 응시한다. 미연으로 추정할 수 있는 앞줄 여성들은 그의 왼쪽 뒤, 군중 속에 있으므로 이 시선은 그들에게 향하지 않는다. 앞줄 여성과 하객들은 전경 또는 카메라 쪽을 바라본다. 따라서 박철진이 무리 속 미연을 노려본다는 핵심 행동이 성립하지 않는다. 명확히 조준하는 총구는 보이지 않는다.",
        "built_space": "왼쪽에 두 봉우리의 천막과 전구 줄, 그 아래 식탁과 의자들이 있고, 왼쪽 아래 드럼통 두 개와 그 뒤 사각 설비가 보인다. 오른쪽에는 덮개 달린 군용 트럭 한 대와 맨 오른쪽 아래 꽃 장식 식탁 일부가 남아 있어 이전 장소의 배치를 강하게 이어간다. 군중은 공터 중경에 밀집하고 박철진은 오른쪽 전경에 크게 배치되어 깊이 간격이 읽힌다. 다만 박철진의 허리 부근까지 보여 주는 미디엄 숏보다는 가슴 중심으로 타이트한 구도다.",
        "entities": "박철진은 짧은 검은 머리의 중년 한국인 남성 외모와 올리브색 야전 재킷을 유지한다. 중경 앞줄에는 갈색 재킷 여성과 회색 후드 위 어두운 재킷을 입은 여성이 함께 있어 어느 쪽이 미연인지 덜 명확하다. 하객 무리는 구체적인 사람들로 표현되어 있다. 트럭 앞 맨 오른쪽의 전술조끼 경비와 박철진 오른쪽 어깨 뒤의 경비는 이전 사진의 동일 위치 경비들의 얼굴·복장 조합을 다시 사용한 것으로 보인다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "이전 사진에서 박철진 외 인물은 재등장시키지 말라는 명시적 제한에도, 트럭 앞 경비 두 명의 얼굴과 군복·전술조끼 조합을 이어서 배치했다."
        ],
        "physics": "중경 하객들의 발과 신발은 흙바닥에 닿아 있고 각자의 몸을 지지한다. 전경 박철진의 하체는 잘렸지만 몸통 자세에 부유나 비정상적인 지지 관계는 없다. 천막은 기둥으로 지지되고 드럼통과 식탁은 바닥에 놓여 있다. 트럭과 인물 사이에도 명백한 관통이나 불가능한 접촉은 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.829,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.579,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 이전 샷 레퍼런스에 있던 경비병 인물들이 얼굴과 복장이 복제되어 재등장함 (금지된 인물 포함 및 복제)",
     "[gpt-high] 이전 사진에서 박철진 외 인물은 재등장시키지 말라는 명시적 제한에도, 트럭 앞 경비 두 명의 얼굴과 군복·전술조끼 조합을 이어서 배치했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "A": 579,
   "B": 2000
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 579,
    "verdict_ko": "레퍼런스 이미지의 배경 인물(민병대원들)을 얼굴과 포즈까지 그대로 복제하여 등장시킨 치명적인 위반이 있으며, 시선 처리도 다소 부정확합니다.  ★위반: [gemini-pro] 이전 샷 레퍼런스에 있던 경비병 인물들이 얼굴과 복장이 복제되어 재등장함 (금지된 인물 포함 및 복제) / [gpt-high] 이전 사진에서 박철진 외 인물은 재등장시키지 말라는 명시적 제한에도, 트럭 앞 경비 두 명의 얼굴과 군복·전술조끼 조합을 이어서 배치했다."
   },
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "박철진이 미연을 노려보는 서늘한 시선과 미연의 놀란 표정이 정확한 심도로 잘 연출되었으며, 요구된 다국적 무리와 배경 세팅을 충실히 반영했습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S25sh19_sel.png",
    "asset_id": "3ecfcb55-eefa-4a9c-a0e2-a00674d50faf",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1401722>",
    "asset_id": "fee7383c-fb61-4b3a-ba7c-79f2555de00b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-28d3-7586-82f1-be875d327bf9",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S25sh19"
  }
 },
 "S26sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:16:00.364900+00:00",
  "fingerprint": "8fcd187b48d13803dadbf928f50f7947fde3b647711f0a414ffb4d77fe1b0de5",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S26sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S26sh7_sel.png",
  "source_sha256": "3ecad0ddf5ef433d531b544e77bd37c7f9aa2d2ba96ae44166c9f01352697a73",
  "file": "S26sh7_cine.png",
  "staged_sha256": "f6705f2b2e5ef8d395e9d239152d97793e2355d692edb351d48a61b5aad103c3",
  "latency_ms": 9300
 },
 "S26sh9::signage": {
  "fp": "9a54d860a83e39c2",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::bbb84cc871e390a4": {
  "subjects": [],
  "subject_text": "인천 난민촌 지하 하수도\n맨홀 아래로 사다리가 이어지는 어두운 지하 배수 통로. 길게 뻗은 통로 중간에 두 갈래로 나뉘는 분기점이 있다.",
  "identity": "canonical",
  "scope_id": "L175",
  "scope_role": "location_interior",
  "scope_sha": "86e38b0060fbf724"
 },
 "S26sh9::bgfirst_bg": {
  "input_fingerprint": "54c6631247e2cdb8",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 어두운 하수도 두 갈래 길 앞에서 단호한 표정으로 페드로의 어깨를 꽉 쥔 현우의 상체.\n\nLOCATION (lock): At a two-way junction inside the refugee settlement's underground sewer, where both branches recede into darkness.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Sewer fork (The passage divides into two routes) — The two branch entrances are visible obliquely behind the interaction; used as Provides the spatial evidence for the decision to separate.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the sewer dark, with restrained contrast preserving the expression and gripping hand without specifying an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 어두운 하수도 두 갈래 길 앞에서 단호한 표정으로 페드로의 어깨를 꽉 쥔 현우의 상체.\n\nLOCATION (lock): At a two-way junction inside the refugee settlement's underground sewer, where both branches recede into darkness.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Sewer fork (The passage divides into two routes) — The two branch entrances are visible obliquely behind the interaction; used as Provides the spatial evidence for the decision to separate.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the sewer dark, with restrained contrast preserving the expression and gripping hand without specifying an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S26sh9__bgfirst_bg.png",
  "asset_id": "ff1b1858-e263-49dc-901d-bd8f3ca0ac13",
  "input_asset_ids": [
   "bad5b160-fb00-401b-90e3-64652fe59154",
   "aacb72f4-e4d3-4331-8b05-48ae8f9cc866"
  ]
 },
 "S26sh9": {
  "input_fingerprint": "f7906e31af313fdc",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 어두운 하수도 두 갈래 길 앞에서 단호한 표정으로 페드로의 어깨를 꽉 쥔 현우의 상체.\n\nLOCATION (lock): At a two-way junction inside the refugee settlement's underground sewer, where both branches recede into darkness. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Sewer fork (The passage divides into two routes) — The two branch entrances are visible obliquely behind the interaction; used as Provides the spatial evidence for the decision to separate.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the sewer dark, with restrained contrast preserving the expression and gripping hand without specifying an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The underground sewer divides into two passages. Charlie remains a worn gorilla-shaped robot in an old coat and hat, with blue-lit eyes and a faded UBIK chest logo. 현우: He is at the sewer fork without his outer garment; his face remains bruised and his dog-bitten leg remains injured. The contact card is concealed in his shoe. 페드로: He has reached the sewer fork and is about to take a separate route.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 어두운 하수도 두 갈래 길 앞에서 단호한 표정으로 페드로의 어깨를 꽉 쥔 현우의 상체.\n\nLOCATION (lock): At a two-way junction inside the refugee settlement's underground sewer, where both branches recede into darkness. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Sewer fork (The passage divides into two routes) — The two branch entrances are visible obliquely behind the interaction; used as Provides the spatial evidence for the decision to separate.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the sewer dark, with restrained contrast preserving the expression and gripping hand without specifying an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The underground sewer divides into two passages. Charlie remains a worn gorilla-shaped robot in an old coat and hat, with blue-lit eyes and a faded UBIK chest logo. 현우: He is at the sewer fork without his outer garment; his face remains bruised and his dog-bitten leg remains injured. The contact card is concealed in his shoe. 페드로: He has reached the sewer fork and is about to take a separate route.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 어두운 하수도 두 갈래 길 앞에서 단호한 표정으로 페드로의 어깨를 꽉 쥔 현우의 상체.\n\nLOCATION (lock): At a two-way junction inside the refugee settlement's underground sewer, where both branches recede into darkness. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Sewer fork (The passage divides into two routes) — The two branch entrances are visible obliquely behind the interaction; used as Provides the spatial evidence for the decision to separate.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the sewer dark, with restrained contrast preserving the expression and gripping hand without specifying an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The underground sewer divides into two passages. Charlie remains a worn gorilla-shaped robot in an old coat and hat, with blue-lit eyes and a faded UBIK chest logo. 현우: He is at the sewer fork without his outer garment; his face remains bruised and his dog-bitten leg remains injured. The contact card is concealed in his shoe. 페드로: He has reached the sewer fork and is about to take a separate route.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S26sh9__bgfirst_bg.png",
     "asset_id": "ff1b1858-e263-49dc-901d-bd8f3ca0ac13",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S26sh9.png",
     "asset_id": "bad5b160-fb00-401b-90e3-64652fe59154",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 페드로: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1278830>",
     "asset_id": "b09df655-64d4-4db4-a1b3-2f0bb5d29c95",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L175B01.png",
     "asset_id": "aacb72f4-e4d3-4331-8b05-48ae8f9cc866",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 페드로: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1278830>",
     "asset_id": "b09df655-64d4-4db4-a1b3-2f0bb5d29c95",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우는 페드로를 향해 시선을 두고 있으며, 왼손으로 페드로의 오른쪽 어깨를 단단히 쥐고 있습니다.",
    "built_space": "하수도 내부입니다. 왼쪽에 사다리가 있고 우측으로 뻗은 터널이 보이나, 좌측 갈래길은 인물에 가려졌거나 명확하지 않습니다.",
    "entities": "현우(겉옷 없음, 멍든 얼굴 등 참조 일치), 페드로(참조 일치) 모두 지시된 외형에 부합합니다.",
    "hard_violations": [],
    "physics": "어깨를 쥐고 서 있는 두 인물의 자세와 무게중심이 자연스럽게 바닥과 몸에 지지되어 있습니다."
   },
   {
    "label": "B",
    "direction": "현우는 렌즈를 직시하고 있으며, 페드로는 화면 우측을 봅니다. 현우의 왼손이 페드로의 왼쪽 어깨를 잡고 있습니다.",
    "built_space": "하수도 내부이며 좌측의 사다리와 배경에 나뉘는 두 갈래 터널 입구가 모두 뚜렷하게 보입니다.",
    "entities": "두 인물의 얼굴은 참조와 일치하나, 지시문과 달리 두 사람 모두 겉옷(재킷)과 배낭을 착용하고 있습니다.",
    "hard_violations": [
     "[gpt-high] 프롬프트와 인물 참고에 없는 배낭을 두 사람에게 각각 추가했습니다."
    ],
    "physics": "인물들이 바닥에 잘 서 있고 손의 접촉도 유지되나 자세가 다소 경직되어 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "겉옷을 입지 않은 복장 상태와 상대방을 향한 단호한 시선 및 동작을 지시문대로 훌륭하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "배경의 두 갈래 길은 잘 표현되었으나, 지시와 달리 겉옷을 착용하고 카메라 렌즈를 정면으로 응시하여 감점이 큽니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 페드로를 향해 시선을 두고 있으며, 왼손으로 페드로의 오른쪽 어깨를 단단히 쥐고 있습니다.",
        "built_space": "하수도 내부입니다. 왼쪽에 사다리가 있고 우측으로 뻗은 터널이 보이나, 좌측 갈래길은 인물에 가려졌거나 명확하지 않습니다.",
        "entities": "현우(겉옷 없음, 멍든 얼굴 등 참조 일치), 페드로(참조 일치) 모두 지시된 외형에 부합합니다.",
        "hard_violations": [],
        "physics": "어깨를 쥐고 서 있는 두 인물의 자세와 무게중심이 자연스럽게 바닥과 몸에 지지되어 있습니다."
       },
       {
        "label": "B",
        "direction": "현우는 렌즈를 직시하고 있으며, 페드로는 화면 우측을 봅니다. 현우의 왼손이 페드로의 왼쪽 어깨를 잡고 있습니다.",
        "built_space": "하수도 내부이며 좌측의 사다리와 배경에 나뉘는 두 갈래 터널 입구가 모두 뚜렷하게 보입니다.",
        "entities": "두 인물의 얼굴은 참조와 일치하나, 지시문과 달리 두 사람 모두 겉옷(재킷)과 배낭을 착용하고 있습니다.",
        "hard_violations": [],
        "physics": "인물들이 바닥에 잘 서 있고 손의 접촉도 유지되나 자세가 다소 경직되어 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "겉옷을 입지 않은 복장 상태와 상대방을 향한 단호한 시선 및 동작을 지시문대로 훌륭하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "배경의 두 갈래 길은 잘 표현되었으나, 지시와 달리 겉옷을 착용하고 카메라 렌즈를 정면으로 응시하여 감점이 큽니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우는 페드로를 향해 시선을 두고 있으며, 왼손으로 페드로의 오른쪽 어깨를 단단히 쥐고 있습니다.",
        "built_space": "하수도 내부입니다. 왼쪽에 사다리가 있고 우측으로 뻗은 터널이 보이나, 좌측 갈래길은 인물에 가려졌거나 명확하지 않습니다.",
        "entities": "현우(겉옷 없음, 멍든 얼굴 등 참조 일치), 페드로(참조 일치) 모두 지시된 외형에 부합합니다.",
        "hard_violations": [],
        "physics": "어깨를 쥐고 서 있는 두 인물의 자세와 무게중심이 자연스럽게 바닥과 몸에 지지되어 있습니다."
       },
       {
        "label": "B",
        "direction": "현우는 렌즈를 직시하고 있으며, 페드로는 화면 우측을 봅니다. 현우의 왼손이 페드로의 왼쪽 어깨를 잡고 있습니다.",
        "built_space": "하수도 내부이며 좌측의 사다리와 배경에 나뉘는 두 갈래 터널 입구가 모두 뚜렷하게 보입니다.",
        "entities": "두 인물의 얼굴은 참조와 일치하나, 지시문과 달리 두 사람 모두 겉옷(재킷)과 배낭을 착용하고 있습니다.",
        "hard_violations": [],
        "physics": "인물들이 바닥에 잘 서 있고 손의 접촉도 유지되나 자세가 다소 경직되어 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "상체와 어깨 접촉은 담았지만 현우가 페드로 대신 카메라를 바라보고 겉옷을 입었으며, 불필요한 배낭까지 추가되어 지시에서 벗어납니다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "겉옷을 벗은 현우의 멍든 얼굴, 페드로를 향한 단호한 시선과 어깨를 움켜쥔 손을 미디엄 숏으로 충실히 담았으나 왼쪽 갈래 입구는 많이 가려집니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 얼굴과 시선은 페드로가 아니라 카메라 쪽을 향합니다. 페드로는 왼쪽의 현우 쪽을 바라봅니다. 현우의 뻗은 팔과 손은 페드로의 가까운 어깨에 정확히 닿지만, 손가락을 넓게 편 접촉이라 꽉 움켜쥐는 힘은 약하게 읽힙니다.",
        "built_space": "왼쪽 벽에 고정 사다리 하나와 그 아래 돌출 받침이 보이며, 낮은 콘크리트 천장과 젖은 중앙 수로가 참고 장소와 대응합니다. 뒤쪽 오른쪽 통로 입구는 보이지만 왼쪽 입구는 현우에게 대부분 가려져 두 갈래의 공간 증거가 약합니다. 두 사람은 분기점 앞에 서 있고 설비와 신체가 충돌하지 않습니다.",
        "entities": "젊은 남성 두 명만 보이며 현우의 동아시아계 외모, 검은 머리와 얼굴 상처, 페드로의 짙은 머리와 옆얼굴은 참고 인물에 대체로 대응합니다. 다만 현우는 벗어야 할 회색 겉옷을 입고 있고 두 사람 모두 지시되지 않은 배낭을 멨습니다. 다리 부상과 신발 속 카드는 화면 밖이므로 확인 대상이 아닙니다. 로봇이나 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "프롬프트와 인물 참고에 없는 배낭을 두 사람에게 각각 추가했습니다."
        ],
        "physics": "현우의 손은 팔과 자연스럽게 이어져 페드로의 어깨에 접촉합니다. 배낭은 어깨끈으로 몸에 지지됩니다. 두 사람의 하체와 발은 화면 밖이지만 상체는 정상적인 직립 자세이며 공중 부양이나 불가능한 관절 배치는 보이지 않습니다."
       },
       {
        "label": "B",
        "direction": "현우는 앞에 선 페드로의 얼굴을 내려다보고, 페드로도 현우 쪽으로 얼굴을 돌립니다. 현우의 손가락은 페드로의 어깨 윗부분을 감싸며 옷감을 잡고 있어 행동의 대상과 힘의 방향이 분명합니다.",
        "built_space": "왼쪽 벽의 고정 사다리 하나, 낮은 콘크리트 천장, 중앙의 젖은 수로와 양옆 좁은 턱이 참고 장소와 대응합니다. 뒤쪽 오른쪽 통로 입구는 뚜렷하고 왼쪽 갈래는 현우의 왼편 틈으로 일부 보입니다. 인물들이 왼쪽 입구를 상당히 가리지만 두 방향으로 갈라지는 구조는 남아 있습니다. 두 사람은 분기점 앞에서 서로 마주보며 고정 설비와 충돌하지 않습니다.",
        "entities": "현우는 참고와 유사한 앳된 얼굴과 헝클어진 검은 머리를 가진 젊은 남성으로, 얼굴에 멍과 찰과상이 있고 겉옷 없이 짙은 반팔 티셔츠를 입었습니다. 페드로는 짙은 머리와 젊은 옆얼굴이 보이지만 뒷모습 위주여서 얼굴 전체의 동일성은 제한적으로만 확인됩니다. 페드로의 티셔츠는 참고의 남색보다 회갈색에 가깝습니다. 추가 인물, 로봇, 배낭과 읽을 수 있는 글자는 없으며 다리와 신발 속 카드는 구도 밖입니다.",
        "hard_violations": [],
        "physics": "현우의 굽힌 팔은 어깨를 잡은 손까지 자연스럽게 연결되고 손가락이 페드로의 어깨 표면과 옷감에 밀착합니다. 접촉 부위의 주름도 잡아당기는 동작과 맞습니다. 발은 보이지 않지만 두 상체는 서 있는 자세로 일관되며 지지 없이 떠 있는 신체나 물체는 없습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "상체와 어깨 접촉은 담았지만 현우가 페드로 대신 카메라를 바라보고 겉옷을 입었으며, 불필요한 배낭까지 추가되어 지시에서 벗어납니다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "겉옷을 벗은 현우의 멍든 얼굴, 페드로를 향한 단호한 시선과 어깨를 움켜쥔 손을 미디엄 숏으로 충실히 담았으나 왼쪽 갈래 입구는 많이 가려집니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 얼굴과 시선은 페드로가 아니라 카메라 쪽을 향합니다. 페드로는 왼쪽의 현우 쪽을 바라봅니다. 현우의 뻗은 팔과 손은 페드로의 가까운 어깨에 정확히 닿지만, 손가락을 넓게 편 접촉이라 꽉 움켜쥐는 힘은 약하게 읽힙니다.",
        "built_space": "왼쪽 벽에 고정 사다리 하나와 그 아래 돌출 받침이 보이며, 낮은 콘크리트 천장과 젖은 중앙 수로가 참고 장소와 대응합니다. 뒤쪽 오른쪽 통로 입구는 보이지만 왼쪽 입구는 현우에게 대부분 가려져 두 갈래의 공간 증거가 약합니다. 두 사람은 분기점 앞에 서 있고 설비와 신체가 충돌하지 않습니다.",
        "entities": "젊은 남성 두 명만 보이며 현우의 동아시아계 외모, 검은 머리와 얼굴 상처, 페드로의 짙은 머리와 옆얼굴은 참고 인물에 대체로 대응합니다. 다만 현우는 벗어야 할 회색 겉옷을 입고 있고 두 사람 모두 지시되지 않은 배낭을 멨습니다. 다리 부상과 신발 속 카드는 화면 밖이므로 확인 대상이 아닙니다. 로봇이나 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "프롬프트와 인물 참고에 없는 배낭을 두 사람에게 각각 추가했습니다."
        ],
        "physics": "현우의 손은 팔과 자연스럽게 이어져 페드로의 어깨에 접촉합니다. 배낭은 어깨끈으로 몸에 지지됩니다. 두 사람의 하체와 발은 화면 밖이지만 상체는 정상적인 직립 자세이며 공중 부양이나 불가능한 관절 배치는 보이지 않습니다."
       },
       {
        "label": "A",
        "direction": "현우는 앞에 선 페드로의 얼굴을 내려다보고, 페드로도 현우 쪽으로 얼굴을 돌립니다. 현우의 손가락은 페드로의 어깨 윗부분을 감싸며 옷감을 잡고 있어 행동의 대상과 힘의 방향이 분명합니다.",
        "built_space": "왼쪽 벽의 고정 사다리 하나, 낮은 콘크리트 천장, 중앙의 젖은 수로와 양옆 좁은 턱이 참고 장소와 대응합니다. 뒤쪽 오른쪽 통로 입구는 뚜렷하고 왼쪽 갈래는 현우의 왼편 틈으로 일부 보입니다. 인물들이 왼쪽 입구를 상당히 가리지만 두 방향으로 갈라지는 구조는 남아 있습니다. 두 사람은 분기점 앞에서 서로 마주보며 고정 설비와 충돌하지 않습니다.",
        "entities": "현우는 참고와 유사한 앳된 얼굴과 헝클어진 검은 머리를 가진 젊은 남성으로, 얼굴에 멍과 찰과상이 있고 겉옷 없이 짙은 반팔 티셔츠를 입었습니다. 페드로는 짙은 머리와 젊은 옆얼굴이 보이지만 뒷모습 위주여서 얼굴 전체의 동일성은 제한적으로만 확인됩니다. 페드로의 티셔츠는 참고의 남색보다 회갈색에 가깝습니다. 추가 인물, 로봇, 배낭과 읽을 수 있는 글자는 없으며 다리와 신발 속 카드는 구도 밖입니다.",
        "hard_violations": [],
        "physics": "현우의 굽힌 팔은 어깨를 잡은 손까지 자연스럽게 연결되고 손가락이 페드로의 어깨 표면과 옷감에 밀착합니다. 접촉 부위의 주름도 잡아당기는 동작과 맞습니다. 발은 보이지 않지만 두 상체는 서 있는 자세로 일관되며 지지 없이 떠 있는 신체나 물체는 없습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.016
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.766
   },
   "violations": {
    "B": [
     "[gpt-high] 프롬프트와 인물 참고에 없는 배낭을 두 사람에게 각각 추가했습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 766
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "겉옷을 입지 않은 복장 상태와 상대방을 향한 단호한 시선 및 동작을 지시문대로 훌륭하게 구현했습니다."
   },
   {
    "label": "B",
    "score": 766,
    "verdict_ko": "배경의 두 갈래 길은 잘 표현되었으나, 지시와 달리 겉옷을 착용하고 카메라 렌즈를 정면으로 응시하여 감점이 큽니다.  ★위반: [gpt-high] 프롬프트와 인물 참고에 없는 배낭을 두 사람에게 각각 추가했습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L175B01.png",
    "asset_id": "aacb72f4-e4d3-4331-8b05-48ae8f9cc866",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 페드로: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1278830>",
    "asset_id": "b09df655-64d4-4db4-a1b3-2f0bb5d29c95",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-2a8f-7e47-a081-d334cebae642",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S26sh9__bgfirst_bg.png",
   "bg_asset_id": "ff1b1858-e263-49dc-901d-bd8f3ca0ac13",
   "bg_record_key": "S26sh9::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S26sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:20:33.531488+00:00",
  "fingerprint": "a9f610ab8319506fa455aaf9572ecb74fc9ddf91820b2596b6775cf61cfeffde",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S26sh9_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S26sh9_sel.png",
  "source_sha256": "ab56bcb5b002867ec1af2bed3bd4ce1951f8f7a72e7f739be4b7d7d8f7fa2a09",
  "file": "S26sh9_cine.png",
  "staged_sha256": "c24037dddf41611a691fdf34fac0fac24e48f21963de46ef2d3c724b73d35005",
  "latency_ms": 18768
 },
 "S26sh11::signage": {
  "fp": "bdc6108b03f52d68",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S26sh11": {
  "input_fingerprint": "e46aacbcc54e3b6b",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 얼굴을 등에 기댄 앰버를 꽉 업은 채 어두운 하수도 터널을 전속력으로 내달리며, 앞으로 몸을 깊게 기울인 현우의 땀 맺힌 mid-action 얼굴.\n\nLOCATION (lock): Inside a dark underground sewer passage beyond the junction, beneath the refugee settlement. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Chosen sewer passage (현우 is running through it carrying 앰버) — The passage recedes behind their shoulders; used as A narrow, softly resolved background preserves the direction of escape without competing with the faces.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the sewer's darkness and controlled tonal separation, allowing the stated perspiration and both faces to remain legible without adding a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The sewer route continues beyond the fork. Charlie remains in the old coat and hat, with his worn metal body and blue-lit eyes unchanged. 현우: He is running in a load-bearing posture, his outer garment removed; his facial bruises and injured leg persist. The contact card remains concealed in his shoe. 앰버: She is being carried piggyback, with worsening coughing and the outer garment covering her mouth.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 얼굴을 등에 기댄 앰버를 꽉 업은 채 어두운 하수도 터널을 전속력으로 내달리며, 앞으로 몸을 깊게 기울인 현우의 땀 맺힌 mid-action 얼굴.\n\nLOCATION (lock): Inside a dark underground sewer passage beyond the junction, beneath the refugee settlement. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Chosen sewer passage (현우 is running through it carrying 앰버) — The passage recedes behind their shoulders; used as A narrow, softly resolved background preserves the direction of escape without competing with the faces.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the sewer's darkness and controlled tonal separation, allowing the stated perspiration and both faces to remain legible without adding a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The sewer route continues beyond the fork. Charlie remains in the old coat and hat, with his worn metal body and blue-lit eyes unchanged. 현우: He is running in a load-bearing posture, his outer garment removed; his facial bruises and injured leg persist. The contact card remains concealed in his shoe. 앰버: She is being carried piggyback, with worsening coughing and the outer garment covering her mouth.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 얼굴을 등에 기댄 앰버를 꽉 업은 채 어두운 하수도 터널을 전속력으로 내달리며, 앞으로 몸을 깊게 기울인 현우의 땀 맺힌 mid-action 얼굴.\n\nLOCATION (lock): Inside a dark underground sewer passage beyond the junction, beneath the refugee settlement. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Chosen sewer passage (현우 is running through it carrying 앰버) — The passage recedes behind their shoulders; used as A narrow, softly resolved background preserves the direction of escape without competing with the faces.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the sewer's darkness and controlled tonal separation, allowing the stated perspiration and both faces to remain legible without adding a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The sewer route continues beyond the fork. Charlie remains in the old coat and hat, with his worn metal body and blue-lit eyes unchanged. 현우: He is running in a load-bearing posture, his outer garment removed; his facial bruises and injured leg persist. The contact card remains concealed in his shoe. 앰버: She is being carried piggyback, with worsening coughing and the outer garment covering her mouth.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선이 카메라 렌즈를 정면으로 향해 있음. 앰버는 비스듬히 아래를 향함.",
    "built_space": "하수도 터널 내부이나, 레퍼런스와 달리 수로가 왼쪽, 보행로가 오른쪽에 위치하며 샷에 고정되어야 할 사다리가 누락됨.",
    "entities": "현우(땀 맺힌 얼굴, 검은 티셔츠), 앰버(금발, 겉옷이 입을 완전히 가리지 못하고 노출됨).",
    "hard_violations": [],
    "physics": "앞으로 기울인 달리기 자세를 취하고 있으며 프레임 밖 바닥에 의해 지탱됨. 양손으로 앰버를 받치고 있음."
   },
   {
    "label": "B",
    "direction": "현우는 터널 진행 방향을 주시하며 앞을 향하고 있음. 앰버는 현우의 등에 얼굴을 깊게 묻고 있음.",
    "built_space": "돌로 된 하수도 터널. 레퍼런스 사진과 동일하게 왼쪽 벽에 사다리가 부착되어 있고 왼쪽이 보행로, 오른쪽이 수로인 구조를 정확히 따름.",
    "entities": "현우(땀 맺힌 얼굴, 타박상, 검은 티셔츠), 앰버(금발, 겉옷으로 입 주변이 가려진 채 업혀 있음) 모두 지시사항과 일치함.",
    "hard_violations": [],
    "physics": "몸을 앞으로 깊게 기울여 내달리는 역동적인 자세이며, 양손으로 앰버의 하체를 단단히 받치고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "레퍼런스의 장소 구조(왼쪽 사다리와 오른쪽 수로)를 정확히 유지했고, 카메라를 의식하지 않는 시선 처리와 앰버의 입이 가려진 묘사까지 프롬프트를 충실히 구현함."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "현우가 렌즈를 정면으로 응시하여 연기 지침을 어겼으며, 레퍼런스의 사다리가 누락되고 수로와 보행로의 위치가 반전되어 장소 일치도가 낮음."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "현우는 터널 진행 방향을 주시하며 앞을 향하고 있음. 앰버는 현우의 등에 얼굴을 깊게 묻고 있음.",
        "built_space": "돌로 된 하수도 터널. 레퍼런스 사진과 동일하게 왼쪽 벽에 사다리가 부착되어 있고 왼쪽이 보행로, 오른쪽이 수로인 구조를 정확히 따름.",
        "entities": "현우(땀 맺힌 얼굴, 타박상, 검은 티셔츠), 앰버(금발, 겉옷으로 입 주변이 가려진 채 업혀 있음) 모두 지시사항과 일치함.",
        "hard_violations": [],
        "physics": "몸을 앞으로 깊게 기울여 내달리는 역동적인 자세이며, 양손으로 앰버의 하체를 단단히 받치고 있음."
       },
       {
        "label": "A",
        "direction": "현우의 시선이 카메라 렌즈를 정면으로 향해 있음. 앰버는 비스듬히 아래를 향함.",
        "built_space": "하수도 터널 내부이나, 레퍼런스와 달리 수로가 왼쪽, 보행로가 오른쪽에 위치하며 샷에 고정되어야 할 사다리가 누락됨.",
        "entities": "현우(땀 맺힌 얼굴, 검은 티셔츠), 앰버(금발, 겉옷이 입을 완전히 가리지 못하고 노출됨).",
        "hard_violations": [],
        "physics": "앞으로 기울인 달리기 자세를 취하고 있으며 프레임 밖 바닥에 의해 지탱됨. 양손으로 앰버를 받치고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "레퍼런스의 장소 구조(왼쪽 사다리와 오른쪽 수로)를 정확히 유지했고, 카메라를 의식하지 않는 시선 처리와 앰버의 입이 가려진 묘사까지 프롬프트를 충실히 구현함."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "현우가 렌즈를 정면으로 응시하여 연기 지침을 어겼으며, 레퍼런스의 사다리가 누락되고 수로와 보행로의 위치가 반전되어 장소 일치도가 낮음."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 터널 진행 방향을 주시하며 앞을 향하고 있음. 앰버는 현우의 등에 얼굴을 깊게 묻고 있음.",
        "built_space": "돌로 된 하수도 터널. 레퍼런스 사진과 동일하게 왼쪽 벽에 사다리가 부착되어 있고 왼쪽이 보행로, 오른쪽이 수로인 구조를 정확히 따름.",
        "entities": "현우(땀 맺힌 얼굴, 타박상, 검은 티셔츠), 앰버(금발, 겉옷으로 입 주변이 가려진 채 업혀 있음) 모두 지시사항과 일치함.",
        "hard_violations": [],
        "physics": "몸을 앞으로 깊게 기울여 내달리는 역동적인 자세이며, 양손으로 앰버의 하체를 단단히 받치고 있음."
       },
       {
        "label": "A",
        "direction": "현우의 시선이 카메라 렌즈를 정면으로 향해 있음. 앰버는 비스듬히 아래를 향함.",
        "built_space": "하수도 터널 내부이나, 레퍼런스와 달리 수로가 왼쪽, 보행로가 오른쪽에 위치하며 샷에 고정되어야 할 사다리가 누락됨.",
        "entities": "현우(땀 맺힌 얼굴, 검은 티셔츠), 앰버(금발, 겉옷이 입을 완전히 가리지 못하고 노출됨).",
        "hard_violations": [],
        "physics": "앞으로 기울인 달리기 자세를 취하고 있으며 프레임 밖 바닥에 의해 지탱됨. 양손으로 앰버를 받치고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "장소와 업힌 접촉은 잘 유지하지만, 상체와 하수도 설비를 넓게 보여 주어 얼굴 클로즈업과 전속력 질주의 순간성이 B보다 약하다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "더 크게 잡힌 땀 맺힌 얼굴, 전방으로 숙인 운반 자세, 흐려지는 통로와 앰버의 입을 덮은 겉옷이 핵심 지시를 더 충실히 구현한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 몸을 앞으로 숙이고 화면 왼쪽 전방의 화면 밖 진행로를 바라본다. 앰버는 고개를 아래로 떨어뜨려 현우의 오른쪽 어깨 뒤에 얼굴을 댄다. 통로는 두 사람 뒤로 이어지며, 무기나 방향을 판정할 휴대 도구는 없다.",
        "built_space": "왼쪽 벽에 고정 사다리 하나, 뒤쪽 천장 부근에 작은 조명 하나, 중앙 수로와 양옆 턱이 보인다. 젖고 거친 콘크리트와 낮은 천장은 이전 장면에 가깝다. 두 사람은 통로 전경을 차지하지만 발이 잘려 어느 턱을 밟는지는 확인할 수 없다. 배경과 몸통이 비교적 넓고 선명하게 들어와 요구된 좁고 부드러운 배경보다 비중이 크다.",
        "entities": "현우와 앰버 두 명만 보인다. 현우의 앳된 동아시아계 남성 얼굴, 헝클어진 검은 머리, 짙은 반소매 티셔츠와 얼굴 상처는 참조에 부합하며 땀도 보인다. 앰버는 금발의 어린 여자아이로, 얼굴을 찡그리고 있어 기침이나 고통이 읽힌다. 겉옷은 어깨 주변에 있으나 입 부위가 현우의 어깨와 겹쳐 겉옷으로 입을 덮었는지는 덜 명확하다. 다리 부상과 신발 속 카드는 프레임 밖이다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "앰버의 몸은 현우의 등에 밀착하고 다리는 그의 양옆으로 내려온다. 현우는 양팔을 뒤로 내려 아이의 하체를 받치는 자세이며 손의 정확한 접촉점은 하단 밖이다. 앰버의 얼굴은 어깨 뒤에 닿아 있어 지지가 읽힌다. 발은 보이지 않지만 공중에 떠 있는 묘사는 없으며, 앞으로 기운 자세는 운반 동작으로 가능하다. 다만 보이는 순간만으로 전속력 달리기의 속도감은 강하지 않다."
       },
       {
        "label": "B",
        "direction": "현우의 얼굴과 상체는 카메라 쪽 전방으로 향하고 시선은 약간 아래의 진행로를 향한다. 앰버는 눈을 내리감고 현우의 오른쪽 어깨 뒤에 얼굴을 기댄다. 뒤로 멀어지는 통로와 주변의 흐림이 두 사람이 전경 쪽으로 달려 나오는 방향을 뒷받침한다.",
        "built_space": "거친 콘크리트 벽 두 면, 낮은 천장, 뒤쪽 조명 하나와 젖은 수로 가장자리의 턱이 보인다. 사다리는 보이지 않지만 분기 이후 통로를 가까이 잡은 구도이므로 누락을 구조 모순으로 볼 근거는 없다. 배경은 어깨 뒤로 후퇴하고 주변부가 흐려져 이전 장면의 어두운 하수도 재질을 유지하면서 얼굴과 덜 경쟁한다. 상체가 여전히 다소 많이 포함되지만 A보다 얼굴 비중이 크다.",
        "entities": "현우와 앰버만 등장한다. 현우는 참조와 가까운 젊은 동아시아계 남성 얼굴, 검은 헝클어진 머리와 짙은 반소매 옷을 유지하며 뺨의 상처와 이마·코의 땀이 뚜렷하다. 앰버는 참조에 가까운 금발과 둥근 어린 얼굴을 가지며, 큰 눈의 형태는 내리감긴 눈꺼풀 때문에 충분히 확인되지 않는다. 올리브색 겉옷이 입을 명확히 덮는다. 신발 속 카드와 다리 부상은 프레임 밖이고 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우는 어깨와 몸통을 앞으로 기울이고 팔을 뒤로 보내 앰버의 하체를 받치는 자세다. 아이의 몸은 등에 실리고 다리가 양옆으로 이어지며, 얼굴과 입을 덮은 옷은 어깨에 접촉한다. 손과 발의 정확한 접촉점은 잘려 있지만 무지지 부유나 불가능한 관절은 보이지 않는다. 하중을 견디는 전경 자세와 배경의 이동 흐림이 달리는 순간으로 자연스럽게 연결된다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "장소와 업힌 접촉은 잘 유지하지만, 상체와 하수도 설비를 넓게 보여 주어 얼굴 클로즈업과 전속력 질주의 순간성이 B보다 약하다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "더 크게 잡힌 땀 맺힌 얼굴, 전방으로 숙인 운반 자세, 흐려지는 통로와 앰버의 입을 덮은 겉옷이 핵심 지시를 더 충실히 구현한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 몸을 앞으로 숙이고 화면 왼쪽 전방의 화면 밖 진행로를 바라본다. 앰버는 고개를 아래로 떨어뜨려 현우의 오른쪽 어깨 뒤에 얼굴을 댄다. 통로는 두 사람 뒤로 이어지며, 무기나 방향을 판정할 휴대 도구는 없다.",
        "built_space": "왼쪽 벽에 고정 사다리 하나, 뒤쪽 천장 부근에 작은 조명 하나, 중앙 수로와 양옆 턱이 보인다. 젖고 거친 콘크리트와 낮은 천장은 이전 장면에 가깝다. 두 사람은 통로 전경을 차지하지만 발이 잘려 어느 턱을 밟는지는 확인할 수 없다. 배경과 몸통이 비교적 넓고 선명하게 들어와 요구된 좁고 부드러운 배경보다 비중이 크다.",
        "entities": "현우와 앰버 두 명만 보인다. 현우의 앳된 동아시아계 남성 얼굴, 헝클어진 검은 머리, 짙은 반소매 티셔츠와 얼굴 상처는 참조에 부합하며 땀도 보인다. 앰버는 금발의 어린 여자아이로, 얼굴을 찡그리고 있어 기침이나 고통이 읽힌다. 겉옷은 어깨 주변에 있으나 입 부위가 현우의 어깨와 겹쳐 겉옷으로 입을 덮었는지는 덜 명확하다. 다리 부상과 신발 속 카드는 프레임 밖이다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "앰버의 몸은 현우의 등에 밀착하고 다리는 그의 양옆으로 내려온다. 현우는 양팔을 뒤로 내려 아이의 하체를 받치는 자세이며 손의 정확한 접촉점은 하단 밖이다. 앰버의 얼굴은 어깨 뒤에 닿아 있어 지지가 읽힌다. 발은 보이지 않지만 공중에 떠 있는 묘사는 없으며, 앞으로 기운 자세는 운반 동작으로 가능하다. 다만 보이는 순간만으로 전속력 달리기의 속도감은 강하지 않다."
       },
       {
        "label": "A",
        "direction": "현우의 얼굴과 상체는 카메라 쪽 전방으로 향하고 시선은 약간 아래의 진행로를 향한다. 앰버는 눈을 내리감고 현우의 오른쪽 어깨 뒤에 얼굴을 기댄다. 뒤로 멀어지는 통로와 주변의 흐림이 두 사람이 전경 쪽으로 달려 나오는 방향을 뒷받침한다.",
        "built_space": "거친 콘크리트 벽 두 면, 낮은 천장, 뒤쪽 조명 하나와 젖은 수로 가장자리의 턱이 보인다. 사다리는 보이지 않지만 분기 이후 통로를 가까이 잡은 구도이므로 누락을 구조 모순으로 볼 근거는 없다. 배경은 어깨 뒤로 후퇴하고 주변부가 흐려져 이전 장면의 어두운 하수도 재질을 유지하면서 얼굴과 덜 경쟁한다. 상체가 여전히 다소 많이 포함되지만 A보다 얼굴 비중이 크다.",
        "entities": "현우와 앰버만 등장한다. 현우는 참조와 가까운 젊은 동아시아계 남성 얼굴, 검은 헝클어진 머리와 짙은 반소매 옷을 유지하며 뺨의 상처와 이마·코의 땀이 뚜렷하다. 앰버는 참조에 가까운 금발과 둥근 어린 얼굴을 가지며, 큰 눈의 형태는 내리감긴 눈꺼풀 때문에 충분히 확인되지 않는다. 올리브색 겉옷이 입을 명확히 덮는다. 신발 속 카드와 다리 부상은 프레임 밖이고 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우는 어깨와 몸통을 앞으로 기울이고 팔을 뒤로 보내 앰버의 하체를 받치는 자세다. 아이의 몸은 등에 실리고 다리가 양옆으로 이어지며, 얼굴과 입을 덮은 옷은 어깨에 접촉한다. 손과 발의 정확한 접촉점은 잘려 있지만 무지지 부유나 불가능한 관절은 보이지 않는다. 하중을 견디는 전경 자세와 배경의 이동 흐림이 달리는 순간으로 자연스럽게 연결된다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.571,
    "B": 1.875
   },
   "adjusted": {
    "A": 1.571,
    "B": 1.875
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1875,
   "A": 1571
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1875,
    "verdict_ko": "레퍼런스의 장소 구조(왼쪽 사다리와 오른쪽 수로)를 정확히 유지했고, 카메라를 의식하지 않는 시선 처리와 앰버의 입이 가려진 묘사까지 프롬프트를 충실히 구현함."
   },
   {
    "label": "A",
    "score": 1571,
    "verdict_ko": "현우가 렌즈를 정면으로 응시하여 연기 지침을 어겼으며, 레퍼런스의 사다리가 누락되고 수로와 보행로의 위치가 반전되어 장소 일치도가 낮음."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S26sh9_sel.png",
    "asset_id": "275e1d10-b182-41dc-8cdf-2fbaada0133f",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-2e41-7259-9920-da3c36a4928e",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S26sh9"
  }
 },
 "S26sh11::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:22:03.186917+00:00",
  "fingerprint": "89cc52f536fddb06cded29552cfd8e78cd288480c73befd4dee75e423fc9e8aa",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S26sh11_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S26sh11_sel.png",
  "source_sha256": "56e32b803b633db64df21139a2f17908087d42e0de27f0f554aa98415a6bb908",
  "file": "S26sh11_cine.png",
  "staged_sha256": "7705ba16914bf09d6cd8baa9446505bfa0639c2a0bfcaa589c95b25a46b1199c",
  "latency_ms": 15038
 },
 "S27sh12::signage": {
  "fp": "dbf48cb7b94e26f8",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S27sh12": {
  "input_fingerprint": "c7e5d0e14c9febe4",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 도로 위를 달리며 한 발이 공중에 뜬 찰리를 향해 빠른 속도로 돌진하는 거대한 자동차의 전경.\n\nLOCATION (lock): On the open roadway in the refugee settlement, in the path of an approaching vehicle's headlights. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nSTRUCTURE LOOK AUTHORITY: the attached STRUCTURE LOOK photograph is the identity of the fixed structure at this location — wherever that structure appears in the frame, its shape, proportions, openings, materials and colors are LOCKED to it. The LOCATION PHOTOGRAPH remains the authority for this shot's sub-space, surroundings, time of day and lighting. If the two conflict on the structure itself, the STRUCTURE LOOK photo wins; for everything else, the LOCATION PHOTOGRAPH wins.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Approaching large vehicle in the middle-left of the frame, foreground, moves toward 찰리's running path in the right midground.\n- KEY BACKGROUND ELEMENTS: Approaching large vehicle (Moving rapidly toward 찰리 before the emergency stop) — Its front and one side are visible diagonally from the road edge; used as Forms the left foreground threat while remaining fully within the frame; Road (찰리 and the approaching vehicle occupy converging paths) — The road extends diagonally from the foreground toward the distance; used as Keeps the collision geometry and remaining separation visible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the approaching headlights and the streetlights' established switching behavior to shape the nighttime threat, keeping the road and 찰리 readable without added atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Streetlights continue switching on and off along Charlie's route, and the approaching vehicle has its headlights on. Charlie retains his old coat and hat; the small bird has already flown away.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 도로 위를 달리며 한 발이 공중에 뜬 찰리를 향해 빠른 속도로 돌진하는 거대한 자동차의 전경.\n\nLOCATION (lock): On the open roadway in the refugee settlement, in the path of an approaching vehicle's headlights. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Approaching large vehicle in the middle-left of the frame, foreground, moves toward 찰리's running path in the right midground.\n- KEY BACKGROUND ELEMENTS: Approaching large vehicle (Moving rapidly toward 찰리 before the emergency stop) — Its front and one side are visible diagonally from the road edge; used as Forms the left foreground threat while remaining fully within the frame; Road (찰리 and the approaching vehicle occupy converging paths) — The road extends diagonally from the foreground toward the distance; used as Keeps the collision geometry and remaining separation visible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the approaching headlights and the streetlights' established switching behavior to shape the nighttime threat, keeping the road and 찰리 readable without added atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Streetlights continue switching on and off along Charlie's route, and the approaching vehicle has its headlights on. Charlie retains his old coat and hat; the small bird has already flown away.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 도로 위를 달리며 한 발이 공중에 뜬 찰리를 향해 빠른 속도로 돌진하는 거대한 자동차의 전경.\n\nLOCATION (lock): On the open roadway in the refugee settlement, in the path of an approaching vehicle's headlights. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nSTRUCTURE LOOK AUTHORITY: the attached STRUCTURE LOOK photograph is the identity of the fixed structure at this location — wherever that structure appears in the frame, its shape, proportions, openings, materials and colors are LOCKED to it. The LOCATION PHOTOGRAPH remains the authority for this shot's sub-space, surroundings, time of day and lighting. If the two conflict on the structure itself, the STRUCTURE LOOK photo wins; for everything else, the LOCATION PHOTOGRAPH wins.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Approaching large vehicle in the middle-left of the frame, foreground, moves toward 찰리's running path in the right midground.\n- KEY BACKGROUND ELEMENTS: Approaching large vehicle (Moving rapidly toward 찰리 before the emergency stop) — Its front and one side are visible diagonally from the road edge; used as Forms the left foreground threat while remaining fully within the frame; Road (찰리 and the approaching vehicle occupy converging paths) — The road extends diagonally from the foreground toward the distance; used as Keeps the collision geometry and remaining separation visible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the approaching headlights and the streetlights' established switching behavior to shape the nighttime threat, keeping the road and 찰리 readable without added atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Streetlights continue switching on and off along Charlie's route, and the approaching vehicle has its headlights on. Charlie retains his old coat and hat; the small bird has already flown away.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S27sh12__bgfirst_bg.png",
     "asset_id": "164bd776-d0dc-43e4-bf32-6ef9f6356b78",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S27sh12.png",
     "asset_id": "29e8ef08-4d4f-4022-ac9f-847c4599b99d",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its spatial layout, surroundings, fixed features, time of day and lighting mood are spatial truth; stage the moment inside this place. If a STRUCTURE LOOK photograph is also attached, that photo wins for the fixed structure itself — this photograph wins for everything around it. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L177B01.png",
     "asset_id": "785d4e92-6642-43a7-a3f4-70e3c4f7a9b3",
     "role": "location_plate"
    },
    {
     "label": "STRUCTURE LOOK — the confirmed photograph of the fixed structure at this location: wherever the structure appears in the frame, its shape, proportions, materials, colors and openings are LOCKED to this photo. Never copy its camera framing, time of day or lighting — the shot text and the LOCATION PHOTOGRAPH are the authorities for those.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_sewer_and_manhole_sel.png",
     "asset_id": "2b4b407b-73bc-405e-8265-11c6b8d30ef7",
     "role": "structure_seed_look"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "트럭은 화면 정면으로 향하고 있으며, 찰리는 트럭을 등지고 화면 우측을 향해 달리고 있음.",
    "built_space": "로케이션 사진과 일치하는 야간 골목길 구조와 가로등 배치를 보여줌.",
    "entities": "트럭과 코트, 모자를 착용한 찰리. 하지만 찰리의 체형이 인간형이며 2D 평면 일러스트로 렌더링됨.",
    "hard_violations": [
     "[gemini-pro] 콜라주/그래픽 오버레이 (3D 실사 배경 위에 2D 평면 캐릭터가 스티커처럼 합성됨)"
    ],
    "physics": "오른발이 지면에 닿아 몸을 지탱하고 있으나, 캐릭터가 평면적으로 합성되어 입체적인 공간감과 물리적 접촉이 성립하지 않음."
   },
   {
    "label": "B",
    "direction": "트럭의 전조등과 주행 방향이 찰리의 이동 경로를 향해 대각선으로 교차함. 찰리는 화면 좌측면으로 몸을 틀어 이동 중.",
    "built_space": "로케이션 사진의 건물 및 가로등 배치와 함께 전경에 구조물 참조 사진의 열린 맨홀이 올바르게 렌더링됨.",
    "entities": "거대한 트럭과 참조 사진의 고릴라형 장갑 및 마스크를 정확히 재현한 찰리. 코트와 모자는 누락됨.",
    "hard_violations": [
     "[gpt-high] 전면 번호판의 '23-80'으로 읽히는 숫자가 노출되어, 어디에도 판독 가능한 글자를 두지 말라는 금지 조건을 위반한다."
    ],
    "physics": "달리기 동작의 도약 단계로 두 발이 지면에서 살짝 떠 있으나, 질주하는 관성과 역동적인 자세로 물리적 움직임이 충분히 설명됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "찰리의 코트와 모자가 누락되었으나, 지정된 구도와 3D 모델링의 질감, 고릴라형 외형 및 트럭의 돌진 상황을 매우 사실적으로 구현했습니다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "3D 배경에 2D 만화 캐릭터가 스티커처럼 합성되어 재질 사실주의 규정을 심각하게 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "트럭은 화면 정면으로 향하고 있으며, 찰리는 트럭을 등지고 화면 우측을 향해 달리고 있음.",
        "built_space": "로케이션 사진과 일치하는 야간 골목길 구조와 가로등 배치를 보여줌.",
        "entities": "트럭과 코트, 모자를 착용한 찰리. 하지만 찰리의 체형이 인간형이며 2D 평면 일러스트로 렌더링됨.",
        "hard_violations": [
         "콜라주/그래픽 오버레이 (3D 실사 배경 위에 2D 평면 캐릭터가 스티커처럼 합성됨)"
        ],
        "physics": "오른발이 지면에 닿아 몸을 지탱하고 있으나, 캐릭터가 평면적으로 합성되어 입체적인 공간감과 물리적 접촉이 성립하지 않음."
       },
       {
        "label": "B",
        "direction": "트럭의 전조등과 주행 방향이 찰리의 이동 경로를 향해 대각선으로 교차함. 찰리는 화면 좌측면으로 몸을 틀어 이동 중.",
        "built_space": "로케이션 사진의 건물 및 가로등 배치와 함께 전경에 구조물 참조 사진의 열린 맨홀이 올바르게 렌더링됨.",
        "entities": "거대한 트럭과 참조 사진의 고릴라형 장갑 및 마스크를 정확히 재현한 찰리. 코트와 모자는 누락됨.",
        "hard_violations": [],
        "physics": "달리기 동작의 도약 단계로 두 발이 지면에서 살짝 떠 있으나, 질주하는 관성과 역동적인 자세로 물리적 움직임이 충분히 설명됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "찰리의 코트와 모자가 누락되었으나, 지정된 구도와 3D 모델링의 질감, 고릴라형 외형 및 트럭의 돌진 상황을 매우 사실적으로 구현했습니다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "3D 배경에 2D 만화 캐릭터가 스티커처럼 합성되어 재질 사실주의 규정을 심각하게 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "트럭은 화면 정면으로 향하고 있으며, 찰리는 트럭을 등지고 화면 우측을 향해 달리고 있음.",
        "built_space": "로케이션 사진과 일치하는 야간 골목길 구조와 가로등 배치를 보여줌.",
        "entities": "트럭과 코트, 모자를 착용한 찰리. 하지만 찰리의 체형이 인간형이며 2D 평면 일러스트로 렌더링됨.",
        "hard_violations": [
         "콜라주/그래픽 오버레이 (3D 실사 배경 위에 2D 평면 캐릭터가 스티커처럼 합성됨)"
        ],
        "physics": "오른발이 지면에 닿아 몸을 지탱하고 있으나, 캐릭터가 평면적으로 합성되어 입체적인 공간감과 물리적 접촉이 성립하지 않음."
       },
       {
        "label": "B",
        "direction": "트럭의 전조등과 주행 방향이 찰리의 이동 경로를 향해 대각선으로 교차함. 찰리는 화면 좌측면으로 몸을 틀어 이동 중.",
        "built_space": "로케이션 사진의 건물 및 가로등 배치와 함께 전경에 구조물 참조 사진의 열린 맨홀이 올바르게 렌더링됨.",
        "entities": "거대한 트럭과 참조 사진의 고릴라형 장갑 및 마스크를 정확히 재현한 찰리. 코트와 모자는 누락됨.",
        "hard_violations": [],
        "physics": "달리기 동작의 도약 단계로 두 발이 지면에서 살짝 떠 있으나, 질주하는 관성과 역동적인 자세로 물리적 움직임이 충분히 설명됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "차량과 찰리의 교차 경로 및 장소 외벽은 더 잘 맞지만, 판독 가능한 번호판이 금지 조건을 위반하고 코트와 모자도 빠졌다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "차량 전체가 들어오는 야간 와이드숏과 코트·모자는 맞지만, 차량이 찰리에게 돌진하는 충돌 방향과 지정 장소의 구조 재현은 부족하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "트럭은 왼쪽에서 화면 오른쪽 아래를 향해 비스듬히 전진하고, 오른쪽 찰리는 왼쪽 아래로 달리며 얼굴도 트럭 쪽으로 돌렸다. 두 진행 경로가 사이의 도로에서 만날 가능성이 보여 충돌 위협은 비교적 명확하다. 전조등은 트럭 앞쪽 노면을 비춘다.",
        "built_space": "중앙의 좁은 골목과 양옆 낡은 건물, 왼쪽 차양 및 셔터, 오른쪽 창과 셔터가 보인다. 가까운 가로등은 왼쪽 주황빛 한 개와 오른쪽 흰빛 한 개다. 전경에는 열린 맨홀 한 개와 분리된 뚜껑 한 개, 왼쪽에는 닫힌 맨홀 한 개가 있다. 열린 맨홀의 내부 방사형 철물은 구조 참조와 유사하지만, 이를 화면 전경의 큰 주제로 삼아 차량과 찰리 사이의 도로보다 강조했다. 도로도 요구한 대각선 원근보다 횡방향으로 읽힌다.",
        "entities": "차량 한 대와 찰리 한 명이 보이며 추가 인물이나 새는 없다. 찰리는 흰 마스크형 얼굴, 샌드 베이지 장갑, 넓은 어깨와 긴 팔로 인물 참조에 가깝다. 얼굴이 기계 마스크로 가려져 인간의 나이·성별·민족성은 판별할 수 없다. 반드시 유지해야 할 낡은 코트와 모자가 모두 없다. 전조등은 켜져 있고 밤이지만, 타이어 주변과 왼쪽 노면에 추가된 연무가 보인다. 번호판에는 숫자가 판독된다.",
        "hard_violations": [
         "전면 번호판의 '23-80'으로 읽히는 숫자가 노출되어, 어디에도 판독 가능한 글자를 두지 말라는 금지 조건을 위반한다."
        ],
        "physics": "찰리의 화면 오른쪽 발은 노면을 지지하고 반대쪽 발은 들려 있다. 몸통의 전방 기울기와 벌어진 팔은 달리는 동작으로 가능하며 무지지 부유는 아니다. 트럭은 타이어로 도로에 지지된다. 맨홀 뚜껑은 구멍 뒤쪽 노면에 놓여 있고 내부 철물은 맨홀 테두리에 연결되어 있다."
       },
       {
        "label": "B",
        "direction": "트럭의 정면과 전조등은 주로 화면 왼쪽 아래의 카메라 쪽을 향한다. 찰리는 오른쪽 아래로 달리며 얼굴도 오른쪽 진행 방향을 본다. 트럭의 현재 진행 축이 찰리의 경로로 꺾이거나 수렴하는 모습은 뚜렷하지 않아, 찰리를 향한 돌진보다 서로 다른 차로에서 전진하는 장면에 가깝다.",
        "built_space": "도로가 화면 아래에서 중앙 원경으로 길게 이어지고 양옆에 낡은 건물과 셔터가 늘어서 있다. 가까운 가로등은 왼쪽 주황빛 한 개와 오른쪽 흰빛 한 개이며 원경에도 여러 등이 이어진다. 왼쪽 전경에는 닫힌 맨홀 한 개가 보이지만 구조 참조의 열린 구멍, 내부 방사형 철물, 분리된 뚜껑은 없다. 장소 참조의 두 건물 사이 좁은 골목을 중심으로 한 배치보다는 긴 직선 도로로 바뀌었다. 트럭 전체는 왼쪽에 들어오지만 찰리는 오른쪽 중경보다 전경에 가까워 보인다.",
        "entities": "대형 화물차 한 대와 찰리 한 명만 보이며 새나 추가 인물은 없다. 찰리는 흰 마스크형 얼굴과 베이지 장갑에 낡은 코트와 챙 있는 모자를 착용했다. 다만 인물 참조보다 몸통과 팔이 가늘고 다리가 길어 고릴라형 체격의 일치도는 낮다. 가려진 얼굴로 인간의 나이·성별·민족성을 판별할 수 없다. 전조등과 가로등이 켜진 밤이며 번호판은 판독되지 않는다. 가로등의 점멸 동작 자체는 정지 이미지로 확인할 수 없다.",
        "hard_violations": [],
        "physics": "찰리의 앞쪽 발끝이 노면에 닿아 체중을 받는 것으로 보이고 뒤쪽 다리는 무릎을 굽혀 들려 있다. 팔의 앞뒤 움직임과 뒤로 날리는 코트는 달리기로 가능한 자세다. 모자는 머리에 얹혀 있고 트럭의 바퀴들은 도로에 접지한다. 지지 없이 떠 있는 인물이나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "차량과 찰리의 교차 경로 및 장소 외벽은 더 잘 맞지만, 판독 가능한 번호판이 금지 조건을 위반하고 코트와 모자도 빠졌다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "차량 전체가 들어오는 야간 와이드숏과 코트·모자는 맞지만, 차량이 찰리에게 돌진하는 충돌 방향과 지정 장소의 구조 재현은 부족하다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "트럭은 왼쪽에서 화면 오른쪽 아래를 향해 비스듬히 전진하고, 오른쪽 찰리는 왼쪽 아래로 달리며 얼굴도 트럭 쪽으로 돌렸다. 두 진행 경로가 사이의 도로에서 만날 가능성이 보여 충돌 위협은 비교적 명확하다. 전조등은 트럭 앞쪽 노면을 비춘다.",
        "built_space": "중앙의 좁은 골목과 양옆 낡은 건물, 왼쪽 차양 및 셔터, 오른쪽 창과 셔터가 보인다. 가까운 가로등은 왼쪽 주황빛 한 개와 오른쪽 흰빛 한 개다. 전경에는 열린 맨홀 한 개와 분리된 뚜껑 한 개, 왼쪽에는 닫힌 맨홀 한 개가 있다. 열린 맨홀의 내부 방사형 철물은 구조 참조와 유사하지만, 이를 화면 전경의 큰 주제로 삼아 차량과 찰리 사이의 도로보다 강조했다. 도로도 요구한 대각선 원근보다 횡방향으로 읽힌다.",
        "entities": "차량 한 대와 찰리 한 명이 보이며 추가 인물이나 새는 없다. 찰리는 흰 마스크형 얼굴, 샌드 베이지 장갑, 넓은 어깨와 긴 팔로 인물 참조에 가깝다. 얼굴이 기계 마스크로 가려져 인간의 나이·성별·민족성은 판별할 수 없다. 반드시 유지해야 할 낡은 코트와 모자가 모두 없다. 전조등은 켜져 있고 밤이지만, 타이어 주변과 왼쪽 노면에 추가된 연무가 보인다. 번호판에는 숫자가 판독된다.",
        "hard_violations": [
         "전면 번호판의 '23-80'으로 읽히는 숫자가 노출되어, 어디에도 판독 가능한 글자를 두지 말라는 금지 조건을 위반한다."
        ],
        "physics": "찰리의 화면 오른쪽 발은 노면을 지지하고 반대쪽 발은 들려 있다. 몸통의 전방 기울기와 벌어진 팔은 달리는 동작으로 가능하며 무지지 부유는 아니다. 트럭은 타이어로 도로에 지지된다. 맨홀 뚜껑은 구멍 뒤쪽 노면에 놓여 있고 내부 철물은 맨홀 테두리에 연결되어 있다."
       },
       {
        "label": "A",
        "direction": "트럭의 정면과 전조등은 주로 화면 왼쪽 아래의 카메라 쪽을 향한다. 찰리는 오른쪽 아래로 달리며 얼굴도 오른쪽 진행 방향을 본다. 트럭의 현재 진행 축이 찰리의 경로로 꺾이거나 수렴하는 모습은 뚜렷하지 않아, 찰리를 향한 돌진보다 서로 다른 차로에서 전진하는 장면에 가깝다.",
        "built_space": "도로가 화면 아래에서 중앙 원경으로 길게 이어지고 양옆에 낡은 건물과 셔터가 늘어서 있다. 가까운 가로등은 왼쪽 주황빛 한 개와 오른쪽 흰빛 한 개이며 원경에도 여러 등이 이어진다. 왼쪽 전경에는 닫힌 맨홀 한 개가 보이지만 구조 참조의 열린 구멍, 내부 방사형 철물, 분리된 뚜껑은 없다. 장소 참조의 두 건물 사이 좁은 골목을 중심으로 한 배치보다는 긴 직선 도로로 바뀌었다. 트럭 전체는 왼쪽에 들어오지만 찰리는 오른쪽 중경보다 전경에 가까워 보인다.",
        "entities": "대형 화물차 한 대와 찰리 한 명만 보이며 새나 추가 인물은 없다. 찰리는 흰 마스크형 얼굴과 베이지 장갑에 낡은 코트와 챙 있는 모자를 착용했다. 다만 인물 참조보다 몸통과 팔이 가늘고 다리가 길어 고릴라형 체격의 일치도는 낮다. 가려진 얼굴로 인간의 나이·성별·민족성을 판별할 수 없다. 전조등과 가로등이 켜진 밤이며 번호판은 판독되지 않는다. 가로등의 점멸 동작 자체는 정지 이미지로 확인할 수 없다.",
        "hard_violations": [],
        "physics": "찰리의 앞쪽 발끝이 노면에 닿아 체중을 받는 것으로 보이고 뒤쪽 다리는 무릎을 굽혀 들려 있다. 팔의 앞뒤 움직임과 뒤로 날리는 코트는 달리기로 가능한 자세다. 모자는 머리에 얹혀 있고 트럭의 바퀴들은 도로에 접지한다. 지지 없이 떠 있는 인물이나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.286,
    "B": 1.6
   },
   "adjusted": {
    "A": 1.036,
    "B": 1.35
   },
   "violations": {
    "A": [
     "[gemini-pro] 콜라주/그래픽 오버레이 (3D 실사 배경 위에 2D 평면 캐릭터가 스티커처럼 합성됨)"
    ],
    "B": [
     "[gpt-high] 전면 번호판의 '23-80'으로 읽히는 숫자가 노출되어, 어디에도 판독 가능한 글자를 두지 말라는 금지 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1350,
   "A": 1036
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1350,
    "verdict_ko": "찰리의 코트와 모자가 누락되었으나, 지정된 구도와 3D 모델링의 질감, 고릴라형 외형 및 트럭의 돌진 상황을 매우 사실적으로 구현했습니다.  ★위반: [gpt-high] 전면 번호판의 '23-80'으로 읽히는 숫자가 노출되어, 어디에도 판독 가능한 글자를 두지 말라는 금지 조건을 위반한다."
   },
   {
    "label": "A",
    "score": 1036,
    "verdict_ko": "3D 배경에 2D 만화 캐릭터가 스티커처럼 합성되어 재질 사실주의 규정을 심각하게 위반했습니다.  ★위반: [gemini-pro] 콜라주/그래픽 오버레이 (3D 실사 배경 위에 2D 평면 캐릭터가 스티커처럼 합성됨)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its spatial layout, surroundings, fixed features, time of day and lighting mood are spatial truth; stage the moment inside this place. If a STRUCTURE LOOK photograph is also attached, that photo wins for the fixed structure itself — this photograph wins for everything around it. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L177B01.png",
    "asset_id": "785d4e92-6642-43a7-a3f4-70e3c4f7a9b3",
    "role": "location_plate"
   },
   {
    "label": "STRUCTURE LOOK — the confirmed photograph of the fixed structure at this location: wherever the structure appears in the frame, its shape, proportions, materials, colors and openings are LOCKED to this photo. Never copy its camera framing, time of day or lighting — the shot text and the LOCATION PHOTOGRAPH are the authorities for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_sewer_and_manhole_sel.png",
    "asset_id": "2b4b407b-73bc-405e-8265-11c6b8d30ef7",
    "role": "structure_seed_look"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-301c-7886-b591-21f84241fa9e",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S27sh12__bgfirst_bg.png",
   "bg_asset_id": "164bd776-d0dc-43e4-bf32-6ef9f6356b78",
   "bg_record_key": "S27sh12::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate",
   "seed_attached": true
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  },
  "lane_policy": "ab_select_ready"
 },
 "S27sh12::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:17:44.631833+00:00",
  "fingerprint": "f1d9d54b12f9f7185daa472ab71b405a9ce3e07e8f7d8ad9509eb0e0b8a71426",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S27sh12_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S27sh12_sel.png",
  "source_sha256": "07b21b12d86c97cb0b43f8768c121a8a4b9980d0dd7f1777b761315674f70e9f",
  "file": "S27sh12_cine.png",
  "staged_sha256": "f60fddb97143c9c640fbf0d66c8f5737949f78c45c0784b18a2c795d80179493",
  "latency_ms": 11927
 },
 "S27sh18::signage": {
  "fp": "df1c5668dd64f889",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S27sh18": {
  "input_fingerprint": "aad2dc4585241023",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 헤드라이트 불빛을 등진 채 찰리 쪽을 향해 한 발을 내디딘 신부의 짙은 실루엣 전신.\n\nLOCATION (lock): On the road immediately in front of a stopped van in the refugee settlement, backlit by its headlights. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Stopped vehicle (Stopped after braking in front of 찰리) — Its front faces obliquely toward the camera behind 신부; used as Anchors the approaching figure to the vehicle from which he emerged; Road between 찰리 and 신부 (신부 is crossing the remaining separation on foot) — The visible ground connects the lower-left foreground to the center-right midground; used as Makes the approach and the natural upward eyeline physically legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The vehicle's headlights backlight 신부 into a dense silhouette, withholding his facial detail at this first step while preserving his full-body outline.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the same roadway, immediate roadside surroundings, vehicle exterior, and nighttime headlight illumination. Exclude any transient motion effects from the vehicle's approach.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The van has stopped abruptly and its headlights remain on. Charlie, still wearing the old coat and hat, stands in front of it; the bird is no longer perched on him. 신부: He has stepped out of the van and is approaching on foot, still wearing his clerical collar.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 헤드라이트 불빛을 등진 채 찰리 쪽을 향해 한 발을 내디딘 신부의 짙은 실루엣 전신.\n\nLOCATION (lock): On the road immediately in front of a stopped van in the refugee settlement, backlit by its headlights. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Stopped vehicle (Stopped after braking in front of 찰리) — Its front faces obliquely toward the camera behind 신부; used as Anchors the approaching figure to the vehicle from which he emerged; Road between 찰리 and 신부 (신부 is crossing the remaining separation on foot) — The visible ground connects the lower-left foreground to the center-right midground; used as Makes the approach and the natural upward eyeline physically legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The vehicle's headlights backlight 신부 into a dense silhouette, withholding his facial detail at this first step while preserving his full-body outline.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the same roadway, immediate roadside surroundings, vehicle exterior, and nighttime headlight illumination. Exclude any transient motion effects from the vehicle's approach.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The van has stopped abruptly and its headlights remain on. Charlie, still wearing the old coat and hat, stands in front of it; the bird is no longer perched on him. 신부: He has stepped out of the van and is approaching on foot, still wearing his clerical collar.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 헤드라이트 불빛을 등진 채 찰리 쪽을 향해 한 발을 내디딘 신부의 짙은 실루엣 전신.\n\nLOCATION (lock): On the road immediately in front of a stopped van in the refugee settlement, backlit by its headlights. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Stopped vehicle (Stopped after braking in front of 찰리) — Its front faces obliquely toward the camera behind 신부; used as Anchors the approaching figure to the vehicle from which he emerged; Road between 찰리 and 신부 (신부 is crossing the remaining separation on foot) — The visible ground connects the lower-left foreground to the center-right midground; used as Makes the approach and the natural upward eyeline physically legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The vehicle's headlights backlight 신부 into a dense silhouette, withholding his facial detail at this first step while preserving his full-body outline.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the same roadway, immediate roadside surroundings, vehicle exterior, and nighttime headlight illumination. Exclude any transient motion effects from the vehicle's approach.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The van has stopped abruptly and its headlights remain on. Charlie, still wearing the old coat and hat, stands in front of it; the bird is no longer perched on him. 신부: He has stepped out of the van and is approaching on foot, still wearing his clerical collar.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "신부는 화면 좌측 전방을 향해 걷고 있으며 시선도 같은 방향임. 찰리는 화면에 없음.",
    "built_space": "우측에 판잣집과 텐트가 있는 도로. 레퍼런스의 셔터 건물과 맨홀 등은 전혀 반영되지 않음.",
    "entities": "신부는 레퍼런스와 인상착의가 일치하며 클레리컬 칼라를 착용함. 차량은 레퍼런스와 완전히 다른 형태임.",
    "hard_violations": [],
    "physics": "신부의 발과 트럭 바퀴 모두 지면에 닿아 자연스럽게 무게를 지탱하고 있음."
   },
   {
    "label": "B",
    "direction": "신부는 정면 카메라 방향을 향해 걸어오고 있음.",
    "built_space": "레퍼런스의 건물, 가로등, 바닥의 맨홀 2개가 정확히 일치함. 좌측에 텍스트가 적힌 표지판이 임의로 추가됨.",
    "entities": "신부는 지시대로 어두운 실루엣으로 표현됨. 트럭은 레퍼런스와 동일한 장갑차량임. 화면 좌측 하단에 정체불명의 거대한 눈과 코가 등장함.",
    "hard_violations": [
     "[gemini-pro] 화면 좌측 하단에 거대한 얼굴 일부가 떠 있는 물리적으로 불가능한 콜라주 (collage/physically impossible staging)",
     "[gemini-pro] 명시적으로 금지된 읽을 수 있는 텍스트 표지판 배치 (leaked text/invented objects)",
     "[gpt-high] 왼쪽 아래에 거대한 눈과 얼굴 일부를 별도 이미지처럼 합성하여 단일 실사 장면을 깨뜨렸다.",
     "[gpt-high] 신부의 전신 외에 허용되지 않은 추가 인물의 얼굴 일부가 등장한다.",
     "[gpt-high] 참고에 없던 안내판에 '난민구역 출입통제'라는 읽을 수 있는 문구를 노출했다."
    ],
    "physics": "신부와 트럭은 지면에 닿아 있으나, 좌측 하단의 거대한 얼굴은 아무런 지지체 없이 허공에 겹쳐져 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "배경을 판자촌으로 왜곡하고 얼굴이 밝게 드러나 조명 지시를 어겼으나, 치명적인 구조적 오류나 콜라주 현상이 없어 차악으로 선정됨."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "배경 위치와 실루엣 조명은 완벽히 구현했으나, 좌측 하단의 거대한 얼굴 콜라주와 금지된 텍스트 표지판으로 인해 실격됨."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "신부는 화면 좌측 전방을 향해 걷고 있으며 시선도 같은 방향임. 찰리는 화면에 없음.",
        "built_space": "우측에 판잣집과 텐트가 있는 도로. 레퍼런스의 셔터 건물과 맨홀 등은 전혀 반영되지 않음.",
        "entities": "신부는 레퍼런스와 인상착의가 일치하며 클레리컬 칼라를 착용함. 차량은 레퍼런스와 완전히 다른 형태임.",
        "hard_violations": [],
        "physics": "신부의 발과 트럭 바퀴 모두 지면에 닿아 자연스럽게 무게를 지탱하고 있음."
       },
       {
        "label": "B",
        "direction": "신부는 정면 카메라 방향을 향해 걸어오고 있음.",
        "built_space": "레퍼런스의 건물, 가로등, 바닥의 맨홀 2개가 정확히 일치함. 좌측에 텍스트가 적힌 표지판이 임의로 추가됨.",
        "entities": "신부는 지시대로 어두운 실루엣으로 표현됨. 트럭은 레퍼런스와 동일한 장갑차량임. 화면 좌측 하단에 정체불명의 거대한 눈과 코가 등장함.",
        "hard_violations": [
         "화면 좌측 하단에 거대한 얼굴 일부가 떠 있는 물리적으로 불가능한 콜라주 (collage/physically impossible staging)",
         "명시적으로 금지된 읽을 수 있는 텍스트 표지판 배치 (leaked text/invented objects)"
        ],
        "physics": "신부와 트럭은 지면에 닿아 있으나, 좌측 하단의 거대한 얼굴은 아무런 지지체 없이 허공에 겹쳐져 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "배경을 판자촌으로 왜곡하고 얼굴이 밝게 드러나 조명 지시를 어겼으나, 치명적인 구조적 오류나 콜라주 현상이 없어 차악으로 선정됨."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "배경 위치와 실루엣 조명은 완벽히 구현했으나, 좌측 하단의 거대한 얼굴 콜라주와 금지된 텍스트 표지판으로 인해 실격됨."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "신부는 화면 좌측 전방을 향해 걷고 있으며 시선도 같은 방향임. 찰리는 화면에 없음.",
        "built_space": "우측에 판잣집과 텐트가 있는 도로. 레퍼런스의 셔터 건물과 맨홀 등은 전혀 반영되지 않음.",
        "entities": "신부는 레퍼런스와 인상착의가 일치하며 클레리컬 칼라를 착용함. 차량은 레퍼런스와 완전히 다른 형태임.",
        "hard_violations": [],
        "physics": "신부의 발과 트럭 바퀴 모두 지면에 닿아 자연스럽게 무게를 지탱하고 있음."
       },
       {
        "label": "B",
        "direction": "신부는 정면 카메라 방향을 향해 걸어오고 있음.",
        "built_space": "레퍼런스의 건물, 가로등, 바닥의 맨홀 2개가 정확히 일치함. 좌측에 텍스트가 적힌 표지판이 임의로 추가됨.",
        "entities": "신부는 지시대로 어두운 실루엣으로 표현됨. 트럭은 레퍼런스와 동일한 장갑차량임. 화면 좌측 하단에 정체불명의 거대한 눈과 코가 등장함.",
        "hard_violations": [
         "화면 좌측 하단에 거대한 얼굴 일부가 떠 있는 물리적으로 불가능한 콜라주 (collage/physically impossible staging)",
         "명시적으로 금지된 읽을 수 있는 텍스트 표지판 배치 (leaked text/invented objects)"
        ],
        "physics": "신부와 트럭은 지면에 닿아 있으나, 좌측 하단의 거대한 얼굴은 아무런 지지체 없이 허공에 겹쳐져 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 1,
        "verdict_ko": "역광 속 신부의 전신과 기존 건물은 살렸지만, 거대한 눈 합성과 추가 인물 신체, 읽히는 안내판이 명백한 실격 요소다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "한 명의 신부가 접지한 채 걸어가는 전신 와이드숏은 성립하지만, 오른쪽으로 향하는 동선과 천막촌으로 바뀐 장소가 지정된 접근 구도 및 장소 연속성을 어긴다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "신부의 얼굴과 몸은 거의 카메라 정면을 향하고 한 발을 앞으로 내민다. 왼쪽 아래 화면 밖 찰리에게 향하는 대각선 접근은 분명하지 않다. 차량 전면과 켜진 조명은 신부 뒤에서 카메라 쪽을 향한다. 왼쪽 아래에는 별도로 합성된 눈이 있지만 이를 공간 속 찰리의 시선으로 볼 수는 없다.",
        "built_space": "양옆의 낡은 건물, 셔터, 줄무늬 차양, 왼쪽 주황색 가로등 한 개와 오른쪽 흰색 가로등 한 개가 참고 장소와 유사하다. 전경에는 열린 맨홀 한 개와 그 뒤에 놓인 뚜껑 한 개가 보인다. 차량 한 대는 신부 뒤 중앙에서 오른쪽을 차지한다. 참고에 없던 큰 안내판이 왼쪽 도로변에 추가되었고, 눈 합성이 왼쪽 아래 도로 공간을 가린다.",
        "entities": "검은 성직자복과 흰 성직자 칼라를 착용한 남성 한 명의 전신이 보인다. 얼굴은 어두워 참고 인물의 한국인 60대 외모와 주름을 정확히 대조하기 어렵지만 짧은 머리와 체격은 대체로 양립한다. 차량은 참고와 비슷한 대형 특장 트럭이다. 별도의 거대한 눈과 얼굴 일부는 허용된 단일 인물 구성에 맞지 않는다. 안내판의 '난민구역 출입통제'가 선명하게 읽힌다. 새는 없다.",
        "hard_violations": [
         "왼쪽 아래에 거대한 눈과 얼굴 일부를 별도 이미지처럼 합성하여 단일 실사 장면을 깨뜨렸다.",
         "신부의 전신 외에 허용되지 않은 추가 인물의 얼굴 일부가 등장한다.",
         "참고에 없던 안내판에 '난민구역 출입통제'라는 읽을 수 있는 문구를 노출했다."
        ],
        "physics": "신부는 뒤쪽 신발로 노면을 지지하고 앞쪽 발을 내미는 보행 순간으로 보인다. 다리가 다소 교차하지만 공중에 뜬 몸은 아니다. 차량은 타이어로, 맨홀 뚜껑은 도로면으로 지지된다. 거대한 눈은 실제 공간에서 지지되는 신체가 아니라 화면 위 합성 요소로 보인다."
       },
       {
        "label": "B",
        "direction": "신부의 얼굴, 몸통과 앞발은 화면 오른쪽을 향한다. 시선도 오른쪽 화면 밖을 향해, 왼쪽 아래 전경의 찰리 쪽으로 다가오는 지정 동선과 반대다. 차량은 신부 뒤에서 약간 비스듬한 전면을 카메라에 보이며 헤드라이트를 켜고 있다.",
        "built_space": "차량 한 대 앞 도로에 신부 한 명이 서서 걷고 있다. 양쪽에는 여러 천막과 판잣집, 전신주 및 도로변 물통들이 이어진다. 참고의 콘크리트 건물, 셔터와 줄무늬 차양 대신 전혀 다른 가로 공간이 제시되어 같은 장소로 읽히지 않는다. 왼쪽 아래에서 중앙 오른쪽으로 이어지는 도로 면은 있으나 인물의 진행 방향과 목표 배치가 요구한 관계를 만들지 못한다.",
        "entities": "짧은 회색 섞인 머리의 나이 든 동아시아계 남성 한 명이 검은 성직자복과 흰 칼라를 착용한다. 참고 인물과 대체로 양립하지만 얼굴 각도와 어둠 때문에 정확한 동일성은 확인하기 어렵다. 얼굴 옆면이 드러나 첫 등장 때 얼굴 세부를 감추라는 요구보다 밝다. 차량은 참고처럼 상자형 후부를 가진 대형 트럭이지만 상부 조명 등 외관 세부가 다르다. 추가 인물과 새는 없고 읽히는 글씨도 보이지 않는다.",
        "hard_violations": [],
        "physics": "앞발은 노면에 닿고 뒷발은 뒤꿈치를 들며 발 앞부분으로 지지되어 자연스러운 보행 체중 이동이 성립한다. 팔은 몸 옆에서 걸음에 맞게 놓여 있고 들고 있는 물건은 없다. 차량은 타이어로 도로에 지지되며 인물의 그림자도 노면 위에 이어진다. 지지 없이 떠 있는 신체나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "역광 속 신부의 전신과 기존 건물은 살렸지만, 거대한 눈 합성과 추가 인물 신체, 읽히는 안내판이 명백한 실격 요소다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "한 명의 신부가 접지한 채 걸어가는 전신 와이드숏은 성립하지만, 오른쪽으로 향하는 동선과 천막촌으로 바뀐 장소가 지정된 접근 구도 및 장소 연속성을 어긴다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "신부의 얼굴과 몸은 거의 카메라 정면을 향하고 한 발을 앞으로 내민다. 왼쪽 아래 화면 밖 찰리에게 향하는 대각선 접근은 분명하지 않다. 차량 전면과 켜진 조명은 신부 뒤에서 카메라 쪽을 향한다. 왼쪽 아래에는 별도로 합성된 눈이 있지만 이를 공간 속 찰리의 시선으로 볼 수는 없다.",
        "built_space": "양옆의 낡은 건물, 셔터, 줄무늬 차양, 왼쪽 주황색 가로등 한 개와 오른쪽 흰색 가로등 한 개가 참고 장소와 유사하다. 전경에는 열린 맨홀 한 개와 그 뒤에 놓인 뚜껑 한 개가 보인다. 차량 한 대는 신부 뒤 중앙에서 오른쪽을 차지한다. 참고에 없던 큰 안내판이 왼쪽 도로변에 추가되었고, 눈 합성이 왼쪽 아래 도로 공간을 가린다.",
        "entities": "검은 성직자복과 흰 성직자 칼라를 착용한 남성 한 명의 전신이 보인다. 얼굴은 어두워 참고 인물의 한국인 60대 외모와 주름을 정확히 대조하기 어렵지만 짧은 머리와 체격은 대체로 양립한다. 차량은 참고와 비슷한 대형 특장 트럭이다. 별도의 거대한 눈과 얼굴 일부는 허용된 단일 인물 구성에 맞지 않는다. 안내판의 '난민구역 출입통제'가 선명하게 읽힌다. 새는 없다.",
        "hard_violations": [
         "왼쪽 아래에 거대한 눈과 얼굴 일부를 별도 이미지처럼 합성하여 단일 실사 장면을 깨뜨렸다.",
         "신부의 전신 외에 허용되지 않은 추가 인물의 얼굴 일부가 등장한다.",
         "참고에 없던 안내판에 '난민구역 출입통제'라는 읽을 수 있는 문구를 노출했다."
        ],
        "physics": "신부는 뒤쪽 신발로 노면을 지지하고 앞쪽 발을 내미는 보행 순간으로 보인다. 다리가 다소 교차하지만 공중에 뜬 몸은 아니다. 차량은 타이어로, 맨홀 뚜껑은 도로면으로 지지된다. 거대한 눈은 실제 공간에서 지지되는 신체가 아니라 화면 위 합성 요소로 보인다."
       },
       {
        "label": "A",
        "direction": "신부의 얼굴, 몸통과 앞발은 화면 오른쪽을 향한다. 시선도 오른쪽 화면 밖을 향해, 왼쪽 아래 전경의 찰리 쪽으로 다가오는 지정 동선과 반대다. 차량은 신부 뒤에서 약간 비스듬한 전면을 카메라에 보이며 헤드라이트를 켜고 있다.",
        "built_space": "차량 한 대 앞 도로에 신부 한 명이 서서 걷고 있다. 양쪽에는 여러 천막과 판잣집, 전신주 및 도로변 물통들이 이어진다. 참고의 콘크리트 건물, 셔터와 줄무늬 차양 대신 전혀 다른 가로 공간이 제시되어 같은 장소로 읽히지 않는다. 왼쪽 아래에서 중앙 오른쪽으로 이어지는 도로 면은 있으나 인물의 진행 방향과 목표 배치가 요구한 관계를 만들지 못한다.",
        "entities": "짧은 회색 섞인 머리의 나이 든 동아시아계 남성 한 명이 검은 성직자복과 흰 칼라를 착용한다. 참고 인물과 대체로 양립하지만 얼굴 각도와 어둠 때문에 정확한 동일성은 확인하기 어렵다. 얼굴 옆면이 드러나 첫 등장 때 얼굴 세부를 감추라는 요구보다 밝다. 차량은 참고처럼 상자형 후부를 가진 대형 트럭이지만 상부 조명 등 외관 세부가 다르다. 추가 인물과 새는 없고 읽히는 글씨도 보이지 않는다.",
        "hard_violations": [],
        "physics": "앞발은 노면에 닿고 뒷발은 뒤꿈치를 들며 발 앞부분으로 지지되어 자연스러운 보행 체중 이동이 성립한다. 팔은 몸 옆에서 걸음에 맞게 놓여 있고 들고 있는 물건은 없다. 차량은 타이어로 도로에 지지되며 인물의 그림자도 노면 위에 이어진다. 지지 없이 떠 있는 신체나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.85
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.6
   },
   "violations": {
    "B": [
     "[gemini-pro] 화면 좌측 하단에 거대한 얼굴 일부가 떠 있는 물리적으로 불가능한 콜라주 (collage/physically impossible staging)",
     "[gemini-pro] 명시적으로 금지된 읽을 수 있는 텍스트 표지판 배치 (leaked text/invented objects)",
     "[gpt-high] 왼쪽 아래에 거대한 눈과 얼굴 일부를 별도 이미지처럼 합성하여 단일 실사 장면을 깨뜨렸다.",
     "[gpt-high] 신부의 전신 외에 허용되지 않은 추가 인물의 얼굴 일부가 등장한다.",
     "[gpt-high] 참고에 없던 안내판에 '난민구역 출입통제'라는 읽을 수 있는 문구를 노출했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 600
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "배경을 판자촌으로 왜곡하고 얼굴이 밝게 드러나 조명 지시를 어겼으나, 치명적인 구조적 오류나 콜라주 현상이 없어 차악으로 선정됨."
   },
   {
    "label": "B",
    "score": 600,
    "verdict_ko": "배경 위치와 실루엣 조명은 완벽히 구현했으나, 좌측 하단의 거대한 얼굴 콜라주와 금지된 텍스트 표지판으로 인해 실격됨.  ★위반: [gemini-pro] 화면 좌측 하단에 거대한 얼굴 일부가 떠 있는 물리적으로 불가능한 콜라주 (collage/physically impossible staging) / [gemini-pro] 명시적으로 금지된 읽을 수 있는 텍스트 표지판 배치 (leaked text/invented objects) / [gpt-high] 왼쪽 아래에 거대한 눈과 얼굴 일부를 별도 이미지처럼 합성하여 단일 실사 장면을 깨뜨렸다. / [gpt-high] 신부의 전신 외에 허용되지 않은 추가 인물의 얼굴 일부가 등장한다. / [gpt-high] 참고에 없던 안내판에 '난민구역 출입통제'라는 읽을 수 있는 문구를 노출했다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S27sh12_sel.png",
    "asset_id": "9e4604d0-d417-4a79-b13f-fd4398dd3fbf",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 신부: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1402213>",
    "asset_id": "8696070d-ac09-4a5c-95f2-3daf1c015c32",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-33b0-765f-b2f2-627538550b76",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S27sh12"
  },
  "lane_policy": "ab_select_bypass:prev"
 },
 "S27sh18::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:18:56.785504+00:00",
  "fingerprint": "09b1615afc8ca5d4a4c6c274cb1c5ba5b9c9ebb82a3f8a6d31a31ffe938adc58",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S27sh18_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S27sh18_sel.png",
  "source_sha256": "46a342f743e23543e69bba00f47822f713dd850dbf37175b27b900016d9d5b14",
  "file": "S27sh18_cine.png",
  "staged_sha256": "b00d94928bbdb8dde10f7d028aa2fd7afa593975d0149451721657e9bd56855b",
  "latency_ms": 11097
 },
 "S27sh21::signage": {
  "fp": "3e5eeffaf2a5c370",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S27sh21": {
  "input_fingerprint": "6665f17613c66213",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 어두운 골목 모퉁이에 숨어 찰리가 타는 자동차를 예의주시하는 구도환의 날카로운 눈매 클로즈업.\n\nLOCATION (lock): At a dark alley corner overlooking the stopped van on the refugee settlement road. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Alley corner (Conceals 구도환 as he watches the vehicle) — The near edge is visible along the left side of the image; used as Provides a restrained foreground occlusion that explains his hidden position.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the alley's stated darkness with enough neutral tonal separation to read 구도환's eyes, without transferring the road's headlight effect onto his face.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The van remains stopped with its headlights on, and Charlie is boarding in the old coat and hat. The vehicle has a crude halogen cross ornament on its roof. 구도환: He is observing the boarding from outside the van.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 어두운 골목 모퉁이에 숨어 찰리가 타는 자동차를 예의주시하는 구도환의 날카로운 눈매 클로즈업.\n\nLOCATION (lock): At a dark alley corner overlooking the stopped van on the refugee settlement road. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Alley corner (Conceals 구도환 as he watches the vehicle) — The near edge is visible along the left side of the image; used as Provides a restrained foreground occlusion that explains his hidden position.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the alley's stated darkness with enough neutral tonal separation to read 구도환's eyes, without transferring the road's headlight effect onto his face.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The van remains stopped with its headlights on, and Charlie is boarding in the old coat and hat. The vehicle has a crude halogen cross ornament on its roof. 구도환: He is observing the boarding from outside the van.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 어두운 골목 모퉁이에 숨어 찰리가 타는 자동차를 예의주시하는 구도환의 날카로운 눈매 클로즈업.\n\nLOCATION (lock): At a dark alley corner overlooking the stopped van on the refugee settlement road. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Alley corner (Conceals 구도환 as he watches the vehicle) — The near edge is visible along the left side of the image; used as Provides a restrained foreground occlusion that explains his hidden position.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the alley's stated darkness with enough neutral tonal separation to read 구도환's eyes, without transferring the road's headlight effect onto his face.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The van remains stopped with its headlights on, and Charlie is boarding in the old coat and hat. The vehicle has a crude halogen cross ornament on its roof. 구도환: He is observing the boarding from outside the van.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "우측 밖을 주시하고 있으나, 프롬프트상 주시 대상이어야 할 밴이 존재하지 않음.",
    "built_space": "화면 좌측에 모퉁이 벽이 있고 배경으로 어두운 판자촌이 보이나 도로의 차량은 없음.",
    "entities": "구도환의 얼굴은 일치하나 프롬프트에서 금지한 이전 샷 인물의 사제복 깃을 그대로 입고 있음. 밴과 찰리가 누락됨.",
    "hard_violations": [
     "[gemini-pro] 이전 샷 인물의 의상(사제복)을 캐릭터에게 그대로 복사하여 적용함"
    ],
    "physics": "벽 뒤에서 얼굴만 내민 자세 자체는 물리적으로 안정적임."
   },
   {
    "label": "B",
    "direction": "시선이 우측 배경에 있는 밴과 탑승객(찰리)을 명확히 향하고 있음.",
    "built_space": "왼쪽에 전경을 가리는 골목 모퉁이 벽이 위치하고, 배경 도로에 밴이 정차할 공간이 알맞게 확보됨.",
    "entities": "구도환의 외모와 정장 의상이 레퍼런스와 일치함. 헤드라이트를 켠 밴, 지붕의 빛나는 십자가 장식, 코트와 모자를 쓴 찰리가 모두 존재함.",
    "hard_violations": [
     "[gpt-high] 구도환 외에는 등장시키지 말라는 인물 제한과 달리 차량 문 옆에 추가 인물 한 명이 보인다."
    ],
    "physics": "모퉁이 뒤에 몸을 숨기고 서서 고개를 내민 자세가 자연스럽고 지지에 문제가 없음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "구도환의 클로즈업과 전경을 가리는 골목 모퉁이 구도를 잘 살렸으며, 배경에 명시된 밴, 십자가 장식, 탑승하는 찰리를 모두 정확히 구현함."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "프롬프트에 명시된 밴과 찰리가 완전히 누락되었으며, 이전 샷 인물의 의상(사제복)을 복사하는 금지 사항을 위반함."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "시선이 우측 배경에 있는 밴과 탑승객(찰리)을 명확히 향하고 있음.",
        "built_space": "왼쪽에 전경을 가리는 골목 모퉁이 벽이 위치하고, 배경 도로에 밴이 정차할 공간이 알맞게 확보됨.",
        "entities": "구도환의 외모와 정장 의상이 레퍼런스와 일치함. 헤드라이트를 켠 밴, 지붕의 빛나는 십자가 장식, 코트와 모자를 쓴 찰리가 모두 존재함.",
        "hard_violations": [],
        "physics": "모퉁이 뒤에 몸을 숨기고 서서 고개를 내민 자세가 자연스럽고 지지에 문제가 없음."
       },
       {
        "label": "A",
        "direction": "우측 밖을 주시하고 있으나, 프롬프트상 주시 대상이어야 할 밴이 존재하지 않음.",
        "built_space": "화면 좌측에 모퉁이 벽이 있고 배경으로 어두운 판자촌이 보이나 도로의 차량은 없음.",
        "entities": "구도환의 얼굴은 일치하나 프롬프트에서 금지한 이전 샷 인물의 사제복 깃을 그대로 입고 있음. 밴과 찰리가 누락됨.",
        "hard_violations": [
         "이전 샷 인물의 의상(사제복)을 캐릭터에게 그대로 복사하여 적용함"
        ],
        "physics": "벽 뒤에서 얼굴만 내민 자세 자체는 물리적으로 안정적임."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "구도환의 클로즈업과 전경을 가리는 골목 모퉁이 구도를 잘 살렸으며, 배경에 명시된 밴, 십자가 장식, 탑승하는 찰리를 모두 정확히 구현함."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "프롬프트에 명시된 밴과 찰리가 완전히 누락되었으며, 이전 샷 인물의 의상(사제복)을 복사하는 금지 사항을 위반함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "시선이 우측 배경에 있는 밴과 탑승객(찰리)을 명확히 향하고 있음.",
        "built_space": "왼쪽에 전경을 가리는 골목 모퉁이 벽이 위치하고, 배경 도로에 밴이 정차할 공간이 알맞게 확보됨.",
        "entities": "구도환의 외모와 정장 의상이 레퍼런스와 일치함. 헤드라이트를 켠 밴, 지붕의 빛나는 십자가 장식, 코트와 모자를 쓴 찰리가 모두 존재함.",
        "hard_violations": [],
        "physics": "모퉁이 뒤에 몸을 숨기고 서서 고개를 내민 자세가 자연스럽고 지지에 문제가 없음."
       },
       {
        "label": "A",
        "direction": "우측 밖을 주시하고 있으나, 프롬프트상 주시 대상이어야 할 밴이 존재하지 않음.",
        "built_space": "화면 좌측에 모퉁이 벽이 있고 배경으로 어두운 판자촌이 보이나 도로의 차량은 없음.",
        "entities": "구도환의 얼굴은 일치하나 프롬프트에서 금지한 이전 샷 인물의 사제복 깃을 그대로 입고 있음. 밴과 찰리가 누락됨.",
        "hard_violations": [
         "이전 샷 인물의 의상(사제복)을 캐릭터에게 그대로 복사하여 적용함"
        ],
        "physics": "벽 뒤에서 얼굴만 내민 자세 자체는 물리적으로 안정적임."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "차량은 오른쪽 뒤에 있는데 시선은 왼쪽을 향하며, 허용되지 않은 추가 인물까지 보여 눈매 중심의 은밀한 관찰 장면을 벗어난다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "왼쪽 모퉁이에 숨은 얼굴과 날카로운 눈매의 클로즈업을 충실히 구현하지만, 성직자 칼라는 구도환의 의상 참조와 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "구도환의 두 눈은 화면 왼쪽 바깥을 향한다. 관찰해야 할 차량과 승차 위치는 그의 오른쪽 뒤에 보여 시선이 목표에 닿지 않는다. 차량 옆 추가 인물은 차문 쪽을 향한다.",
        "built_space": "왼쪽의 거친 벽 모퉁이 하나가 화면 약 3분의 1을 가리고, 구도환은 그 오른쪽으로 머리와 상체를 내민다. 오른쪽 뒤에는 정차 차량 한 대, 지붕 십자가 하나와 켜진 전조등들이 보인다. 도로와 연석은 이전 장소에 어울리지만, 차량이 화면에서 매우 큰 비중을 차지하고 얼굴은 요청한 눈매 클로즈업보다 작다.",
        "entities": "검은 머리의 한국인 중년 남성 얼굴은 구도환 참조와 대체로 닮았다. 어두운 재킷은 보이나 참조의 흰 셔츠는 확인되지 않는다. 차량과 발광 십자가는 존재한다. 차량 문 옆에 별도의 사람 한 명이 있으며, 흐려서 낡은 코트와 모자는 확인하기 어렵다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "구도환 외에는 등장시키지 말라는 인물 제한과 달리 차량 문 옆에 추가 인물 한 명이 보인다."
        ],
        "physics": "구도환은 벽 뒤에서 상체를 기울인 자세로, 목과 몸통이 자연스럽게 연결된다. 발은 프레임 밖이므로 접지는 확인할 수 없지만 공중에 떠 있는 형상은 아니다. 추가 인물은 도로에 발을 딛고 있으며 차량은 바퀴로 도로에 지지된다. 십자가는 차량 지붕에 부착되어 있다."
       },
       {
        "label": "B",
        "direction": "구도환은 렌즈를 응시하지 않고 눈을 화면 왼쪽 바깥으로 돌려 집중한다. 차량은 프레임 밖이라 실제 목표와의 정렬까지 확인할 수 없지만, 화면 안에 시선과 모순되는 차량 위치는 없다.",
        "built_space": "왼쪽에 거친 벽 모퉁이 하나가 전경을 가리고, 바로 오른쪽에서 구도환의 얼굴이 크게 드러난다. 뒤에는 젖은 도로, 연석과 낮은 임시 건물의 지붕들이 흐리게 보여 이전 정착촌의 재료와 야간 분위기가 이어진다. 얼굴과 눈매가 중심인 클로즈업이며 중복 시설이나 불가능한 반사는 보이지 않는다.",
        "entities": "검은 머리, 중년의 주름과 얼굴 윤곽은 한국인 50대 남성 구도환의 참조와 대체로 일치한다. 한 사람만 등장하고 눈은 정상적인 홍채와 동공을 갖는다. 다만 목의 흰 성직자 칼라와 검은 셔츠는 참조의 흰 셔츠 및 남색 재킷과 다르며 이전 장면 인물의 복장을 연상시킨다. 차량, 십자가와 승차자는 이 클로즈업 밖에 있어 확인 대상이 아니다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리는 목과 어깨에 자연스럽게 연결되고, 몸을 벽 뒤에 둔 채 얼굴을 조금 내민 자세가 가능하다. 하체와 발은 클로즈업 밖이므로 지면 접촉은 확인할 수 없다. 떠 있는 신체나 지지 없이 매달린 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "차량은 오른쪽 뒤에 있는데 시선은 왼쪽을 향하며, 허용되지 않은 추가 인물까지 보여 눈매 중심의 은밀한 관찰 장면을 벗어난다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "왼쪽 모퉁이에 숨은 얼굴과 날카로운 눈매의 클로즈업을 충실히 구현하지만, 성직자 칼라는 구도환의 의상 참조와 다르다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "구도환의 두 눈은 화면 왼쪽 바깥을 향한다. 관찰해야 할 차량과 승차 위치는 그의 오른쪽 뒤에 보여 시선이 목표에 닿지 않는다. 차량 옆 추가 인물은 차문 쪽을 향한다.",
        "built_space": "왼쪽의 거친 벽 모퉁이 하나가 화면 약 3분의 1을 가리고, 구도환은 그 오른쪽으로 머리와 상체를 내민다. 오른쪽 뒤에는 정차 차량 한 대, 지붕 십자가 하나와 켜진 전조등들이 보인다. 도로와 연석은 이전 장소에 어울리지만, 차량이 화면에서 매우 큰 비중을 차지하고 얼굴은 요청한 눈매 클로즈업보다 작다.",
        "entities": "검은 머리의 한국인 중년 남성 얼굴은 구도환 참조와 대체로 닮았다. 어두운 재킷은 보이나 참조의 흰 셔츠는 확인되지 않는다. 차량과 발광 십자가는 존재한다. 차량 문 옆에 별도의 사람 한 명이 있으며, 흐려서 낡은 코트와 모자는 확인하기 어렵다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "구도환 외에는 등장시키지 말라는 인물 제한과 달리 차량 문 옆에 추가 인물 한 명이 보인다."
        ],
        "physics": "구도환은 벽 뒤에서 상체를 기울인 자세로, 목과 몸통이 자연스럽게 연결된다. 발은 프레임 밖이므로 접지는 확인할 수 없지만 공중에 떠 있는 형상은 아니다. 추가 인물은 도로에 발을 딛고 있으며 차량은 바퀴로 도로에 지지된다. 십자가는 차량 지붕에 부착되어 있다."
       },
       {
        "label": "A",
        "direction": "구도환은 렌즈를 응시하지 않고 눈을 화면 왼쪽 바깥으로 돌려 집중한다. 차량은 프레임 밖이라 실제 목표와의 정렬까지 확인할 수 없지만, 화면 안에 시선과 모순되는 차량 위치는 없다.",
        "built_space": "왼쪽에 거친 벽 모퉁이 하나가 전경을 가리고, 바로 오른쪽에서 구도환의 얼굴이 크게 드러난다. 뒤에는 젖은 도로, 연석과 낮은 임시 건물의 지붕들이 흐리게 보여 이전 정착촌의 재료와 야간 분위기가 이어진다. 얼굴과 눈매가 중심인 클로즈업이며 중복 시설이나 불가능한 반사는 보이지 않는다.",
        "entities": "검은 머리, 중년의 주름과 얼굴 윤곽은 한국인 50대 남성 구도환의 참조와 대체로 일치한다. 한 사람만 등장하고 눈은 정상적인 홍채와 동공을 갖는다. 다만 목의 흰 성직자 칼라와 검은 셔츠는 참조의 흰 셔츠 및 남색 재킷과 다르며 이전 장면 인물의 복장을 연상시킨다. 차량, 십자가와 승차자는 이 클로즈업 밖에 있어 확인 대상이 아니다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리는 목과 어깨에 자연스럽게 연결되고, 몸을 벽 뒤에 둔 채 얼굴을 조금 내민 자세가 가능하다. 하체와 발은 클로즈업 밖이므로 지면 접촉은 확인할 수 없다. 떠 있는 신체나 지지 없이 매달린 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.429,
    "B": 1.25
   },
   "adjusted": {
    "A": 1.179,
    "B": 1.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 이전 샷 인물의 의상(사제복)을 캐릭터에게 그대로 복사하여 적용함"
    ],
    "B": [
     "[gpt-high] 구도환 외에는 등장시키지 말라는 인물 제한과 달리 차량 문 옆에 추가 인물 한 명이 보인다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1000,
   "A": 1179
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1000,
    "verdict_ko": "구도환의 클로즈업과 전경을 가리는 골목 모퉁이 구도를 잘 살렸으며, 배경에 명시된 밴, 십자가 장식, 탑승하는 찰리를 모두 정확히 구현함.  ★위반: [gpt-high] 구도환 외에는 등장시키지 말라는 인물 제한과 달리 차량 문 옆에 추가 인물 한 명이 보인다."
   },
   {
    "label": "A",
    "score": 1179,
    "verdict_ko": "프롬프트에 명시된 밴과 찰리가 완전히 누락되었으며, 이전 샷 인물의 의상(사제복)을 복사하는 금지 사항을 위반함.  ★위반: [gemini-pro] 이전 샷 인물의 의상(사제복)을 캐릭터에게 그대로 복사하여 적용함"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S27sh18_sel.png",
    "asset_id": "831b8e87-50bb-486b-9dbf-7161a937b6be",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1311816>",
    "asset_id": "623f0421-67dc-4592-a009-148d8e957276",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-3578-7bac-bb93-c847b3385911",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S27sh18"
  }
 },
 "S27sh21::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:19:53.221658+00:00",
  "fingerprint": "1675bd906367539a3732fd65b0869dac4d8894847035b7dbfda70f9205b6df99",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S27sh21_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S27sh21_sel.png",
  "source_sha256": "e6e9bc04af9f30394dc1ee1756b028e60038fc91d0e4192a6d3d6a83ee7c4085",
  "file": "S27sh21_cine.png",
  "staged_sha256": "b57748a906e3e34cd0e38c098e8ac123379f08fdf618043313209a0675aa354e",
  "latency_ms": 9076
 },
 "S28sh2::signage": {
  "fp": "9f143c5992eeaeae",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::22ae2a4d0af72fb5": {
  "subjects": [],
  "subject_text": "신부의 밴 내부\n운전석 뒤로 뒷좌석과 적재 공간이 이어지는 밴 실내. 물품 상자들이 듬성듬성 놓여 있고 상자 사이에 빈 공간이 남아 있다.",
  "identity": "canonical",
  "scope_id": "L179",
  "scope_role": "location_interior",
  "scope_sha": "b707d6881f7a896b"
 },
 "S28sh2::bgfirst_bg": {
  "input_fingerprint": "03f810a54ad37b98",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 차 안 뒷좌석, 낡은 종이상자 더미 사이사이에 몸을 욱여넣고 웅크린 현우, 앰버, 찰리의 굳은 전신.\n\nLOCATION (lock): Inside the van's rear passenger-and-cargo area, among scattered cardboard boxes in the nighttime darkness.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Worn paper boxes (Distributed through the rear seating area with gaps occupied by the hiding group) — Their tops and differently angled sides are visible from above; used as Create separate pockets of concealment without obscuring the three complete crouched poses; Rear seating area (Occupied by boxes and the concealed group) — Seen diagonally from the front passenger side; used as Establishes the limited shared space and realistic scale of the boxes.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued ambient illumination appropriate to the nighttime van interior, retaining readable bodies and boxes without introducing a cabin light or exterior color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 차 안 뒷좌석, 낡은 종이상자 더미 사이사이에 몸을 욱여넣고 웅크린 현우, 앰버, 찰리의 굳은 전신.\n\nLOCATION (lock): Inside the van's rear passenger-and-cargo area, among scattered cardboard boxes in the nighttime darkness.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Worn paper boxes (Distributed through the rear seating area with gaps occupied by the hiding group) — Their tops and differently angled sides are visible from above; used as Create separate pockets of concealment without obscuring the three complete crouched poses; Rear seating area (Occupied by boxes and the concealed group) — Seen diagonally from the front passenger side; used as Establishes the limited shared space and realistic scale of the boxes.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued ambient illumination appropriate to the nighttime van interior, retaining readable bodies and boxes without introducing a cabin light or exterior color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S28sh2__bgfirst_bg.png",
  "asset_id": "bc133181-b478-4d81-83c2-7aa1d682c287",
  "input_asset_ids": [
   "999f65d7-3a65-487a-8bd7-c6b9ebc6b8a4",
   "f6306063-3d19-4618-946d-060af53e1589"
  ]
 },
 "S28sh2": {
  "input_fingerprint": "1b614e380db77321",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 차 안 뒷좌석, 낡은 종이상자 더미 사이사이에 몸을 욱여넣고 웅크린 현우, 앰버, 찰리의 굳은 전신.\n\nLOCATION (lock): Inside the van's rear passenger-and-cargo area, among scattered cardboard boxes in the nighttime darkness. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Worn paper boxes (Distributed through the rear seating area with gaps occupied by the hiding group) — Their tops and differently angled sides are visible from above; used as Create separate pockets of concealment without obscuring the three complete crouched poses; Rear seating area (Occupied by boxes and the concealed group) — Seen diagonally from the front passenger side; used as Establishes the limited shared space and realistic scale of the boxes.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued ambient illumination appropriate to the nighttime van interior, retaining readable bodies and boxes without introducing a cabin light or exterior color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Scattered cargo boxes occupy the van's rear seating area, and a crude halogen cross is mounted on the roof. Charlie remains concealed among the boxes in his old coat and hat. 현우: He is concealed among the rear cargo boxes, with facial bruises and an injured leg; his outer garment remains removed. The contact card remains concealed in his shoe. 앰버: She is concealed among the rear boxes, still exhausted, with the outer garment used as a mouth covering.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 차 안 뒷좌석, 낡은 종이상자 더미 사이사이에 몸을 욱여넣고 웅크린 현우, 앰버, 찰리의 굳은 전신.\n\nLOCATION (lock): Inside the van's rear passenger-and-cargo area, among scattered cardboard boxes in the nighttime darkness. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Worn paper boxes (Distributed through the rear seating area with gaps occupied by the hiding group) — Their tops and differently angled sides are visible from above; used as Create separate pockets of concealment without obscuring the three complete crouched poses; Rear seating area (Occupied by boxes and the concealed group) — Seen diagonally from the front passenger side; used as Establishes the limited shared space and realistic scale of the boxes.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued ambient illumination appropriate to the nighttime van interior, retaining readable bodies and boxes without introducing a cabin light or exterior color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Scattered cargo boxes occupy the van's rear seating area, and a crude halogen cross is mounted on the roof. Charlie remains concealed among the boxes in his old coat and hat. 현우: He is concealed among the rear cargo boxes, with facial bruises and an injured leg; his outer garment remains removed. The contact card remains concealed in his shoe. 앰버: She is concealed among the rear boxes, still exhausted, with the outer garment used as a mouth covering.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 차 안 뒷좌석, 낡은 종이상자 더미 사이사이에 몸을 욱여넣고 웅크린 현우, 앰버, 찰리의 굳은 전신.\n\nLOCATION (lock): Inside the van's rear passenger-and-cargo area, among scattered cardboard boxes in the nighttime darkness. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Worn paper boxes (Distributed through the rear seating area with gaps occupied by the hiding group) — Their tops and differently angled sides are visible from above; used as Create separate pockets of concealment without obscuring the three complete crouched poses; Rear seating area (Occupied by boxes and the concealed group) — Seen diagonally from the front passenger side; used as Establishes the limited shared space and realistic scale of the boxes.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued ambient illumination appropriate to the nighttime van interior, retaining readable bodies and boxes without introducing a cabin light or exterior color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Scattered cargo boxes occupy the van's rear seating area, and a crude halogen cross is mounted on the roof. Charlie remains concealed among the boxes in his old coat and hat. 현우: He is concealed among the rear cargo boxes, with facial bruises and an injured leg; his outer garment remains removed. The contact card remains concealed in his shoe. 앰버: She is concealed among the rear boxes, still exhausted, with the outer garment used as a mouth covering.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S28sh2__bgfirst_bg.png",
     "asset_id": "bc133181-b478-4d81-83c2-7aa1d682c287",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S28sh2.png",
     "asset_id": "999f65d7-3a65-487a-8bd7-c6b9ebc6b8a4",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L179B01.png",
     "asset_id": "f6306063-3d19-4618-946d-060af53e1589",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우와 앰버는 시선을 아래로 향하고 웅크려 있으며, 찰리는 정면을 응시함.",
    "built_space": "밴 뒷좌석 공간과 상자 배치는 레퍼런스를 따르나, 전체적인 환경이 실사가 아닌 2D 일러스트 스타일로 표현됨.",
    "entities": "현우(얼굴 타박상 묘사), 앰버(무릎에 얼굴을 묻고 있으나 겉옷으로 입을 가린 형태가 아님), 찰리(로봇 형태는 일치하나 지정된 낡은 코트와 모자 없음).",
    "hard_violations": [],
    "physics": "바닥과 종이상자에 몸을 기대어 안정적으로 웅크리고 있음."
   },
   {
    "label": "B",
    "direction": "세 캐릭터 모두 밴 앞쪽 대각선 방향으로 시선을 두고 웅크려 있음.",
    "built_space": "조수석 대각선 뷰에서 본 밴 내부 구조가 레퍼런스와 일치하며 사실적인 질감을 가짐. 단, 전경 상자에 읽을 수 있는 글자(취급주의)가 그대로 노출됨.",
    "entities": "앰버(지시대로 겉옷을 이용해 입을 가림), 현우(가슴을 열고 겉옷을 벗은 상태이며 얼굴과 맨다리에 타박상 있음), 찰리(로봇 형태는 일치하나 낡은 코트와 모자 없음).",
    "hard_violations": [
     "[gpt-high] 왼쪽 전경 상자에 '취급주의'라는 읽을 수 있는 글자가 있어 이미지 내 모든 가독성 있는 글자 금지 조건을 위반한다."
    ],
    "physics": "밴 바닥에 앉아 상자들 사이에 기대어 신체를 자연스럽게 지탱하고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "찰리의 의상이 누락되고 상표 텍스트가 노출되었으나, 완벽한 실사 질감과 앰버의 입 가림, 현우의 맨다리 상처 등 세부 묘사를 훌륭히 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "실사 영화 스틸컷 지침을 명백히 위반한 일러스트 스타일이며, 앰버의 입 가림 액션과 찰리의 의상 지시를 모두 누락했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우와 앰버는 시선을 아래로 향하고 웅크려 있으며, 찰리는 정면을 응시함.",
        "built_space": "밴 뒷좌석 공간과 상자 배치는 레퍼런스를 따르나, 전체적인 환경이 실사가 아닌 2D 일러스트 스타일로 표현됨.",
        "entities": "현우(얼굴 타박상 묘사), 앰버(무릎에 얼굴을 묻고 있으나 겉옷으로 입을 가린 형태가 아님), 찰리(로봇 형태는 일치하나 지정된 낡은 코트와 모자 없음).",
        "hard_violations": [],
        "physics": "바닥과 종이상자에 몸을 기대어 안정적으로 웅크리고 있음."
       },
       {
        "label": "B",
        "direction": "세 캐릭터 모두 밴 앞쪽 대각선 방향으로 시선을 두고 웅크려 있음.",
        "built_space": "조수석 대각선 뷰에서 본 밴 내부 구조가 레퍼런스와 일치하며 사실적인 질감을 가짐. 단, 전경 상자에 읽을 수 있는 글자(취급주의)가 그대로 노출됨.",
        "entities": "앰버(지시대로 겉옷을 이용해 입을 가림), 현우(가슴을 열고 겉옷을 벗은 상태이며 얼굴과 맨다리에 타박상 있음), 찰리(로봇 형태는 일치하나 낡은 코트와 모자 없음).",
        "hard_violations": [],
        "physics": "밴 바닥에 앉아 상자들 사이에 기대어 신체를 자연스럽게 지탱하고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "찰리의 의상이 누락되고 상표 텍스트가 노출되었으나, 완벽한 실사 질감과 앰버의 입 가림, 현우의 맨다리 상처 등 세부 묘사를 훌륭히 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "실사 영화 스틸컷 지침을 명백히 위반한 일러스트 스타일이며, 앰버의 입 가림 액션과 찰리의 의상 지시를 모두 누락했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우와 앰버는 시선을 아래로 향하고 웅크려 있으며, 찰리는 정면을 응시함.",
        "built_space": "밴 뒷좌석 공간과 상자 배치는 레퍼런스를 따르나, 전체적인 환경이 실사가 아닌 2D 일러스트 스타일로 표현됨.",
        "entities": "현우(얼굴 타박상 묘사), 앰버(무릎에 얼굴을 묻고 있으나 겉옷으로 입을 가린 형태가 아님), 찰리(로봇 형태는 일치하나 지정된 낡은 코트와 모자 없음).",
        "hard_violations": [],
        "physics": "바닥과 종이상자에 몸을 기대어 안정적으로 웅크리고 있음."
       },
       {
        "label": "B",
        "direction": "세 캐릭터 모두 밴 앞쪽 대각선 방향으로 시선을 두고 웅크려 있음.",
        "built_space": "조수석 대각선 뷰에서 본 밴 내부 구조가 레퍼런스와 일치하며 사실적인 질감을 가짐. 단, 전경 상자에 읽을 수 있는 글자(취급주의)가 그대로 노출됨.",
        "entities": "앰버(지시대로 겉옷을 이용해 입을 가림), 현우(가슴을 열고 겉옷을 벗은 상태이며 얼굴과 맨다리에 타박상 있음), 찰리(로봇 형태는 일치하나 낡은 코트와 모자 없음).",
        "hard_violations": [],
        "physics": "밴 바닥에 앉아 상자들 사이에 기대어 신체를 자연스럽게 지탱하고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "읽을 수 있는 상자 글자가 금지 조건을 위반하며, 뒤쪽에서 앞을 보는 구도와 전경 좌석이 세 인물의 웅크린 전신을 가린다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "상자 사이의 개별 은신 공간과 현우·앰버의 웅크린 전신은 더 충실하지만, 촬영 방향이 다르고 찰리의 하체와 낡은 코트·모자가 구현되지 않았다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 화면 오른쪽 앞을, 앰버는 거의 정면에서 약간 오른쪽을 본다. 찰리의 얼굴도 화면 오른쪽 아래로 향한다. 지정된 응시 대상은 없으며 무기나 이동 동작도 없다. 카메라는 조수석 쪽에서 뒤를 보는 것이 아니라 차량 뒤쪽에서 앞유리와 운전석을 바라본다.",
        "built_space": "앞좌석 두 개와 각각의 머리받침, 운전대 하나, 중앙 룸미러 하나, 전경의 뒷좌석 등받이 하나가 보인다. 오른쪽 큰 측창과 안전벨트, 회색 내장재는 장소 사진과 가깝다. 세 인물은 앞좌석과 전경 등받이 사이에 밀집해 있고, 앰버와 현우는 같은 좁은 틈을 공유한다. 상자 윗면은 보이지만 큰 등받이가 현우의 몸 일부와 찰리의 하체를 가려 세 명의 완전한 웅크린 자세를 보여주지 못한다. 불가능한 반사나 중복된 고정 설비는 보이지 않는다.",
        "entities": "등장 인물은 세 명이다. 현우는 헝클어진 검은 머리의 젊은 동아시아계 남성으로 얼굴 멍과 다리 상처가 보이지만, 열린 회색 셔츠 차림은 참고의 남색 티셔츠와 다르다. 앰버는 금발 여자아이이며 짙은 천으로 입을 가렸다. 찰리는 베이지 장갑판과 흰 기계형 얼굴로 참고의 정체성을 따르지만 낡은 코트와 모자가 없다. 낡은 종이상자, 녹색 수납함과 가방이 보인다. 왼쪽 큰 상자의 '취급주의'는 읽을 수 있다. 지붕 바깥의 십자가와 신발 속 카드는 확인할 수 없다. 밤의 실내이며 켜진 실내등은 없다.",
        "hard_violations": [
         "왼쪽 전경 상자에 '취급주의'라는 읽을 수 있는 글자가 있어 이미지 내 모든 가독성 있는 글자 금지 조건을 위반한다."
        ],
        "physics": "현우는 바닥에 앉아 무릎을 세우고 드러난 발을 바닥에 댄다. 앰버도 다리를 접어 바닥 가까이에 앉아 있으며, 몸이 공중에 떠 있다는 징후는 없다. 찰리의 하체 지지점은 등받이와 상자에 가려 확인되지 않지만 상체는 낮게 접혀 있다. 찰리 앞 상자는 주변 짐 사이에 끼어 있고 양손은 그 아래에서 모인다. 상자들은 바닥이나 다른 상자 위에 놓여 있으며 명백히 무지지 상태인 물체는 없다."
       },
       {
        "label": "B",
        "direction": "현우는 고개를 숙여 자신의 무릎과 발 쪽을 보고, 앰버는 입을 가린 천과 무릎 쪽으로 얼굴을 묻는다. 찰리도 얼굴을 아래로 숙여 앞쪽 상자 방향을 향한다. 세 인물 모두 움츠린 채 숨는 행동으로 읽힌다. 다만 카메라는 이 후보에서도 차량 뒤에서 앞유리를 바라보므로 조수석 쪽에서 후방을 대각선으로 보는 배치와 다르다.",
        "built_space": "앞좌석 두 개와 머리받침 두 개, 운전대 하나, 중앙 룸미러 하나가 보이고 전경 오른쪽에 뒷좌석 등받이 일부가 있다. 측창, 천장과 회색 내장재는 장소 사진을 대체로 유지한다. 현우는 왼쪽, 앰버는 중앙, 찰리는 오른쪽에서 상자 사이의 서로 다른 틈을 차지한다. 여러 상자의 윗면과 비스듬한 옆면이 보여 요구된 은신 공간 분리가 A보다 명확하다. 현우와 앰버는 발까지 보이지만 찰리의 하체는 앞쪽 상자에 가려진다. 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "세 인물만 등장한다. 현우는 검은 헝클어진 머리, 앳된 동아시아계 얼굴과 볼의 멍을 갖추고 겉옷 없이 반소매 상의를 입었다. 바지로 가려진 다리의 부상은 확인되지 않는다. 앰버는 금발 여자아이로 겉옷으로 보이는 천을 입에 대고 지친 표정으로 웅크린다. 찰리는 육중한 베이지 기계 몸체와 흰 얼굴을 유지하지만 코트와 모자가 없다. 낡고 찌그러진 종이상자가 다수 있으며 확실히 읽히는 글자는 보이지 않는다. 지붕 외부 십자가와 신발 속 카드는 보이지 않는다. 어두운 야간 실내이고 켜진 실내등이나 강한 외부 색조는 없다.",
        "hard_violations": [],
        "physics": "현우는 엉덩이를 낮게 두고 발을 바닥에 놓은 채 접은 다리를 양손으로 붙든다. 앰버도 바닥에 앉아 두 무릎을 가슴에 모으고 팔로 감싸며 신발이 바닥에 닿는다. 입을 가린 천은 얼굴과 무릎 사이에 받쳐져 있다. 찰리는 상체를 앞으로 접었고 하체의 실제 접점은 상자에 가려 확인되지 않는다. 떠 있는 자세로 보이지는 않는다. 상자들은 바닥이나 다른 짐에 받쳐져 있으며 뚜렷한 부유나 불가능한 관절 자세는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "읽을 수 있는 상자 글자가 금지 조건을 위반하며, 뒤쪽에서 앞을 보는 구도와 전경 좌석이 세 인물의 웅크린 전신을 가린다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "상자 사이의 개별 은신 공간과 현우·앰버의 웅크린 전신은 더 충실하지만, 촬영 방향이 다르고 찰리의 하체와 낡은 코트·모자가 구현되지 않았다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 화면 오른쪽 앞을, 앰버는 거의 정면에서 약간 오른쪽을 본다. 찰리의 얼굴도 화면 오른쪽 아래로 향한다. 지정된 응시 대상은 없으며 무기나 이동 동작도 없다. 카메라는 조수석 쪽에서 뒤를 보는 것이 아니라 차량 뒤쪽에서 앞유리와 운전석을 바라본다.",
        "built_space": "앞좌석 두 개와 각각의 머리받침, 운전대 하나, 중앙 룸미러 하나, 전경의 뒷좌석 등받이 하나가 보인다. 오른쪽 큰 측창과 안전벨트, 회색 내장재는 장소 사진과 가깝다. 세 인물은 앞좌석과 전경 등받이 사이에 밀집해 있고, 앰버와 현우는 같은 좁은 틈을 공유한다. 상자 윗면은 보이지만 큰 등받이가 현우의 몸 일부와 찰리의 하체를 가려 세 명의 완전한 웅크린 자세를 보여주지 못한다. 불가능한 반사나 중복된 고정 설비는 보이지 않는다.",
        "entities": "등장 인물은 세 명이다. 현우는 헝클어진 검은 머리의 젊은 동아시아계 남성으로 얼굴 멍과 다리 상처가 보이지만, 열린 회색 셔츠 차림은 참고의 남색 티셔츠와 다르다. 앰버는 금발 여자아이이며 짙은 천으로 입을 가렸다. 찰리는 베이지 장갑판과 흰 기계형 얼굴로 참고의 정체성을 따르지만 낡은 코트와 모자가 없다. 낡은 종이상자, 녹색 수납함과 가방이 보인다. 왼쪽 큰 상자의 '취급주의'는 읽을 수 있다. 지붕 바깥의 십자가와 신발 속 카드는 확인할 수 없다. 밤의 실내이며 켜진 실내등은 없다.",
        "hard_violations": [
         "왼쪽 전경 상자에 '취급주의'라는 읽을 수 있는 글자가 있어 이미지 내 모든 가독성 있는 글자 금지 조건을 위반한다."
        ],
        "physics": "현우는 바닥에 앉아 무릎을 세우고 드러난 발을 바닥에 댄다. 앰버도 다리를 접어 바닥 가까이에 앉아 있으며, 몸이 공중에 떠 있다는 징후는 없다. 찰리의 하체 지지점은 등받이와 상자에 가려 확인되지 않지만 상체는 낮게 접혀 있다. 찰리 앞 상자는 주변 짐 사이에 끼어 있고 양손은 그 아래에서 모인다. 상자들은 바닥이나 다른 상자 위에 놓여 있으며 명백히 무지지 상태인 물체는 없다."
       },
       {
        "label": "A",
        "direction": "현우는 고개를 숙여 자신의 무릎과 발 쪽을 보고, 앰버는 입을 가린 천과 무릎 쪽으로 얼굴을 묻는다. 찰리도 얼굴을 아래로 숙여 앞쪽 상자 방향을 향한다. 세 인물 모두 움츠린 채 숨는 행동으로 읽힌다. 다만 카메라는 이 후보에서도 차량 뒤에서 앞유리를 바라보므로 조수석 쪽에서 후방을 대각선으로 보는 배치와 다르다.",
        "built_space": "앞좌석 두 개와 머리받침 두 개, 운전대 하나, 중앙 룸미러 하나가 보이고 전경 오른쪽에 뒷좌석 등받이 일부가 있다. 측창, 천장과 회색 내장재는 장소 사진을 대체로 유지한다. 현우는 왼쪽, 앰버는 중앙, 찰리는 오른쪽에서 상자 사이의 서로 다른 틈을 차지한다. 여러 상자의 윗면과 비스듬한 옆면이 보여 요구된 은신 공간 분리가 A보다 명확하다. 현우와 앰버는 발까지 보이지만 찰리의 하체는 앞쪽 상자에 가려진다. 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "세 인물만 등장한다. 현우는 검은 헝클어진 머리, 앳된 동아시아계 얼굴과 볼의 멍을 갖추고 겉옷 없이 반소매 상의를 입었다. 바지로 가려진 다리의 부상은 확인되지 않는다. 앰버는 금발 여자아이로 겉옷으로 보이는 천을 입에 대고 지친 표정으로 웅크린다. 찰리는 육중한 베이지 기계 몸체와 흰 얼굴을 유지하지만 코트와 모자가 없다. 낡고 찌그러진 종이상자가 다수 있으며 확실히 읽히는 글자는 보이지 않는다. 지붕 외부 십자가와 신발 속 카드는 보이지 않는다. 어두운 야간 실내이고 켜진 실내등이나 강한 외부 색조는 없다.",
        "hard_violations": [],
        "physics": "현우는 엉덩이를 낮게 두고 발을 바닥에 놓은 채 접은 다리를 양손으로 붙든다. 앰버도 바닥에 앉아 두 무릎을 가슴에 모으고 팔로 감싸며 신발이 바닥에 닿는다. 입을 가린 천은 얼굴과 무릎 사이에 받쳐져 있다. 찰리는 상체를 앞으로 접었고 하체의 실제 접점은 상자에 가려 확인되지 않는다. 떠 있는 자세로 보이지는 않는다. 상자들은 바닥이나 다른 짐에 받쳐져 있으며 뚜렷한 부유나 불가능한 관절 자세는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.429,
    "B": 1.4
   },
   "adjusted": {
    "A": 1.429,
    "B": 1.15
   },
   "violations": {
    "B": [
     "[gpt-high] 왼쪽 전경 상자에 '취급주의'라는 읽을 수 있는 글자가 있어 이미지 내 모든 가독성 있는 글자 금지 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1150,
   "A": 1429
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1150,
    "verdict_ko": "찰리의 의상이 누락되고 상표 텍스트가 노출되었으나, 완벽한 실사 질감과 앰버의 입 가림, 현우의 맨다리 상처 등 세부 묘사를 훌륭히 구현했습니다.  ★위반: [gpt-high] 왼쪽 전경 상자에 '취급주의'라는 읽을 수 있는 글자가 있어 이미지 내 모든 가독성 있는 글자 금지 조건을 위반한다."
   },
   {
    "label": "A",
    "score": 1429,
    "verdict_ko": "실사 영화 스틸컷 지침을 명백히 위반한 일러스트 스타일이며, 앰버의 입 가림 액션과 찰리의 의상 지시를 모두 누락했습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L179B01.png",
    "asset_id": "f6306063-3d19-4618-946d-060af53e1589",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-3735-75bc-8af7-df9afccda0e0",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S28sh2__bgfirst_bg.png",
   "bg_asset_id": "bc133181-b478-4d81-83c2-7aa1d682c287",
   "bg_record_key": "S28sh2::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S28sh2::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:27:39.426743+00:00",
  "fingerprint": "115d490c6b90f4bcbc25306e7f1fa24665aa0b47447485676c4cb31c6c5efd1c",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S28sh2_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S28sh2_sel.png",
  "source_sha256": "20dab48b1d8e518653dc12c39ab4109d9a0ce0ad9d02805bddfe913e3898e33d",
  "file": "S28sh2_cine.png",
  "staged_sha256": "39ae19454b02e43392c88450077c8116a26d619daf94e42908cc6ca0581c5445",
  "latency_ms": 13784
 },
 "S28sh5::signage": {
  "fp": "0227ca4fd6663347",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S28sh5::bgfirst_bg": {
  "input_fingerprint": "ee9bf078ea28e12d",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 자동차 앞유리창 너머로 무기를 든 채 차의 앞길을 막아선 민병대원들의 위협적인 전경.\n\nLOCATION (lock): On the nighttime road directly ahead of the van at a militia checkpoint, seen through its windshield.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Front windshield (The armed militia are visible through it) — Seen from inside the van, with the blocked road and militia beyond; used as Frames the threat as a direct exterior view from the protected cabin rather than a reflection or screen image; Road ahead (Blocked by the armed militia) — Extends forward beyond the windshield; used as Shows that the van's route is physically obstructed; Militia weapons (Held by the figures blocking the van) — Visible at varied oblique angles alongside their holders; used as Supply the immediate threat while remaining subordinate in size to the people.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the nighttime exterior and cabin in restrained tonal contrast, preserving a direct view through the windshield without invented reflections, colored light, or glare.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 자동차 앞유리창 너머로 무기를 든 채 차의 앞길을 막아선 민병대원들의 위협적인 전경.\n\nLOCATION (lock): On the nighttime road directly ahead of the van at a militia checkpoint, seen through its windshield.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Front windshield (The armed militia are visible through it) — Seen from inside the van, with the blocked road and militia beyond; used as Frames the threat as a direct exterior view from the protected cabin rather than a reflection or screen image; Road ahead (Blocked by the armed militia) — Extends forward beyond the windshield; used as Shows that the van's route is physically obstructed; Militia weapons (Held by the figures blocking the van) — Visible at varied oblique angles alongside their holders; used as Supply the immediate threat while remaining subordinate in size to the people.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the nighttime exterior and cabin in restrained tonal contrast, preserving a direct view through the windshield without invented reflections, colored light, or glare.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S28sh5__bgfirst_bg.png",
  "asset_id": "2fb1b93e-74b7-4cf0-85b0-9fb67e335dcd",
  "input_asset_ids": [
   "029b4a24-f56c-4f1f-aeee-61bd3bdf819c",
   "f6306063-3d19-4618-946d-060af53e1589"
  ]
 },
 "S28sh5": {
  "input_fingerprint": "198ac8749577c253",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 자동차 앞유리창 너머로 무기를 든 채 차의 앞길을 막아선 민병대원들의 위협적인 전경.\n\nLOCATION (lock): On the nighttime road directly ahead of the van at a militia checkpoint, seen through its windshield. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Front windshield (The armed militia are visible through it) — Seen from inside the van, with the blocked road and militia beyond; used as Frames the threat as a direct exterior view from the protected cabin rather than a reflection or screen image; Road ahead (Blocked by the armed militia) — Extends forward beyond the windshield; used as Shows that the van's route is physically obstructed; Militia weapons (Held by the figures blocking the van) — Visible at varied oblique angles alongside their holders; used as Supply the immediate threat while remaining subordinate in size to the people.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the nighttime exterior and cabin in restrained tonal contrast, preserving a direct view through the windshield without invented reflections, colored light, or glare.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rear cargo boxes and roof-mounted halogen cross remain in place. Charlie is concealed among the boxes, still wearing his old coat and hat.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 민병대원들 right now, so 민병대원들's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 민병대원들: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 자동차 앞유리창 너머로 무기를 든 채 차의 앞길을 막아선 민병대원들의 위협적인 전경.\n\nLOCATION (lock): On the nighttime road directly ahead of the van at a militia checkpoint, seen through its windshield. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Front windshield (The armed militia are visible through it) — Seen from inside the van, with the blocked road and militia beyond; used as Frames the threat as a direct exterior view from the protected cabin rather than a reflection or screen image; Road ahead (Blocked by the armed militia) — Extends forward beyond the windshield; used as Shows that the van's route is physically obstructed; Militia weapons (Held by the figures blocking the van) — Visible at varied oblique angles alongside their holders; used as Supply the immediate threat while remaining subordinate in size to the people.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the nighttime exterior and cabin in restrained tonal contrast, preserving a direct view through the windshield without invented reflections, colored light, or glare.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rear cargo boxes and roof-mounted halogen cross remain in place. Charlie is concealed among the boxes, still wearing his old coat and hat.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 민병대원들 right now, so 민병대원들's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 민병대원들: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 자동차 앞유리창 너머로 무기를 든 채 차의 앞길을 막아선 민병대원들의 위협적인 전경.\n\nLOCATION (lock): On the nighttime road directly ahead of the van at a militia checkpoint, seen through its windshield. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Front windshield (The armed militia are visible through it) — Seen from inside the van, with the blocked road and militia beyond; used as Frames the threat as a direct exterior view from the protected cabin rather than a reflection or screen image; Road ahead (Blocked by the armed militia) — Extends forward beyond the windshield; used as Shows that the van's route is physically obstructed; Militia weapons (Held by the figures blocking the van) — Visible at varied oblique angles alongside their holders; used as Supply the immediate threat while remaining subordinate in size to the people.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the nighttime exterior and cabin in restrained tonal contrast, preserving a direct view through the windshield without invented reflections, colored light, or glare.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rear cargo boxes and roof-mounted halogen cross remain in place. Charlie is concealed among the boxes, still wearing his old coat and hat.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 민병대원들 right now, so 민병대원들's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 민병대원들: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S28sh5__bgfirst_bg.png",
     "asset_id": "2fb1b93e-74b7-4cf0-85b0-9fb67e335dcd",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S28sh5.png",
     "asset_id": "029b4a24-f56c-4f1f-aeee-61bd3bdf819c",
     "role": "conti_light"
    },
    {
     "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:842741>",
     "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
     "role": "prop_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L179B01.png",
     "asset_id": "f6306063-3d19-4618-946d-060af53e1589",
     "role": "location_plate"
    },
    {
     "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:842741>",
     "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
     "role": "prop_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "중앙의 민병대원이 손을 들어 밴을 멈춰 세우고 있으며, 운전자는 정면을 주시함.",
    "built_space": "승합차 내부. 대시보드와 앞좌석, 뒤편의 화물 상자 구역이 보임.",
    "entities": "운전석에 앉은 지시문에 없는 인물, 상자 사이에 숨은 찰리, 외부의 민병대원 3명. 대원들 주변 허공에 흰색 화살표 마커가 떠 있음. 지정된 형태의 총기는 구현되지 않음.",
    "hard_violations": [
     "[gemini-pro] 지시문에 없는 인물(운전자) 추가",
     "[gemini-pro] 화면에 흰색 화살표 마커 유출",
     "[gpt-high] 실제 장면의 물체가 아닌 흰 화살표 도해가 화면에 노출되어 있다.",
     "[gpt-high] 샷 텍스트에 등장하지 않는 운전자를 앞좌석에 추가하여 등장인물 제한을 위반한다."
    ],
    "physics": "운전자와 찰리는 차량 내부에 자리 잡고 지탱되며, 외부 대원들은 지면에 두 발로 서 있음."
   },
   {
    "label": "B",
    "direction": "중앙의 민병대원이 밴 정면(카메라)을 향해 권총을 정조준하고 있음.",
    "built_space": "승합차 앞좌석 사이에서 밖을 내다보는 시점. 빈 앞좌석과 대시보드, 전방 도로의 바리케이드가 보임.",
    "entities": "도로를 막아선 4명의 민병대원. 찰리와 화물 상자는 프레임 밖에 있어 보이지 않음. 참조된 특수 총기는 정확히 묘사되지 않고 일반 권총과 소총들로 대체됨.",
    "hard_violations": [
     "[gpt-high] 운전대에 식별 가능한 현대 로고가 노출되어 명시적인 로고 금지 조건을 위반한다."
    ],
    "physics": "대원들은 지면에 안정적으로 서서 손으로 무기를 파지하고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "구도의 한계로 찰리가 프레임에서 제외되었으나, 앞유리창 너머로 위협을 가하는 민병대의 전경을 규칙 위반 없이 성공적으로 묘사했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지시문에 명시되지 않은 운전자가 프레임에 추가되었고, 화면 양측에 화살표 형태의 마커가 유출되어 치명적인 위반이 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "중앙의 민병대원이 손을 들어 밴을 멈춰 세우고 있으며, 운전자는 정면을 주시함.",
        "built_space": "승합차 내부. 대시보드와 앞좌석, 뒤편의 화물 상자 구역이 보임.",
        "entities": "운전석에 앉은 지시문에 없는 인물, 상자 사이에 숨은 찰리, 외부의 민병대원 3명. 대원들 주변 허공에 흰색 화살표 마커가 떠 있음. 지정된 형태의 총기는 구현되지 않음.",
        "hard_violations": [
         "지시문에 없는 인물(운전자) 추가",
         "화면에 흰색 화살표 마커 유출"
        ],
        "physics": "운전자와 찰리는 차량 내부에 자리 잡고 지탱되며, 외부 대원들은 지면에 두 발로 서 있음."
       },
       {
        "label": "B",
        "direction": "중앙의 민병대원이 밴 정면(카메라)을 향해 권총을 정조준하고 있음.",
        "built_space": "승합차 앞좌석 사이에서 밖을 내다보는 시점. 빈 앞좌석과 대시보드, 전방 도로의 바리케이드가 보임.",
        "entities": "도로를 막아선 4명의 민병대원. 찰리와 화물 상자는 프레임 밖에 있어 보이지 않음. 참조된 특수 총기는 정확히 묘사되지 않고 일반 권총과 소총들로 대체됨.",
        "hard_violations": [],
        "physics": "대원들은 지면에 안정적으로 서서 손으로 무기를 파지하고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "구도의 한계로 찰리가 프레임에서 제외되었으나, 앞유리창 너머로 위협을 가하는 민병대의 전경을 규칙 위반 없이 성공적으로 묘사했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지시문에 명시되지 않은 운전자가 프레임에 추가되었고, 화면 양측에 화살표 형태의 마커가 유출되어 치명적인 위반이 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "중앙의 민병대원이 손을 들어 밴을 멈춰 세우고 있으며, 운전자는 정면을 주시함.",
        "built_space": "승합차 내부. 대시보드와 앞좌석, 뒤편의 화물 상자 구역이 보임.",
        "entities": "운전석에 앉은 지시문에 없는 인물, 상자 사이에 숨은 찰리, 외부의 민병대원 3명. 대원들 주변 허공에 흰색 화살표 마커가 떠 있음. 지정된 형태의 총기는 구현되지 않음.",
        "hard_violations": [
         "지시문에 없는 인물(운전자) 추가",
         "화면에 흰색 화살표 마커 유출"
        ],
        "physics": "운전자와 찰리는 차량 내부에 자리 잡고 지탱되며, 외부 대원들은 지면에 두 발로 서 있음."
       },
       {
        "label": "B",
        "direction": "중앙의 민병대원이 밴 정면(카메라)을 향해 권총을 정조준하고 있음.",
        "built_space": "승합차 앞좌석 사이에서 밖을 내다보는 시점. 빈 앞좌석과 대시보드, 전방 도로의 바리케이드가 보임.",
        "entities": "도로를 막아선 4명의 민병대원. 찰리와 화물 상자는 프레임 밖에 있어 보이지 않음. 참조된 특수 총기는 정확히 묘사되지 않고 일반 권총과 소총들로 대체됨.",
        "hard_violations": [],
        "physics": "대원들은 지면에 안정적으로 서서 손으로 무기를 파지하고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "앞유리 너머 차량을 가로막는 민병대의 위협이라는 핵심 구도는 더 충실하지만, 노출된 로고와 한국인 설정에 맞지 않는 인물 구성, 참조 총기와의 차이가 남는다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "흰 화살표가 화면에 노출되고 지시되지 않은 운전자가 추가되었으며, 화물칸까지 넓힌 구도가 앞유리 밖 민병대의 위협을 약화한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "중앙 남성은 차 안을 똑바로 보며 권총 총구를 앞유리와 카메라 쪽으로 향한다. 나머지 네 명은 대체로 차량을 향해 서 있고, 긴 총들은 몸 앞에서 아래쪽 또는 비스듬한 옆쪽을 향한다. 왼쪽에서 두 번째 인물의 총은 거의 수직으로 내려가 있다. 모든 총이 차량을 겨눠야 한다는 지시는 없으므로 이 방향 배치는 위협적인 길막 장면과 양립한다.",
        "built_space": "밴 내부에서 앞유리 하나를 통해 도로를 직접 본다. 앞좌석 두 개의 일부, 왼쪽 운전대 하나, 중앙 룸미러 하나, 선바이저 두 개, 앞유리 아래 와이퍼 두 개가 보인다. 참조와 유사한 낡은 대시보드와 중앙 조작부가 유지된다. 밖에는 다섯 명이 도로 폭을 가로질러 서 있고 그 뒤로 가로 차단봉 하나가 놓여 있다. 낮은 건물과 가로등이 밤길을 따라 이어지며, 위협을 반사상으로 대체하지 않았다.",
        "entities": "민병대는 성인 남성 다섯 명이며 낡은 야전 재킷과 모자, 총기를 착용하거나 들고 있다. 중앙 인물은 백인으로, 오른쪽 끝 인물 등은 흑인으로 보이므로 별도 설명 없는 현지인은 한국인이라는 설정과 맞지 않는다. 중앙의 짧은 권총과 주변의 긴 소총들은 참조의 개머리판이 결합된 권총형 총기와 다르다. 운전대에는 식별 가능한 현대 로고가 있다. 후방 상자, 찰리, 지붕 십자가는 이 구도에서 보이지 않으므로 부재로 판정하지 않는다.",
        "hard_violations": [
         "운전대에 식별 가능한 현대 로고가 노출되어 명시적인 로고 금지 조건을 위반한다."
        ],
        "physics": "중앙 인물은 두 손으로 권총을 받쳐 들고, 나머지 인물들도 손과 팔로 총을 지지한다. 인물들의 하체 일부는 대시보드와 창틀에 가리지만 지면에 서 있는 자세로 자연스럽게 이어지며 공중에 떠 있는 징후는 없다. 차단봉은 수직 기둥으로 지지되고 실내 부품들도 정상적으로 장착되어 있다."
       },
       {
        "label": "B",
        "direction": "중앙 민병대원은 차량 내부를 바라보면서 손바닥을 차량 쪽으로 들어 정지를 지시한다. 그의 총구는 화면 오른쪽 아래를 향한다. 왼쪽 대원의 얼굴과 소총은 도로 중앙 및 오른쪽 아래를 향하고, 오른쪽 대원은 오른쪽으로 이동하며 총을 낮추고 있다. 운전자는 앞길을 보고 있으며, 후방의 모자 쓴 인물은 상자 사이로 몸을 숙인다.",
        "built_space": "카메라는 화물 공간 쪽에서 앞좌석 두 개와 앞유리 하나를 본다. 왼쪽 운전대 하나, 중앙 룸미러 하나, 선바이저 두 개, 중앙 화면 하나가 있으며 참조의 라디오 중심 조작부와는 다르다. 왼쪽 큰 상자 하나와 중앙·오른쪽의 여러 상자 및 운반함이 화면 아래를 크게 차지한다. 밖에는 민병대원 세 명, 도로 양옆의 콘크리트 방벽, 교차형 장애물과 오른쪽 초소가 보인다. 참조에서 확인되는 장소보다 검문소 시설을 훨씬 구체적으로 추가했고, 실내 비중이 커져 외부 위협의 크기가 줄었다.",
        "entities": "외부의 세 대원은 동아시아계 성인 남성으로 보이며 군복과 방탄 장비를 착용한다. 총들은 참조의 개머리판 부착 권총보다는 긴 총몸과 전방 구조를 가진 소총형으로 보인다. 내부에는 샷 텍스트에 없는 운전자 한 명이 추가되어 있다. 후방에는 낡은 외투와 모자를 쓴 찰리로 해석되는 인물이 있지만 상체가 상당히 드러나 은폐 상태가 약하다. 화면 밖 지붕 십자가는 확인할 수 없다. 외부 인물과 건물 주변에 흰 화살표 표식 여러 개가 붙어 있다.",
        "hard_violations": [
         "실제 장면의 물체가 아닌 흰 화살표 도해가 화면에 노출되어 있다.",
         "샷 텍스트에 등장하지 않는 운전자를 앞좌석에 추가하여 등장인물 제한을 위반한다."
        ],
        "physics": "중앙 대원은 한 손으로 총을 잡고 다른 손을 들어 정지 신호를 하며, 총은 몸에 붙어 지지된다. 양옆 대원들도 손으로 총을 잡고 있고 오른쪽 대원은 다리를 벌린 보행 자세로 지면을 딛는다. 운전자는 좌석에 앉아 운전대를 잡는다. 후방 인물은 상자 뒤로 숙여 하체 지지가 가려져 있으나 떠 있다고 볼 근거는 없다. 상자와 운반함은 바닥 또는 다른 상자 위에 놓여 있다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "앞유리 너머 차량을 가로막는 민병대의 위협이라는 핵심 구도는 더 충실하지만, 노출된 로고와 한국인 설정에 맞지 않는 인물 구성, 참조 총기와의 차이가 남는다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "흰 화살표가 화면에 노출되고 지시되지 않은 운전자가 추가되었으며, 화물칸까지 넓힌 구도가 앞유리 밖 민병대의 위협을 약화한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "중앙 남성은 차 안을 똑바로 보며 권총 총구를 앞유리와 카메라 쪽으로 향한다. 나머지 네 명은 대체로 차량을 향해 서 있고, 긴 총들은 몸 앞에서 아래쪽 또는 비스듬한 옆쪽을 향한다. 왼쪽에서 두 번째 인물의 총은 거의 수직으로 내려가 있다. 모든 총이 차량을 겨눠야 한다는 지시는 없으므로 이 방향 배치는 위협적인 길막 장면과 양립한다.",
        "built_space": "밴 내부에서 앞유리 하나를 통해 도로를 직접 본다. 앞좌석 두 개의 일부, 왼쪽 운전대 하나, 중앙 룸미러 하나, 선바이저 두 개, 앞유리 아래 와이퍼 두 개가 보인다. 참조와 유사한 낡은 대시보드와 중앙 조작부가 유지된다. 밖에는 다섯 명이 도로 폭을 가로질러 서 있고 그 뒤로 가로 차단봉 하나가 놓여 있다. 낮은 건물과 가로등이 밤길을 따라 이어지며, 위협을 반사상으로 대체하지 않았다.",
        "entities": "민병대는 성인 남성 다섯 명이며 낡은 야전 재킷과 모자, 총기를 착용하거나 들고 있다. 중앙 인물은 백인으로, 오른쪽 끝 인물 등은 흑인으로 보이므로 별도 설명 없는 현지인은 한국인이라는 설정과 맞지 않는다. 중앙의 짧은 권총과 주변의 긴 소총들은 참조의 개머리판이 결합된 권총형 총기와 다르다. 운전대에는 식별 가능한 현대 로고가 있다. 후방 상자, 찰리, 지붕 십자가는 이 구도에서 보이지 않으므로 부재로 판정하지 않는다.",
        "hard_violations": [
         "운전대에 식별 가능한 현대 로고가 노출되어 명시적인 로고 금지 조건을 위반한다."
        ],
        "physics": "중앙 인물은 두 손으로 권총을 받쳐 들고, 나머지 인물들도 손과 팔로 총을 지지한다. 인물들의 하체 일부는 대시보드와 창틀에 가리지만 지면에 서 있는 자세로 자연스럽게 이어지며 공중에 떠 있는 징후는 없다. 차단봉은 수직 기둥으로 지지되고 실내 부품들도 정상적으로 장착되어 있다."
       },
       {
        "label": "A",
        "direction": "중앙 민병대원은 차량 내부를 바라보면서 손바닥을 차량 쪽으로 들어 정지를 지시한다. 그의 총구는 화면 오른쪽 아래를 향한다. 왼쪽 대원의 얼굴과 소총은 도로 중앙 및 오른쪽 아래를 향하고, 오른쪽 대원은 오른쪽으로 이동하며 총을 낮추고 있다. 운전자는 앞길을 보고 있으며, 후방의 모자 쓴 인물은 상자 사이로 몸을 숙인다.",
        "built_space": "카메라는 화물 공간 쪽에서 앞좌석 두 개와 앞유리 하나를 본다. 왼쪽 운전대 하나, 중앙 룸미러 하나, 선바이저 두 개, 중앙 화면 하나가 있으며 참조의 라디오 중심 조작부와는 다르다. 왼쪽 큰 상자 하나와 중앙·오른쪽의 여러 상자 및 운반함이 화면 아래를 크게 차지한다. 밖에는 민병대원 세 명, 도로 양옆의 콘크리트 방벽, 교차형 장애물과 오른쪽 초소가 보인다. 참조에서 확인되는 장소보다 검문소 시설을 훨씬 구체적으로 추가했고, 실내 비중이 커져 외부 위협의 크기가 줄었다.",
        "entities": "외부의 세 대원은 동아시아계 성인 남성으로 보이며 군복과 방탄 장비를 착용한다. 총들은 참조의 개머리판 부착 권총보다는 긴 총몸과 전방 구조를 가진 소총형으로 보인다. 내부에는 샷 텍스트에 없는 운전자 한 명이 추가되어 있다. 후방에는 낡은 외투와 모자를 쓴 찰리로 해석되는 인물이 있지만 상체가 상당히 드러나 은폐 상태가 약하다. 화면 밖 지붕 십자가는 확인할 수 없다. 외부 인물과 건물 주변에 흰 화살표 표식 여러 개가 붙어 있다.",
        "hard_violations": [
         "실제 장면의 물체가 아닌 흰 화살표 도해가 화면에 노출되어 있다.",
         "샷 텍스트에 등장하지 않는 운전자를 앞좌석에 추가하여 등장인물 제한을 위반한다."
        ],
        "physics": "중앙 대원은 한 손으로 총을 잡고 다른 손을 들어 정지 신호를 하며, 총은 몸에 붙어 지지된다. 양옆 대원들도 손으로 총을 잡고 있고 오른쪽 대원은 다리를 벌린 보행 자세로 지면을 딛는다. 운전자는 좌석에 앉아 운전대를 잡는다. 후방 인물은 상자 뒤로 숙여 하체 지지가 가려져 있으나 떠 있다고 볼 근거는 없다. 상자와 운반함은 바닥 또는 다른 상자 위에 놓여 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.762,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.512,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 지시문에 없는 인물(운전자) 추가",
     "[gemini-pro] 화면에 흰색 화살표 마커 유출",
     "[gpt-high] 실제 장면의 물체가 아닌 흰 화살표 도해가 화면에 노출되어 있다.",
     "[gpt-high] 샷 텍스트에 등장하지 않는 운전자를 앞좌석에 추가하여 등장인물 제한을 위반한다."
    ],
    "B": [
     "[gpt-high] 운전대에 식별 가능한 현대 로고가 노출되어 명시적인 로고 금지 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 1750,
   "A": 512
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "구도의 한계로 찰리가 프레임에서 제외되었으나, 앞유리창 너머로 위협을 가하는 민병대의 전경을 규칙 위반 없이 성공적으로 묘사했습니다.  ★위반: [gpt-high] 운전대에 식별 가능한 현대 로고가 노출되어 명시적인 로고 금지 조건을 위반한다."
   },
   {
    "label": "A",
    "score": 512,
    "verdict_ko": "지시문에 명시되지 않은 운전자가 프레임에 추가되었고, 화면 양측에 화살표 형태의 마커가 유출되어 치명적인 위반이 발생했습니다.  ★위반: [gemini-pro] 지시문에 없는 인물(운전자) 추가 / [gemini-pro] 화면에 흰색 화살표 마커 유출 / [gpt-high] 실제 장면의 물체가 아닌 흰 화살표 도해가 화면에 노출되어 있다. / [gpt-high] 샷 텍스트에 등장하지 않는 운전자를 앞좌석에 추가하여 등장인물 제한을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L179B01.png",
    "asset_id": "f6306063-3d19-4618-946d-060af53e1589",
    "role": "location_plate"
   },
   {
    "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
    "path": "<bytes:842741>",
    "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
    "role": "prop_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-3ac0-7379-9c72-359a1a99124e",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S28sh5__bgfirst_bg.png",
   "bg_asset_id": "2fb1b93e-74b7-4cf0-85b0-9fb67e335dcd",
   "bg_record_key": "S28sh5::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S28sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:29:41.624724+00:00",
  "fingerprint": "1b838b6bbcd292dc6e45f3ba0c6f7848667a51cd292343588a51fa89870ed32d",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S28sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S28sh5_sel.png",
  "source_sha256": "6115c6a9dc5bb761a567220ea203835422021d49f560fa5eb86bfe98756071de",
  "file": "S28sh5_cine.png",
  "staged_sha256": "38740b888248d9b956da3b795a54c63c33ad44b3a63011d92ca3c83d72712eb5",
  "latency_ms": 10970
 },
 "S28sh7::confined_fp_apt": {
  "applies": true,
  "reason_ko": "차량(밴)의 운전석 내부를 배경으로 하는 숏으로, 운전대 및 열린 운전석 창문과 외부 인물(민병대원들) 간의 공간적 위치 관계를 정확히 묘사해야 하므로 평면도 레이아웃 가이드가 필요합니다.",
  "input_fingerprint": "3be4f6485ed8fa97"
 },
 "S28sh7::signage": {
  "fp": "231d94694e2e23f2",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "confinedfp::9fd979762448": {
  "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/confinedfp_base_9fd979762448.png",
  "place_text": "At the driver's seat inside the van's compact front cab, beside the lowered side window during the nighttime checkpoint stop.",
  "input_fingerprint": "5612a230c9bd725a"
 },
 "S28sh7::confined_fp": {
  "reads": {
   "controls": "The steering wheel is located at the front-left station (driver's seat).",
   "mirrors": "No mirrors are depicted in the diagram.",
   "camera": "The camera is positioned in the middle of the cabin, forward of the seats, pointing directly to the left towards the driver's seat and the adjacent side window.",
   "occupants": "신부 is seated in the left (driver's) seat. The right (passenger) seat is marked but unoccupied."
  },
  "mismatches": [],
  "scene_description_en": "The camera is positioned inside the center of the cabin, pointing directly left at the driver's seat in a close-up framing. In the center foreground, 신부 is seated, seen in profile facing left. On the left edge of the frame, slightly in front of him, the right side of the steering wheel is visible. Beyond his profile on the far left, the driver's side window is entirely lowered, opening to the exterior night. The view focuses entirely on the driver and the window, with the rest of the cabin behind the camera or off-screen to the right.",
  "fixed": true,
  "input_fingerprint": "25f45d1e1eff4916"
 },
 "S28sh7": {
  "input_fingerprint": "98908b64918c5925",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 열린 운전석 창문 너머의 민병대원들을 향해 활짝 웃으며 찡긋 눈웃음을 짓는 신부의 여유로운 얼굴.\n\nLOCATION (lock): At the driver's seat inside the van's compact front cab, beside the lowered side window during the nighttime checkpoint stop. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Open driver's window (Lowered for the exchange with the militia) — Its open aperture is seen obliquely beyond 신부's profile; no glass lies across the opening; used as Establishes the direction of his smile and offscreen attention; Driver's seat (Occupied by 신부) — A limited portion is visible behind his shoulder; used as Keeps the close portrait physically grounded inside the van.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained nighttime ambient light across the cabin and open window, keeping the smile and wink legible without attributing facial illumination to the roof ornament.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The driver's window is lowered for the inspection, with the halogen cross fixed on the roof. The rear boxes continue to conceal Charlie in his old coat and hat. 신부: He remains seated at the wheel in his clerical collar, smiling through the lowered window.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera is positioned inside the center of the cabin, pointing directly left at the driver's seat in a close-up framing. In the center foreground, 신부 is seated, seen in profile facing left. On the left edge of the frame, slightly in front of him, the right side of the steering wheel is visible. Beyond his profile on the far left, the driver's side window is entirely lowered, opening to the exterior night. The view focuses entirely on the driver and the window, with the rest of the cabin behind the camera or off-screen to the right.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 열린 운전석 창문 너머의 민병대원들을 향해 활짝 웃으며 찡긋 눈웃음을 짓는 신부의 여유로운 얼굴.\n\nLOCATION (lock): At the driver's seat inside the van's compact front cab, beside the lowered side window during the nighttime checkpoint stop. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained nighttime ambient light across the cabin and open window, keeping the smile and wink legible without attributing facial illumination to the roof ornament.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The driver's window is lowered for the inspection, with the halogen cross fixed on the roof. The rear boxes continue to conceal Charlie in his old coat and hat. 신부: He remains seated at the wheel in his clerical collar, smiling through the lowered window.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera is positioned inside the center of the cabin, pointing directly left at the driver's seat in a close-up framing. In the center foreground, 신부 is seated, seen in profile facing left. On the left edge of the frame, slightly in front of him, the right side of the steering wheel is visible. Beyond his profile on the far left, the driver's side window is entirely lowered, opening to the exterior night. The view focuses entirely on the driver and the window, with the rest of the cabin behind the camera or off-screen to the right.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 열린 운전석 창문 너머의 민병대원들을 향해 활짝 웃으며 찡긋 눈웃음을 짓는 신부의 여유로운 얼굴.\n\nLOCATION (lock): At the driver's seat inside the van's compact front cab, beside the lowered side window during the nighttime checkpoint stop. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained nighttime ambient light across the cabin and open window, keeping the smile and wink legible without attributing facial illumination to the roof ornament.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The driver's window is lowered for the inspection, with the halogen cross fixed on the roof. The rear boxes continue to conceal Charlie in his old coat and hat. 신부: He remains seated at the wheel in his clerical collar, smiling through the lowered window.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S28sh7_confinedfp.png",
     "asset_id": null,
     "role": null
    },
    {
     "label": "신부",
     "path": "<bytes:1402213>",
     "asset_id": "8696070d-ac09-4a5c-95f2-3daf1c015c32",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S28sh7_confinedfp.png",
     "asset_id": null,
     "role": null
    },
    {
     "label": "신부",
     "path": "<bytes:1402213>",
     "asset_id": "8696070d-ac09-4a5c-95f2-3daf1c015c32",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "신부의 시선과 미소가 창밖의 민병대원들을 정확히 향하고 있음.",
    "built_space": "차량 내부 조수석에서 운전석을 바라보는 구도. 지시대로 창문에 유리가 없음.",
    "entities": "신부의 외모와 복장이 레퍼런스와 일치하며, 창밖 인물들도 한국인 민병대원으로 묘사됨.",
    "hard_violations": [
     "[gpt-high] 화면에 허용된 신부 외에 창밖 인물 일곱 명을 추가했다."
    ],
    "physics": "신부가 운전석 좌석에 자연스럽게 기대어 착석해 있음."
   },
   {
    "label": "B",
    "direction": "신부의 시선이 창밖의 군인들을 향하고 있음.",
    "built_space": "차량 내부 구도이나, 열려 있어야 할 창문에 유리가 묘사되어 스티어링 휠이 반사됨.",
    "entities": "신부는 레퍼런스와 일치하나, 창밖 인물들이 성조기 패치를 단 다인종 군인으로 묘사됨.",
    "hard_violations": [
     "[gemini-pro] 창문에 유리가 없어야 한다는 명시적 지시 위반 (유리 반사가 선명함)",
     "[gpt-high] 화면에 허용된 신부 외에 창밖 군복 차림 인물 세 명을 추가했다.",
     "[gpt-high] 운전대에 식별 가능한 차량 로고가 노출되어 로고 금지 조건을 위반한다."
    ],
    "physics": "신부가 운전석 좌석에 자연스럽게 착석해 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "유리가 없는 열린 창문과 한국인 민병대원 설정을 명확히 구현하여 프롬프트 지시를 잘 따랐습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "창문에 유리가 묘사되어 반사가 발생했고, 민병대원을 성조기를 단 미군으로 묘사하여 인물 및 배경 설정을 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "신부의 시선과 미소가 창밖의 민병대원들을 정확히 향하고 있음.",
        "built_space": "차량 내부 조수석에서 운전석을 바라보는 구도. 지시대로 창문에 유리가 없음.",
        "entities": "신부의 외모와 복장이 레퍼런스와 일치하며, 창밖 인물들도 한국인 민병대원으로 묘사됨.",
        "hard_violations": [],
        "physics": "신부가 운전석 좌석에 자연스럽게 기대어 착석해 있음."
       },
       {
        "label": "B",
        "direction": "신부의 시선이 창밖의 군인들을 향하고 있음.",
        "built_space": "차량 내부 구도이나, 열려 있어야 할 창문에 유리가 묘사되어 스티어링 휠이 반사됨.",
        "entities": "신부는 레퍼런스와 일치하나, 창밖 인물들이 성조기 패치를 단 다인종 군인으로 묘사됨.",
        "hard_violations": [
         "창문에 유리가 없어야 한다는 명시적 지시 위반 (유리 반사가 선명함)"
        ],
        "physics": "신부가 운전석 좌석에 자연스럽게 착석해 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "유리가 없는 열린 창문과 한국인 민병대원 설정을 명확히 구현하여 프롬프트 지시를 잘 따랐습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "창문에 유리가 묘사되어 반사가 발생했고, 민병대원을 성조기를 단 미군으로 묘사하여 인물 및 배경 설정을 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "신부의 시선과 미소가 창밖의 민병대원들을 정확히 향하고 있음.",
        "built_space": "차량 내부 조수석에서 운전석을 바라보는 구도. 지시대로 창문에 유리가 없음.",
        "entities": "신부의 외모와 복장이 레퍼런스와 일치하며, 창밖 인물들도 한국인 민병대원으로 묘사됨.",
        "hard_violations": [],
        "physics": "신부가 운전석 좌석에 자연스럽게 기대어 착석해 있음."
       },
       {
        "label": "B",
        "direction": "신부의 시선이 창밖의 군인들을 향하고 있음.",
        "built_space": "차량 내부 구도이나, 열려 있어야 할 창문에 유리가 묘사되어 스티어링 휠이 반사됨.",
        "entities": "신부는 레퍼런스와 일치하나, 창밖 인물들이 성조기 패치를 단 다인종 군인으로 묘사됨.",
        "hard_violations": [
         "창문에 유리가 없어야 한다는 명시적 지시 위반 (유리 반사가 선명함)"
        ],
        "physics": "신부가 운전석 좌석에 자연스럽게 착석해 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "창문에 유리가 없고 미소와 찡긋한 눈웃음이 더 선명하지만, 허용되지 않은 외부 인물과 차량 로고가 등장하며 얼굴 클로즈업보다 넓고 민병대를 향한 시선도 불명확하다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "신부의 착석과 야간 분위기는 맞지만, 외부 군중을 추가하고 창문 아래쪽에 유리를 남겼으며 얼굴보다 운전석과 군중을 크게 보여 주어 지정 클로즈업을 벗어난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "신부의 얼굴과 가늘게 뜬 눈은 화면 왼쪽 차량 전방 쪽을 향한다. 창밖 민병대는 얼굴보다 뒤쪽 배경에 있어 그들에게 직접 눈을 맞추는 방향으로 명확히 읽히지 않는다. 외부 인물들은 운전석 안쪽을 바라본다. 몸 앞에 든 총기류는 일부 가려져 총구의 정확한 방향과 표적을 판독하기 어렵다.",
        "built_space": "운전대 하나, 운전석 등받이와 머리받침 하나, 운전석 측면 창문 하나, 사이드미러 하나, 기둥 손잡이 하나와 문 안쪽 손잡이가 보인다. 신부는 운전대 뒤 좌석에 있고 등받이는 어깨 뒤에 있어 좌석 배치는 자연스럽다. 창문 개구부에는 유리가 보이지 않는다. 다만 운전대와 문, 상체를 넓게 포함하여 요구한 얼굴 클로즈업과 제한적인 좌석 노출보다 넓다. 사이드미러에는 흐린 빛만 보여 불가능한 반사를 확인할 근거는 없다.",
        "entities": "신부는 희끗한 짧은 머리, 이마와 눈가 주름이 있는 한국인 60대 남성의 인상으로 참조와 대체로 부합한다. 검은 성직자복과 흰 성직자 칼라를 착용하고 치아를 드러내며 웃고 한쪽 눈을 더 감고 있다. 창밖에는 군복 차림 인물 세 명이 보이며 한 명은 상당 부분 가려져 있다. 이는 신부만 화면에 허용한 인물 제한에 어긋난다. 운전대에는 식별 가능한 현대차 로고가 있다. 지붕 십자가와 뒤쪽 상자, 찰리는 이 구도 밖이므로 미노출 자체는 결함이 아니다.",
        "hard_violations": [
         "화면에 허용된 신부 외에 창밖 군복 차림 인물 세 명을 추가했다.",
         "운전대에 식별 가능한 차량 로고가 노출되어 로고 금지 조건을 위반한다."
        ],
        "physics": "신부의 몸은 좌석에 놓이고 등과 어깨 뒤에 등받이가 있으며 안전벨트가 몸통을 가로지른다. 착석 자세와 고개 회전은 물리적으로 가능하다. 외부 인물들의 하체는 문에 가려졌지만 서 있는 상체 배치에 공중부양 징후는 없다. 총기류는 손과 몸 앞의 장비에 겹쳐 있고 명백히 무지지 상태인 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "신부는 화면 왼쪽 전방을 향해 웃으며 눈을 좁힌다. 창밖 군중은 신부의 측후방 배경에 놓여 있어 미소와 시선이 그들에게 직접 향한다고 확신하기 어렵다. 군중은 대체로 운전석 안을 바라본다. 명확히 식별되는 총구나 겨냥 동작은 없다.",
        "built_space": "운전대 하나의 일부, 운전석 등받이와 머리받침 하나, 측면 창문 하나, 사이드미러 하나, 실내 거울 일부, 기둥 손잡이 하나와 천장 손잡이 하나, 문 손잡이와 수동 창문 조작부가 보인다. 신부는 운전대 뒤 좌석에 정상적으로 앉아 있다. 창문 아래쪽에는 올라온 유리의 곡선 가장자리가 보여 유리가 개구부를 가로막지 않아야 한다는 조건과 다르다. 상체와 문 전체에 가까운 영역을 포함해 얼굴 클로즈업보다 넓다. 거울에서 불가능한 반사는 확인되지 않는다.",
        "entities": "주인공은 참조와 비슷한 희끗한 머리와 주름을 가진 한국인 60대 남성으로 보이며 검은 성직자복과 흰 칼라를 착용한다. 치아를 드러낸 미소는 분명하지만 한쪽 눈을 찡긋하는 동작은 A보다 덜 뚜렷하다. 창밖에는 군복 또는 어두운 옷을 입은 인물 일곱 명이 보여 신부만 허용한 화면 인물 조건을 위반한다. 읽을 수 있는 문자나 뚜렷한 로고는 보이지 않는다. 십자가, 상자와 찰리는 프레임 밖이므로 평가상 누락으로 보지 않는다.",
        "hard_violations": [
         "화면에 허용된 신부 외에 창밖 인물 일곱 명을 추가했다."
        ],
        "physics": "신부의 등과 몸통은 좌석 등받이에 맞게 놓여 있고 하단에 보이는 팔과 손도 운전대 주변에서 자연스럽게 이어진다. 창문 유리는 문 내부 기구에 지지되는 형태다. 외부 인물들은 하체가 가려진 채 서 있는 것으로 읽히며, 지지 없이 떠 있는 몸이나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "창문에 유리가 없고 미소와 찡긋한 눈웃음이 더 선명하지만, 허용되지 않은 외부 인물과 차량 로고가 등장하며 얼굴 클로즈업보다 넓고 민병대를 향한 시선도 불명확하다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "신부의 착석과 야간 분위기는 맞지만, 외부 군중을 추가하고 창문 아래쪽에 유리를 남겼으며 얼굴보다 운전석과 군중을 크게 보여 주어 지정 클로즈업을 벗어난다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "신부의 얼굴과 가늘게 뜬 눈은 화면 왼쪽 차량 전방 쪽을 향한다. 창밖 민병대는 얼굴보다 뒤쪽 배경에 있어 그들에게 직접 눈을 맞추는 방향으로 명확히 읽히지 않는다. 외부 인물들은 운전석 안쪽을 바라본다. 몸 앞에 든 총기류는 일부 가려져 총구의 정확한 방향과 표적을 판독하기 어렵다.",
        "built_space": "운전대 하나, 운전석 등받이와 머리받침 하나, 운전석 측면 창문 하나, 사이드미러 하나, 기둥 손잡이 하나와 문 안쪽 손잡이가 보인다. 신부는 운전대 뒤 좌석에 있고 등받이는 어깨 뒤에 있어 좌석 배치는 자연스럽다. 창문 개구부에는 유리가 보이지 않는다. 다만 운전대와 문, 상체를 넓게 포함하여 요구한 얼굴 클로즈업과 제한적인 좌석 노출보다 넓다. 사이드미러에는 흐린 빛만 보여 불가능한 반사를 확인할 근거는 없다.",
        "entities": "신부는 희끗한 짧은 머리, 이마와 눈가 주름이 있는 한국인 60대 남성의 인상으로 참조와 대체로 부합한다. 검은 성직자복과 흰 성직자 칼라를 착용하고 치아를 드러내며 웃고 한쪽 눈을 더 감고 있다. 창밖에는 군복 차림 인물 세 명이 보이며 한 명은 상당 부분 가려져 있다. 이는 신부만 화면에 허용한 인물 제한에 어긋난다. 운전대에는 식별 가능한 현대차 로고가 있다. 지붕 십자가와 뒤쪽 상자, 찰리는 이 구도 밖이므로 미노출 자체는 결함이 아니다.",
        "hard_violations": [
         "화면에 허용된 신부 외에 창밖 군복 차림 인물 세 명을 추가했다.",
         "운전대에 식별 가능한 차량 로고가 노출되어 로고 금지 조건을 위반한다."
        ],
        "physics": "신부의 몸은 좌석에 놓이고 등과 어깨 뒤에 등받이가 있으며 안전벨트가 몸통을 가로지른다. 착석 자세와 고개 회전은 물리적으로 가능하다. 외부 인물들의 하체는 문에 가려졌지만 서 있는 상체 배치에 공중부양 징후는 없다. 총기류는 손과 몸 앞의 장비에 겹쳐 있고 명백히 무지지 상태인 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "신부는 화면 왼쪽 전방을 향해 웃으며 눈을 좁힌다. 창밖 군중은 신부의 측후방 배경에 놓여 있어 미소와 시선이 그들에게 직접 향한다고 확신하기 어렵다. 군중은 대체로 운전석 안을 바라본다. 명확히 식별되는 총구나 겨냥 동작은 없다.",
        "built_space": "운전대 하나의 일부, 운전석 등받이와 머리받침 하나, 측면 창문 하나, 사이드미러 하나, 실내 거울 일부, 기둥 손잡이 하나와 천장 손잡이 하나, 문 손잡이와 수동 창문 조작부가 보인다. 신부는 운전대 뒤 좌석에 정상적으로 앉아 있다. 창문 아래쪽에는 올라온 유리의 곡선 가장자리가 보여 유리가 개구부를 가로막지 않아야 한다는 조건과 다르다. 상체와 문 전체에 가까운 영역을 포함해 얼굴 클로즈업보다 넓다. 거울에서 불가능한 반사는 확인되지 않는다.",
        "entities": "주인공은 참조와 비슷한 희끗한 머리와 주름을 가진 한국인 60대 남성으로 보이며 검은 성직자복과 흰 칼라를 착용한다. 치아를 드러낸 미소는 분명하지만 한쪽 눈을 찡긋하는 동작은 A보다 덜 뚜렷하다. 창밖에는 군복 또는 어두운 옷을 입은 인물 일곱 명이 보여 신부만 허용한 화면 인물 조건을 위반한다. 읽을 수 있는 문자나 뚜렷한 로고는 보이지 않는다. 십자가, 상자와 찰리는 프레임 밖이므로 평가상 누락으로 보지 않는다.",
        "hard_violations": [
         "화면에 허용된 신부 외에 창밖 인물 일곱 명을 추가했다."
        ],
        "physics": "신부의 등과 몸통은 좌석 등받이에 맞게 놓여 있고 하단에 보이는 팔과 손도 운전대 주변에서 자연스럽게 이어진다. 창문 유리는 문 내부 기구에 지지되는 형태다. 외부 인물들은 하체가 가려진 채 서 있는 것으로 읽히며, 지지 없이 떠 있는 몸이나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.667,
    "B": 1.429
   },
   "adjusted": {
    "A": 1.417,
    "B": 1.179
   },
   "violations": {
    "B": [
     "[gemini-pro] 창문에 유리가 없어야 한다는 명시적 지시 위반 (유리 반사가 선명함)",
     "[gpt-high] 화면에 허용된 신부 외에 창밖 군복 차림 인물 세 명을 추가했다.",
     "[gpt-high] 운전대에 식별 가능한 차량 로고가 노출되어 로고 금지 조건을 위반한다."
    ],
    "A": [
     "[gpt-high] 화면에 허용된 신부 외에 창밖 인물 일곱 명을 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1417,
   "B": 1179
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1417,
    "verdict_ko": "유리가 없는 열린 창문과 한국인 민병대원 설정을 명확히 구현하여 프롬프트 지시를 잘 따랐습니다.  ★위반: [gpt-high] 화면에 허용된 신부 외에 창밖 인물 일곱 명을 추가했다."
   },
   {
    "label": "B",
    "score": 1179,
    "verdict_ko": "창문에 유리가 묘사되어 반사가 발생했고, 민병대원을 성조기를 단 미군으로 묘사하여 인물 및 배경 설정을 위반했습니다.  ★위반: [gemini-pro] 창문에 유리가 없어야 한다는 명시적 지시 위반 (유리 반사가 선명함) / [gpt-high] 화면에 허용된 신부 외에 창밖 군복 차림 인물 세 명을 추가했다. / [gpt-high] 운전대에 식별 가능한 차량 로고가 노출되어 로고 금지 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S28sh7_confinedfp.png",
    "asset_id": null,
    "role": null
   },
   {
    "label": "신부",
    "path": "<bytes:1402213>",
    "asset_id": "8696070d-ac09-4a5c-95f2-3daf1c015c32",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-3e32-7d7f-a986-495d18ee8fa4",
  "confined_fp": {
   "base_key": "confinedfp::9fd979762448",
   "apt_reason": "차량(밴)의 운전석 내부를 배경으로 하는 숏으로, 운전대 및 열린 운전석 창문과 외부 인물(민병대원들) 간의 공간적 위치 관계를 정확히 묘사해야 하므로 평면도 레이아웃 가이드가 필요합니다.",
   "fixed": true,
   "mismatches": []
  },
  "ref_mode": "confined_fp: 도면+장면설명+엔티티",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S28sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:32:37.339435+00:00",
  "fingerprint": "d6cf74a879009bdf035b071d522ebc418e6a8d33af0024b8ea78d68ca81df9fd",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S28sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S28sh7_sel.png",
  "source_sha256": "87000210a9a8815c0a1580c02378679f3e1eccc5534117999bcc7bda8ec5c53c",
  "file": "S28sh7_cine.png",
  "staged_sha256": "9b17ded46cf02f9ad7f547543d2c4d1364e7cce3dc45134228499cc7e3ec4e7b",
  "latency_ms": 12564
 },
 "S29sh6::signage": {
  "fp": "0b03705f58a7f802",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::b7ea15bbe8cf36a0": {
  "subjects": [],
  "subject_text": "인천 성당 지하 기도실\n십자가와 작은 책상, 의자, 작은 풍금이 놓인 단출한 지하 공간. 구석에는 낡은 라디오가 있고 촛불과 랜턴이 실내를 밝힌다.",
  "identity": "canonical",
  "scope_id": "L180",
  "scope_role": "location_interior",
  "scope_sha": "e2c03ad952b754fc"
 },
 "S29sh6::bgfirst_bg": {
  "input_fingerprint": "211f365f6147facf",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 1층에서 들려오는 둔탁한 타격음에 일제히 어깨를 움츠린 채 위를 올려다보며 얼어붙은 일행의 전경.\n\nLOCATION (lock): Inside the church's sparsely furnished basement prayer room, lit by the priest's candle beside a small organ, table, and chairs.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Stairs to the upper floor (Connect the prayer room to the source of the knocking) — The lower steps enter along the left foreground and continue out of view upward; used as Anchor the camera position and explain the group's common upward attention; Small desk and chair (Present in the sparsely furnished prayer room) — Seen obliquely beyond the group; used as Provide modest human-scale context without crowding the reaction tableau; Small reed organ (Present among the room's few furnishings) — Partially visible at an oblique angle in the rear of the frame; used as Preserves the room's identity while remaining secondary to the startled figures; Cross (Present in the prayer room) — Its recognizable form remains visible in the background; used as Adds a restrained location cue without becoming the group's gaze target.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the established candlelight supply selective warmth within the subdued prayer room, preserving the upward-looking faces with controlled contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 1층에서 들려오는 둔탁한 타격음에 일제히 어깨를 움츠린 채 위를 올려다보며 얼어붙은 일행의 전경.\n\nLOCATION (lock): Inside the church's sparsely furnished basement prayer room, lit by the priest's candle beside a small organ, table, and chairs.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Stairs to the upper floor (Connect the prayer room to the source of the knocking) — The lower steps enter along the left foreground and continue out of view upward; used as Anchor the camera position and explain the group's common upward attention; Small desk and chair (Present in the sparsely furnished prayer room) — Seen obliquely beyond the group; used as Provide modest human-scale context without crowding the reaction tableau; Small reed organ (Present among the room's few furnishings) — Partially visible at an oblique angle in the rear of the frame; used as Preserves the room's identity while remaining secondary to the startled figures; Cross (Present in the prayer room) — Its recognizable form remains visible in the background; used as Adds a restrained location cue without becoming the group's gaze target.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the established candlelight supply selective warmth within the subdued prayer room, preserving the upward-looking faces with controlled contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S29sh6__bgfirst_bg.png",
  "asset_id": "4df963a5-17a9-4cdd-a2c1-6834d045d891",
  "input_asset_ids": [
   "51f9964d-9eec-4e5f-a76a-7feb0e6f9752",
   "973d81a2-fcfc-4412-82af-c60138fde795"
  ]
 },
 "S29sh6": {
  "input_fingerprint": "69cac38fa859f751",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 1층에서 들려오는 둔탁한 타격음에 일제히 어깨를 움츠린 채 위를 올려다보며 얼어붙은 일행의 전경.\n\nLOCATION (lock): Inside the church's sparsely furnished basement prayer room, lit by the priest's candle beside a small organ, table, and chairs. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Stairs to the upper floor (Connect the prayer room to the source of the knocking) — The lower steps enter along the left foreground and continue out of view upward; used as Anchor the camera position and explain the group's common upward attention; Small desk and chair (Present in the sparsely furnished prayer room) — Seen obliquely beyond the group; used as Provide modest human-scale context without crowding the reaction tableau; Small reed organ (Present among the room's few furnishings) — Partially visible at an oblique angle in the rear of the frame; used as Preserves the room's identity while remaining secondary to the startled figures; Cross (Present in the prayer room) — Its recognizable form remains visible in the background; used as Adds a restrained location cue without becoming the group's gaze target.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the established candlelight supply selective warmth within the subdued prayer room, preserving the upward-looking faces with controlled contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement prayer room contains a cross, a small desk, chairs and a small organ, with candlelight brought down the stairs. Charlie retains his old coat and hat. 현우: He is in the basement prayer room with facial bruises and an injured leg, without his outer garment. The contact card remains concealed in his shoe. 앰버: She is in the basement, tired and hungry, retaining the outer garment used as a mouth covering. 신부: He is in the basement wearing his clerical collar and carrying a lit candle.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 1층에서 들려오는 둔탁한 타격음에 일제히 어깨를 움츠린 채 위를 올려다보며 얼어붙은 일행의 전경.\n\nLOCATION (lock): Inside the church's sparsely furnished basement prayer room, lit by the priest's candle beside a small organ, table, and chairs. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Stairs to the upper floor (Connect the prayer room to the source of the knocking) — The lower steps enter along the left foreground and continue out of view upward; used as Anchor the camera position and explain the group's common upward attention; Small desk and chair (Present in the sparsely furnished prayer room) — Seen obliquely beyond the group; used as Provide modest human-scale context without crowding the reaction tableau; Small reed organ (Present among the room's few furnishings) — Partially visible at an oblique angle in the rear of the frame; used as Preserves the room's identity while remaining secondary to the startled figures; Cross (Present in the prayer room) — Its recognizable form remains visible in the background; used as Adds a restrained location cue without becoming the group's gaze target.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the established candlelight supply selective warmth within the subdued prayer room, preserving the upward-looking faces with controlled contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement prayer room contains a cross, a small desk, chairs and a small organ, with candlelight brought down the stairs. Charlie retains his old coat and hat. 현우: He is in the basement prayer room with facial bruises and an injured leg, without his outer garment. The contact card remains concealed in his shoe. 앰버: She is in the basement, tired and hungry, retaining the outer garment used as a mouth covering. 신부: He is in the basement wearing his clerical collar and carrying a lit candle.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 1층에서 들려오는 둔탁한 타격음에 일제히 어깨를 움츠린 채 위를 올려다보며 얼어붙은 일행의 전경.\n\nLOCATION (lock): Inside the church's sparsely furnished basement prayer room, lit by the priest's candle beside a small organ, table, and chairs. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Stairs to the upper floor (Connect the prayer room to the source of the knocking) — The lower steps enter along the left foreground and continue out of view upward; used as Anchor the camera position and explain the group's common upward attention; Small desk and chair (Present in the sparsely furnished prayer room) — Seen obliquely beyond the group; used as Provide modest human-scale context without crowding the reaction tableau; Small reed organ (Present among the room's few furnishings) — Partially visible at an oblique angle in the rear of the frame; used as Preserves the room's identity while remaining secondary to the startled figures; Cross (Present in the prayer room) — Its recognizable form remains visible in the background; used as Adds a restrained location cue without becoming the group's gaze target.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the established candlelight supply selective warmth within the subdued prayer room, preserving the upward-looking faces with controlled contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement prayer room contains a cross, a small desk, chairs and a small organ, with candlelight brought down the stairs. Charlie retains his old coat and hat. 현우: He is in the basement prayer room with facial bruises and an injured leg, without his outer garment. The contact card remains concealed in his shoe. 앰버: She is in the basement, tired and hungry, retaining the outer garment used as a mouth covering. 신부: He is in the basement wearing his clerical collar and carrying a lit candle.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S29sh6__bgfirst_bg.png",
     "asset_id": "4df963a5-17a9-4cdd-a2c1-6834d045d891",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S29sh6.png",
     "asset_id": "51f9964d-9eec-4e5f-a76a-7feb0e6f9752",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 신부: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1402213>",
     "asset_id": "8696070d-ac09-4a5c-95f2-3daf1c015c32",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L180B02.png",
     "asset_id": "973d81a2-fcfc-4412-82af-c60138fde795",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 신부: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1402213>",
     "asset_id": "8696070d-ac09-4a5c-95f2-3daf1c015c32",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "신부, 현우, 앰버는 둔탁한 소리가 나는 화면 좌측의 계단 위쪽을 일제히 올려다보고 있다. 찰리(로봇)의 시선은 정면 우측을 향하고 있다.",
    "built_space": "지하실 기도실의 레퍼런스와 일치한다. 좌측 전경에 위로 향하는 계단이 있고, 배경 중앙 벽면에 십자가가 보이며, 촛불이 놓인 책상과 의자, 우측 뒤편에 작은 오르간이 정확한 스케일과 위치에 배치되어 있다.",
    "entities": "지시된 4명의 캐릭터만 존재한다. 신부(참조 이미지 일치, 로만 칼라 착용, 양초 들고 있음), 현우(참조 이미지 일치, 얼굴에 멍, 겉옷 없음), 앰버(참조 이미지 일치, 겉옷으로 입을 가림), 찰리(참조 이미지의 로봇 외형 일치, 단 코트와 모자는 생략됨).",
    "hard_violations": [],
    "physics": "모든 인물은 지하실 바닥에 체중을 싣고 서 있으며, 어깨를 움츠리고 굳어있는 자세가 물리적으로 자연스럽다. 신부의 손은 촛대를 안정적으로 쥐고 있다."
   },
   {
    "label": "B",
    "direction": "등장인물 대부분이 천장이나 위쪽을 올려다보고 있다.",
    "built_space": "지하실 기도실의 구조가 잘 반영되었다. 좌측 계단, 중앙의 책상과 의자, 배경의 십자가, 우측 뒤편의 오르간이 적절히 배치되어 있다.",
    "entities": "참조 이미지의 신부와 똑같이 생긴 남성, 촛불을 들고 로만 칼라를 입은 또 다른 남성, 현우, 앰버, 찰리(로봇), 그리고 우측 끝에 신원 미상의 여성까지 총 6명이 등장한다. 프롬프트가 지시한 4명 제한을 어겼다.",
    "hard_violations": [
     "[gemini-pro] 프롬프트에 명시되지 않은 발명된 인물(신원 미상의 여성) 추가",
     "[gemini-pro] 신부 캐릭터가 두 명의 별개 인물(참조 얼굴을 한 남성과 칼라를 입은 남성)로 복제 및 분열됨",
     "[gpt-high] 허용된 네 인물 대신 여섯 인물이 등장하여 지정되지 않은 인물 두 명이 추가됐다."
    ],
    "physics": "인물들은 바닥에 잘 서 있으며 촛불을 쥔 손이나 신체 자세에 물리적 오류는 없다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "명시된 4명의 인물만 정확하게 배치하고 위층을 올려다보는 긴장된 순간과 캐릭터별 복장 및 상태 지침(앰버의 입 가림, 현우의 상처 및 겉옷 없음, 신부의 촛불)을 매우 충실히 구현했습니다."
       },
       {
        "label": "B",
        "score": 0,
        "verdict_ko": "지시된 4명의 인물 외에 프롬프트에 없는 인물들이 추가로 등장하고, 신부 캐릭터가 두 명으로 분열되어 나타나는 치명적인 오류가 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "신부, 현우, 앰버는 둔탁한 소리가 나는 화면 좌측의 계단 위쪽을 일제히 올려다보고 있다. 찰리(로봇)의 시선은 정면 우측을 향하고 있다.",
        "built_space": "지하실 기도실의 레퍼런스와 일치한다. 좌측 전경에 위로 향하는 계단이 있고, 배경 중앙 벽면에 십자가가 보이며, 촛불이 놓인 책상과 의자, 우측 뒤편에 작은 오르간이 정확한 스케일과 위치에 배치되어 있다.",
        "entities": "지시된 4명의 캐릭터만 존재한다. 신부(참조 이미지 일치, 로만 칼라 착용, 양초 들고 있음), 현우(참조 이미지 일치, 얼굴에 멍, 겉옷 없음), 앰버(참조 이미지 일치, 겉옷으로 입을 가림), 찰리(참조 이미지의 로봇 외형 일치, 단 코트와 모자는 생략됨).",
        "hard_violations": [],
        "physics": "모든 인물은 지하실 바닥에 체중을 싣고 서 있으며, 어깨를 움츠리고 굳어있는 자세가 물리적으로 자연스럽다. 신부의 손은 촛대를 안정적으로 쥐고 있다."
       },
       {
        "label": "B",
        "direction": "등장인물 대부분이 천장이나 위쪽을 올려다보고 있다.",
        "built_space": "지하실 기도실의 구조가 잘 반영되었다. 좌측 계단, 중앙의 책상과 의자, 배경의 십자가, 우측 뒤편의 오르간이 적절히 배치되어 있다.",
        "entities": "참조 이미지의 신부와 똑같이 생긴 남성, 촛불을 들고 로만 칼라를 입은 또 다른 남성, 현우, 앰버, 찰리(로봇), 그리고 우측 끝에 신원 미상의 여성까지 총 6명이 등장한다. 프롬프트가 지시한 4명 제한을 어겼다.",
        "hard_violations": [
         "프롬프트에 명시되지 않은 발명된 인물(신원 미상의 여성) 추가",
         "신부 캐릭터가 두 명의 별개 인물(참조 얼굴을 한 남성과 칼라를 입은 남성)로 복제 및 분열됨"
        ],
        "physics": "인물들은 바닥에 잘 서 있으며 촛불을 쥔 손이나 신체 자세에 물리적 오류는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "명시된 4명의 인물만 정확하게 배치하고 위층을 올려다보는 긴장된 순간과 캐릭터별 복장 및 상태 지침(앰버의 입 가림, 현우의 상처 및 겉옷 없음, 신부의 촛불)을 매우 충실히 구현했습니다."
       },
       {
        "label": "B",
        "score": 0,
        "verdict_ko": "지시된 4명의 인물 외에 프롬프트에 없는 인물들이 추가로 등장하고, 신부 캐릭터가 두 명으로 분열되어 나타나는 치명적인 오류가 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "신부, 현우, 앰버는 둔탁한 소리가 나는 화면 좌측의 계단 위쪽을 일제히 올려다보고 있다. 찰리(로봇)의 시선은 정면 우측을 향하고 있다.",
        "built_space": "지하실 기도실의 레퍼런스와 일치한다. 좌측 전경에 위로 향하는 계단이 있고, 배경 중앙 벽면에 십자가가 보이며, 촛불이 놓인 책상과 의자, 우측 뒤편에 작은 오르간이 정확한 스케일과 위치에 배치되어 있다.",
        "entities": "지시된 4명의 캐릭터만 존재한다. 신부(참조 이미지 일치, 로만 칼라 착용, 양초 들고 있음), 현우(참조 이미지 일치, 얼굴에 멍, 겉옷 없음), 앰버(참조 이미지 일치, 겉옷으로 입을 가림), 찰리(참조 이미지의 로봇 외형 일치, 단 코트와 모자는 생략됨).",
        "hard_violations": [],
        "physics": "모든 인물은 지하실 바닥에 체중을 싣고 서 있으며, 어깨를 움츠리고 굳어있는 자세가 물리적으로 자연스럽다. 신부의 손은 촛대를 안정적으로 쥐고 있다."
       },
       {
        "label": "B",
        "direction": "등장인물 대부분이 천장이나 위쪽을 올려다보고 있다.",
        "built_space": "지하실 기도실의 구조가 잘 반영되었다. 좌측 계단, 중앙의 책상과 의자, 배경의 십자가, 우측 뒤편의 오르간이 적절히 배치되어 있다.",
        "entities": "참조 이미지의 신부와 똑같이 생긴 남성, 촛불을 들고 로만 칼라를 입은 또 다른 남성, 현우, 앰버, 찰리(로봇), 그리고 우측 끝에 신원 미상의 여성까지 총 6명이 등장한다. 프롬프트가 지시한 4명 제한을 어겼다.",
        "hard_violations": [
         "프롬프트에 명시되지 않은 발명된 인물(신원 미상의 여성) 추가",
         "신부 캐릭터가 두 명의 별개 인물(참조 얼굴을 한 남성과 칼라를 입은 남성)로 복제 및 분열됨"
        ],
        "physics": "인물들은 바닥에 잘 서 있으며 촛불을 쥔 손이나 신체 자세에 물리적 오류는 없다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "지하 기도실의 와이드 구도는 맞지만 일행이 여섯으로 늘어나는 치명적 오류가 있으며, 앰버의 입가 가림과 찰리의 외투·모자도 빠졌다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "정확한 네 인물과 왼쪽 계단, 움츠린 인간 인물들의 상향 시선을 구현했으나 찰리의 상향 반응과 외투·모자가 누락됐다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "인간 다섯 명은 대체로 화면 왼쪽 위, 계단 너머의 상층 방향을 올려다본다. 찰리의 얼굴은 왼쪽을 향하지만 거의 수평이어서 같은 상층을 주시한다고 보기 어렵다. 배경 십자가를 바라보는 인물은 없다.",
        "built_space": "왼쪽 전경에 위로 이어져 화면 밖으로 나가는 계단 하나가 있다. 뒤쪽에는 책상 하나, 식별 가능한 의자 하나, 오른쪽 오르간 하나, 벽 십자가 하나, 왼쪽 책장 하나가 보인다. 벽등 두 개와 천장등 하나가 있으며, 낡은 벽과 바닥도 장소 참조와 가깝다. 일행은 가구 앞 바닥에 서 있다. 와이드 구도이지만 참조 사진의 정면 배치를 거의 그대로 따르며 책상과 오르간의 사선 노출은 약하다.",
        "entities": "총 여섯 인물이 보인다. 중앙의 검은 머리 청년과 금발 여자아이는 현우와 앰버에 대체로 대응하고, 베이지 장갑판과 흰 얼굴의 찰리도 있다. 그러나 왼쪽 노년 남성, 가운데 성직자 칼라를 한 촛불 소지자, 오른쪽 성인 남성이 함께 있어 허용된 네 명보다 두 명이 많다. 신부 참조와 닮은 왼쪽 남성과 성직자 역할이 분리되어 있다. 현우의 얼굴 멍과 다리 부상은 뚜렷하지 않고, 앰버의 입을 가리는 외투와 찰리의 낡은 외투·모자는 없다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "허용된 네 인물 대신 여섯 인물이 등장하여 지정되지 않은 인물 두 명이 추가됐다."
        ],
        "physics": "모든 인물은 발로 바닥을 딛고 있으며 무릎을 굽힌 자세도 지지 가능하다. 가운데 성직자가 촛불을 손으로 받쳐 들고 있다. 책상 위 초와 가구는 각각 상판과 바닥에 지지된다. 부유하거나 지지 없이 매달린 대상은 없다."
       },
       {
        "label": "B",
        "direction": "신부, 현우, 앰버는 몸을 움츠리고 왼쪽 위 계단의 상층 방향을 바라본다. 세 사람의 주시 대상은 배경 십자가가 아니다. 찰리는 왼쪽을 향하지만 얼굴이 거의 수평으로 서 있어, 다른 세 명처럼 위층을 올려다보는 반응은 분명하지 않다.",
        "built_space": "왼쪽 전경의 계단 하나가 왼쪽 위 화면 밖으로 이어진다. 일행 뒤에는 책상 하나와 의자 두 개가 있으며, 오른쪽 뒤에는 오르간 하나와 그 앞 작은 좌석이 부분적으로 보인다. 뒤 벽 십자가 하나, 왼쪽 책장 하나, 천장등 하나와 벽등이 원래 방의 배치를 유지한다. 네 인물은 계단 아래 기도실 바닥에 있고, 가구와 충돌하지 않는다. 와이드 구도와 전경 계단은 맞지만 배경 가구의 각도는 요청보다 정면에 가깝다.",
        "entities": "신부, 현우, 앰버, 찰리에 대응하는 네 인물만 있다. 신부는 한국계 중노년 남성으로 성직자 칼라를 착용하고 켜진 초를 들었다. 현우는 헝클어진 검은 머리의 젊은 남성이며 겉옷 없이 반소매 차림이고 얼굴에 상처 자국이 보이지만 다리 부상은 명확하지 않다. 앰버는 금발 여자아이로 외투와 천을 입가에 붙들고 있어 가림 상태를 구현한다. 찰리는 육중한 베이지 장갑판, 긴 팔, 흰 마스크형 얼굴이 참조와 잘 맞지만 낡은 외투와 모자는 없다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "신부와 현우는 양발로 바닥을 딛고, 앰버는 무릎을 굽혀 발에 체중을 싣는다. 찰리도 넓은 두 발로 바닥에 서 있다. 신부의 손이 초를 잡고 앰버의 손은 입가의 천을 붙든다. 책상 위 초와 오르간 등은 정상적으로 지지되어 있다. 찰리의 자세가 덜 움츠러들었을 뿐 물리적으로 불가능한 자세는 아니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "지하 기도실의 와이드 구도는 맞지만 일행이 여섯으로 늘어나는 치명적 오류가 있으며, 앰버의 입가 가림과 찰리의 외투·모자도 빠졌다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "정확한 네 인물과 왼쪽 계단, 움츠린 인간 인물들의 상향 시선을 구현했으나 찰리의 상향 반응과 외투·모자가 누락됐다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "인간 다섯 명은 대체로 화면 왼쪽 위, 계단 너머의 상층 방향을 올려다본다. 찰리의 얼굴은 왼쪽을 향하지만 거의 수평이어서 같은 상층을 주시한다고 보기 어렵다. 배경 십자가를 바라보는 인물은 없다.",
        "built_space": "왼쪽 전경에 위로 이어져 화면 밖으로 나가는 계단 하나가 있다. 뒤쪽에는 책상 하나, 식별 가능한 의자 하나, 오른쪽 오르간 하나, 벽 십자가 하나, 왼쪽 책장 하나가 보인다. 벽등 두 개와 천장등 하나가 있으며, 낡은 벽과 바닥도 장소 참조와 가깝다. 일행은 가구 앞 바닥에 서 있다. 와이드 구도이지만 참조 사진의 정면 배치를 거의 그대로 따르며 책상과 오르간의 사선 노출은 약하다.",
        "entities": "총 여섯 인물이 보인다. 중앙의 검은 머리 청년과 금발 여자아이는 현우와 앰버에 대체로 대응하고, 베이지 장갑판과 흰 얼굴의 찰리도 있다. 그러나 왼쪽 노년 남성, 가운데 성직자 칼라를 한 촛불 소지자, 오른쪽 성인 남성이 함께 있어 허용된 네 명보다 두 명이 많다. 신부 참조와 닮은 왼쪽 남성과 성직자 역할이 분리되어 있다. 현우의 얼굴 멍과 다리 부상은 뚜렷하지 않고, 앰버의 입을 가리는 외투와 찰리의 낡은 외투·모자는 없다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "허용된 네 인물 대신 여섯 인물이 등장하여 지정되지 않은 인물 두 명이 추가됐다."
        ],
        "physics": "모든 인물은 발로 바닥을 딛고 있으며 무릎을 굽힌 자세도 지지 가능하다. 가운데 성직자가 촛불을 손으로 받쳐 들고 있다. 책상 위 초와 가구는 각각 상판과 바닥에 지지된다. 부유하거나 지지 없이 매달린 대상은 없다."
       },
       {
        "label": "A",
        "direction": "신부, 현우, 앰버는 몸을 움츠리고 왼쪽 위 계단의 상층 방향을 바라본다. 세 사람의 주시 대상은 배경 십자가가 아니다. 찰리는 왼쪽을 향하지만 얼굴이 거의 수평으로 서 있어, 다른 세 명처럼 위층을 올려다보는 반응은 분명하지 않다.",
        "built_space": "왼쪽 전경의 계단 하나가 왼쪽 위 화면 밖으로 이어진다. 일행 뒤에는 책상 하나와 의자 두 개가 있으며, 오른쪽 뒤에는 오르간 하나와 그 앞 작은 좌석이 부분적으로 보인다. 뒤 벽 십자가 하나, 왼쪽 책장 하나, 천장등 하나와 벽등이 원래 방의 배치를 유지한다. 네 인물은 계단 아래 기도실 바닥에 있고, 가구와 충돌하지 않는다. 와이드 구도와 전경 계단은 맞지만 배경 가구의 각도는 요청보다 정면에 가깝다.",
        "entities": "신부, 현우, 앰버, 찰리에 대응하는 네 인물만 있다. 신부는 한국계 중노년 남성으로 성직자 칼라를 착용하고 켜진 초를 들었다. 현우는 헝클어진 검은 머리의 젊은 남성이며 겉옷 없이 반소매 차림이고 얼굴에 상처 자국이 보이지만 다리 부상은 명확하지 않다. 앰버는 금발 여자아이로 외투와 천을 입가에 붙들고 있어 가림 상태를 구현한다. 찰리는 육중한 베이지 장갑판, 긴 팔, 흰 마스크형 얼굴이 참조와 잘 맞지만 낡은 외투와 모자는 없다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "신부와 현우는 양발로 바닥을 딛고, 앰버는 무릎을 굽혀 발에 체중을 싣는다. 찰리도 넓은 두 발로 바닥에 서 있다. 신부의 손이 초를 잡고 앰버의 손은 입가의 천을 붙든다. 책상 위 초와 오르간 등은 정상적으로 지지되어 있다. 찰리의 자세가 덜 움츠러들었을 뿐 물리적으로 불가능한 자세는 아니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.286
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.036
   },
   "violations": {
    "B": [
     "[gemini-pro] 프롬프트에 명시되지 않은 발명된 인물(신원 미상의 여성) 추가",
     "[gemini-pro] 신부 캐릭터가 두 명의 별개 인물(참조 얼굴을 한 남성과 칼라를 입은 남성)로 복제 및 분열됨",
     "[gpt-high] 허용된 네 인물 대신 여섯 인물이 등장하여 지정되지 않은 인물 두 명이 추가됐다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 36
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "명시된 4명의 인물만 정확하게 배치하고 위층을 올려다보는 긴장된 순간과 캐릭터별 복장 및 상태 지침(앰버의 입 가림, 현우의 상처 및 겉옷 없음, 신부의 촛불)을 매우 충실히 구현했습니다."
   },
   {
    "label": "B",
    "score": 36,
    "verdict_ko": "지시된 4명의 인물 외에 프롬프트에 없는 인물들이 추가로 등장하고, 신부 캐릭터가 두 명으로 분열되어 나타나는 치명적인 오류가 발생했습니다.  ★위반: [gemini-pro] 프롬프트에 명시되지 않은 발명된 인물(신원 미상의 여성) 추가 / [gemini-pro] 신부 캐릭터가 두 명의 별개 인물(참조 얼굴을 한 남성과 칼라를 입은 남성)로 복제 및 분열됨 / [gpt-high] 허용된 네 인물 대신 여섯 인물이 등장하여 지정되지 않은 인물 두 명이 추가됐다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L180B02.png",
    "asset_id": "973d81a2-fcfc-4412-82af-c60138fde795",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 신부: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1402213>",
    "asset_id": "8696070d-ac09-4a5c-95f2-3daf1c015c32",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-4184-732c-bf1b-0d6d2ad0fc86",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S29sh6__bgfirst_bg.png",
   "bg_asset_id": "4df963a5-17a9-4cdd-a2c1-6834d045d891",
   "bg_record_key": "S29sh6::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S29sh6::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:34:22.493435+00:00",
  "fingerprint": "d933a90ed8ca412ef14b17eb2c4bac930ee9e6d052776d3bb3e8af4b0a63b82b",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S29sh6_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S29sh6_sel.png",
  "source_sha256": "c7d8887314f6d9384080015e98a875fa7d14deb572a6bc0afb0aff0d855b988c",
  "file": "S29sh6_cine.png",
  "staged_sha256": "2b1015029e675319d57b4c5f197d9107245089549265aab5756fa40aacd10db1",
  "latency_ms": 12668
 },
 "S29sh10::signage": {
  "fp": "48ac4b585de92106",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S29sh10::bgfirst_bg": {
  "input_fingerprint": "dc99fe6c7d163142",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 살짝 열린 문틈으로 신부와 낯선 남자가 마주 서 있는 모습을 훔쳐보는 현우의 시점 쇼트.\n\nLOCATION (lock): Outside the church's ground-floor entrance at night, seen through the door opening from inside.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Open viewing window in the entrance door in the middle-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Entrance door's small viewing window (Open sufficiently for 현우 to look outside) — The opening is viewed from inside, with 신부 and the visitor directly visible beyond its edges; used as Restricts the field of view and establishes the concealed first-person perspective.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Render the nighttime conversation with subdued ambient tonal separation through the opening, without carrying the basement candlelight outside or applying a flashback treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 살짝 열린 문틈으로 신부와 낯선 남자가 마주 서 있는 모습을 훔쳐보는 현우의 시점 쇼트.\n\nLOCATION (lock): Outside the church's ground-floor entrance at night, seen through the door opening from inside.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Open viewing window in the entrance door in the middle-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Entrance door's small viewing window (Open sufficiently for 현우 to look outside) — The opening is viewed from inside, with 신부 and the visitor directly visible beyond its edges; used as Restricts the field of view and establishes the concealed first-person perspective.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Render the nighttime conversation with subdued ambient tonal separation through the opening, without carrying the basement candlelight outside or applying a flashback treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S29sh10__bgfirst_bg.png",
  "asset_id": "bd7dfe1d-fc8f-4e44-ac36-c3cfb142c309",
  "input_asset_ids": [
   "552fe877-7268-4883-a1c6-0d965dfe77b6",
   "db270a75-23c6-4def-abb3-d877c7750bce"
  ]
 },
 "S29sh10": {
  "input_fingerprint": "acf5b45f9c209c2f",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 살짝 열린 문틈으로 신부와 낯선 남자가 마주 서 있는 모습을 훔쳐보는 현우의 시점 쇼트.\n\nLOCATION (lock): Outside the church's ground-floor entrance at night, seen through the door opening from inside. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Open viewing window in the entrance door in the middle-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Entrance door's small viewing window (Open sufficiently for 현우 to look outside) — The opening is viewed from inside, with 신부 and the visitor directly visible beyond its edges; used as Restricts the field of view and establishes the concealed first-person perspective.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Render the nighttime conversation with subdued ambient tonal separation through the opening, without carrying the basement candlelight outside or applying a flashback treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The church entrance has been opened, and its small viewing window is open. The basement furnishings remain unchanged, with Charlie still downstairs in the old coat and hat. 신부: He is outside the church entrance, wearing his clerical collar. 구도환: He stands outside the church entrance, engaged in conversation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름); 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 살짝 열린 문틈으로 신부와 낯선 남자가 마주 서 있는 모습을 훔쳐보는 현우의 시점 쇼트.\n\nLOCATION (lock): Outside the church's ground-floor entrance at night, seen through the door opening from inside. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Open viewing window in the entrance door in the middle-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Entrance door's small viewing window (Open sufficiently for 현우 to look outside) — The opening is viewed from inside, with 신부 and the visitor directly visible beyond its edges; used as Restricts the field of view and establishes the concealed first-person perspective.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Render the nighttime conversation with subdued ambient tonal separation through the opening, without carrying the basement candlelight outside or applying a flashback treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The church entrance has been opened, and its small viewing window is open. The basement furnishings remain unchanged, with Charlie still downstairs in the old coat and hat. 신부: He is outside the church entrance, wearing his clerical collar. 구도환: He stands outside the church entrance, engaged in conversation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름); 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 살짝 열린 문틈으로 신부와 낯선 남자가 마주 서 있는 모습을 훔쳐보는 현우의 시점 쇼트.\n\nLOCATION (lock): Outside the church's ground-floor entrance at night, seen through the door opening from inside. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Open viewing window in the entrance door in the middle-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Entrance door's small viewing window (Open sufficiently for 현우 to look outside) — The opening is viewed from inside, with 신부 and the visitor directly visible beyond its edges; used as Restricts the field of view and establishes the concealed first-person perspective.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Render the nighttime conversation with subdued ambient tonal separation through the opening, without carrying the basement candlelight outside or applying a flashback treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The church entrance has been opened, and its small viewing window is open. The basement furnishings remain unchanged, with Charlie still downstairs in the old coat and hat. 신부: He is outside the church entrance, wearing his clerical collar. 구도환: He stands outside the church entrance, engaged in conversation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름); 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S29sh10__bgfirst_bg.png",
     "asset_id": "bd7dfe1d-fc8f-4e44-ac36-c3cfb142c309",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S29sh10.png",
     "asset_id": "552fe877-7268-4883-a1c6-0d965dfe77b6",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 신부: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1402213>",
     "asset_id": "8696070d-ac09-4a5c-95f2-3daf1c015c32",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1311816>",
     "asset_id": "623f0421-67dc-4592-a009-148d8e957276",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L182B01.png",
     "asset_id": "db270a75-23c6-4def-abb3-d877c7750bce",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 신부: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1402213>",
     "asset_id": "8696070d-ac09-4a5c-95f2-3daf1c015c32",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1311816>",
     "asset_id": "623f0421-67dc-4592-a009-148d8e957276",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "신부와 구도환이 서로의 얼굴을 마주보고 시선을 교환하고 있음.",
    "built_space": "카메라가 금속 문 안쪽에 위치하며, 중앙에 열린 직사각형의 작은 관측용 창틀이 프레임의 전경을 둘러싸고 있음.",
    "entities": "신부(60대 남성, 로만 칼라 착용)와 구도환(50대 남성, 정장 착용) 모두 레퍼런스의 얼굴 및 체형과 정확히 일치함.",
    "hard_violations": [],
    "physics": "두 인물 모두 야외 바닥에 안정적인 자세로 서 있음."
   },
   {
    "label": "B",
    "direction": "두 인물이 서로를 마주보며 대화하는 시선을 유지함.",
    "built_space": "카메라가 살짝 열린 주 출입문 틈새에 위치하며, 프레임 우측 문 패널에 관측용 창문이 닫힌 채로 배치되어 있음.",
    "entities": "신부와 구도환 모두 주어진 캐릭터 레퍼런스와 외형 및 복장이 잘 일치함.",
    "hard_violations": [],
    "physics": "두 인물 모두 땅을 디디고 자연스럽게 서 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "문 중앙의 열린 관측용 창문을 통해 두 인물을 바라보는 지시된 프레이밍을 정확하게 구현했으며, 캐릭터의 외모와 조명 묘사도 우수합니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "레퍼런스와 인물은 잘 일치하나, 중앙의 관측용 창문을 통해 훔쳐보는 구도를 지시한 프레임 레이아웃을 따르지 않고 문틈 사이로 구도를 잡아 우선순위가 높은 지시사항을 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "신부와 구도환이 서로의 얼굴을 마주보고 시선을 교환하고 있음.",
        "built_space": "카메라가 금속 문 안쪽에 위치하며, 중앙에 열린 직사각형의 작은 관측용 창틀이 프레임의 전경을 둘러싸고 있음.",
        "entities": "신부(60대 남성, 로만 칼라 착용)와 구도환(50대 남성, 정장 착용) 모두 레퍼런스의 얼굴 및 체형과 정확히 일치함.",
        "hard_violations": [],
        "physics": "두 인물 모두 야외 바닥에 안정적인 자세로 서 있음."
       },
       {
        "label": "B",
        "direction": "두 인물이 서로를 마주보며 대화하는 시선을 유지함.",
        "built_space": "카메라가 살짝 열린 주 출입문 틈새에 위치하며, 프레임 우측 문 패널에 관측용 창문이 닫힌 채로 배치되어 있음.",
        "entities": "신부와 구도환 모두 주어진 캐릭터 레퍼런스와 외형 및 복장이 잘 일치함.",
        "hard_violations": [],
        "physics": "두 인물 모두 땅을 디디고 자연스럽게 서 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "문 중앙의 열린 관측용 창문을 통해 두 인물을 바라보는 지시된 프레이밍을 정확하게 구현했으며, 캐릭터의 외모와 조명 묘사도 우수합니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "레퍼런스와 인물은 잘 일치하나, 중앙의 관측용 창문을 통해 훔쳐보는 구도를 지시한 프레임 레이아웃을 따르지 않고 문틈 사이로 구도를 잡아 우선순위가 높은 지시사항을 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "신부와 구도환이 서로의 얼굴을 마주보고 시선을 교환하고 있음.",
        "built_space": "카메라가 금속 문 안쪽에 위치하며, 중앙에 열린 직사각형의 작은 관측용 창틀이 프레임의 전경을 둘러싸고 있음.",
        "entities": "신부(60대 남성, 로만 칼라 착용)와 구도환(50대 남성, 정장 착용) 모두 레퍼런스의 얼굴 및 체형과 정확히 일치함.",
        "hard_violations": [],
        "physics": "두 인물 모두 야외 바닥에 안정적인 자세로 서 있음."
       },
       {
        "label": "B",
        "direction": "두 인물이 서로를 마주보며 대화하는 시선을 유지함.",
        "built_space": "카메라가 살짝 열린 주 출입문 틈새에 위치하며, 프레임 우측 문 패널에 관측용 창문이 닫힌 채로 배치되어 있음.",
        "entities": "신부와 구도환 모두 주어진 캐릭터 레퍼런스와 외형 및 복장이 잘 일치함.",
        "hard_violations": [],
        "physics": "두 인물 모두 땅을 디디고 자연스럽게 서 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "야간 장소와 두 사람의 마주 보는 관계는 맞지만, 작은 확인창이 아닌 크게 열린 출입문 사이로 허벅지까지 보여 핵심인 중앙 확인창 너머의 은밀한 미디엄 시점을 놓쳤다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "전경 중앙의 열린 확인창으로 시야를 제한하고 그 너머 대화하는 두 사람을 미디엄 구도로 담아, 현우의 숨은 시점과 야간 장소를 가장 충실하게 구현했다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 신부는 오른쪽 구도환의 얼굴을 보고, 구도환은 왼쪽 신부의 얼굴을 바라본다. 서로 마주 서서 대화하는 방향은 정확하며 카메라를 응시하지 않는다. 무기나 방향을 판정할 휴대 물건은 없다.",
        "built_space": "왼쪽에 벗겨진 회벽과 문설주, 오른쪽에 낡은 금속 문짝 하나가 보인다. 문짝에는 덮개와 하단 걸쇠가 달린 확인창 하나가 있지만 화면 오른쪽으로 밀려 있다. 두 사람은 확인창 너머가 아니라 문짝과 왼쪽 문설주 사이의 넓은 출입구에 보인다. 외부 콘크리트 담, 철제 난간, 젖은 바닥, 가로등과 주거 건물은 장소 참조와 잘 연결된다. 내부에서 밖을 보는 방향은 맞지만 중앙 확인창을 통한 제한된 시야가 아니다.",
        "entities": "인물은 신부와 구도환으로 읽히는 남성 두 명뿐이다. 신부는 주름진 얼굴과 회색 섞인 머리의 고령 한국인 남성으로 보이며 참조 인상에 가깝고, 흰 성직자 칼라와 검은 성직복을 착용한다. 구도환은 검은 머리의 중년 한국인 남성으로 보이나 참조의 남색 재킷 대신 어두운 긴 외투를 입었다. 얼굴과 머리의 전반적인 인상은 참조에 가깝다. 현우나 지하의 찰리는 추가되지 않았고 읽을 수 있는 글자도 없다.",
        "hard_violations": [],
        "physics": "두 사람의 몸통과 팔은 자연스럽게 아래로 이어지고 손은 몸 옆에 내려와 있다. 발은 화면 밖이지만 출입구 바닥에 서 있는 자세로 읽히며 공중에 떠 있다는 징후는 없다. 문짝은 문틀에 연결된 구조이고 확인창 덮개와 걸쇠도 금속 문에 부착되어 있다. 다만 두 사람 모두 팔을 거의 곧게 내린 자세여서 대화 중의 자연스러운 움직임은 약하다."
       },
       {
        "label": "B",
        "direction": "왼쪽 신부가 고개를 조금 숙여 오른쪽 구도환의 얼굴을 보고, 구도환은 고개를 들어 신부의 눈 쪽을 바라본다. 서로에게 향한 시선과 몸의 방향이 대화 장면에 맞으며 관찰자인 카메라를 보지 않는다. 별도의 휴대 물건이나 지향성 소품은 없다.",
        "built_space": "화면 전경 중앙에 직사각형 확인창 하나가 있고, 위로 열린 금속 덮개 하나와 아래쪽 걸쇠 하나가 보인다. 확인창 테두리가 두 사람의 상반신 주변 시야를 직접 제한한다. 그 너머 출입구 양옆의 문틀과 낡은 회벽, 외부 담장과 철제 난간, 젖은 바닥, 가로등 및 건물 창들이 보인다. 확인창이 크게 보이는 것은 카메라 바로 앞의 전경 배치로 설명된다. 두 사람은 창 바깥쪽 출입구에 서 있으며, 참조의 금속·회벽 재질과 야간 외부 공간을 유지한다.",
        "entities": "신부와 구도환으로 읽히는 남성 두 명만 등장한다. 신부는 회색 섞인 머리와 이마·눈가 주름이 있는 고령 한국인 남성으로 보이고 성직자 칼라를 착용한다. 참조보다 얼굴이 다소 길고 마른 인상이지만 주요 연령·머리 특징은 맞는다. 구도환은 검은 머리의 중년 한국인 남성으로 보이며 남색 재킷과 흰 셔츠가 참조 복장에 가깝다. 두 인물 모두 구체적인 얼굴과 정상적인 눈을 가진 실사 인물로 표현되며, 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 사람은 상반신만 확인창 안에 보이고 하체와 발은 아래 테두리에 가려져 있다. 몸통의 직립과 서로를 향한 가벼운 고개 기울임은 바깥 바닥에 서서 대화하는 자세로 자연스럽고, 부유나 비정상적인 신체 연결은 보이지 않는다. 확인창 덮개는 상단 연결부를 축으로 열려 있고 하단 걸쇠와 테두리는 문짝에 고정되어 있다. 손이나 휴대 물건이 보이지 않는 것은 제한된 구도에 따른 가림이다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "야간 장소와 두 사람의 마주 보는 관계는 맞지만, 작은 확인창이 아닌 크게 열린 출입문 사이로 허벅지까지 보여 핵심인 중앙 확인창 너머의 은밀한 미디엄 시점을 놓쳤다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "전경 중앙의 열린 확인창으로 시야를 제한하고 그 너머 대화하는 두 사람을 미디엄 구도로 담아, 현우의 숨은 시점과 야간 장소를 가장 충실하게 구현했다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽 신부는 오른쪽 구도환의 얼굴을 보고, 구도환은 왼쪽 신부의 얼굴을 바라본다. 서로 마주 서서 대화하는 방향은 정확하며 카메라를 응시하지 않는다. 무기나 방향을 판정할 휴대 물건은 없다.",
        "built_space": "왼쪽에 벗겨진 회벽과 문설주, 오른쪽에 낡은 금속 문짝 하나가 보인다. 문짝에는 덮개와 하단 걸쇠가 달린 확인창 하나가 있지만 화면 오른쪽으로 밀려 있다. 두 사람은 확인창 너머가 아니라 문짝과 왼쪽 문설주 사이의 넓은 출입구에 보인다. 외부 콘크리트 담, 철제 난간, 젖은 바닥, 가로등과 주거 건물은 장소 참조와 잘 연결된다. 내부에서 밖을 보는 방향은 맞지만 중앙 확인창을 통한 제한된 시야가 아니다.",
        "entities": "인물은 신부와 구도환으로 읽히는 남성 두 명뿐이다. 신부는 주름진 얼굴과 회색 섞인 머리의 고령 한국인 남성으로 보이며 참조 인상에 가깝고, 흰 성직자 칼라와 검은 성직복을 착용한다. 구도환은 검은 머리의 중년 한국인 남성으로 보이나 참조의 남색 재킷 대신 어두운 긴 외투를 입었다. 얼굴과 머리의 전반적인 인상은 참조에 가깝다. 현우나 지하의 찰리는 추가되지 않았고 읽을 수 있는 글자도 없다.",
        "hard_violations": [],
        "physics": "두 사람의 몸통과 팔은 자연스럽게 아래로 이어지고 손은 몸 옆에 내려와 있다. 발은 화면 밖이지만 출입구 바닥에 서 있는 자세로 읽히며 공중에 떠 있다는 징후는 없다. 문짝은 문틀에 연결된 구조이고 확인창 덮개와 걸쇠도 금속 문에 부착되어 있다. 다만 두 사람 모두 팔을 거의 곧게 내린 자세여서 대화 중의 자연스러운 움직임은 약하다."
       },
       {
        "label": "A",
        "direction": "왼쪽 신부가 고개를 조금 숙여 오른쪽 구도환의 얼굴을 보고, 구도환은 고개를 들어 신부의 눈 쪽을 바라본다. 서로에게 향한 시선과 몸의 방향이 대화 장면에 맞으며 관찰자인 카메라를 보지 않는다. 별도의 휴대 물건이나 지향성 소품은 없다.",
        "built_space": "화면 전경 중앙에 직사각형 확인창 하나가 있고, 위로 열린 금속 덮개 하나와 아래쪽 걸쇠 하나가 보인다. 확인창 테두리가 두 사람의 상반신 주변 시야를 직접 제한한다. 그 너머 출입구 양옆의 문틀과 낡은 회벽, 외부 담장과 철제 난간, 젖은 바닥, 가로등 및 건물 창들이 보인다. 확인창이 크게 보이는 것은 카메라 바로 앞의 전경 배치로 설명된다. 두 사람은 창 바깥쪽 출입구에 서 있으며, 참조의 금속·회벽 재질과 야간 외부 공간을 유지한다.",
        "entities": "신부와 구도환으로 읽히는 남성 두 명만 등장한다. 신부는 회색 섞인 머리와 이마·눈가 주름이 있는 고령 한국인 남성으로 보이고 성직자 칼라를 착용한다. 참조보다 얼굴이 다소 길고 마른 인상이지만 주요 연령·머리 특징은 맞는다. 구도환은 검은 머리의 중년 한국인 남성으로 보이며 남색 재킷과 흰 셔츠가 참조 복장에 가깝다. 두 인물 모두 구체적인 얼굴과 정상적인 눈을 가진 실사 인물로 표현되며, 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 사람은 상반신만 확인창 안에 보이고 하체와 발은 아래 테두리에 가려져 있다. 몸통의 직립과 서로를 향한 가벼운 고개 기울임은 바깥 바닥에 서서 대화하는 자세로 자연스럽고, 부유나 비정상적인 신체 연결은 보이지 않는다. 확인창 덮개는 상단 연결부를 축으로 열려 있고 하단 걸쇠와 테두리는 문짝에 고정되어 있다. 손이나 휴대 물건이 보이지 않는 것은 제한된 구도에 따른 가림이다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.167
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.167
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1167
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "문 중앙의 열린 관측용 창문을 통해 두 인물을 바라보는 지시된 프레이밍을 정확하게 구현했으며, 캐릭터의 외모와 조명 묘사도 우수합니다."
   },
   {
    "label": "B",
    "score": 1167,
    "verdict_ko": "레퍼런스와 인물은 잘 일치하나, 중앙의 관측용 창문을 통해 훔쳐보는 구도를 지시한 프레임 레이아웃을 따르지 않고 문틈 사이로 구도를 잡아 우선순위가 높은 지시사항을 위반했습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L182B01.png",
    "asset_id": "db270a75-23c6-4def-abb3-d877c7750bce",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 신부: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1402213>",
    "asset_id": "8696070d-ac09-4a5c-95f2-3daf1c015c32",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1311816>",
    "asset_id": "623f0421-67dc-4592-a009-148d8e957276",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-452f-7b09-aa4b-157747ff73d1",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S29sh10__bgfirst_bg.png",
   "bg_asset_id": "bd7dfe1d-fc8f-4e44-ac36-c3cfb142c309",
   "bg_record_key": "S29sh10::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S29sh10::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:35:52.155602+00:00",
  "fingerprint": "ea0f0c9cf9899da4a80469a4475f2126494724210db738d6dcaa3701e0a2b4db",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S29sh10_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S29sh10_sel.png",
  "source_sha256": "e4c90d46cfd38b0a9b120c16867d01b4433935614be552b9af20b6fbaf267588",
  "file": "S29sh10_cine.png",
  "staged_sha256": "7eb6713cbefc77d52bec7c75c5e5c10a2770d024fc5a8b79ec1ba5a0e5a5556e",
  "latency_ms": 13815
 },
 "S29sh12::signage": {
  "fp": "4d57f2d70cc579d7",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S29sh12": {
  "input_fingerprint": "92a79f19580a8fb2",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 서치라이트 불빛 주변의 어두운 하늘을 벌떼처럼 빽빽하게 덮은 채 앞으로 향한 비행 자세로 허공에 뜬 소형 드론들의 실루엣.\n\nLOCATION (lock): In the night sky above the refugee settlement, where small drones fly around sweeping helicopter searchlights. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Night sky (Dark behind the airborne drones); used as Negative space between clusters preserves the silhouettes' readability; Small drone swarm (Densely clustered and flying forward along the searchlight) — Undersides and oblique side profiles are visible, with their forward axes directed toward screen right; used as Distributed silhouettes convey overwhelming numbers without enlarging any single drone.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The helicopter's searchlight separates the drone silhouettes from the dark night sky with controlled contrast, without adding a dreamlike treatment to the flashback.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Helicopter searchlights sweep the night sky, with several drones flying along the beams. The church basement still contains the cross, desk, chairs and organ, and Charlie remains there in his old coat and hat.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 서치라이트 불빛 주변의 어두운 하늘을 벌떼처럼 빽빽하게 덮은 채 앞으로 향한 비행 자세로 허공에 뜬 소형 드론들의 실루엣.\n\nLOCATION (lock): In the night sky above the refugee settlement, where small drones fly around sweeping helicopter searchlights. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Night sky (Dark behind the airborne drones); used as Negative space between clusters preserves the silhouettes' readability; Small drone swarm (Densely clustered and flying forward along the searchlight) — Undersides and oblique side profiles are visible, with their forward axes directed toward screen right; used as Distributed silhouettes convey overwhelming numbers without enlarging any single drone.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The helicopter's searchlight separates the drone silhouettes from the dark night sky with controlled contrast, without adding a dreamlike treatment to the flashback.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Helicopter searchlights sweep the night sky, with several drones flying along the beams. The church basement still contains the cross, desk, chairs and organ, and Charlie remains there in his old coat and hat.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 서치라이트 불빛 주변의 어두운 하늘을 벌떼처럼 빽빽하게 덮은 채 앞으로 향한 비행 자세로 허공에 뜬 소형 드론들의 실루엣.\n\nLOCATION (lock): In the night sky above the refugee settlement, where small drones fly around sweeping helicopter searchlights. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Night sky (Dark behind the airborne drones); used as Negative space between clusters preserves the silhouettes' readability; Small drone swarm (Densely clustered and flying forward along the searchlight) — Undersides and oblique side profiles are visible, with their forward axes directed toward screen right; used as Distributed silhouettes convey overwhelming numbers without enlarging any single drone.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The helicopter's searchlight separates the drone silhouettes from the dark night sky with controlled contrast, without adding a dreamlike treatment to the flashback.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Helicopter searchlights sweep the night sky, with several drones flying along the beams. The church basement still contains the cross, desk, chairs and organ, and Charlie remains there in his old coat and hat.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "헬리콥터 탐조등이 우측 하단을 향해 빛을 비추고 있으며, 드론들의 전면 축이 화면 우측을 향해 비행하고 있음.",
    "built_space": "화면 최하단에 난민촌의 구조물 지붕들이 배치되어 있음.",
    "entities": "1대의 헬리콥터, 1개의 탐조등 불빛, 벌떼처럼 빽빽하게 밀집된 다수의 소형 쿼드코프터 드론 실루엣이 보임.",
    "hard_violations": [],
    "physics": "헬리콥터와 드론들 모두 로터의 회전을 통한 양력으로 공중에 안정적으로 떠 있음."
   },
   {
    "label": "B",
    "direction": "화면 상단 양쪽에서 두 개의 탐조등이 중앙 하단을 향해 교차하며 빛을 비추고 있으며, 드론들은 대체로 우측이나 정면을 향함.",
    "built_space": "화면 하단에 넓은 면적의 난민촌 건물과 골목들이 위치해 있음.",
    "entities": "화면 상단 양 끝에 2대의 헬리콥터(일부), 2개의 탐조등 불빛, 다수의 소형 쿼드코프터 드론들이 존재함.",
    "hard_violations": [
     "[gemini-pro] 레퍼런스와 프롬프트(The helicopter's searchlight)에 명시된 단일 광원을 위반하고 헬리콥터와 조명을 하나 더 복제하여 대칭으로 만들어냄 (Duplicated or extra objects)."
    ],
    "physics": "헬리콥터와 드론들 모두 양력을 받아 허공에 떠 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "단일 헬리콥터 조명과 우측을 향해 빽빽하게 밀집된 드론 떼를 프롬프트와 레퍼런스에 맞게 정확히 구현했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "레퍼런스와 설정에 없는 두 번째 헬리콥터와 교차하는 탐조등 불빛을 임의로 복제하여 추가한 점이 치명적인 오류입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "헬리콥터 탐조등이 우측 하단을 향해 빛을 비추고 있으며, 드론들의 전면 축이 화면 우측을 향해 비행하고 있음.",
        "built_space": "화면 최하단에 난민촌의 구조물 지붕들이 배치되어 있음.",
        "entities": "1대의 헬리콥터, 1개의 탐조등 불빛, 벌떼처럼 빽빽하게 밀집된 다수의 소형 쿼드코프터 드론 실루엣이 보임.",
        "hard_violations": [],
        "physics": "헬리콥터와 드론들 모두 로터의 회전을 통한 양력으로 공중에 안정적으로 떠 있음."
       },
       {
        "label": "B",
        "direction": "화면 상단 양쪽에서 두 개의 탐조등이 중앙 하단을 향해 교차하며 빛을 비추고 있으며, 드론들은 대체로 우측이나 정면을 향함.",
        "built_space": "화면 하단에 넓은 면적의 난민촌 건물과 골목들이 위치해 있음.",
        "entities": "화면 상단 양 끝에 2대의 헬리콥터(일부), 2개의 탐조등 불빛, 다수의 소형 쿼드코프터 드론들이 존재함.",
        "hard_violations": [
         "레퍼런스와 프롬프트(The helicopter's searchlight)에 명시된 단일 광원을 위반하고 헬리콥터와 조명을 하나 더 복제하여 대칭으로 만들어냄 (Duplicated or extra objects)."
        ],
        "physics": "헬리콥터와 드론들 모두 양력을 받아 허공에 떠 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "단일 헬리콥터 조명과 우측을 향해 빽빽하게 밀집된 드론 떼를 프롬프트와 레퍼런스에 맞게 정확히 구현했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "레퍼런스와 설정에 없는 두 번째 헬리콥터와 교차하는 탐조등 불빛을 임의로 복제하여 추가한 점이 치명적인 오류입니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "헬리콥터 탐조등이 우측 하단을 향해 빛을 비추고 있으며, 드론들의 전면 축이 화면 우측을 향해 비행하고 있음.",
        "built_space": "화면 최하단에 난민촌의 구조물 지붕들이 배치되어 있음.",
        "entities": "1대의 헬리콥터, 1개의 탐조등 불빛, 벌떼처럼 빽빽하게 밀집된 다수의 소형 쿼드코프터 드론 실루엣이 보임.",
        "hard_violations": [],
        "physics": "헬리콥터와 드론들 모두 로터의 회전을 통한 양력으로 공중에 안정적으로 떠 있음."
       },
       {
        "label": "B",
        "direction": "화면 상단 양쪽에서 두 개의 탐조등이 중앙 하단을 향해 교차하며 빛을 비추고 있으며, 드론들은 대체로 우측이나 정면을 향함.",
        "built_space": "화면 하단에 넓은 면적의 난민촌 건물과 골목들이 위치해 있음.",
        "entities": "화면 상단 양 끝에 2대의 헬리콥터(일부), 2개의 탐조등 불빛, 다수의 소형 쿼드코프터 드론들이 존재함.",
        "hard_violations": [
         "레퍼런스와 프롬프트(The helicopter's searchlight)에 명시된 단일 광원을 위반하고 헬리콥터와 조명을 하나 더 복제하여 대칭으로 만들어냄 (Duplicated or extra objects)."
        ],
        "physics": "헬리콥터와 드론들 모두 양력을 받아 허공에 떠 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "하단의 정착촌이 화면 약 3분의 1을 차지해 하늘과 드론 군집 중심의 구도를 약화하며, 드론들도 화면 오른쪽으로 전진하기보다 카메라 쪽을 향해 떠 있는 모습이다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "어두운 하늘을 넓게 채운 군집과 대각선 서치라이트가 요구 구도 및 참조에 더 충실하지만, 드론들의 오른쪽 전진 방향은 불명확하고 우상단 한 기가 유독 크게 보인다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "좌상단 서치라이트는 오른쪽 아래로, 우상단 서치라이트는 왼쪽 아래로 향해 중앙 군집과 그 아래 공간을 비춘다. 드론들은 대체로 좌우로 펼쳐진 팔과 정면의 하부 장치를 보여 카메라 쪽을 향한 자세로 읽힌다. 일부 사선 기체는 있으나 군집 전체의 전방 축이 화면 오른쪽을 향하지 않으며, 두 광선에도 단일한 진행 방향이 없다.",
        "built_space": "화면 하단 약 3분의 1에 골함석 지붕, 방수포, 창문과 통로가 있는 임시 주거 건물 수십 채가 보인다. 상단 양쪽에는 잘린 헬리콥터 동체 일부와 서치라이트가 각각 하나씩 있다. 참조 사진에는 정착촌 건축물이 보이지 않아 이 건물들의 정확한 장소 일치 여부는 검증할 수 없다. 사람이나 반사는 없으며, 하늘 중심이어야 할 화면에서 건축물의 비중이 크다.",
        "entities": "소형 다중회전익 드론 수십 기, 헬리콥터 일부 두 대, 서치라이트 두 줄기, 어두운 밤하늘과 정착촌이 보인다. 드론의 몸체와 하부 장치는 실제 기계 형태로 표현됐고 빛 앞에서 실루엣이 드러난다. 사람, 얼굴, 읽을 수 있는 글자는 없다. 교회 지하실과 그 안의 인물·물건은 이 하늘 장면의 프레임 밖이므로 누락으로 보지 않는다.",
        "hard_violations": [],
        "physics": "드론들은 회전익과 일부 회전 흐림이 보여 공중 체류를 지탱할 추진 장치가 있다. 수평에 가까운 자세는 호버링으로 가능하지만 오른쪽으로 전진하는 기울기는 뚜렷하지 않다. 잘린 헬리콥터의 로터 전체는 확인하기 어렵지만 기체 일부만 보이는 구도 자체가 무지지 부유를 뜻하지는 않는다. 건물은 지면에 놓여 있고 광선은 상단의 실제 광원에서 퍼진다."
       },
       {
        "label": "B",
        "direction": "좌상단 헬리콥터의 서치라이트가 화면 오른쪽 아래로 뻗어 중앙의 드론 무리를 비춘다. 군집 분포는 광선의 대각선 흐름을 따르지만, 개별 기체 대부분은 하부 카메라와 좌우 대칭에 가까운 팔을 정면으로 보여 오른쪽을 향한 전방 축이 확인되지 않는다. 우상단의 큰 드론 역시 주로 카메라 쪽 아래를 드러낸다.",
        "built_space": "정착촌 지붕은 화면 맨 아래의 좁은 띠에만 있고, 나머지 대부분은 드론과 구름 낀 밤하늘이다. 좌상단에 일부 잘린 헬리콥터 한 대와 서치라이트 하나가 보인다. 참조의 헬리콥터·광선·하늘 관계를 유지하면서 하늘을 주된 공간으로 삼는다. 참조에 건축 세부가 없어 하단 지붕의 정확한 일치는 검증할 수 없으며, 사람이나 반사는 없다.",
        "entities": "소형 다중회전익 드론 수십 기가 화면 전반에 분산되고, 헬리콥터 한 대와 밝은 서치라이트, 어두운 구름, 하단의 임시 주거 지붕들이 보인다. 드론은 참조처럼 팔, 회전익, 하부 장치를 갖춘 기계로 읽힌다. 다만 우상단 한 기는 다른 기체보다 현저히 커서 개별 드론을 강조하지 말라는 요구에 덜 맞는다. 사람, 얼굴, 판독 가능한 글자는 없고 지하실의 인물과 소품은 적절히 화면 밖에 있다.",
        "hard_violations": [],
        "physics": "드론의 회전익과 회전 흐림, 헬리콥터의 주회전익이 보여 각 항공기의 부양 수단을 확인할 수 있다. 드론의 거의 수평인 자세는 물리적으로 가능한 비행 또는 호버링이지만 요구된 오른쪽 전진 동작을 명확히 보여주지는 않는다. 지붕은 건물 위에 놓여 있고 서치라이트는 헬리콥터 하부에서 시작해 대기 중으로 확산된다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "하단의 정착촌이 화면 약 3분의 1을 차지해 하늘과 드론 군집 중심의 구도를 약화하며, 드론들도 화면 오른쪽으로 전진하기보다 카메라 쪽을 향해 떠 있는 모습이다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "어두운 하늘을 넓게 채운 군집과 대각선 서치라이트가 요구 구도 및 참조에 더 충실하지만, 드론들의 오른쪽 전진 방향은 불명확하고 우상단 한 기가 유독 크게 보인다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "좌상단 서치라이트는 오른쪽 아래로, 우상단 서치라이트는 왼쪽 아래로 향해 중앙 군집과 그 아래 공간을 비춘다. 드론들은 대체로 좌우로 펼쳐진 팔과 정면의 하부 장치를 보여 카메라 쪽을 향한 자세로 읽힌다. 일부 사선 기체는 있으나 군집 전체의 전방 축이 화면 오른쪽을 향하지 않으며, 두 광선에도 단일한 진행 방향이 없다.",
        "built_space": "화면 하단 약 3분의 1에 골함석 지붕, 방수포, 창문과 통로가 있는 임시 주거 건물 수십 채가 보인다. 상단 양쪽에는 잘린 헬리콥터 동체 일부와 서치라이트가 각각 하나씩 있다. 참조 사진에는 정착촌 건축물이 보이지 않아 이 건물들의 정확한 장소 일치 여부는 검증할 수 없다. 사람이나 반사는 없으며, 하늘 중심이어야 할 화면에서 건축물의 비중이 크다.",
        "entities": "소형 다중회전익 드론 수십 기, 헬리콥터 일부 두 대, 서치라이트 두 줄기, 어두운 밤하늘과 정착촌이 보인다. 드론의 몸체와 하부 장치는 실제 기계 형태로 표현됐고 빛 앞에서 실루엣이 드러난다. 사람, 얼굴, 읽을 수 있는 글자는 없다. 교회 지하실과 그 안의 인물·물건은 이 하늘 장면의 프레임 밖이므로 누락으로 보지 않는다.",
        "hard_violations": [],
        "physics": "드론들은 회전익과 일부 회전 흐림이 보여 공중 체류를 지탱할 추진 장치가 있다. 수평에 가까운 자세는 호버링으로 가능하지만 오른쪽으로 전진하는 기울기는 뚜렷하지 않다. 잘린 헬리콥터의 로터 전체는 확인하기 어렵지만 기체 일부만 보이는 구도 자체가 무지지 부유를 뜻하지는 않는다. 건물은 지면에 놓여 있고 광선은 상단의 실제 광원에서 퍼진다."
       },
       {
        "label": "A",
        "direction": "좌상단 헬리콥터의 서치라이트가 화면 오른쪽 아래로 뻗어 중앙의 드론 무리를 비춘다. 군집 분포는 광선의 대각선 흐름을 따르지만, 개별 기체 대부분은 하부 카메라와 좌우 대칭에 가까운 팔을 정면으로 보여 오른쪽을 향한 전방 축이 확인되지 않는다. 우상단의 큰 드론 역시 주로 카메라 쪽 아래를 드러낸다.",
        "built_space": "정착촌 지붕은 화면 맨 아래의 좁은 띠에만 있고, 나머지 대부분은 드론과 구름 낀 밤하늘이다. 좌상단에 일부 잘린 헬리콥터 한 대와 서치라이트 하나가 보인다. 참조의 헬리콥터·광선·하늘 관계를 유지하면서 하늘을 주된 공간으로 삼는다. 참조에 건축 세부가 없어 하단 지붕의 정확한 일치는 검증할 수 없으며, 사람이나 반사는 없다.",
        "entities": "소형 다중회전익 드론 수십 기가 화면 전반에 분산되고, 헬리콥터 한 대와 밝은 서치라이트, 어두운 구름, 하단의 임시 주거 지붕들이 보인다. 드론은 참조처럼 팔, 회전익, 하부 장치를 갖춘 기계로 읽힌다. 다만 우상단 한 기는 다른 기체보다 현저히 커서 개별 드론을 강조하지 말라는 요구에 덜 맞는다. 사람, 얼굴, 판독 가능한 글자는 없고 지하실의 인물과 소품은 적절히 화면 밖에 있다.",
        "hard_violations": [],
        "physics": "드론의 회전익과 회전 흐림, 헬리콥터의 주회전익이 보여 각 항공기의 부양 수단을 확인할 수 있다. 드론의 거의 수평인 자세는 물리적으로 가능한 비행 또는 호버링이지만 요구된 오른쪽 전진 동작을 명확히 보여주지는 않는다. 지붕은 건물 위에 놓여 있고 서치라이트는 헬리콥터 하부에서 시작해 대기 중으로 확산된다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.143
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.893
   },
   "violations": {
    "B": [
     "[gemini-pro] 레퍼런스와 프롬프트(The helicopter's searchlight)에 명시된 단일 광원을 위반하고 헬리콥터와 조명을 하나 더 복제하여 대칭으로 만들어냄 (Duplicated or extra objects)."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 893
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "단일 헬리콥터 조명과 우측을 향해 빽빽하게 밀집된 드론 떼를 프롬프트와 레퍼런스에 맞게 정확히 구현했습니다."
   },
   {
    "label": "B",
    "score": 893,
    "verdict_ko": "레퍼런스와 설정에 없는 두 번째 헬리콥터와 교차하는 탐조등 불빛을 임의로 복제하여 추가한 점이 치명적인 오류입니다.  ★위반: [gemini-pro] 레퍼런스와 프롬프트(The helicopter's searchlight)에 명시된 단일 광원을 위반하고 헬리콥터와 조명을 하나 더 복제하여 대칭으로 만들어냄 (Duplicated or extra objects)."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L183B01.png",
    "asset_id": "b50a2636-aacd-45fa-8152-e123b37e1b11",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-48a2-7d1f-a283-b01b15207a1e",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S29sh12::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:36:48.950919+00:00",
  "fingerprint": "0df4cd58968d7973cc869fd23b287a49a48d67b9f32c63b3690620a212cc7ddc",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S29sh12_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S29sh12_sel.png",
  "source_sha256": "d3655d86403ed2185c74bf9024ed511d1ee019d3182c1d374a3690e5f07df0d1",
  "file": "S29sh12_cine.png",
  "staged_sha256": "f07f8b48e7650b56c17a35555d0d8e1807e6dfe6564b91434c33bc0ae11a072f",
  "latency_ms": 17487
 },
 "S30sh5::signage": {
  "fp": "02a1ddaf0c2a122a",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S30sh5": {
  "input_fingerprint": "1751bca49d6a24a6",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 찰리의 금속 가슴 장갑 틈새로 눈부신 푸른 불빛이 강하게 뿜어져 나오는 찰나의 클로즈업.\n\nLOCATION (lock): Beside the small organ and old radio inside the church basement prayer room, as lanterns brighten around the robot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Brilliant blue light bursts through the chest armor gaps, retaining readable metal edges around the brightest emission.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The small organ and the old broken radio remain in the basement, along with the cross, desk and chairs. Light is emerging from Charlie's chest; his old coat and hat, worn metal body and faded chest logo persist.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 찰리의 금속 가슴 장갑 틈새로 눈부신 푸른 불빛이 강하게 뿜어져 나오는 찰나의 클로즈업.\n\nLOCATION (lock): Beside the small organ and old radio inside the church basement prayer room, as lanterns brighten around the robot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Brilliant blue light bursts through the chest armor gaps, retaining readable metal edges around the brightest emission.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The small organ and the old broken radio remain in the basement, along with the cross, desk and chairs. Light is emerging from Charlie's chest; his old coat and hat, worn metal body and faded chest logo persist.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 찰리의 금속 가슴 장갑 틈새로 눈부신 푸른 불빛이 강하게 뿜어져 나오는 찰나의 클로즈업.\n\nLOCATION (lock): Beside the small organ and old radio inside the church basement prayer room, as lanterns brighten around the robot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Brilliant blue light bursts through the chest armor gaps, retaining readable metal edges around the brightest emission.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The small organ and the old broken radio remain in the basement, along with the cross, desk and chairs. Light is emerging from Charlie's chest; his old coat and hat, worn metal body and faded chest logo persist.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "로봇의 가슴 장갑 틈새에서 푸른 불빛이 밖을 향해 강하게 뿜어져 나옵니다.",
    "built_space": "양쪽 벽면에 랜턴이 있으며, 오른쪽 배경에 오르간이 실제 비율에 맞게 배치되어 있습니다.",
    "entities": "지시된 형태의 찰리(몸통 부분), 벽걸이 랜턴들, 오르간.",
    "hard_violations": [],
    "physics": "로봇은 화면 밖 바닥에 안정적으로 서 있는 상태입니다."
   },
   {
    "label": "B",
    "direction": "로봇의 가슴에서 푸른 불빛이 사방으로 뿜어져 나옵니다.",
    "built_space": "배경 벽에 십자가와 랜턴이 있고, 로봇 바로 앞에 책상이 위치합니다.",
    "entities": "지시된 형태의 찰리, 미니어처 크기로 축소된 오르간과 라디오, 십자가, 랜턴.",
    "hard_violations": [
     "[gemini-pro] 오르간과 라디오가 장난감 크기로 축소되어 물리적으로 불가능한 스케일/무대 연출을 보여줌 (비율 유지 지시 위반)",
     "[gpt-high] 기존 세로형 건반악기를 오른쪽 뒤에 남겨둔 채 왼쪽 아래에 별도의 작은 오르간을 추가하여, 장소에 고정된 건반악기를 중복시켰습니다."
    ],
    "physics": "로봇은 서 있고 축소된 소품들은 앞쪽 책상 위에 놓여 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "가슴 장갑 틈새로 뿜어져 나오는 푸른 불빛과 클로즈업 샷을 정확히 구현했으며, 배경 사물의 물리적 비율을 적절히 유지했습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "배경의 오르간과 라디오가 로봇 크기에 맞춰 미니어처처럼 축소되어, 실제 물리적 비율을 유지하라는 지시를 심각하게 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "로봇의 가슴 장갑 틈새에서 푸른 불빛이 밖을 향해 강하게 뿜어져 나옵니다.",
        "built_space": "양쪽 벽면에 랜턴이 있으며, 오른쪽 배경에 오르간이 실제 비율에 맞게 배치되어 있습니다.",
        "entities": "지시된 형태의 찰리(몸통 부분), 벽걸이 랜턴들, 오르간.",
        "hard_violations": [],
        "physics": "로봇은 화면 밖 바닥에 안정적으로 서 있는 상태입니다."
       },
       {
        "label": "B",
        "direction": "로봇의 가슴에서 푸른 불빛이 사방으로 뿜어져 나옵니다.",
        "built_space": "배경 벽에 십자가와 랜턴이 있고, 로봇 바로 앞에 책상이 위치합니다.",
        "entities": "지시된 형태의 찰리, 미니어처 크기로 축소된 오르간과 라디오, 십자가, 랜턴.",
        "hard_violations": [
         "오르간과 라디오가 장난감 크기로 축소되어 물리적으로 불가능한 스케일/무대 연출을 보여줌 (비율 유지 지시 위반)"
        ],
        "physics": "로봇은 서 있고 축소된 소품들은 앞쪽 책상 위에 놓여 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "가슴 장갑 틈새로 뿜어져 나오는 푸른 불빛과 클로즈업 샷을 정확히 구현했으며, 배경 사물의 물리적 비율을 적절히 유지했습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "배경의 오르간과 라디오가 로봇 크기에 맞춰 미니어처처럼 축소되어, 실제 물리적 비율을 유지하라는 지시를 심각하게 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "로봇의 가슴 장갑 틈새에서 푸른 불빛이 밖을 향해 강하게 뿜어져 나옵니다.",
        "built_space": "양쪽 벽면에 랜턴이 있으며, 오른쪽 배경에 오르간이 실제 비율에 맞게 배치되어 있습니다.",
        "entities": "지시된 형태의 찰리(몸통 부분), 벽걸이 랜턴들, 오르간.",
        "hard_violations": [],
        "physics": "로봇은 화면 밖 바닥에 안정적으로 서 있는 상태입니다."
       },
       {
        "label": "B",
        "direction": "로봇의 가슴에서 푸른 불빛이 사방으로 뿜어져 나옵니다.",
        "built_space": "배경 벽에 십자가와 랜턴이 있고, 로봇 바로 앞에 책상이 위치합니다.",
        "entities": "지시된 형태의 찰리, 미니어처 크기로 축소된 오르간과 라디오, 십자가, 랜턴.",
        "hard_violations": [
         "오르간과 라디오가 장난감 크기로 축소되어 물리적으로 불가능한 스케일/무대 연출을 보여줌 (비율 유지 지시 위반)"
        ],
        "physics": "로봇은 서 있고 축소된 소품들은 앞쪽 책상 위에 놓여 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "가슴 장갑 틈의 강렬한 청색 발광은 구현했지만, 기존 건반악기 외에 작은 오르간을 하나 더 만들어 장소의 고정 설비를 중복시켰습니다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "가슴 중심 클로즈업에서 장갑 틈으로 터지는 푸른빛과 금속 가장자리를 선명하게 보여주며, 찰리의 외형과 기존 건반악기가 있는 지하실을 유지했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 가슴은 거의 정면을 향하며, 푸른빛은 중앙 가슴판의 이음새와 아래쪽 장갑 틈에서 카메라 쪽과 좌우로 퍼집니다. 눈은 화면 밖이므로 시선 방향은 확인할 수 없습니다. 무기나 조준 대상은 없습니다.",
        "built_space": "양쪽 가장자리에 등불이 하나씩, 뒤쪽 왼편에 십자가 하나가 보입니다. 오른쪽 뒤에는 이전 장면의 세로형 건반악기가 있고, 왼쪽 아래에는 별도의 작은 오르간과 연주용 의자가 추가되어 건반악기가 두 대입니다. 오른쪽 아래 탁자 위에는 라디오 한 대가 있습니다. 기존 지하실의 낡은 벽과 따뜻한 조명은 이어지지만 악기 수가 맞지 않습니다.",
        "entities": "찰리 한 개체만 보이며 다른 사람은 없습니다. 샌드 베이지색 각진 흉부·어깨 장갑, 검은 관절과 배관, 마모 흔적은 참조와 대체로 일치합니다. 얼굴과 모자는 화면 밖입니다. 보이는 몸통에는 외투가 없으나 이전 장면에서도 외투는 보이지 않습니다. 낡은 라디오는 식별되지만 고장 여부는 확인할 수 없습니다. 식별 가능한 글자는 없습니다.",
        "hard_violations": [
         "기존 세로형 건반악기를 오른쪽 뒤에 남겨둔 채 왼쪽 아래에 별도의 작은 오르간을 추가하여, 장소에 고정된 건반악기를 중복시켰습니다."
        ],
        "physics": "가슴판과 어깨·팔 장갑은 내부 프레임 및 관절에 연결되어 있고, 팔은 어깨와 팔꿈치 관절에서 아래로 이어집니다. 하체와 발의 접지는 화면 밖이지만 공중에 떠 있다는 징후는 없습니다. 라디오는 탁자에, 악기와 의자는 바닥에 놓여 있습니다. 청색광은 장갑 틈에서 발생하며 주변 금속에도 반사됩니다."
       },
       {
        "label": "B",
        "direction": "찰리의 몸통을 비스듬한 전면에서 보며, 빛은 가슴 중앙과 흉부 장갑 하단의 틈에서 전방 및 화면 오른쪽으로 강하게 퍼집니다. 턱 일부만 보이고 눈은 잘려 있어 시선은 판단할 수 없습니다. 발광의 출발점은 요구된 가슴 장갑 틈입니다.",
        "built_space": "오른쪽 배경에 세로형 건반악기 한 대가 있고, 왼쪽 벽과 오른쪽 악기 위쪽에 등불이 각각 하나씩 보입니다. 낡은 회벽과 노출 배관이 이전 지하실의 재질을 이어갑니다. 악기는 찰리 뒤에 있으며 중복되지 않습니다. 십자가, 라디오, 책상과 의자는 이 좁은 구도에서 식별되지 않아 배치 변경 여부를 단정할 수 없습니다.",
        "entities": "찰리만 등장합니다. 흰 마스크의 턱 일부, 육중한 어깨와 팔, 샌드 베이지 장갑, 검은 배관과 관절, 긁힌 표면이 참조 외형에 부합합니다. 얼굴 전체와 다리는 잘려 있어 눈과 다리 비율은 평가할 수 없습니다. 몸통에 외투가 없는 모습은 이전 장면과 같습니다. 배경 건반악기는 이전 장면의 목제 악기와 유사하며, 읽을 수 있는 글자나 추가 인물은 없습니다.",
        "hard_violations": [],
        "physics": "장갑판은 몸통 프레임에, 팔은 어깨와 팔꿈치 기계 관절에 연결되어 있습니다. 발은 화면 밖이지만 몸통이 떠 있거나 부품이 지지 없이 분리된 모습은 없습니다. 왼쪽 등불은 벽 부착부에, 오른쪽 등불은 악기 윗면에 지지됩니다. 내부에서 새어 나오는 청색광이 금속 가장자리와 배관을 비추어 물리적 발광으로 읽힙니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "가슴 장갑 틈의 강렬한 청색 발광은 구현했지만, 기존 건반악기 외에 작은 오르간을 하나 더 만들어 장소의 고정 설비를 중복시켰습니다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "가슴 중심 클로즈업에서 장갑 틈으로 터지는 푸른빛과 금속 가장자리를 선명하게 보여주며, 찰리의 외형과 기존 건반악기가 있는 지하실을 유지했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 가슴은 거의 정면을 향하며, 푸른빛은 중앙 가슴판의 이음새와 아래쪽 장갑 틈에서 카메라 쪽과 좌우로 퍼집니다. 눈은 화면 밖이므로 시선 방향은 확인할 수 없습니다. 무기나 조준 대상은 없습니다.",
        "built_space": "양쪽 가장자리에 등불이 하나씩, 뒤쪽 왼편에 십자가 하나가 보입니다. 오른쪽 뒤에는 이전 장면의 세로형 건반악기가 있고, 왼쪽 아래에는 별도의 작은 오르간과 연주용 의자가 추가되어 건반악기가 두 대입니다. 오른쪽 아래 탁자 위에는 라디오 한 대가 있습니다. 기존 지하실의 낡은 벽과 따뜻한 조명은 이어지지만 악기 수가 맞지 않습니다.",
        "entities": "찰리 한 개체만 보이며 다른 사람은 없습니다. 샌드 베이지색 각진 흉부·어깨 장갑, 검은 관절과 배관, 마모 흔적은 참조와 대체로 일치합니다. 얼굴과 모자는 화면 밖입니다. 보이는 몸통에는 외투가 없으나 이전 장면에서도 외투는 보이지 않습니다. 낡은 라디오는 식별되지만 고장 여부는 확인할 수 없습니다. 식별 가능한 글자는 없습니다.",
        "hard_violations": [
         "기존 세로형 건반악기를 오른쪽 뒤에 남겨둔 채 왼쪽 아래에 별도의 작은 오르간을 추가하여, 장소에 고정된 건반악기를 중복시켰습니다."
        ],
        "physics": "가슴판과 어깨·팔 장갑은 내부 프레임 및 관절에 연결되어 있고, 팔은 어깨와 팔꿈치 관절에서 아래로 이어집니다. 하체와 발의 접지는 화면 밖이지만 공중에 떠 있다는 징후는 없습니다. 라디오는 탁자에, 악기와 의자는 바닥에 놓여 있습니다. 청색광은 장갑 틈에서 발생하며 주변 금속에도 반사됩니다."
       },
       {
        "label": "A",
        "direction": "찰리의 몸통을 비스듬한 전면에서 보며, 빛은 가슴 중앙과 흉부 장갑 하단의 틈에서 전방 및 화면 오른쪽으로 강하게 퍼집니다. 턱 일부만 보이고 눈은 잘려 있어 시선은 판단할 수 없습니다. 발광의 출발점은 요구된 가슴 장갑 틈입니다.",
        "built_space": "오른쪽 배경에 세로형 건반악기 한 대가 있고, 왼쪽 벽과 오른쪽 악기 위쪽에 등불이 각각 하나씩 보입니다. 낡은 회벽과 노출 배관이 이전 지하실의 재질을 이어갑니다. 악기는 찰리 뒤에 있으며 중복되지 않습니다. 십자가, 라디오, 책상과 의자는 이 좁은 구도에서 식별되지 않아 배치 변경 여부를 단정할 수 없습니다.",
        "entities": "찰리만 등장합니다. 흰 마스크의 턱 일부, 육중한 어깨와 팔, 샌드 베이지 장갑, 검은 배관과 관절, 긁힌 표면이 참조 외형에 부합합니다. 얼굴 전체와 다리는 잘려 있어 눈과 다리 비율은 평가할 수 없습니다. 몸통에 외투가 없는 모습은 이전 장면과 같습니다. 배경 건반악기는 이전 장면의 목제 악기와 유사하며, 읽을 수 있는 글자나 추가 인물은 없습니다.",
        "hard_violations": [],
        "physics": "장갑판은 몸통 프레임에, 팔은 어깨와 팔꿈치 기계 관절에 연결되어 있습니다. 발은 화면 밖이지만 몸통이 떠 있거나 부품이 지지 없이 분리된 모습은 없습니다. 왼쪽 등불은 벽 부착부에, 오른쪽 등불은 악기 윗면에 지지됩니다. 내부에서 새어 나오는 청색광이 금속 가장자리와 배관을 비추어 물리적 발광으로 읽힙니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.619
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.369
   },
   "violations": {
    "B": [
     "[gemini-pro] 오르간과 라디오가 장난감 크기로 축소되어 물리적으로 불가능한 스케일/무대 연출을 보여줌 (비율 유지 지시 위반)",
     "[gpt-high] 기존 세로형 건반악기를 오른쪽 뒤에 남겨둔 채 왼쪽 아래에 별도의 작은 오르간을 추가하여, 장소에 고정된 건반악기를 중복시켰습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 369
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "가슴 장갑 틈새로 뿜어져 나오는 푸른 불빛과 클로즈업 샷을 정확히 구현했으며, 배경 사물의 물리적 비율을 적절히 유지했습니다."
   },
   {
    "label": "B",
    "score": 369,
    "verdict_ko": "배경의 오르간과 라디오가 로봇 크기에 맞춰 미니어처처럼 축소되어, 실제 물리적 비율을 유지하라는 지시를 심각하게 위반했습니다.  ★위반: [gemini-pro] 오르간과 라디오가 장난감 크기로 축소되어 물리적으로 불가능한 스케일/무대 연출을 보여줌 (비율 유지 지시 위반) / [gpt-high] 기존 세로형 건반악기를 오른쪽 뒤에 남겨둔 채 왼쪽 아래에 별도의 작은 오르간을 추가하여, 장소에 고정된 건반악기를 중복시켰습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S29sh6_sel.png",
    "asset_id": "dc45cd18-8d8a-4dfc-97ac-b46ab9e2e5c8",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-4a5d-74e8-bbde-8ada72b4d04f",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S29sh6"
  }
 },
 "S30sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:20:47.480762+00:00",
  "fingerprint": "f3629fcb77829123107cf8e1d70e596e27595b57080fe1bdeba6dd37696c0fc4",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S30sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S30sh5_sel.png",
  "source_sha256": "e822b0de83822f34bf71e1466593f2182e12b0529d2cb62cea2d4d32794281d9",
  "file": "S30sh5_cine.png",
  "staged_sha256": "49b734495dfcc9208bb7a1ff91e91ca782540f8f22f449591f95e5cb6ed23946",
  "latency_ms": 10628
 },
 "S30sh10::signage": {
  "fp": "8c35ec84c871a895",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S30sh10": {
  "input_fingerprint": "c8c2427b73ef7f71",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 찰리의 거대한 금속 몸통을 향해 양팔을 뻗어 덥석 끌어안은 채 밀착된 구도환의 환한 상체.\n\nLOCATION (lock): Inside the sparse church basement prayer room, near the foot of the stairs and the organ, under brightened lantern light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The now-bright lantern illumination gives the embrace clear facial and bodily separation with restrained contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same basement prayer-room surfaces and its cross, small desk, chairs, and organ. Exclude the upstairs doorway and exterior surveillance equipment.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement lanterns are now brightly lit, and the previously broken radio is operating. Charlie retains his old coat and hat beside the small organ. 구도환: He has entered the basement and holds an excited, forward-leaning embrace posture.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 찰리의 거대한 금속 몸통을 향해 양팔을 뻗어 덥석 끌어안은 채 밀착된 구도환의 환한 상체.\n\nLOCATION (lock): Inside the sparse church basement prayer room, near the foot of the stairs and the organ, under brightened lantern light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The now-bright lantern illumination gives the embrace clear facial and bodily separation with restrained contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same basement prayer-room surfaces and its cross, small desk, chairs, and organ. Exclude the upstairs doorway and exterior surveillance equipment.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement lanterns are now brightly lit, and the previously broken radio is operating. Charlie retains his old coat and hat beside the small organ. 구도환: He has entered the basement and holds an excited, forward-leaning embrace posture.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 찰리의 거대한 금속 몸통을 향해 양팔을 뻗어 덥석 끌어안은 채 밀착된 구도환의 환한 상체.\n\nLOCATION (lock): Inside the sparse church basement prayer room, near the foot of the stairs and the organ, under brightened lantern light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The now-bright lantern illumination gives the embrace clear facial and bodily separation with restrained contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same basement prayer-room surfaces and its cross, small desk, chairs, and organ. Exclude the upstairs doorway and exterior surveillance equipment.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement lanterns are now brightly lit, and the previously broken radio is operating. Charlie retains his old coat and hat beside the small organ. 구도환: He has entered the basement and holds an excited, forward-leaning embrace posture.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "구도환의 시선과 양팔이 찰리를 향하고 있으며, 자연스럽게 몸통을 껴안고 있음.",
    "built_space": "교회 지하실 배경으로 계단, 랜턴이 놓인 책상, 십자가, 오르간이 적절히 배치되었으며, 오르간 주변에 코트와 모자가 있음.",
    "entities": "구도환과 찰리의 몸통 모두 레퍼런스와 일치함.",
    "hard_violations": [],
    "physics": "구도환이 찰리를 단단히 껴안고 체중을 지탱하는 자세가 물리적으로 타당함."
   },
   {
    "label": "B",
    "direction": "구도환이 찰리를 향해 미소 지으며 팔을 뻗고 있음.",
    "built_space": "계단, 책상, 십자가, 오르간 등의 지하실 배경 요소들이 존재함.",
    "entities": "구도환의 얼굴과 찰리의 외형은 레퍼런스와 일치함.",
    "hard_violations": [
     "[gemini-pro] 구도환의 등 뒤쪽으로 정체불명의 잉여 로봇 팔이 허공에 뻗어 있는 해부학적/물리적 불가능 오류"
    ],
    "physics": "구도환의 자세는 찰리에 닿아 있으나, 등 뒤에 떠 있는 여분의 로봇 팔은 어떠한 논리적 지지 기반도 없음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "구도환이 찰리의 금속 몸통을 밀착하여 껴안는 모습을 미디엄 샷으로 자연스럽게 구현했으며, 소품과 배경 설정도 프롬프트를 잘 따랐습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "구도환의 등 뒤에 정체불명의 로봇 팔이 추가로 생성되는 치명적인 해부학적 오류가 발생하여 실격 사유에 해당합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "구도환의 시선과 양팔이 찰리를 향하고 있으며, 자연스럽게 몸통을 껴안고 있음.",
        "built_space": "교회 지하실 배경으로 계단, 랜턴이 놓인 책상, 십자가, 오르간이 적절히 배치되었으며, 오르간 주변에 코트와 모자가 있음.",
        "entities": "구도환과 찰리의 몸통 모두 레퍼런스와 일치함.",
        "hard_violations": [],
        "physics": "구도환이 찰리를 단단히 껴안고 체중을 지탱하는 자세가 물리적으로 타당함."
       },
       {
        "label": "B",
        "direction": "구도환이 찰리를 향해 미소 지으며 팔을 뻗고 있음.",
        "built_space": "계단, 책상, 십자가, 오르간 등의 지하실 배경 요소들이 존재함.",
        "entities": "구도환의 얼굴과 찰리의 외형은 레퍼런스와 일치함.",
        "hard_violations": [
         "구도환의 등 뒤쪽으로 정체불명의 잉여 로봇 팔이 허공에 뻗어 있는 해부학적/물리적 불가능 오류"
        ],
        "physics": "구도환의 자세는 찰리에 닿아 있으나, 등 뒤에 떠 있는 여분의 로봇 팔은 어떠한 논리적 지지 기반도 없음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "구도환이 찰리의 금속 몸통을 밀착하여 껴안는 모습을 미디엄 샷으로 자연스럽게 구현했으며, 소품과 배경 설정도 프롬프트를 잘 따랐습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "구도환의 등 뒤에 정체불명의 로봇 팔이 추가로 생성되는 치명적인 해부학적 오류가 발생하여 실격 사유에 해당합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "구도환의 시선과 양팔이 찰리를 향하고 있으며, 자연스럽게 몸통을 껴안고 있음.",
        "built_space": "교회 지하실 배경으로 계단, 랜턴이 놓인 책상, 십자가, 오르간이 적절히 배치되었으며, 오르간 주변에 코트와 모자가 있음.",
        "entities": "구도환과 찰리의 몸통 모두 레퍼런스와 일치함.",
        "hard_violations": [],
        "physics": "구도환이 찰리를 단단히 껴안고 체중을 지탱하는 자세가 물리적으로 타당함."
       },
       {
        "label": "B",
        "direction": "구도환이 찰리를 향해 미소 지으며 팔을 뻗고 있음.",
        "built_space": "계단, 책상, 십자가, 오르간 등의 지하실 배경 요소들이 존재함.",
        "entities": "구도환의 얼굴과 찰리의 외형은 레퍼런스와 일치함.",
        "hard_violations": [
         "구도환의 등 뒤쪽으로 정체불명의 잉여 로봇 팔이 허공에 뻗어 있는 해부학적/물리적 불가능 오류"
        ],
        "physics": "구도환의 자세는 찰리에 닿아 있으나, 등 뒤에 떠 있는 여분의 로봇 팔은 어떠한 논리적 지지 기반도 없음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "환한 표정과 밀착은 맞지만, 허벅지와 넓은 배경까지 담아 상체 중심 미디엄 숏보다 넓고 양팔로 몸통을 끌어안는 동작도 덜 명확하다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "구도환의 환한 상체와 찰리의 거대한 몸통을 중심으로 양팔 포옹을 명확히 담아, 지정된 미디엄 숏과 밀착 순간을 더 충실히 구현한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "구도환은 찰리 쪽으로 상체를 기울이고 뺨을 가슴에 붙인다. 보이는 손은 찰리의 가슴과 어깨 경계에 닿지만 반대쪽 팔의 감싸는 경로는 가려져 있다. 웃는 얼굴은 카메라 쪽으로 돌아와 있고, 찰리의 얼굴은 구도환 쪽으로 약간 내려가 있다.",
        "built_space": "왼쪽에 목재 계단 한 구간, 뒤쪽에 십자가 하나와 작은 책상 하나, 의자 두 개가 보인다. 오른쪽에는 목재 건반악기 하나와 벤치 하나가 있으며 랜턴은 책상과 악기 위에 각각 하나씩 있다. 두 인물은 계단 아래와 악기 사이에 서 있다. 낡은 벽과 목재 악기는 장소에 부합하지만, 넓어진 화면에서 계단과 가구가 상체 포옹만큼 큰 비중을 차지한다.",
        "entities": "등장인물은 구도환과 찰리뿐이다. 구도환은 검은 머리의 한국인 중년 남성으로 보이고 참조 얼굴과 유사하지만, 재킷은 참조의 남색보다 회색에 가깝다. 찰리는 샌드 베이지 장갑판, 육중한 팔, 흰 마스크형 얼굴과 주황색 눈을 유지하며 이전 숏의 청색 몸통 발광도 이어진다. 외투·모자·라디오는 보이는 범위에서 확인되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "구도환의 몸은 화면 아래로 이어지는 다리로 지탱되는 서 있는 자세이며, 앞으로 기울인 가슴과 손이 찰리에 접촉한다. 발은 화면 밖이지만 공중에 뜬 모습은 아니다. 찰리의 들어 올린 팔은 어깨와 팔꿈치 관절로 연결되어 있고 다른 팔은 아래로 내려와 있다. 가구와 랜턴도 각각 바닥과 가구 상판의 지지를 받는다."
       },
       {
        "label": "B",
        "direction": "구도환은 찰리의 가슴에 몸을 밀착하고 양팔을 좌우로 벌려 감싼다. 가까운 손은 찰리의 위팔 장갑에 닿고 반대쪽 팔은 몸통 뒤쪽으로 돌아간다. 웃는 얼굴과 시선은 화면 왼쪽을 향한다. 찰리의 눈은 프레임 밖이라 시선을 판정할 수 없으며, 두 기계 손은 구도환의 등과 허리 쪽을 감싼다.",
        "built_space": "왼쪽 뒤로 올라가는 콘크리트 계단 한 구간, 오른쪽 벽의 십자가 하나, 오른쪽의 목재 건반악기 하나와 벤치 하나가 보인다. 왼쪽 작은 탁자 위에는 밝은 랜턴 하나가 있고 의자 일부가 아래에 보인다. 두 인물은 계단 발치와 악기 바로 옆에 밀착해 있다. 낡은 회벽과 목재 악기의 재질이 이전 숏에 잘 이어지며, 배경은 포옹을 방해하지 않는 크기로 남아 있다.",
        "entities": "구도환은 참조와 유사한 검은 머리와 중년 얼굴을 가진 한국인 남성으로 보이며 흰 셔츠와 짙은 재킷을 입었다. 재킷은 참조보다 검게 보인다. 찰리의 거대한 베이지 금속 몸통, 긴 팔, 마모된 장갑판과 청색 발광이 참조에 부합한다. 흰 얼굴은 하단만 보이며 짧은 다리 비율은 이 구도에서 판정할 수 없다. 악기 옆에는 낡은 외투와 모자가 있고, 작동 중인 라디오는 확인되지 않는다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "구도환의 양팔은 어깨에서 자연스럽게 이어져 찰리를 감싸며, 가슴은 금속 몸통에 접촉한다. 찰리의 두 팔도 관절을 굽혀 구도환의 등과 허리를 받치는 구조로 읽힌다. 두 몸의 하체는 화면 아래로 이어지고 부유를 나타내는 자세는 없다. 외투는 악기 가장자리에 걸쳐 있고 모자와 랜턴은 상판 위에 놓여 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "환한 표정과 밀착은 맞지만, 허벅지와 넓은 배경까지 담아 상체 중심 미디엄 숏보다 넓고 양팔로 몸통을 끌어안는 동작도 덜 명확하다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "구도환의 환한 상체와 찰리의 거대한 몸통을 중심으로 양팔 포옹을 명확히 담아, 지정된 미디엄 숏과 밀착 순간을 더 충실히 구현한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "구도환은 찰리 쪽으로 상체를 기울이고 뺨을 가슴에 붙인다. 보이는 손은 찰리의 가슴과 어깨 경계에 닿지만 반대쪽 팔의 감싸는 경로는 가려져 있다. 웃는 얼굴은 카메라 쪽으로 돌아와 있고, 찰리의 얼굴은 구도환 쪽으로 약간 내려가 있다.",
        "built_space": "왼쪽에 목재 계단 한 구간, 뒤쪽에 십자가 하나와 작은 책상 하나, 의자 두 개가 보인다. 오른쪽에는 목재 건반악기 하나와 벤치 하나가 있으며 랜턴은 책상과 악기 위에 각각 하나씩 있다. 두 인물은 계단 아래와 악기 사이에 서 있다. 낡은 벽과 목재 악기는 장소에 부합하지만, 넓어진 화면에서 계단과 가구가 상체 포옹만큼 큰 비중을 차지한다.",
        "entities": "등장인물은 구도환과 찰리뿐이다. 구도환은 검은 머리의 한국인 중년 남성으로 보이고 참조 얼굴과 유사하지만, 재킷은 참조의 남색보다 회색에 가깝다. 찰리는 샌드 베이지 장갑판, 육중한 팔, 흰 마스크형 얼굴과 주황색 눈을 유지하며 이전 숏의 청색 몸통 발광도 이어진다. 외투·모자·라디오는 보이는 범위에서 확인되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "구도환의 몸은 화면 아래로 이어지는 다리로 지탱되는 서 있는 자세이며, 앞으로 기울인 가슴과 손이 찰리에 접촉한다. 발은 화면 밖이지만 공중에 뜬 모습은 아니다. 찰리의 들어 올린 팔은 어깨와 팔꿈치 관절로 연결되어 있고 다른 팔은 아래로 내려와 있다. 가구와 랜턴도 각각 바닥과 가구 상판의 지지를 받는다."
       },
       {
        "label": "A",
        "direction": "구도환은 찰리의 가슴에 몸을 밀착하고 양팔을 좌우로 벌려 감싼다. 가까운 손은 찰리의 위팔 장갑에 닿고 반대쪽 팔은 몸통 뒤쪽으로 돌아간다. 웃는 얼굴과 시선은 화면 왼쪽을 향한다. 찰리의 눈은 프레임 밖이라 시선을 판정할 수 없으며, 두 기계 손은 구도환의 등과 허리 쪽을 감싼다.",
        "built_space": "왼쪽 뒤로 올라가는 콘크리트 계단 한 구간, 오른쪽 벽의 십자가 하나, 오른쪽의 목재 건반악기 하나와 벤치 하나가 보인다. 왼쪽 작은 탁자 위에는 밝은 랜턴 하나가 있고 의자 일부가 아래에 보인다. 두 인물은 계단 발치와 악기 바로 옆에 밀착해 있다. 낡은 회벽과 목재 악기의 재질이 이전 숏에 잘 이어지며, 배경은 포옹을 방해하지 않는 크기로 남아 있다.",
        "entities": "구도환은 참조와 유사한 검은 머리와 중년 얼굴을 가진 한국인 남성으로 보이며 흰 셔츠와 짙은 재킷을 입었다. 재킷은 참조보다 검게 보인다. 찰리의 거대한 베이지 금속 몸통, 긴 팔, 마모된 장갑판과 청색 발광이 참조에 부합한다. 흰 얼굴은 하단만 보이며 짧은 다리 비율은 이 구도에서 판정할 수 없다. 악기 옆에는 낡은 외투와 모자가 있고, 작동 중인 라디오는 확인되지 않는다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "구도환의 양팔은 어깨에서 자연스럽게 이어져 찰리를 감싸며, 가슴은 금속 몸통에 접촉한다. 찰리의 두 팔도 관절을 굽혀 구도환의 등과 허리를 받치는 구조로 읽힌다. 두 몸의 하체는 화면 아래로 이어지고 부유를 나타내는 자세는 없다. 외투는 악기 가장자리에 걸쳐 있고 모자와 랜턴은 상판 위에 놓여 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.206
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.956
   },
   "violations": {
    "B": [
     "[gemini-pro] 구도환의 등 뒤쪽으로 정체불명의 잉여 로봇 팔이 허공에 뻗어 있는 해부학적/물리적 불가능 오류"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 956
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "구도환이 찰리의 금속 몸통을 밀착하여 껴안는 모습을 미디엄 샷으로 자연스럽게 구현했으며, 소품과 배경 설정도 프롬프트를 잘 따랐습니다."
   },
   {
    "label": "B",
    "score": 956,
    "verdict_ko": "구도환의 등 뒤에 정체불명의 로봇 팔이 추가로 생성되는 치명적인 해부학적 오류가 발생하여 실격 사유에 해당합니다.  ★위반: [gemini-pro] 구도환의 등 뒤쪽으로 정체불명의 잉여 로봇 팔이 허공에 뻗어 있는 해부학적/물리적 불가능 오류"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S30sh5_sel.png",
    "asset_id": "1fe1c2fc-a535-4cf7-b8e4-c1a9f7557858",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1311816>",
    "asset_id": "623f0421-67dc-4592-a009-148d8e957276",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-4c1c-7c10-a6bc-d1555b5378d4",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S30sh5"
  }
 },
 "S30sh10::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:21:46.974413+00:00",
  "fingerprint": "9b721130f3ba48c36cfd2e13842d7ea89e11410742276860e33ab7fb0ebadebd",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S30sh10_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S30sh10_sel.png",
  "source_sha256": "7de28f99c3a975acf79889e66882653689b220fd02158e5665d50eb348ea2bbe",
  "file": "S30sh10_cine.png",
  "staged_sha256": "990f2cd4da0e5571a86e15773e25be259342b9f1164189d9825fa44fdfe9addb",
  "latency_ms": 9545
 },
 "S30sh17::signage": {
  "fp": "b36907ac3be0935c",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S30sh17": {
  "input_fingerprint": "c0bea28847cd3191",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 유빅사라는 말에 충격을 받은 듯 놀란 눈으로 구도환을 일제히 쳐다보는 앰버와 신부의 상체.\n\nLOCATION (lock): In the lantern-lit gathering area of the church basement prayer room, near the small table and organ. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established bright lantern illumination and controlled contrast, letting widened eyes carry the shock without a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the basement prayer room's fixed surfaces and sparse furnishings, including the small organ. Exclude the upstairs entrance and outdoor searchlights or drones.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement lanterns remain lit and the old radio remains functional, with the organ and other sparse furnishings unchanged. Charlie still wears the old coat and hat. 앰버: She remains in the basement, startled, with the outer garment used as a mouth covering. 신부: He stands in the basement wearing his clerical collar, visibly startled.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 유빅사라는 말에 충격을 받은 듯 놀란 눈으로 구도환을 일제히 쳐다보는 앰버와 신부의 상체.\n\nLOCATION (lock): In the lantern-lit gathering area of the church basement prayer room, near the small table and organ. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established bright lantern illumination and controlled contrast, letting widened eyes carry the shock without a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the basement prayer room's fixed surfaces and sparse furnishings, including the small organ. Exclude the upstairs entrance and outdoor searchlights or drones.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement lanterns remain lit and the old radio remains functional, with the organ and other sparse furnishings unchanged. Charlie still wears the old coat and hat. 앰버: She remains in the basement, startled, with the outer garment used as a mouth covering. 신부: He stands in the basement wearing his clerical collar, visibly startled.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 유빅사라는 말에 충격을 받은 듯 놀란 눈으로 구도환을 일제히 쳐다보는 앰버와 신부의 상체.\n\nLOCATION (lock): In the lantern-lit gathering area of the church basement prayer room, near the small table and organ. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established bright lantern illumination and controlled contrast, letting widened eyes carry the shock without a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the basement prayer room's fixed surfaces and sparse furnishings, including the small organ. Exclude the upstairs entrance and outdoor searchlights or drones.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement lanterns remain lit and the old radio remains functional, with the organ and other sparse furnishings unchanged. Charlie still wears the old coat and hat. 앰버: She remains in the basement, startled, with the outer garment used as a mouth covering. 신부: He stands in the basement wearing his clerical collar, visibly startled.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "B",
    "direction": "앰버와 신부 모두 눈을 크게 뜨고 화면 오른쪽 바깥을 바라본다. 앰버의 시선은 신부보다 조금 위로 향해 키 차이가 있는 두 사람이 같은 상대를 보는 상황과 양립한다. 구도환 자체는 보이지 않지만, 두 시선은 화면 밖 구도환을 향한다는 지시와 맞는다.",
    "built_space": "왼쪽 뒤에 콘크리트 계단 한 구간과 벽 배선, 왼쪽 아래에 작은 탁자 하나와 그 위 켜진 랜턴 하나가 보인다. 오른쪽에는 목재 오르간 하나와 벽 십자가 하나가 있으며, 오르간 위 모자와 걸쳐진 외투도 이전 공간의 배치를 따른다. 위쪽에는 매달린 랜턴 하나와 천장 조명이 추가로 보인다. 두 사람은 탁자와 오르간 사이 전경에 서 있고 상체 중심으로 잘린다. 위층 출입구나 외부 장비는 보이지 않는다.",
    "entities": "보이는 인물은 금발 여자아이 앰버와 나이 든 동아시아계 남성 신부 두 명뿐이다. 앰버의 둥근 얼굴, 큰 눈과 금발은 참조와 대체로 맞고 입은 짙은 겉옷으로 가려져 있다. 신부는 참조와 유사한 회색 섞인 머리와 이마·눈가 주름을 지녔으며 흰 성직 칼라를 착용했다. 작은 목재 오르간과 켜진 랜턴이 확인되지만 오래된 라디오의 정체와 작동 상태는 이 화면만으로 확정하기 어렵다. 읽을 수 있는 글자는 없다.",
    "hard_violations": [],
    "physics": "앰버가 손으로 겉옷을 움켜쥐어 입까지 올리고 있어 천의 지지가 명확하다. 두 사람의 상체는 자연스럽게 수직으로 이어지며 발이 화면 밖이라는 이유만으로 부유로 볼 근거는 없다. 탁상 랜턴은 탁자에 놓이고, 공중의 랜턴은 위로 이어지는 매달림 부재로 지지된다. 오르간과 그 위 물건도 접촉면이 자연스럽다."
   },
   {
    "label": "A",
    "direction": "앰버는 오른쪽 위 전경의 남성을 올려다보고, 신부도 고개와 눈을 같은 남성 쪽으로 돌린다. 시선의 실제 도착점은 명확하지만, 그 대상을 화면에 인물로 추가한 것은 허용된 출연자 범위를 벗어난다.",
    "built_space": "왼쪽에 콘크리트 계단 한 구간, 중앙 뒤에 작은 목재 오르간 하나와 그 앞 벤치 하나, 벽 십자가 하나, 오른쪽 위에 창 일부가 보인다. 두 주인공은 오르간 앞에 서 있다. 작은 탁자와 랜턴은 프레임 밖이라 점등 여부를 확인할 수 없다. 오른쪽 전경 남성의 등과 어깨가 화면을 크게 차지해 두 사람만의 상체 반응 숏보다 어깨너머 대화 숏으로 읽힌다.",
    "entities": "앰버와 신부 외에 검은 머리의 성인 남성 한 명이 오른쪽 전경에 추가되어 있다. 앰버의 금발, 어린 얼굴과 큰 눈, 신부의 나이 든 얼굴·회색 섞인 머리·성직 칼라는 대체로 지시에 맞는다. 앰버는 카키색 겉옷을 입 가까이 들었지만 입술이 드러나 입 가림이 불완전하다. 오르간은 확인되며 라디오와 랜턴은 보이지 않는다. 읽을 수 있는 글자는 없다.",
    "hard_violations": [
     "출연이 허용된 앰버와 신부 외에 세 번째 성인 남성의 머리와 상체를 오른쪽 전경에 추가했다."
    ],
    "physics": "앰버가 손으로 겉옷 가장자리를 잡아 들어 올리므로 천이 떠 있지 않다. 신부는 몸을 조금 앞으로 기울였지만 직립 상태에서 가능한 자세이며, 아래쪽 손에도 불가능한 접촉은 보이지 않는다. 전경 남성 역시 서 있는 상체로 읽힌다. 오르간과 벤치는 바닥에 놓인 가구이며 지지 없는 몸이나 물체는 보이지 않는다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": null,
     "normalized": null,
     "ok": false
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "앰버와 신부만의 상체 미디엄 숏에서 같은 화면 밖 대상을 향한 놀란 시선, 겉옷 입 가림, 성직 칼라와 지하실 공간을 충실히 구현했다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "놀란 시선의 대상은 명확하지만, 허용되지 않은 세 번째 남성을 전경에 넣어 두 사람의 상체 반응 숏을 어깨너머 구도로 바꾼 중대 위반이 있다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버와 신부 모두 눈을 크게 뜨고 화면 오른쪽 바깥을 바라본다. 앰버의 시선은 신부보다 조금 위로 향해 키 차이가 있는 두 사람이 같은 상대를 보는 상황과 양립한다. 구도환 자체는 보이지 않지만, 두 시선은 화면 밖 구도환을 향한다는 지시와 맞는다.",
        "built_space": "왼쪽 뒤에 콘크리트 계단 한 구간과 벽 배선, 왼쪽 아래에 작은 탁자 하나와 그 위 켜진 랜턴 하나가 보인다. 오른쪽에는 목재 오르간 하나와 벽 십자가 하나가 있으며, 오르간 위 모자와 걸쳐진 외투도 이전 공간의 배치를 따른다. 위쪽에는 매달린 랜턴 하나와 천장 조명이 추가로 보인다. 두 사람은 탁자와 오르간 사이 전경에 서 있고 상체 중심으로 잘린다. 위층 출입구나 외부 장비는 보이지 않는다.",
        "entities": "보이는 인물은 금발 여자아이 앰버와 나이 든 동아시아계 남성 신부 두 명뿐이다. 앰버의 둥근 얼굴, 큰 눈과 금발은 참조와 대체로 맞고 입은 짙은 겉옷으로 가려져 있다. 신부는 참조와 유사한 회색 섞인 머리와 이마·눈가 주름을 지녔으며 흰 성직 칼라를 착용했다. 작은 목재 오르간과 켜진 랜턴이 확인되지만 오래된 라디오의 정체와 작동 상태는 이 화면만으로 확정하기 어렵다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "앰버가 손으로 겉옷을 움켜쥐어 입까지 올리고 있어 천의 지지가 명확하다. 두 사람의 상체는 자연스럽게 수직으로 이어지며 발이 화면 밖이라는 이유만으로 부유로 볼 근거는 없다. 탁상 랜턴은 탁자에 놓이고, 공중의 랜턴은 위로 이어지는 매달림 부재로 지지된다. 오르간과 그 위 물건도 접촉면이 자연스럽다."
       },
       {
        "label": "B",
        "direction": "앰버는 오른쪽 위 전경의 남성을 올려다보고, 신부도 고개와 눈을 같은 남성 쪽으로 돌린다. 시선의 실제 도착점은 명확하지만, 그 대상을 화면에 인물로 추가한 것은 허용된 출연자 범위를 벗어난다.",
        "built_space": "왼쪽에 콘크리트 계단 한 구간, 중앙 뒤에 작은 목재 오르간 하나와 그 앞 벤치 하나, 벽 십자가 하나, 오른쪽 위에 창 일부가 보인다. 두 주인공은 오르간 앞에 서 있다. 작은 탁자와 랜턴은 프레임 밖이라 점등 여부를 확인할 수 없다. 오른쪽 전경 남성의 등과 어깨가 화면을 크게 차지해 두 사람만의 상체 반응 숏보다 어깨너머 대화 숏으로 읽힌다.",
        "entities": "앰버와 신부 외에 검은 머리의 성인 남성 한 명이 오른쪽 전경에 추가되어 있다. 앰버의 금발, 어린 얼굴과 큰 눈, 신부의 나이 든 얼굴·회색 섞인 머리·성직 칼라는 대체로 지시에 맞는다. 앰버는 카키색 겉옷을 입 가까이 들었지만 입술이 드러나 입 가림이 불완전하다. 오르간은 확인되며 라디오와 랜턴은 보이지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "출연이 허용된 앰버와 신부 외에 세 번째 성인 남성의 머리와 상체를 오른쪽 전경에 추가했다."
        ],
        "physics": "앰버가 손으로 겉옷 가장자리를 잡아 들어 올리므로 천이 떠 있지 않다. 신부는 몸을 조금 앞으로 기울였지만 직립 상태에서 가능한 자세이며, 아래쪽 손에도 불가능한 접촉은 보이지 않는다. 전경 남성 역시 서 있는 상체로 읽힌다. 오르간과 벤치는 바닥에 놓인 가구이며 지지 없는 몸이나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "앰버와 신부만의 상체 미디엄 숏에서 같은 화면 밖 대상을 향한 놀란 시선, 겉옷 입 가림, 성직 칼라와 지하실 공간을 충실히 구현했다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "놀란 시선의 대상은 명확하지만, 허용되지 않은 세 번째 남성을 전경에 넣어 두 사람의 상체 반응 숏을 어깨너머 구도로 바꾼 중대 위반이 있다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "앰버와 신부 모두 눈을 크게 뜨고 화면 오른쪽 바깥을 바라본다. 앰버의 시선은 신부보다 조금 위로 향해 키 차이가 있는 두 사람이 같은 상대를 보는 상황과 양립한다. 구도환 자체는 보이지 않지만, 두 시선은 화면 밖 구도환을 향한다는 지시와 맞는다.",
        "built_space": "왼쪽 뒤에 콘크리트 계단 한 구간과 벽 배선, 왼쪽 아래에 작은 탁자 하나와 그 위 켜진 랜턴 하나가 보인다. 오른쪽에는 목재 오르간 하나와 벽 십자가 하나가 있으며, 오르간 위 모자와 걸쳐진 외투도 이전 공간의 배치를 따른다. 위쪽에는 매달린 랜턴 하나와 천장 조명이 추가로 보인다. 두 사람은 탁자와 오르간 사이 전경에 서 있고 상체 중심으로 잘린다. 위층 출입구나 외부 장비는 보이지 않는다.",
        "entities": "보이는 인물은 금발 여자아이 앰버와 나이 든 동아시아계 남성 신부 두 명뿐이다. 앰버의 둥근 얼굴, 큰 눈과 금발은 참조와 대체로 맞고 입은 짙은 겉옷으로 가려져 있다. 신부는 참조와 유사한 회색 섞인 머리와 이마·눈가 주름을 지녔으며 흰 성직 칼라를 착용했다. 작은 목재 오르간과 켜진 랜턴이 확인되지만 오래된 라디오의 정체와 작동 상태는 이 화면만으로 확정하기 어렵다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "앰버가 손으로 겉옷을 움켜쥐어 입까지 올리고 있어 천의 지지가 명확하다. 두 사람의 상체는 자연스럽게 수직으로 이어지며 발이 화면 밖이라는 이유만으로 부유로 볼 근거는 없다. 탁상 랜턴은 탁자에 놓이고, 공중의 랜턴은 위로 이어지는 매달림 부재로 지지된다. 오르간과 그 위 물건도 접촉면이 자연스럽다."
       },
       {
        "label": "A",
        "direction": "앰버는 오른쪽 위 전경의 남성을 올려다보고, 신부도 고개와 눈을 같은 남성 쪽으로 돌린다. 시선의 실제 도착점은 명확하지만, 그 대상을 화면에 인물로 추가한 것은 허용된 출연자 범위를 벗어난다.",
        "built_space": "왼쪽에 콘크리트 계단 한 구간, 중앙 뒤에 작은 목재 오르간 하나와 그 앞 벤치 하나, 벽 십자가 하나, 오른쪽 위에 창 일부가 보인다. 두 주인공은 오르간 앞에 서 있다. 작은 탁자와 랜턴은 프레임 밖이라 점등 여부를 확인할 수 없다. 오른쪽 전경 남성의 등과 어깨가 화면을 크게 차지해 두 사람만의 상체 반응 숏보다 어깨너머 대화 숏으로 읽힌다.",
        "entities": "앰버와 신부 외에 검은 머리의 성인 남성 한 명이 오른쪽 전경에 추가되어 있다. 앰버의 금발, 어린 얼굴과 큰 눈, 신부의 나이 든 얼굴·회색 섞인 머리·성직 칼라는 대체로 지시에 맞는다. 앰버는 카키색 겉옷을 입 가까이 들었지만 입술이 드러나 입 가림이 불완전하다. 오르간은 확인되며 라디오와 랜턴은 보이지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "출연이 허용된 앰버와 신부 외에 세 번째 성인 남성의 머리와 상체를 오른쪽 전경에 추가했다."
        ],
        "physics": "앰버가 손으로 겉옷 가장자리를 잡아 들어 올리므로 천이 떠 있지 않다. 신부는 몸을 조금 앞으로 기울였지만 직립 상태에서 가능한 자세이며, 아래쪽 손에도 불가능한 접촉은 보이지 않는다. 전경 남성 역시 서 있는 상체로 읽힌다. 오르간과 벤치는 바닥에 놓인 가구이며 지지 없는 몸이나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gemini-pro"
   ],
   "route": "single_reverse"
  },
  "totals": {
   "B": 9,
   "A": 3
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 9,
    "verdict_ko": "앰버와 신부만의 상체 미디엄 숏에서 같은 화면 밖 대상을 향한 놀란 시선, 겉옷 입 가림, 성직 칼라와 지하실 공간을 충실히 구현했다."
   },
   {
    "label": "A",
    "score": 3,
    "verdict_ko": "놀란 시선의 대상은 명확하지만, 허용되지 않은 세 번째 남성을 전경에 넣어 두 사람의 상체 반응 숏을 어깨너머 구도로 바꾼 중대 위반이 있다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S30sh10_sel.png",
    "asset_id": "2857dac8-c656-447e-97ca-1e3c0e786e64",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 신부: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1402213>",
    "asset_id": "8696070d-ac09-4a5c-95f2-3daf1c015c32",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-4de7-7e30-ac3e-515c3dae465f",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S30sh10"
  }
 },
 "S30sh17::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:22:44.425890+00:00",
  "fingerprint": "ab9b60b83531450a6ee1e82ecdd1daeecc96fb510361bff0634d6aef7b6ce241",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S30sh17_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S30sh17_sel.png",
  "source_sha256": "1f8891e9e0196213ef214412f836ca25e9c75bc3fe07b6285ff1736ee452c418",
  "file": "S30sh17_cine.png",
  "staged_sha256": "6f7c8c24f1328721e9f5c0353d0eccaeaf844439202516e548c3d9489278c3d8",
  "latency_ms": 9946
 },
 "S31sh2::signage": {
  "fp": "72f50f3f97715f02",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::c1fe96fc846da4bd": {
  "subjects": [],
  "subject_text": "인천 난민촌 관리사무소 3층 감금실\n문고리가 달린 잠금식 출입문과 의자가 있는 폐쇄된 방. 출입문은 바깥 복도와 연결돼 있다.",
  "identity": "canonical",
  "scope_id": "L189",
  "scope_role": "location_interior",
  "scope_sha": "e56880f19b804921"
 },
 "S31sh2::bgfirst_bg": {
  "input_fingerprint": "22441b1de99bf4a4",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 퉁퉁 부은 얼굴의 미연 앞에 바짝 의자를 당겨 앉은 박철진의 거만한 상체.\n\nLOCATION (lock): Inside a third-floor detention room in the brightly lit refugee administration building, at the facing interrogation chairs.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 박철진's chair (Occupied and pulled close to 미연) — A partial side edge is visible beneath his seated torso; used as The cropped chair establishes his seated position and intrusive proximity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The brightly lit office exposes 미연's facial swelling and 박철진's expression with restrained, unsentimental contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 퉁퉁 부은 얼굴의 미연 앞에 바짝 의자를 당겨 앉은 박철진의 거만한 상체.\n\nLOCATION (lock): Inside a third-floor detention room in the brightly lit refugee administration building, at the facing interrogation chairs.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 박철진's chair (Occupied and pulled close to 미연) — A partial side edge is visible beneath his seated torso; used as The cropped chair establishes his seated position and intrusive proximity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The brightly lit office exposes 미연's facial swelling and 박철진's expression with restrained, unsentimental contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S31sh2__bgfirst_bg.png",
  "asset_id": "6d0e961f-cfc1-4ccb-921e-40cb43da0593",
  "input_asset_ids": [
   "c9bf8684-260f-4c7f-bb67-b37f8e71578e",
   "d31b4d4c-5488-4514-8d64-f7f3972f4a12"
  ]
 },
 "S31sh2": {
  "input_fingerprint": "fc740966efade74c",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 퉁퉁 부은 얼굴의 미연 앞에 바짝 의자를 당겨 앉은 박철진의 거만한 상체.\n\nLOCATION (lock): Inside a third-floor detention room in the brightly lit refugee administration building, at the facing interrogation chairs. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 박철진's chair (Occupied and pulled close to 미연) — A partial side edge is visible beneath his seated torso; used as The cropped chair establishes his seated position and intrusive proximity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The brightly lit office exposes 미연's facial swelling and 박철진's expression with restrained, unsentimental contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The management office is brightly lit, with the interrogation taking place on the third floor. 미연: Her face is badly swollen from the beating, and she remains in the interrogation room. 박철진: He is seated in the interrogation room after dismissing the subordinate from the beating.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리); 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 퉁퉁 부은 얼굴의 미연 앞에 바짝 의자를 당겨 앉은 박철진의 거만한 상체.\n\nLOCATION (lock): Inside a third-floor detention room in the brightly lit refugee administration building, at the facing interrogation chairs. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 박철진's chair (Occupied and pulled close to 미연) — A partial side edge is visible beneath his seated torso; used as The cropped chair establishes his seated position and intrusive proximity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The brightly lit office exposes 미연's facial swelling and 박철진's expression with restrained, unsentimental contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The management office is brightly lit, with the interrogation taking place on the third floor. 미연: Her face is badly swollen from the beating, and she remains in the interrogation room. 박철진: He is seated in the interrogation room after dismissing the subordinate from the beating.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리); 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 퉁퉁 부은 얼굴의 미연 앞에 바짝 의자를 당겨 앉은 박철진의 거만한 상체.\n\nLOCATION (lock): Inside a third-floor detention room in the brightly lit refugee administration building, at the facing interrogation chairs. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 박철진's chair (Occupied and pulled close to 미연) — A partial side edge is visible beneath his seated torso; used as The cropped chair establishes his seated position and intrusive proximity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The brightly lit office exposes 미연's facial swelling and 박철진's expression with restrained, unsentimental contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The management office is brightly lit, with the interrogation taking place on the third floor. 미연: Her face is badly swollen from the beating, and she remains in the interrogation room. 박철진: He is seated in the interrogation room after dismissing the subordinate from the beating.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리); 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S31sh2__bgfirst_bg.png",
     "asset_id": "6d0e961f-cfc1-4ccb-921e-40cb43da0593",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S31sh2.png",
     "asset_id": "c9bf8684-260f-4c7f-bb67-b37f8e71578e",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1113064>",
     "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1401722>",
     "asset_id": "fee7383c-fb61-4b3a-ba7c-79f2555de00b",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L189B01.png",
     "asset_id": "d31b4d4c-5488-4514-8d64-f7f3972f4a12",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1113064>",
     "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1401722>",
     "asset_id": "fee7383c-fb61-4b3a-ba7c-79f2555de00b",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "박철진의 시선은 미연을 향함.",
    "built_space": "참조된 방 구조가 구현됨. 단, 지정된 2개의 의자 외에 박철진 우측에 빈 의자가 추가되어 총 3개의 의자가 보임.",
    "entities": "박철진과 미연의 인상착의가 참조와 일치함. 미연의 측면 얼굴 부기가 확인됨.",
    "hard_violations": [
     "[gemini-pro] 공간 내 의자 개수 초과 (2개여야 할 의자가 3개로 복제됨)",
     "[gpt-high] 두 개로 정해진 심문 의자 외에 화면 오른쪽에 세 번째 빈 의자를 추가했다."
    ],
    "physics": "인물들은 의자에 앉아 체중을 자연스럽게 지탱하고 있음."
   },
   {
    "label": "B",
    "direction": "박철진은 미연을 응시하고, 미연은 아래를 향함.",
    "built_space": "방과 문이 구현됨. 인물들이 앉은 의자 외에 화면 전경에 추가 의자 등받이가 크게 배치되어 총 3개의 의자가 존재함.",
    "entities": "미연의 심하게 부은 얼굴이 명확히 표현됨. 박철진의 측면 일치함.",
    "hard_violations": [
     "[gemini-pro] 공간 내 의자 개수 초과 (2개여야 할 의자가 3개로 복제됨)"
    ],
    "physics": "인물들은 앉은 자세를 안정적으로 유지함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "박철진의 거만한 자세를 잘 포착했으나, 원본 공간에 없는 세 번째 의자가 우측에 생성되어 실격입니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "미연의 부은 얼굴 묘사는 훌륭하지만, 전경에 세 번째 의자가 잘못 추가되어 구도를 해치므로 실격입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진의 시선은 미연을 향함.",
        "built_space": "참조된 방 구조가 구현됨. 단, 지정된 2개의 의자 외에 박철진 우측에 빈 의자가 추가되어 총 3개의 의자가 보임.",
        "entities": "박철진과 미연의 인상착의가 참조와 일치함. 미연의 측면 얼굴 부기가 확인됨.",
        "hard_violations": [
         "공간 내 의자 개수 초과 (2개여야 할 의자가 3개로 복제됨)"
        ],
        "physics": "인물들은 의자에 앉아 체중을 자연스럽게 지탱하고 있음."
       },
       {
        "label": "B",
        "direction": "박철진은 미연을 응시하고, 미연은 아래를 향함.",
        "built_space": "방과 문이 구현됨. 인물들이 앉은 의자 외에 화면 전경에 추가 의자 등받이가 크게 배치되어 총 3개의 의자가 존재함.",
        "entities": "미연의 심하게 부은 얼굴이 명확히 표현됨. 박철진의 측면 일치함.",
        "hard_violations": [
         "공간 내 의자 개수 초과 (2개여야 할 의자가 3개로 복제됨)"
        ],
        "physics": "인물들은 앉은 자세를 안정적으로 유지함."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "박철진의 거만한 자세를 잘 포착했으나, 원본 공간에 없는 세 번째 의자가 우측에 생성되어 실격입니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "미연의 부은 얼굴 묘사는 훌륭하지만, 전경에 세 번째 의자가 잘못 추가되어 구도를 해치므로 실격입니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "박철진의 시선은 미연을 향함.",
        "built_space": "참조된 방 구조가 구현됨. 단, 지정된 2개의 의자 외에 박철진 우측에 빈 의자가 추가되어 총 3개의 의자가 보임.",
        "entities": "박철진과 미연의 인상착의가 참조와 일치함. 미연의 측면 얼굴 부기가 확인됨.",
        "hard_violations": [
         "공간 내 의자 개수 초과 (2개여야 할 의자가 3개로 복제됨)"
        ],
        "physics": "인물들은 의자에 앉아 체중을 자연스럽게 지탱하고 있음."
       },
       {
        "label": "B",
        "direction": "박철진은 미연을 응시하고, 미연은 아래를 향함.",
        "built_space": "방과 문이 구현됨. 인물들이 앉은 의자 외에 화면 전경에 추가 의자 등받이가 크게 배치되어 총 3개의 의자가 존재함.",
        "entities": "미연의 심하게 부은 얼굴이 명확히 표현됨. 박철진의 측면 일치함.",
        "hard_violations": [
         "공간 내 의자 개수 초과 (2개여야 할 의자가 3개로 복제됨)"
        ],
        "physics": "인물들은 앉은 자세를 안정적으로 유지함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "의상 불일치와 전경을 가리는 의자 등받이가 문제지만, 부은 미연 앞에 접근한 박철진을 보여주며 B의 명백한 의자 추가 오류는 없다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "박철진의 상체 중심 구도와 정장은 더 충실하지만, 두 심문 의자 외에 빈 의자를 하나 추가해 고정된 공간 구성을 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 박철진은 몸과 얼굴을 오른쪽 미연에게 기울이고 미연의 얼굴을 바라본다. 미연은 고개를 숙여 두 사람 사이 아래쪽을 본다. 상대를 압박하는 시선 관계는 성립하며, 무기나 방향을 확인할 휴대 소품은 없다.",
        "built_space": "밝은 천장등 하나, 사각 환기구 하나, 작은 원형 천장 장치 하나, 오른쪽 열린 회색 문 하나와 복도가 보인다. 회백색 이중 도장 벽과 문의 하단 환기구는 장소 참조와 부합한다. 의자 등받이는 전경 중앙과 미연 뒤 오른쪽에 각각 하나씩 보인다. 다만 중앙 등받이가 화면 하단을 크게 가로질러, 요구된 박철진 몸 아래의 작은 의자 측면 대신 큰 장애물이 된다. 박철진이 의자에 비스듬히 앉은 것인지 정확한 좌석 관계는 가림 때문에 불명확하다. 외부가 보이지 않아 밤과 층수는 독립적으로 확인되지 않는다.",
        "entities": "중년 동아시아인 남녀 두 명만 보이며 인원과 대략적인 연령·성별은 부합한다. 박철진의 짧은 검은 머리와 귀걸이, 미연의 검은 단발은 참조와 유사하다. 미연의 눈 주위와 볼에는 심한 부종과 멍이 보인다. 그러나 박철진은 참조의 남색 정장 대신 어두운 캐주얼 재킷을, 미연은 남색 상의 대신 회갈색 겉셔츠를 입었다. 박철진의 굳은 표정과 전진한 상체는 위압적이지만 거만함은 강하지 않다. 판독 가능한 글자나 추가 인물은 보이지 않는다.",
        "hard_violations": [],
        "physics": "미연은 뒤쪽 등받이가 있는 의자에 앉아 상체를 숙이고 있다. 박철진의 골반과 좌판 접촉은 전경 등받이와 화면 경계에 가려져 지지 위치를 직접 확인하기 어렵다. 옆으로 틀어 앉아 앞으로 기울인 자세로는 가능하며, 공중에 떠 있다고 단정할 근거는 없다. 명백한 관절 이상이나 지지 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "박철진은 왼쪽 전경의 미연 얼굴을 똑바로 바라보며 상체를 가까이 기울인다. 미연도 얼굴을 박철진 쪽으로 돌리고 있지만 눈은 측면과 가림 때문에 정확한 초점을 확인하기 어렵다. 두 사람의 대면 방향은 명확하다.",
        "built_space": "회백색 이중 도장 벽, 천장등 하나, 사각 환기구 하나, 작은 원형 천장 장치 하나, 오른쪽 열린 회색 문 하나와 복도가 보이며 장소의 주요 재료와 구조는 참조에 가깝다. 그러나 의자는 미연이 앉은 왼쪽 전경 의자, 박철진 뒤의 의자, 오른쪽 빈 의자로 총 세 개다. 참조의 두 심문 의자보다 하나 많다. 오른쪽 빈 의자는 좌판과 등받이까지 크게 노출된다. 박철진은 미연 앞에 가까이 앉아 있지만 프레임은 상체뿐 아니라 손과 허벅지까지 비교적 넓게 담는다. 외부가 없어 밤과 층수는 확인되지 않는다.",
        "entities": "중년 동아시아인 남녀 두 명이며 추가 인물은 없다. 박철진의 얼굴, 짧게 빗어 넘긴 검은 머리, 귀걸이, 남색 정장·흰 셔츠·사선무늬 넥타이는 참조와 잘 맞는다. 미연은 검은 단발의 중년 여성으로 보이고 드러난 볼에 부종과 멍이 있다. 다만 얼굴 대부분이 옆모습이라 전체 부종과 동일인 여부 확인은 제한되며, 상의는 참조의 남색이 아닌 밝은 회색이다. 박철진의 표정은 집중하거나 추궁하는 인상으로, 거만함은 다소 약하다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "두 개로 정해진 심문 의자 외에 화면 오른쪽에 세 번째 빈 의자를 추가했다."
        ],
        "physics": "두 사람 모두 각자의 의자에 골반을 두고 등받이를 뒤로 향한 정상적인 착석 관계를 보인다. 박철진은 다리를 벌리고 몸을 앞으로 기울였으며, 양손은 무릎 사이에 자연스럽게 모여 있다. 보이는 의자 프레임과 다리는 바닥 쪽으로 이어지고, 지지 없이 떠 있는 몸이나 소품은 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "의상 불일치와 전경을 가리는 의자 등받이가 문제지만, 부은 미연 앞에 접근한 박철진을 보여주며 B의 명백한 의자 추가 오류는 없다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "박철진의 상체 중심 구도와 정장은 더 충실하지만, 두 심문 의자 외에 빈 의자를 하나 추가해 고정된 공간 구성을 위반한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽 박철진은 몸과 얼굴을 오른쪽 미연에게 기울이고 미연의 얼굴을 바라본다. 미연은 고개를 숙여 두 사람 사이 아래쪽을 본다. 상대를 압박하는 시선 관계는 성립하며, 무기나 방향을 확인할 휴대 소품은 없다.",
        "built_space": "밝은 천장등 하나, 사각 환기구 하나, 작은 원형 천장 장치 하나, 오른쪽 열린 회색 문 하나와 복도가 보인다. 회백색 이중 도장 벽과 문의 하단 환기구는 장소 참조와 부합한다. 의자 등받이는 전경 중앙과 미연 뒤 오른쪽에 각각 하나씩 보인다. 다만 중앙 등받이가 화면 하단을 크게 가로질러, 요구된 박철진 몸 아래의 작은 의자 측면 대신 큰 장애물이 된다. 박철진이 의자에 비스듬히 앉은 것인지 정확한 좌석 관계는 가림 때문에 불명확하다. 외부가 보이지 않아 밤과 층수는 독립적으로 확인되지 않는다.",
        "entities": "중년 동아시아인 남녀 두 명만 보이며 인원과 대략적인 연령·성별은 부합한다. 박철진의 짧은 검은 머리와 귀걸이, 미연의 검은 단발은 참조와 유사하다. 미연의 눈 주위와 볼에는 심한 부종과 멍이 보인다. 그러나 박철진은 참조의 남색 정장 대신 어두운 캐주얼 재킷을, 미연은 남색 상의 대신 회갈색 겉셔츠를 입었다. 박철진의 굳은 표정과 전진한 상체는 위압적이지만 거만함은 강하지 않다. 판독 가능한 글자나 추가 인물은 보이지 않는다.",
        "hard_violations": [],
        "physics": "미연은 뒤쪽 등받이가 있는 의자에 앉아 상체를 숙이고 있다. 박철진의 골반과 좌판 접촉은 전경 등받이와 화면 경계에 가려져 지지 위치를 직접 확인하기 어렵다. 옆으로 틀어 앉아 앞으로 기울인 자세로는 가능하며, 공중에 떠 있다고 단정할 근거는 없다. 명백한 관절 이상이나 지지 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "박철진은 왼쪽 전경의 미연 얼굴을 똑바로 바라보며 상체를 가까이 기울인다. 미연도 얼굴을 박철진 쪽으로 돌리고 있지만 눈은 측면과 가림 때문에 정확한 초점을 확인하기 어렵다. 두 사람의 대면 방향은 명확하다.",
        "built_space": "회백색 이중 도장 벽, 천장등 하나, 사각 환기구 하나, 작은 원형 천장 장치 하나, 오른쪽 열린 회색 문 하나와 복도가 보이며 장소의 주요 재료와 구조는 참조에 가깝다. 그러나 의자는 미연이 앉은 왼쪽 전경 의자, 박철진 뒤의 의자, 오른쪽 빈 의자로 총 세 개다. 참조의 두 심문 의자보다 하나 많다. 오른쪽 빈 의자는 좌판과 등받이까지 크게 노출된다. 박철진은 미연 앞에 가까이 앉아 있지만 프레임은 상체뿐 아니라 손과 허벅지까지 비교적 넓게 담는다. 외부가 없어 밤과 층수는 확인되지 않는다.",
        "entities": "중년 동아시아인 남녀 두 명이며 추가 인물은 없다. 박철진의 얼굴, 짧게 빗어 넘긴 검은 머리, 귀걸이, 남색 정장·흰 셔츠·사선무늬 넥타이는 참조와 잘 맞는다. 미연은 검은 단발의 중년 여성으로 보이고 드러난 볼에 부종과 멍이 있다. 다만 얼굴 대부분이 옆모습이라 전체 부종과 동일인 여부 확인은 제한되며, 상의는 참조의 남색이 아닌 밝은 회색이다. 박철진의 표정은 집중하거나 추궁하는 인상으로, 거만함은 다소 약하다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "두 개로 정해진 심문 의자 외에 화면 오른쪽에 세 번째 빈 의자를 추가했다."
        ],
        "physics": "두 사람 모두 각자의 의자에 골반을 두고 등받이를 뒤로 향한 정상적인 착석 관계를 보인다. 박철진은 다리를 벌리고 몸을 앞으로 기울였으며, 양손은 무릎 사이에 자연스럽게 모여 있다. 보이는 의자 프레임과 다리는 바닥 쪽으로 이어지고, 지지 없이 떠 있는 몸이나 소품은 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.6,
    "B": 1.75
   },
   "adjusted": {
    "A": 1.35,
    "B": 1.5
   },
   "violations": {
    "A": [
     "[gemini-pro] 공간 내 의자 개수 초과 (2개여야 할 의자가 3개로 복제됨)",
     "[gpt-high] 두 개로 정해진 심문 의자 외에 화면 오른쪽에 세 번째 빈 의자를 추가했다."
    ],
    "B": [
     "[gemini-pro] 공간 내 의자 개수 초과 (2개여야 할 의자가 3개로 복제됨)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1350,
   "B": 1500
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1350,
    "verdict_ko": "박철진의 거만한 자세를 잘 포착했으나, 원본 공간에 없는 세 번째 의자가 우측에 생성되어 실격입니다.  ★위반: [gemini-pro] 공간 내 의자 개수 초과 (2개여야 할 의자가 3개로 복제됨) / [gpt-high] 두 개로 정해진 심문 의자 외에 화면 오른쪽에 세 번째 빈 의자를 추가했다."
   },
   {
    "label": "B",
    "score": 1500,
    "verdict_ko": "미연의 부은 얼굴 묘사는 훌륭하지만, 전경에 세 번째 의자가 잘못 추가되어 구도를 해치므로 실격입니다.  ★위반: [gemini-pro] 공간 내 의자 개수 초과 (2개여야 할 의자가 3개로 복제됨)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L189B01.png",
    "asset_id": "d31b4d4c-5488-4514-8d64-f7f3972f4a12",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1113064>",
    "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1401722>",
    "asset_id": "fee7383c-fb61-4b3a-ba7c-79f2555de00b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-4fb5-7610-a087-4ecc0174cac2",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S31sh2__bgfirst_bg.png",
   "bg_asset_id": "6d0e961f-cfc1-4ccb-921e-40cb43da0593",
   "bg_record_key": "S31sh2::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S31sh2::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:42:07.497511+00:00",
  "fingerprint": "080e548904893a87fe88e990e3b07b5c0abf51a9ba2740457aa3836de63b8626",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S31sh2_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S31sh2_sel.png",
  "source_sha256": "cb5b09049c0be79e4edcb606002327f80fe42635d10445c58ebd4e373850cd3d",
  "file": "S31sh2_cine.png",
  "staged_sha256": "339aa81c902de00b3741b1a65a8e87525d17baa70697f8a977132be65b77e745",
  "latency_ms": 11325
 },
 "S31sh8::signage": {
  "fp": "193145b90fc8ec5f",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S31sh8": {
  "input_fingerprint": "4c27f623b2073a30",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 수하 1의 귓속말에 놀라 눈이 커진 채 자리에서 반쯤 일어선 박철진의 엉거주춤한 자세.\n\nLOCATION (lock): Beside the interrogation chair inside the illuminated third-floor room of the refugee administration building. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 박철진's chair (Partly vacated as he interrupts his rise) — The seat is visible obliquely below and behind his hips; used as The visible separation between hips and seat proves that he has not fully stood.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the office illumination consistent with the questioning shot, preserving facial detail through the incomplete rise.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the same confinement-room surfaces, interrogation chair, and nighttime interior lighting. Exclude equipment or furniture from the separate militia office.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The third-floor interrogation room remains brightly lit, and its door has been opened for the report. 박철진: He remains at the interrogation position and reacts angrily to the report.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 수하 1의 귓속말에 놀라 눈이 커진 채 자리에서 반쯤 일어선 박철진의 엉거주춤한 자세.\n\nLOCATION (lock): Beside the interrogation chair inside the illuminated third-floor room of the refugee administration building. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 박철진's chair (Partly vacated as he interrupts his rise) — The seat is visible obliquely below and behind his hips; used as The visible separation between hips and seat proves that he has not fully stood.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the office illumination consistent with the questioning shot, preserving facial detail through the incomplete rise.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the same confinement-room surfaces, interrogation chair, and nighttime interior lighting. Exclude equipment or furniture from the separate militia office.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The third-floor interrogation room remains brightly lit, and its door has been opened for the report. 박철진: He remains at the interrogation position and reacts angrily to the report.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 수하 1의 귓속말에 놀라 눈이 커진 채 자리에서 반쯤 일어선 박철진의 엉거주춤한 자세.\n\nLOCATION (lock): Beside the interrogation chair inside the illuminated third-floor room of the refugee administration building. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 박철진's chair (Partly vacated as he interrupts his rise) — The seat is visible obliquely below and behind his hips; used as The visible separation between hips and seat proves that he has not fully stood.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the office illumination consistent with the questioning shot, preserving facial detail through the incomplete rise.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the same confinement-room surfaces, interrogation chair, and nighttime interior lighting. Exclude equipment or furniture from the separate militia office.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The third-floor interrogation room remains brightly lit, and its door has been opened for the report. 박철진: He remains at the interrogation position and reacts angrily to the report.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "박철진은 눈이 커진 채 화면 밖을 응시하고, 우측의 여성이 그의 귀에 대고 속삭이고 있음.",
    "built_space": "이전 샷의 빈 공간과 달리 중앙에 큰 탁자가 배치되어 있고, 의자들은 배경에 놓여 있음.",
    "entities": "박철진(참조 이미지 일치), 이어폰을 낀 낯선 여성(목록에 없는 인물), 탁자(이전 샷에 없던 사물).",
    "hard_violations": [
     "[gemini-pro] 허가되지 않은 인물 등장",
     "[gemini-pro] 허가되지 않은 사물(탁자) 추가 및 공간 왜곡",
     "[gemini-pro] 의자에서 일어나는 구도 위반 (엉덩이 아래 의자 좌석이 보이지 않음)",
     "[gpt-high] 박철진만 화면에 등장할 수 있다는 명시적 인물 제한과 달리, 귓속말하는 별도 인물을 추가했다."
    ],
    "physics": "박철진은 의자에서 일어나는 중이 아니라 탁자에 손을 짚고 기대어 서 있음."
   },
   {
    "label": "B",
    "direction": "박철진은 놀란 표정으로 화면 우측을 향하고, 이전 샷의 여성이 그의 귀에 속삭이고 있음.",
    "built_space": "이전 샷과 동일한 구조의 방이며, 전경에 의자 등받이가 보임.",
    "entities": "박철진, 이전 샷의 여성(상처난 얼굴과 옷차림 동일 - 프롬프트에서 재등장 금지됨).",
    "hard_violations": [
     "[gemini-pro] 이전 샷 인물 재등장 금지 지시 완벽 위반",
     "[gemini-pro] 의자에서 일어나는 구도 위반 (엉덩이 아래 의자 좌석이 명확히 보이지 않음)",
     "[gpt-high] 이전 장면에서 제외하라고 명시한 여성을 같은 머리와 의상으로 재등장시켰으며, 박철진만 허용한 인물 제한도 위반했다."
    ],
    "physics": "박철진은 허리를 굽히고 엉거주춤 서 있으나, 일어나는 동작을 증명할 하단 지지물(의자)과의 관계가 단절됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "의자에서 반쯤 일어나는 구도를 무시하고 임의로 탁자를 추가해 기대고 있으며, 화이트리스트에 없는 인물을 등장시켜 지시를 위반했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "이전 샷의 인물을 절대 재등장시키지 말라는 강력한 금지 지항을 무시하고 그대로 복사해 넣었으며, 요구된 카메라 구도(엉덩이 뒤 의자 노출)도 충족하지 못했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진은 눈이 커진 채 화면 밖을 응시하고, 우측의 여성이 그의 귀에 대고 속삭이고 있음.",
        "built_space": "이전 샷의 빈 공간과 달리 중앙에 큰 탁자가 배치되어 있고, 의자들은 배경에 놓여 있음.",
        "entities": "박철진(참조 이미지 일치), 이어폰을 낀 낯선 여성(목록에 없는 인물), 탁자(이전 샷에 없던 사물).",
        "hard_violations": [
         "허가되지 않은 인물 등장",
         "허가되지 않은 사물(탁자) 추가 및 공간 왜곡",
         "의자에서 일어나는 구도 위반 (엉덩이 아래 의자 좌석이 보이지 않음)"
        ],
        "physics": "박철진은 의자에서 일어나는 중이 아니라 탁자에 손을 짚고 기대어 서 있음."
       },
       {
        "label": "B",
        "direction": "박철진은 놀란 표정으로 화면 우측을 향하고, 이전 샷의 여성이 그의 귀에 속삭이고 있음.",
        "built_space": "이전 샷과 동일한 구조의 방이며, 전경에 의자 등받이가 보임.",
        "entities": "박철진, 이전 샷의 여성(상처난 얼굴과 옷차림 동일 - 프롬프트에서 재등장 금지됨).",
        "hard_violations": [
         "이전 샷 인물 재등장 금지 지시 완벽 위반",
         "의자에서 일어나는 구도 위반 (엉덩이 아래 의자 좌석이 명확히 보이지 않음)"
        ],
        "physics": "박철진은 허리를 굽히고 엉거주춤 서 있으나, 일어나는 동작을 증명할 하단 지지물(의자)과의 관계가 단절됨."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "의자에서 반쯤 일어나는 구도를 무시하고 임의로 탁자를 추가해 기대고 있으며, 화이트리스트에 없는 인물을 등장시켜 지시를 위반했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "이전 샷의 인물을 절대 재등장시키지 말라는 강력한 금지 지항을 무시하고 그대로 복사해 넣었으며, 요구된 카메라 구도(엉덩이 뒤 의자 노출)도 충족하지 못했습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "박철진은 눈이 커진 채 화면 밖을 응시하고, 우측의 여성이 그의 귀에 대고 속삭이고 있음.",
        "built_space": "이전 샷의 빈 공간과 달리 중앙에 큰 탁자가 배치되어 있고, 의자들은 배경에 놓여 있음.",
        "entities": "박철진(참조 이미지 일치), 이어폰을 낀 낯선 여성(목록에 없는 인물), 탁자(이전 샷에 없던 사물).",
        "hard_violations": [
         "허가되지 않은 인물 등장",
         "허가되지 않은 사물(탁자) 추가 및 공간 왜곡",
         "의자에서 일어나는 구도 위반 (엉덩이 아래 의자 좌석이 보이지 않음)"
        ],
        "physics": "박철진은 의자에서 일어나는 중이 아니라 탁자에 손을 짚고 기대어 서 있음."
       },
       {
        "label": "B",
        "direction": "박철진은 놀란 표정으로 화면 우측을 향하고, 이전 샷의 여성이 그의 귀에 속삭이고 있음.",
        "built_space": "이전 샷과 동일한 구조의 방이며, 전경에 의자 등받이가 보임.",
        "entities": "박철진, 이전 샷의 여성(상처난 얼굴과 옷차림 동일 - 프롬프트에서 재등장 금지됨).",
        "hard_violations": [
         "이전 샷 인물 재등장 금지 지시 완벽 위반",
         "의자에서 일어나는 구도 위반 (엉덩이 아래 의자 좌석이 명확히 보이지 않음)"
        ],
        "physics": "박철진은 허리를 굽히고 엉거주춤 서 있으나, 일어나는 동작을 증명할 하단 지지물(의자)과의 관계가 단절됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "놀라며 몸을 일으키는 반응은 보이지만, 명시적으로 제외된 이전 장면의 여성을 그대로 재등장시켰고 좌판과 엉덩이 사이의 간격도 가렸다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "박철진의 얼굴과 귓속말에 놀란 반응은 더 충실하지만, 박철진만 허용한다는 인물 제한을 어겼으며 반쯤 일어섰음을 입증할 좌판이 명확히 보이지 않는다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진은 눈을 크게 뜨고 화면 오른쪽 바깥을 바라보며, 시선의 대상은 보이지 않는다. 오른쪽 여성은 입을 박철진의 귀 가까이에 대어 귓속말하는 방향은 맞는다. 무기나 방향성 있는 소품은 없다.",
        "built_space": "전경 중앙과 오른쪽에 의자 등받이가 하나씩 보여 의자는 두 개다. 중앙의 큰 등받이가 박철진의 엉덩이 아래와 좌판을 가려, 요구된 좌판과 엉덩이의 분리가 보이지 않는다. 오른쪽 열린 문 하나, 천장 직사각형 조명 하나, 사각 환기구와 작은 원형 설비, 밝은 상부와 회색 하부 벽은 이전 장소와 대체로 일치한다.",
        "entities": "박철진은 중년 한국인 남성으로 보이며 짧게 넘긴 검은 머리와 어두운 재킷이 이전 장면에 가깝다. 놀란 눈은 정상적인 인간 눈으로 표현됐다. 그러나 옆 인물은 이전 장면 여성의 검은 단발과 회갈색 셔츠를 그대로 이어받아, 해당 인물의 외형과 의상을 재사용하지 말라는 지시를 위반한다. 전경 의자의 작은 표시는 흐려 읽을 수 없다.",
        "hard_violations": [
         "이전 장면에서 제외하라고 명시한 여성을 같은 머리와 의상으로 재등장시켰으며, 박철진만 허용한 인물 제한도 위반했다."
        ],
        "physics": "박철진은 골반을 뒤로 두고 상체를 앞으로 숙여 일어나는 중간 자세를 취한다. 하체는 아래로 이어지지만 발과 손의 접촉점은 화면에 가려져 직접 확인되지 않는다. 공중에 떠 있는 형태는 아니다. 여성 역시 하체가 화면 밖으로 이어져 있으며, 귀에 가까이 몸을 기울이는 동작 자체는 가능하다."
       },
       {
        "label": "B",
        "direction": "박철진은 눈을 크게 뜬 채 카메라 가까운 앞쪽을 바라보고, 구체적인 시선 대상은 화면에 없다. 오른쪽 인물의 입은 박철진의 왼쪽 귀를 향해 있어 귓속말 관계가 분명하다. 그 인물의 손은 박철진의 위팔을 잡고 있다.",
        "built_space": "박철진 뒤 왼쪽의 의자 하나와 중앙 뒤쪽의 의자 하나, 중앙의 금속 탁자 하나가 보인다. 박철진 뒤 의자는 등받이가 주로 드러나고 좌판은 몸과 화면 하단에 가려져, 엉덩이와 좌판 사이의 틈을 명확하게 확인할 수 없다. 오른쪽 열린 문 하나와 천장 조명 하나, 환기구 및 이색 벽면은 이전 장소를 대체로 유지한다. 새로 드러난 탁자의 연속성은 참고 사진만으로 확인하기 어렵다.",
        "entities": "박철진의 중년 얼굴 윤곽, 검은 머리, 작은 귀걸이와 어두운 재킷은 참고 인물과 잘 맞는다. 눈을 크게 뜨고 입을 벌린 표정은 놀람에 가깝고 분노는 약하다. 오른쪽에는 짧은 갈색 머리, 회색 셔츠, 검은 바지와 이어피스를 착용한 별도 성인이 등장한다. 이전 여성의 얼굴을 그대로 복제하지는 않았지만, 박철진만 허용한다는 명시적 제한에는 맞지 않는다. 탁자 위에는 펼친 서류철이 있으며 글자는 읽히지 않는다.",
        "hard_violations": [
         "박철진만 화면에 등장할 수 있다는 명시적 인물 제한과 달리, 귓속말하는 별도 인물을 추가했다."
        ],
        "physics": "박철진은 상체를 앞으로 기울이고 골반을 뒤에 둔 자세이며 하체는 화면 아래로 이어진다. 발은 잘렸지만 부유하거나 불가능하게 지지된 모습은 아니다. 상대 인물은 손으로 박철진의 위팔을 잡고 허리를 굽혀 귀에 접근한다. 서류철은 탁자 위에 놓여 있어 지지가 분명하다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "놀라며 몸을 일으키는 반응은 보이지만, 명시적으로 제외된 이전 장면의 여성을 그대로 재등장시켰고 좌판과 엉덩이 사이의 간격도 가렸다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "박철진의 얼굴과 귓속말에 놀란 반응은 더 충실하지만, 박철진만 허용한다는 인물 제한을 어겼으며 반쯤 일어섰음을 입증할 좌판이 명확히 보이지 않는다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "박철진은 눈을 크게 뜨고 화면 오른쪽 바깥을 바라보며, 시선의 대상은 보이지 않는다. 오른쪽 여성은 입을 박철진의 귀 가까이에 대어 귓속말하는 방향은 맞는다. 무기나 방향성 있는 소품은 없다.",
        "built_space": "전경 중앙과 오른쪽에 의자 등받이가 하나씩 보여 의자는 두 개다. 중앙의 큰 등받이가 박철진의 엉덩이 아래와 좌판을 가려, 요구된 좌판과 엉덩이의 분리가 보이지 않는다. 오른쪽 열린 문 하나, 천장 직사각형 조명 하나, 사각 환기구와 작은 원형 설비, 밝은 상부와 회색 하부 벽은 이전 장소와 대체로 일치한다.",
        "entities": "박철진은 중년 한국인 남성으로 보이며 짧게 넘긴 검은 머리와 어두운 재킷이 이전 장면에 가깝다. 놀란 눈은 정상적인 인간 눈으로 표현됐다. 그러나 옆 인물은 이전 장면 여성의 검은 단발과 회갈색 셔츠를 그대로 이어받아, 해당 인물의 외형과 의상을 재사용하지 말라는 지시를 위반한다. 전경 의자의 작은 표시는 흐려 읽을 수 없다.",
        "hard_violations": [
         "이전 장면에서 제외하라고 명시한 여성을 같은 머리와 의상으로 재등장시켰으며, 박철진만 허용한 인물 제한도 위반했다."
        ],
        "physics": "박철진은 골반을 뒤로 두고 상체를 앞으로 숙여 일어나는 중간 자세를 취한다. 하체는 아래로 이어지지만 발과 손의 접촉점은 화면에 가려져 직접 확인되지 않는다. 공중에 떠 있는 형태는 아니다. 여성 역시 하체가 화면 밖으로 이어져 있으며, 귀에 가까이 몸을 기울이는 동작 자체는 가능하다."
       },
       {
        "label": "A",
        "direction": "박철진은 눈을 크게 뜬 채 카메라 가까운 앞쪽을 바라보고, 구체적인 시선 대상은 화면에 없다. 오른쪽 인물의 입은 박철진의 왼쪽 귀를 향해 있어 귓속말 관계가 분명하다. 그 인물의 손은 박철진의 위팔을 잡고 있다.",
        "built_space": "박철진 뒤 왼쪽의 의자 하나와 중앙 뒤쪽의 의자 하나, 중앙의 금속 탁자 하나가 보인다. 박철진 뒤 의자는 등받이가 주로 드러나고 좌판은 몸과 화면 하단에 가려져, 엉덩이와 좌판 사이의 틈을 명확하게 확인할 수 없다. 오른쪽 열린 문 하나와 천장 조명 하나, 환기구 및 이색 벽면은 이전 장소를 대체로 유지한다. 새로 드러난 탁자의 연속성은 참고 사진만으로 확인하기 어렵다.",
        "entities": "박철진의 중년 얼굴 윤곽, 검은 머리, 작은 귀걸이와 어두운 재킷은 참고 인물과 잘 맞는다. 눈을 크게 뜨고 입을 벌린 표정은 놀람에 가깝고 분노는 약하다. 오른쪽에는 짧은 갈색 머리, 회색 셔츠, 검은 바지와 이어피스를 착용한 별도 성인이 등장한다. 이전 여성의 얼굴을 그대로 복제하지는 않았지만, 박철진만 허용한다는 명시적 제한에는 맞지 않는다. 탁자 위에는 펼친 서류철이 있으며 글자는 읽히지 않는다.",
        "hard_violations": [
         "박철진만 화면에 등장할 수 있다는 명시적 인물 제한과 달리, 귓속말하는 별도 인물을 추가했다."
        ],
        "physics": "박철진은 상체를 앞으로 기울이고 골반을 뒤에 둔 자세이며 하체는 화면 아래로 이어진다. 발은 잘렸지만 부유하거나 불가능하게 지지된 모습은 아니다. 상대 인물은 손으로 박철진의 위팔을 잡고 허리를 굽혀 귀에 접근한다. 서류철은 탁자 위에 놓여 있어 지지가 분명하다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.5
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.25
   },
   "violations": {
    "A": [
     "[gemini-pro] 허가되지 않은 인물 등장",
     "[gemini-pro] 허가되지 않은 사물(탁자) 추가 및 공간 왜곡",
     "[gemini-pro] 의자에서 일어나는 구도 위반 (엉덩이 아래 의자 좌석이 보이지 않음)",
     "[gpt-high] 박철진만 화면에 등장할 수 있다는 명시적 인물 제한과 달리, 귓속말하는 별도 인물을 추가했다."
    ],
    "B": [
     "[gemini-pro] 이전 샷 인물 재등장 금지 지시 완벽 위반",
     "[gemini-pro] 의자에서 일어나는 구도 위반 (엉덩이 아래 의자 좌석이 명확히 보이지 않음)",
     "[gpt-high] 이전 장면에서 제외하라고 명시한 여성을 같은 머리와 의상으로 재등장시켰으며, 박철진만 허용한 인물 제한도 위반했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 1250
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "의자에서 반쯤 일어나는 구도를 무시하고 임의로 탁자를 추가해 기대고 있으며, 화이트리스트에 없는 인물을 등장시켜 지시를 위반했습니다.  ★위반: [gemini-pro] 허가되지 않은 인물 등장 / [gemini-pro] 허가되지 않은 사물(탁자) 추가 및 공간 왜곡 / [gemini-pro] 의자에서 일어나는 구도 위반 (엉덩이 아래 의자 좌석이 보이지 않음) / [gpt-high] 박철진만 화면에 등장할 수 있다는 명시적 인물 제한과 달리, 귓속말하는 별도 인물을 추가했다."
   },
   {
    "label": "B",
    "score": 1250,
    "verdict_ko": "이전 샷의 인물을 절대 재등장시키지 말라는 강력한 금지 지항을 무시하고 그대로 복사해 넣었으며, 요구된 카메라 구도(엉덩이 뒤 의자 노출)도 충족하지 못했습니다.  ★위반: [gemini-pro] 이전 샷 인물 재등장 금지 지시 완벽 위반 / [gemini-pro] 의자에서 일어나는 구도 위반 (엉덩이 아래 의자 좌석이 명확히 보이지 않음) / [gpt-high] 이전 장면에서 제외하라고 명시한 여성을 같은 머리와 의상으로 재등장시켰으며, 박철진만 허용한 인물 제한도 위반했다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S31sh2_sel.png",
    "asset_id": "209fb114-415e-429b-857e-54ebdb3054c2",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1401722>",
    "asset_id": "fee7383c-fb61-4b3a-ba7c-79f2555de00b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-5325-7529-afe0-2f17f975d5ca",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S31sh2"
  }
 },
 "S31sh8::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:43:33.105822+00:00",
  "fingerprint": "89b1e6b154b99080c61c6c0f270baa17de4ca99cc966a7a5bd629c3a101d1f27",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S31sh8_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S31sh8_sel.png",
  "source_sha256": "3fb52668ad8f6113e6c88342af4cabd268a27edcf1055f994f079d8d7860b7fd",
  "file": "S31sh8_cine.png",
  "staged_sha256": "da3bf319d83ba00919bee40e6ad4559252a5525540158456a43802151d6eee10",
  "latency_ms": 10620
 },
 "S31sh9::signage": {
  "fp": "42982951c83501c8",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S31sh9": {
  "input_fingerprint": "d9b354eb0f5c2176",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 성당이라는 말에 충격을 받은 듯 두 눈이 동그라진 미연의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): At the detainee's seat inside the illuminated third-floor interrogation room of the refugee administration building. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: 미연 alone in the close frame in the middle-left of the frame, foreground.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Unchanged office illumination keeps the swelling and suddenly widened eyes legible without glamorizing the injury.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same confinement-room surfaces, chair, and interior lighting. Exclude the command-office desk, speakerphone, and wall-mounted knife from the separate office scene.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The third-floor interrogation room remains brightly lit, with the door opened during the report. 미연: Her face remains badly swollen from the beating, and she is visibly startled.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 성당이라는 말에 충격을 받은 듯 두 눈이 동그라진 미연의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): At the detainee's seat inside the illuminated third-floor interrogation room of the refugee administration building. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: 미연 alone in the close frame in the middle-left of the frame, foreground.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Unchanged office illumination keeps the swelling and suddenly widened eyes legible without glamorizing the injury.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same confinement-room surfaces, chair, and interior lighting. Exclude the command-office desk, speakerphone, and wall-mounted knife from the separate office scene.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The third-floor interrogation room remains brightly lit, with the door opened during the report. 미연: Her face remains badly swollen from the beating, and she is visibly startled.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 성당이라는 말에 충격을 받은 듯 두 눈이 동그라진 미연의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): At the detainee's seat inside the illuminated third-floor interrogation room of the refugee administration building. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: 미연 alone in the close frame in the middle-left of the frame, foreground.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Unchanged office illumination keeps the swelling and suddenly widened eyes legible without glamorizing the injury.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same confinement-room surfaces, chair, and interior lighting. Exclude the command-office desk, speakerphone, and wall-mounted knife from the separate office scene.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The third-floor interrogation room remains brightly lit, with the door opened during the report. 미연: Her face remains badly swollen from the beating, and she is visibly startled.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "시선은 화면 우측 밖을 향함.",
    "built_space": "투톤 벽면, 책상, 나무 의자가 참조 이미지의 취조실 환경과 일치함.",
    "entities": "미연과 일치하는 인물, 남색 상의, 멍들고 부은 얼굴 묘사 확인됨.",
    "hard_violations": [],
    "physics": "손이 책상 위에 자연스럽게 놓여 있음."
   },
   {
    "label": "B",
    "direction": "카메라 렌즈를 정면으로 응시함.",
    "built_space": "투톤 벽면과 조명은 일치하나, 책상과 의자가 배경에 멀리 떨어져 있음.",
    "entities": "미연과 일치하는 인물, 어두운 상의, 멍든 얼굴 묘사 확인됨.",
    "hard_violations": [
     "[gemini-pro] 지정된 취조실 좌석이 아닌 책상과 멀리 떨어진 위치에 인물이 배치됨",
     "[gpt-high] 고정된 심문실의 의자를 참고 장면의 두 개에서 세 개로 늘려 가구를 추가했습니다.",
     "[gpt-high] 피구금자의 심문 좌석이어야 하는 인물을 심문 탁자에서 떨어진 전경에 배치하고 탁자를 뒤쪽으로 옮겼습니다."
    ],
    "physics": "상반신만 보여 특별한 지지 상태는 확인되지 않음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "프레임 중좌측 배치와 놀란 표정, 부은 얼굴 등 프롬프트의 구도 및 묘사 지시를 훌륭히 따랐습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "인물을 프레임 중앙에 배치하고 지정된 좌석이 아닌 공간에 세워두어 구도와 장소 설정 지시를 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 화면 우측 밖을 향함.",
        "built_space": "투톤 벽면, 책상, 나무 의자가 참조 이미지의 취조실 환경과 일치함.",
        "entities": "미연과 일치하는 인물, 남색 상의, 멍들고 부은 얼굴 묘사 확인됨.",
        "hard_violations": [],
        "physics": "손이 책상 위에 자연스럽게 놓여 있음."
       },
       {
        "label": "B",
        "direction": "카메라 렌즈를 정면으로 응시함.",
        "built_space": "투톤 벽면과 조명은 일치하나, 책상과 의자가 배경에 멀리 떨어져 있음.",
        "entities": "미연과 일치하는 인물, 어두운 상의, 멍든 얼굴 묘사 확인됨.",
        "hard_violations": [
         "지정된 취조실 좌석이 아닌 책상과 멀리 떨어진 위치에 인물이 배치됨"
        ],
        "physics": "상반신만 보여 특별한 지지 상태는 확인되지 않음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "프레임 중좌측 배치와 놀란 표정, 부은 얼굴 등 프롬프트의 구도 및 묘사 지시를 훌륭히 따랐습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "인물을 프레임 중앙에 배치하고 지정된 좌석이 아닌 공간에 세워두어 구도와 장소 설정 지시를 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 화면 우측 밖을 향함.",
        "built_space": "투톤 벽면, 책상, 나무 의자가 참조 이미지의 취조실 환경과 일치함.",
        "entities": "미연과 일치하는 인물, 남색 상의, 멍들고 부은 얼굴 묘사 확인됨.",
        "hard_violations": [],
        "physics": "손이 책상 위에 자연스럽게 놓여 있음."
       },
       {
        "label": "B",
        "direction": "카메라 렌즈를 정면으로 응시함.",
        "built_space": "투톤 벽면과 조명은 일치하나, 책상과 의자가 배경에 멀리 떨어져 있음.",
        "entities": "미연과 일치하는 인물, 어두운 상의, 멍든 얼굴 묘사 확인됨.",
        "hard_violations": [
         "지정된 취조실 좌석이 아닌 책상과 멀리 떨어진 위치에 인물이 배치됨"
        ],
        "physics": "상반신만 보여 특별한 지지 상태는 확인되지 않음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "얼굴 클로즈업과 중간 왼쪽 배치는 잘 맞지만, 의자를 세 개로 늘리고 심문 탁자를 인물 뒤로 옮겨 장소의 연속성을 깨뜨렸습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "부어오른 얼굴과 화면 밖 상대를 향한 놀란 눈, 참고 의상과 심문실 재질은 충실하지만, 손과 탁자까지 담아 요구한 얼굴 클로즈업보다 넓습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴과 두 눈이 거의 정면의 카메라를 향합니다. 충격으로 눈을 크게 뜬 모습은 보이지만, 말을 들려준 화면 밖 상대를 바라보는 관계는 드러나지 않습니다. 방향을 판정할 무기나 휴대 물체는 없습니다.",
        "built_space": "밝은 상부 벽과 회색 하부 벽, 천장 조명 한 개, 환기구 한 개, 오른쪽 뒤 금속 탁자 한 개가 보입니다. 목제 의자 등받이는 인물 바로 뒤와 그보다 먼 중앙, 오른쪽 탁자 뒤에 각각 하나씩 총 세 개가 보이며, 참고 장면의 두 의자와 다릅니다. 인물이 탁자 앞 심문 자리라기보다 탁자에서 떨어진 전경에 놓여 있습니다. 문은 화면에 없어 개방 여부를 확인할 수 없습니다.",
        "entities": "중년 한국인 여성으로 보이는 인물 한 명만 있습니다. 검은 머리와 얼굴 윤곽은 참고 인물과 대체로 닮았지만 머리가 더 길고 흐트러져 있으며, 보이는 상의는 참고의 남색보다 짙은 회흑색에 가깝습니다. 볼과 눈가에 멍이 있으나 심하게 부었다는 상태는 비교적 약합니다. 탁자 위 열린 서류철에는 읽을 수 있는 글자가 없고, 전화기나 벽걸이 칼은 없습니다.",
        "hard_violations": [
         "고정된 심문실의 의자를 참고 장면의 두 개에서 세 개로 늘려 가구를 추가했습니다.",
         "피구금자의 심문 좌석이어야 하는 인물을 심문 탁자에서 떨어진 전경에 배치하고 탁자를 뒤쪽으로 옮겼습니다."
        ],
        "physics": "머리와 목, 어깨의 연결은 자연스럽습니다. 하체와 좌판 접촉은 클로즈업 밖이므로 지지 상태를 확정할 수 없지만 공중에 떠 있다는 증거는 없습니다. 보이는 가구는 바닥에 서 있고 서류철은 탁자에 놓여 있습니다."
       },
       {
        "label": "B",
        "direction": "고개와 두 눈이 화면 오른쪽의 보이지 않는 상대를 향합니다. 눈꺼풀을 크게 열고 입을 조금 벌린 채 굳은 표정이라 말을 듣고 놀라는 반응으로 읽힙니다. 손은 가슴 아래에서 구부러져 있으며 특정 대상을 가리키지는 않습니다.",
        "built_space": "참고와 같은 밝은 상부 벽과 회색 하부 벽, 전경의 금속 테두리 탁자 한 개, 오른쪽의 낡은 목제 의자 등받이 한 개가 보입니다. 여성은 탁자 너머에 자리하며 심문 자리의 관계가 성립합니다. 나머지 의자와 천장 조명, 문은 화면 밖이므로 수량이나 문 개방 상태를 확인할 수 없습니다. 얼굴뿐 아니라 상체와 손, 탁자까지 보여 요구된 클로즈업보다 넓습니다.",
        "entities": "검은 단발머리의 중년 한국인 여성 한 명이며 얼굴 윤곽과 남색 둥근목 상의가 참고 인물에 가깝습니다. 양쪽 눈 밑과 볼에 멍과 뚜렷한 부기가 있고, 눈은 정상적인 홍채와 동공을 유지한 채 크게 떠 있습니다. 열린 서류철은 참고 장면의 소품과 대응하며 판독 가능한 글자는 없습니다. 다른 사람이나 제외 대상인 전화기, 벽걸이 칼은 없습니다.",
        "hard_violations": [],
        "physics": "구부린 손은 팔과 자연스럽게 이어지고 손목과 팔 아래쪽은 탁자 가장자리 부근에 놓여 있습니다. 앉은 하체와 좌판의 접촉은 가려져 있지만 상체 자세에 부유나 불가능한 관절의 징후는 없습니다. 서류철은 탁자 면에 지지되며 종이의 처짐도 자연스럽습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "얼굴 클로즈업과 중간 왼쪽 배치는 잘 맞지만, 의자를 세 개로 늘리고 심문 탁자를 인물 뒤로 옮겨 장소의 연속성을 깨뜨렸습니다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "부어오른 얼굴과 화면 밖 상대를 향한 놀란 눈, 참고 의상과 심문실 재질은 충실하지만, 손과 탁자까지 담아 요구한 얼굴 클로즈업보다 넓습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴과 두 눈이 거의 정면의 카메라를 향합니다. 충격으로 눈을 크게 뜬 모습은 보이지만, 말을 들려준 화면 밖 상대를 바라보는 관계는 드러나지 않습니다. 방향을 판정할 무기나 휴대 물체는 없습니다.",
        "built_space": "밝은 상부 벽과 회색 하부 벽, 천장 조명 한 개, 환기구 한 개, 오른쪽 뒤 금속 탁자 한 개가 보입니다. 목제 의자 등받이는 인물 바로 뒤와 그보다 먼 중앙, 오른쪽 탁자 뒤에 각각 하나씩 총 세 개가 보이며, 참고 장면의 두 의자와 다릅니다. 인물이 탁자 앞 심문 자리라기보다 탁자에서 떨어진 전경에 놓여 있습니다. 문은 화면에 없어 개방 여부를 확인할 수 없습니다.",
        "entities": "중년 한국인 여성으로 보이는 인물 한 명만 있습니다. 검은 머리와 얼굴 윤곽은 참고 인물과 대체로 닮았지만 머리가 더 길고 흐트러져 있으며, 보이는 상의는 참고의 남색보다 짙은 회흑색에 가깝습니다. 볼과 눈가에 멍이 있으나 심하게 부었다는 상태는 비교적 약합니다. 탁자 위 열린 서류철에는 읽을 수 있는 글자가 없고, 전화기나 벽걸이 칼은 없습니다.",
        "hard_violations": [
         "고정된 심문실의 의자를 참고 장면의 두 개에서 세 개로 늘려 가구를 추가했습니다.",
         "피구금자의 심문 좌석이어야 하는 인물을 심문 탁자에서 떨어진 전경에 배치하고 탁자를 뒤쪽으로 옮겼습니다."
        ],
        "physics": "머리와 목, 어깨의 연결은 자연스럽습니다. 하체와 좌판 접촉은 클로즈업 밖이므로 지지 상태를 확정할 수 없지만 공중에 떠 있다는 증거는 없습니다. 보이는 가구는 바닥에 서 있고 서류철은 탁자에 놓여 있습니다."
       },
       {
        "label": "A",
        "direction": "고개와 두 눈이 화면 오른쪽의 보이지 않는 상대를 향합니다. 눈꺼풀을 크게 열고 입을 조금 벌린 채 굳은 표정이라 말을 듣고 놀라는 반응으로 읽힙니다. 손은 가슴 아래에서 구부러져 있으며 특정 대상을 가리키지는 않습니다.",
        "built_space": "참고와 같은 밝은 상부 벽과 회색 하부 벽, 전경의 금속 테두리 탁자 한 개, 오른쪽의 낡은 목제 의자 등받이 한 개가 보입니다. 여성은 탁자 너머에 자리하며 심문 자리의 관계가 성립합니다. 나머지 의자와 천장 조명, 문은 화면 밖이므로 수량이나 문 개방 상태를 확인할 수 없습니다. 얼굴뿐 아니라 상체와 손, 탁자까지 보여 요구된 클로즈업보다 넓습니다.",
        "entities": "검은 단발머리의 중년 한국인 여성 한 명이며 얼굴 윤곽과 남색 둥근목 상의가 참고 인물에 가깝습니다. 양쪽 눈 밑과 볼에 멍과 뚜렷한 부기가 있고, 눈은 정상적인 홍채와 동공을 유지한 채 크게 떠 있습니다. 열린 서류철은 참고 장면의 소품과 대응하며 판독 가능한 글자는 없습니다. 다른 사람이나 제외 대상인 전화기, 벽걸이 칼은 없습니다.",
        "hard_violations": [],
        "physics": "구부린 손은 팔과 자연스럽게 이어지고 손목과 팔 아래쪽은 탁자 가장자리 부근에 놓여 있습니다. 앉은 하체와 좌판의 접촉은 가려져 있지만 상체 자세에 부유나 불가능한 관절의 징후는 없습니다. 서류철은 탁자 면에 지지되며 종이의 처짐도 자연스럽습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.0
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.75
   },
   "violations": {
    "B": [
     "[gemini-pro] 지정된 취조실 좌석이 아닌 책상과 멀리 떨어진 위치에 인물이 배치됨",
     "[gpt-high] 고정된 심문실의 의자를 참고 장면의 두 개에서 세 개로 늘려 가구를 추가했습니다.",
     "[gpt-high] 피구금자의 심문 좌석이어야 하는 인물을 심문 탁자에서 떨어진 전경에 배치하고 탁자를 뒤쪽으로 옮겼습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 750
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "프레임 중좌측 배치와 놀란 표정, 부은 얼굴 등 프롬프트의 구도 및 묘사 지시를 훌륭히 따랐습니다."
   },
   {
    "label": "B",
    "score": 750,
    "verdict_ko": "인물을 프레임 중앙에 배치하고 지정된 좌석이 아닌 공간에 세워두어 구도와 장소 설정 지시를 위반했습니다.  ★위반: [gemini-pro] 지정된 취조실 좌석이 아닌 책상과 멀리 떨어진 위치에 인물이 배치됨 / [gpt-high] 고정된 심문실의 의자를 참고 장면의 두 개에서 세 개로 늘려 가구를 추가했습니다. / [gpt-high] 피구금자의 심문 좌석이어야 하는 인물을 심문 탁자에서 떨어진 전경에 배치하고 탁자를 뒤쪽으로 옮겼습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S31sh8_sel.png",
    "asset_id": "6984b446-c0be-40f5-a67e-c4a7caaa91e2",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1113064>",
    "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-54f3-7094-ad98-9a441cf70e6e",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S31sh8"
  }
 },
 "S31sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:44:26.109041+00:00",
  "fingerprint": "5455582fa9d64d4a8409f9061aef1703d5f8d405cdcea8229332a3781858922b",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S31sh9_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S31sh9_sel.png",
  "source_sha256": "b2030a6f102337a412e0683be00e1712cc71785e9b8320f1437d1c28ae63b5b4",
  "file": "S31sh9_cine.png",
  "staged_sha256": "2f4fb8b1e22140c3394a995f0f5473e5c2b607e4303fa9168ecc92149de0101c",
  "latency_ms": 9838
 },
 "S32sh5::signage": {
  "fp": "920671277827e0c3",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S32sh5": {
  "input_fingerprint": "a177022f4009a5df",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 구도환을 향해 고개를 한쪽으로 강하게 꺾은 채 굳은 표정으로 거절하는 현우의 단호한 얼굴.\n\nLOCATION (lock): Near the table and wall inside the church basement prayer room, in the lantern light established earlier. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Basement wall (Behind 현우, supporting his lean); used as A narrow visible area anchors his resistant posture without introducing additional detail.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral ambient illumination appropriate to the basement, with controlled contrast keeping his firm expression readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the basement prayer room's walls and sparse fixed furnishings, including the cross and small organ. Exclude the upstairs entrance and exterior surveillance lights.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement lanterns remain lit, with the small table, chairs, organ and cross unchanged. Charlie now holds the functioning radio and still wears his old coat and hat. 현우: He stands against the wall with folded arms, facial bruises and an injured leg, without his outer garment. The contact card remains concealed in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 구도환을 향해 고개를 한쪽으로 강하게 꺾은 채 굳은 표정으로 거절하는 현우의 단호한 얼굴.\n\nLOCATION (lock): Near the table and wall inside the church basement prayer room, in the lantern light established earlier. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Basement wall (Behind 현우, supporting his lean); used as A narrow visible area anchors his resistant posture without introducing additional detail.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral ambient illumination appropriate to the basement, with controlled contrast keeping his firm expression readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the basement prayer room's walls and sparse fixed furnishings, including the cross and small organ. Exclude the upstairs entrance and exterior surveillance lights.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement lanterns remain lit, with the small table, chairs, organ and cross unchanged. Charlie now holds the functioning radio and still wears his old coat and hat. 현우: He stands against the wall with folded arms, facial bruises and an injured leg, without his outer garment. The contact card remains concealed in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 구도환을 향해 고개를 한쪽으로 강하게 꺾은 채 굳은 표정으로 거절하는 현우의 단호한 얼굴.\n\nLOCATION (lock): Near the table and wall inside the church basement prayer room, in the lantern light established earlier. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Basement wall (Behind 현우, supporting his lean); used as A narrow visible area anchors his resistant posture without introducing additional detail.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral ambient illumination appropriate to the basement, with controlled contrast keeping his firm expression readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the basement prayer room's walls and sparse fixed furnishings, including the cross and small organ. Exclude the upstairs entrance and exterior surveillance lights.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement lanterns remain lit, with the small table, chairs, organ and cross unchanged. Charlie now holds the functioning radio and still wears his old coat and hat. 현우: He stands against the wall with folded arms, facial bruises and an injured leg, without his outer garment. The contact card remains concealed in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "시선과 고개가 카메라 밖의 타겟을 향하고 있음.",
    "built_space": "인물 뒤로 콘크리트 벽면만 배경으로 보임.",
    "entities": "현우의 얼굴이 일치하며, 멍 자국과 겉옷을 입지 않은 상태가 지시와 부합함.",
    "hard_violations": [],
    "physics": "벽에 몸을 기대고 팔짱을 낀 자연스러운 자세."
   },
   {
    "label": "B",
    "direction": "카메라 렌즈를 정면으로 응시하고 있음.",
    "built_space": "콘크리트 벽면에 십자가가 있으나, 천장 랜턴이 벽 부착형으로 변형됨.",
    "entities": "현우의 얼굴과 멍 자국은 일치하나, 지시와 달리 재킷(겉옷)을 입고 있음.",
    "hard_violations": [],
    "physics": "벽에 기대어 팔짱을 낀 안정적인 자세."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지시된 대로 겉옷을 벗은 차림과 화면 밖 대상을 향한 단호한 시선을 잘 구현했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "겉옷을 입지 말라는 지시를 위반하여 재킷을 착용했으며, 랜턴의 부착 위치가 벽면으로 왜곡되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선과 고개가 카메라 밖의 타겟을 향하고 있음.",
        "built_space": "인물 뒤로 콘크리트 벽면만 배경으로 보임.",
        "entities": "현우의 얼굴이 일치하며, 멍 자국과 겉옷을 입지 않은 상태가 지시와 부합함.",
        "hard_violations": [],
        "physics": "벽에 몸을 기대고 팔짱을 낀 자연스러운 자세."
       },
       {
        "label": "B",
        "direction": "카메라 렌즈를 정면으로 응시하고 있음.",
        "built_space": "콘크리트 벽면에 십자가가 있으나, 천장 랜턴이 벽 부착형으로 변형됨.",
        "entities": "현우의 얼굴과 멍 자국은 일치하나, 지시와 달리 재킷(겉옷)을 입고 있음.",
        "hard_violations": [],
        "physics": "벽에 기대어 팔짱을 낀 안정적인 자세."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지시된 대로 겉옷을 벗은 차림과 화면 밖 대상을 향한 단호한 시선을 잘 구현했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "겉옷을 입지 말라는 지시를 위반하여 재킷을 착용했으며, 랜턴의 부착 위치가 벽면으로 왜곡되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선과 고개가 카메라 밖의 타겟을 향하고 있음.",
        "built_space": "인물 뒤로 콘크리트 벽면만 배경으로 보임.",
        "entities": "현우의 얼굴이 일치하며, 멍 자국과 겉옷을 입지 않은 상태가 지시와 부합함.",
        "hard_violations": [],
        "physics": "벽에 몸을 기대고 팔짱을 낀 자연스러운 자세."
       },
       {
        "label": "B",
        "direction": "카메라 렌즈를 정면으로 응시하고 있음.",
        "built_space": "콘크리트 벽면에 십자가가 있으나, 천장 랜턴이 벽 부착형으로 변형됨.",
        "entities": "현우의 얼굴과 멍 자국은 일치하나, 지시와 달리 재킷(겉옷)을 입고 있음.",
        "hard_violations": [],
        "physics": "벽에 기대어 팔짱을 낀 안정적인 자세."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "얼굴 중심의 밀착 구도와 화면 밖 상대를 향한 시선, 벽에 기댄 팔짱 자세는 더 충실하지만, 고개를 한쪽으로 강하게 꺾어 거절하는 핵심 동작은 부족하다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "벽에 기댄 부상자의 굳은 표정은 맞지만, 배경을 넓게 드러내고 겉옷을 입혔으며 고개를 강하게 꺾는 거절 동작도 약하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴은 화면 왼쪽으로 조금 돌아가 있지만 눈은 카메라 가까운 왼쪽을 곁눈질한다. 구도환은 화면 밖에 있어 실제 시선 도착점은 확인되지 않으며, 상대를 향한 강한 고개 꺾기보다는 정면에 가까운 경계로 보인다. 무기나 방향성 있는 소품은 없다.",
        "built_space": "현우의 등과 뒤통수 뒤에 거친 콘크리트 벽이 있고, 왼쪽 배경에 켜진 랜턴 하나, 십자가 하나, 목제 가구 일부와 그 위의 모자 하나가 보인다. 가구는 흐려 오르간인지 확정하기 어렵다. 벽과 기존 소품의 재질은 이전 장면에 대체로 맞지만, 좁은 벽만 남기는 얼굴 중심 구도보다 배경 설명이 많고 랜턴도 눈에 띄게 크다. 중복 설비나 반사는 없다.",
        "entities": "한 명의 앳된 동아시아계 남성이 보이며 검은 헝클어진 머리와 얼굴 형태는 현우 참고 이미지에 대체로 가깝다. 한국계 미국인이라는 국적 배경은 외형만으로 확인할 수 없다. 얼굴의 멍과 찰과상, 팔짱은 보이나 회갈색 지퍼 겉옷을 입어 겉옷을 벗었다는 조건과 다르다. 다리 부상과 신발 속 카드는 프레임 밖이다. 다른 인물과 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "등과 어깨가 벽에 닿아 기대는 자세를 지지한다. 접힌 팔과 손의 접촉도 자연스럽고 머리는 목에 정상적으로 지지된다. 발은 잘려 있어 서 있는 자세의 하체 지지는 확인할 수 없지만, 공중에 뜬 몸이라는 증거는 없다. 랜턴은 위쪽 현수부가 일부 잘려 있고 목제 가구 위 모자는 표면에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "얼굴과 두 눈이 화면 왼쪽의 프레임 밖 상대를 향해 있어 구도환을 보는 설정에 더 잘 맞는다. 다만 고개는 거의 세워진 채 약간 돌아가 있을 뿐, 한쪽으로 강하게 꺾인 순간은 아니다. 다문 입과 굳은 눈매는 단호함을 전달하지만 거절 동작 자체는 약하다.",
        "built_space": "뒤통수와 오른쪽 어깨 뒤의 거친 콘크리트 벽이 주된 배경이며 왼쪽 가장자리에 어두운 구조물 일부만 걸린다. 랜턴, 십자가, 오르간, 탁자와 의자는 이 밀착 구도에서 식별되지 않으며, 이를 없어진 것으로 판단할 근거는 없다. 벽의 재질과 인물의 기대는 위치가 이전 지하실 설정에 부합한다. 중복 설비나 반사는 없다.",
        "entities": "한 명의 앳된 동아시아계 남성이며 헝클어진 검은 머리, 얼굴 윤곽과 이목구비가 현우 참고 이미지에 비교적 가깝다. 얼굴의 멍과 찰과상, 팔짱이 보이고 별도의 두꺼운 겉옷은 보이지 않는다. 다만 검은 칼라 셔츠는 참고 이미지의 남색 둥근 목 티셔츠와 다르다. 다리와 신발 속 카드는 구도 밖이며 다른 인물, 무전기, 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "뒤통수와 등·어깨가 벽에 맞닿고, 양팔은 몸 앞에서 접혀 서로 지지된다. 손가락과 소매의 접촉도 가능한 형태다. 하체는 프레임 밖이므로 발의 접지나 부상 상태는 판단할 수 없으나, 보이는 상체에 지지 없는 부유나 불가능한 관절 자세는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "얼굴 중심의 밀착 구도와 화면 밖 상대를 향한 시선, 벽에 기댄 팔짱 자세는 더 충실하지만, 고개를 한쪽으로 강하게 꺾어 거절하는 핵심 동작은 부족하다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "벽에 기댄 부상자의 굳은 표정은 맞지만, 배경을 넓게 드러내고 겉옷을 입혔으며 고개를 강하게 꺾는 거절 동작도 약하다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴은 화면 왼쪽으로 조금 돌아가 있지만 눈은 카메라 가까운 왼쪽을 곁눈질한다. 구도환은 화면 밖에 있어 실제 시선 도착점은 확인되지 않으며, 상대를 향한 강한 고개 꺾기보다는 정면에 가까운 경계로 보인다. 무기나 방향성 있는 소품은 없다.",
        "built_space": "현우의 등과 뒤통수 뒤에 거친 콘크리트 벽이 있고, 왼쪽 배경에 켜진 랜턴 하나, 십자가 하나, 목제 가구 일부와 그 위의 모자 하나가 보인다. 가구는 흐려 오르간인지 확정하기 어렵다. 벽과 기존 소품의 재질은 이전 장면에 대체로 맞지만, 좁은 벽만 남기는 얼굴 중심 구도보다 배경 설명이 많고 랜턴도 눈에 띄게 크다. 중복 설비나 반사는 없다.",
        "entities": "한 명의 앳된 동아시아계 남성이 보이며 검은 헝클어진 머리와 얼굴 형태는 현우 참고 이미지에 대체로 가깝다. 한국계 미국인이라는 국적 배경은 외형만으로 확인할 수 없다. 얼굴의 멍과 찰과상, 팔짱은 보이나 회갈색 지퍼 겉옷을 입어 겉옷을 벗었다는 조건과 다르다. 다리 부상과 신발 속 카드는 프레임 밖이다. 다른 인물과 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "등과 어깨가 벽에 닿아 기대는 자세를 지지한다. 접힌 팔과 손의 접촉도 자연스럽고 머리는 목에 정상적으로 지지된다. 발은 잘려 있어 서 있는 자세의 하체 지지는 확인할 수 없지만, 공중에 뜬 몸이라는 증거는 없다. 랜턴은 위쪽 현수부가 일부 잘려 있고 목제 가구 위 모자는 표면에 놓여 있다."
       },
       {
        "label": "A",
        "direction": "얼굴과 두 눈이 화면 왼쪽의 프레임 밖 상대를 향해 있어 구도환을 보는 설정에 더 잘 맞는다. 다만 고개는 거의 세워진 채 약간 돌아가 있을 뿐, 한쪽으로 강하게 꺾인 순간은 아니다. 다문 입과 굳은 눈매는 단호함을 전달하지만 거절 동작 자체는 약하다.",
        "built_space": "뒤통수와 오른쪽 어깨 뒤의 거친 콘크리트 벽이 주된 배경이며 왼쪽 가장자리에 어두운 구조물 일부만 걸린다. 랜턴, 십자가, 오르간, 탁자와 의자는 이 밀착 구도에서 식별되지 않으며, 이를 없어진 것으로 판단할 근거는 없다. 벽의 재질과 인물의 기대는 위치가 이전 지하실 설정에 부합한다. 중복 설비나 반사는 없다.",
        "entities": "한 명의 앳된 동아시아계 남성이며 헝클어진 검은 머리, 얼굴 윤곽과 이목구비가 현우 참고 이미지에 비교적 가깝다. 얼굴의 멍과 찰과상, 팔짱이 보이고 별도의 두꺼운 겉옷은 보이지 않는다. 다만 검은 칼라 셔츠는 참고 이미지의 남색 둥근 목 티셔츠와 다르다. 다리와 신발 속 카드는 구도 밖이며 다른 인물, 무전기, 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "뒤통수와 등·어깨가 벽에 맞닿고, 양팔은 몸 앞에서 접혀 서로 지지된다. 손가락과 소매의 접촉도 가능한 형태다. 하체는 프레임 밖이므로 발의 접지나 부상 상태는 판단할 수 없으나, 보이는 상체에 지지 없는 부유나 불가능한 관절 자세는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.286
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.286
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1286
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지시된 대로 겉옷을 벗은 차림과 화면 밖 대상을 향한 단호한 시선을 잘 구현했습니다."
   },
   {
    "label": "B",
    "score": 1286,
    "verdict_ko": "겉옷을 입지 말라는 지시를 위반하여 재킷을 착용했으며, 랜턴의 부착 위치가 벽면으로 왜곡되었습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S30sh17_sel.png",
    "asset_id": "b541fcbe-926f-4ef3-b778-3535e2331d78",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-56ac-76bc-a899-ae7440595f20",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S30sh17"
  }
 },
 "S32sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:23:40.446046+00:00",
  "fingerprint": "fb24efdd2ae913cc66b942114f8130bbcd80e9cf524e28d213121cb2e035053f",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S32sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S32sh5_sel.png",
  "source_sha256": "0aec9fd4f9860ac3f7e0da649ac70545608811c74ec99174f6eab3f0f5010600",
  "file": "S32sh5_cine.png",
  "staged_sha256": "feb28b38b5b4e355762b0fec886049124a9cae47547299555364ee9242dbe704",
  "latency_ms": 9380
 },
 "S32sh8::signage": {
  "fp": "9854909145f5a1d7",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S32sh8": {
  "input_fingerprint": "14046b8125b23cf0",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 낡은 손전등을 꽉 쥔 채 지하실 구석의 비밀 통로 쪽을 검지손가락으로 가리키는 신부의 역동적인 상체.\n\nLOCATION (lock): At the concealed passage entrance in a corner of the church basement prayer room, lit by lanterns and a carried flashlight. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: 신부 in the middle-left of the frame, midground, points to secret passage approach; secret passage approach in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Secret passage (The escape route indicated by 신부) — The approach to the passage is seen obliquely beyond his pointing hand; its continuation lies outside the frame; used as Provides a spatial destination for the gesture without specifying an unsupported door mechanism; Old flashlight (Gripped tightly by 신부) — Seen side-on near his lower torso; used as A small practical object supporting the urgency of departure, kept subordinate to the pointing hand.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain neutral basement ambient illumination and readable hand contours without assuming that the flashlight has been switched on.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement furnishings and lit lanterns remain unchanged, with a secret exit passage available. Charlie retains the radio, old coat and hat. 신부: He has returned with the warning and carries a flashlight as he leads toward the secret passage, still wearing his clerical collar.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 신부 right now, so 신부's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 신부: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 낡은 손전등을 꽉 쥔 채 지하실 구석의 비밀 통로 쪽을 검지손가락으로 가리키는 신부의 역동적인 상체.\n\nLOCATION (lock): At the concealed passage entrance in a corner of the church basement prayer room, lit by lanterns and a carried flashlight. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: 신부 in the middle-left of the frame, midground, points to secret passage approach; secret passage approach in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Secret passage (The escape route indicated by 신부) — The approach to the passage is seen obliquely beyond his pointing hand; its continuation lies outside the frame; used as Provides a spatial destination for the gesture without specifying an unsupported door mechanism; Old flashlight (Gripped tightly by 신부) — Seen side-on near his lower torso; used as A small practical object supporting the urgency of departure, kept subordinate to the pointing hand.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain neutral basement ambient illumination and readable hand contours without assuming that the flashlight has been switched on.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement furnishings and lit lanterns remain unchanged, with a secret exit passage available. Charlie retains the radio, old coat and hat. 신부: He has returned with the warning and carries a flashlight as he leads toward the secret passage, still wearing his clerical collar.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 신부 right now, so 신부's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 신부: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 낡은 손전등을 꽉 쥔 채 지하실 구석의 비밀 통로 쪽을 검지손가락으로 가리키는 신부의 역동적인 상체.\n\nLOCATION (lock): At the concealed passage entrance in a corner of the church basement prayer room, lit by lanterns and a carried flashlight. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: 신부 in the middle-left of the frame, midground, points to secret passage approach; secret passage approach in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Secret passage (The escape route indicated by 신부) — The approach to the passage is seen obliquely beyond his pointing hand; its continuation lies outside the frame; used as Provides a spatial destination for the gesture without specifying an unsupported door mechanism; Old flashlight (Gripped tightly by 신부) — Seen side-on near his lower torso; used as A small practical object supporting the urgency of departure, kept subordinate to the pointing hand.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain neutral basement ambient illumination and readable hand contours without assuming that the flashlight has been switched on.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The basement furnishings and lit lanterns remain unchanged, with a secret exit passage available. Charlie retains the radio, old coat and hat. 신부: He has returned with the warning and carries a flashlight as he leads toward the secret passage, still wearing his clerical collar.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 신부 right now, so 신부's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 신부: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "신부의 시선은 화면 밖 왼쪽을 향하고 있으며, 왼손 검지손가락은 우측 배경의 어두운 통로 입구를 가리키고 있습니다.",
    "built_space": "이전 샷과 일치하는 둥근 홈이 파인 노출 콘크리트 벽면이 있습니다. 화면 우측에는 커튼이 쳐진 통로가 있고, 종이 상자들이 쌓여 있습니다. 좌측에는 랜턴이 매달려 있습니다.",
    "entities": "신부는 레퍼런스와 얼굴, 연령이 일치하며 클레릭 칼라가 있는 검은 셔츠를 입고 있습니다. 오른손에는 낡은 손전등을 쥐고 있습니다.",
    "hard_violations": [
     "[gpt-high] 오른쪽 상자에 읽을 수 있는 한글 표기가 보여, 어디에도 읽을 수 있는 문자를 두지 말라는 명시적 조건을 위반한다."
    ],
    "physics": "오른손으로 손전등을 자연스럽게 쥐고 있으며, 왼팔을 들어 올린 자세가 물리적으로 안정적입니다."
   },
   {
    "label": "B",
    "direction": "신부의 시선은 정면 약간 왼쪽을 향하고 있으며, 왼손 검지손가락으로 우측의 돌벽으로 된 통로를 가리키고 있습니다.",
    "built_space": "나무틀이 있는 문과 거친 돌을 쌓아 만든 벽면이 보이며, 우측 통로 안쪽에는 나무 지지대가 있습니다. 랜턴이 우측 벽에 걸려 있습니다.",
    "entities": "신부는 레퍼런스의 인물과 일치하며 사제복을 입고 손전등을 들고 있습니다.",
    "hard_violations": [
     "[gemini-pro] 장소 고정(Location lock) 위반: 레퍼런스에 제시된 노출 콘크리트 벽 대신 거친 돌벽을 렌더링함."
    ],
    "physics": "손전등을 쥔 오른손과 통로를 가리키는 왼팔의 자세 및 중력 지탱이 자연스럽게 표현되었습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "이전 샷의 콘크리트 벽면 질감을 정확히 유지하며 프롬프트의 지시대로 신부가 손전등을 쥔 채 통로를 가리키는 모습을 잘 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "인물의 행동과 묘사는 좋으나, 이전 샷에서 확립된 콘크리트 벽면 대신 돌벽을 렌더링하여 장소 설정(Location Lock)을 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "신부의 시선은 화면 밖 왼쪽을 향하고 있으며, 왼손 검지손가락은 우측 배경의 어두운 통로 입구를 가리키고 있습니다.",
        "built_space": "이전 샷과 일치하는 둥근 홈이 파인 노출 콘크리트 벽면이 있습니다. 화면 우측에는 커튼이 쳐진 통로가 있고, 종이 상자들이 쌓여 있습니다. 좌측에는 랜턴이 매달려 있습니다.",
        "entities": "신부는 레퍼런스와 얼굴, 연령이 일치하며 클레릭 칼라가 있는 검은 셔츠를 입고 있습니다. 오른손에는 낡은 손전등을 쥐고 있습니다.",
        "hard_violations": [],
        "physics": "오른손으로 손전등을 자연스럽게 쥐고 있으며, 왼팔을 들어 올린 자세가 물리적으로 안정적입니다."
       },
       {
        "label": "B",
        "direction": "신부의 시선은 정면 약간 왼쪽을 향하고 있으며, 왼손 검지손가락으로 우측의 돌벽으로 된 통로를 가리키고 있습니다.",
        "built_space": "나무틀이 있는 문과 거친 돌을 쌓아 만든 벽면이 보이며, 우측 통로 안쪽에는 나무 지지대가 있습니다. 랜턴이 우측 벽에 걸려 있습니다.",
        "entities": "신부는 레퍼런스의 인물과 일치하며 사제복을 입고 손전등을 들고 있습니다.",
        "hard_violations": [
         "장소 고정(Location lock) 위반: 레퍼런스에 제시된 노출 콘크리트 벽 대신 거친 돌벽을 렌더링함."
        ],
        "physics": "손전등을 쥔 오른손과 통로를 가리키는 왼팔의 자세 및 중력 지탱이 자연스럽게 표현되었습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "이전 샷의 콘크리트 벽면 질감을 정확히 유지하며 프롬프트의 지시대로 신부가 손전등을 쥔 채 통로를 가리키는 모습을 잘 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "인물의 행동과 묘사는 좋으나, 이전 샷에서 확립된 콘크리트 벽면 대신 돌벽을 렌더링하여 장소 설정(Location Lock)을 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "신부의 시선은 화면 밖 왼쪽을 향하고 있으며, 왼손 검지손가락은 우측 배경의 어두운 통로 입구를 가리키고 있습니다.",
        "built_space": "이전 샷과 일치하는 둥근 홈이 파인 노출 콘크리트 벽면이 있습니다. 화면 우측에는 커튼이 쳐진 통로가 있고, 종이 상자들이 쌓여 있습니다. 좌측에는 랜턴이 매달려 있습니다.",
        "entities": "신부는 레퍼런스와 얼굴, 연령이 일치하며 클레릭 칼라가 있는 검은 셔츠를 입고 있습니다. 오른손에는 낡은 손전등을 쥐고 있습니다.",
        "hard_violations": [],
        "physics": "오른손으로 손전등을 자연스럽게 쥐고 있으며, 왼팔을 들어 올린 자세가 물리적으로 안정적입니다."
       },
       {
        "label": "B",
        "direction": "신부의 시선은 정면 약간 왼쪽을 향하고 있으며, 왼손 검지손가락으로 우측의 돌벽으로 된 통로를 가리키고 있습니다.",
        "built_space": "나무틀이 있는 문과 거친 돌을 쌓아 만든 벽면이 보이며, 우측 통로 안쪽에는 나무 지지대가 있습니다. 랜턴이 우측 벽에 걸려 있습니다.",
        "entities": "신부는 레퍼런스의 인물과 일치하며 사제복을 입고 손전등을 들고 있습니다.",
        "hard_violations": [
         "장소 고정(Location lock) 위반: 레퍼런스에 제시된 노출 콘크리트 벽 대신 거친 돌벽을 렌더링함."
        ],
        "physics": "손전등을 쥔 오른손과 통로를 가리키는 왼팔의 자세 및 중력 지탱이 자연스럽게 표현되었습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "중간 왼쪽 신부의 상체와 오른쪽 입구를 가리키는 손, 허리의 손전등은 지시대로지만, 석조·목재 공간과 판자로 가려진 입구는 장소 및 탈출 통로 설정에 덜 충실하다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "콘크리트 지하실과 열린 통로, 역동적인 상체 배치는 더 충실하지만, 오른쪽 상자에 읽을 수 있는 한글 표기가 노출되어 명시적인 문자 금지 조건을 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "뻗은 검지는 오른쪽 배경의 목재 테두리 입구 안쪽 판자를 향한다. 입구라는 목적지는 분명하지만, 손끝 너머가 열린 이동 경로보다는 판자와 돌로 막힌 공간처럼 보인다. 신부의 눈은 통로가 아닌 카메라 쪽을 향한다. 손전등의 앞부분은 오른쪽으로 향하며 뚜렷한 조사광은 없다.",
        "built_space": "신부는 중간 왼쪽에 허리 부근까지 보이고, 오른쪽에는 목재 테두리 입구 하나가 있다. 켜진 등불은 신부 뒤와 오른쪽 가장자리에 하나씩, 총 두 개 보인다. 입구 안에는 비스듬한 판자 여러 개와 돌벽이 있으며 왼쪽 가장자리에는 목재 가구 일부가 보인다. 참고 장소의 거푸집 흔적이 있는 콘크리트 벽과 달리 거친 석재와 목재가 공간을 지배하고, 통로의 실제 접근 가능성도 불분명하다.",
        "entities": "등장인물은 신부 한 명뿐이다. 회색이 섞인 뒤로 넘긴 머리, 이마와 눈가 주름, 얼굴 윤곽은 참고 인물과 대체로 일치하는 한국인 60대 남성으로 보인다. 검은 성직자 옷과 흰 성직 칼라가 있으며, 아래쪽 손에는 낡은 금속 손전등 하나가 옆면으로 보인다. 추가 인물이나 읽을 수 있는 문자는 없다.",
        "hard_violations": [],
        "physics": "손전등 손잡이는 허리 앞의 손가락과 손바닥에 잡혀 있다. 반대쪽 팔은 어깨에서 자연스럽게 이어져 입구를 가리키며 상체 회전도 가능한 자세다. 하체는 프레임 밖이지만 몸이 떠 있다는 징후는 없다. 등불은 벽 쪽 부착부에 걸려 있고 판자는 벽과 입구에 기대어 있어 지지 없는 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "검지는 오른쪽 배경의 어두운 통로 개구부 안쪽을 향해, 손 너머의 탈출 방향을 명확하게 지정한다. 신부는 고개와 눈을 화면 왼쪽으로 돌려 통로 반대편을 보고 있다. 허리 앞 손전등은 오른쪽을 향하고 있으며 켜진 광선은 보이지 않는다.",
        "built_space": "신부는 중간 왼쪽에 허리까지 보이고, 오른쪽 배경에는 열린 입구 하나가 있다. 입구 오른쪽에는 접힌 주름 형태의 패널이 보이며 내부의 연속 공간은 어둠에 가려진다. 위쪽 왼편에는 등불 두 개, 뒤쪽 아래에는 낮은 목재 가구, 오른쪽에는 상자 세 개가 보인다. 벽의 콘크리트 질감과 이음선은 장소 참고에 A보다 가깝고, 신부와 입구의 위치도 요구한 배치에 맞는다.",
        "entities": "등장인물은 신부 한 명이다. 나이 든 한국인 남성의 얼굴, 희끗한 머리, 이마 주름과 수염 자국이 참고 인물에 대체로 부합한다. 어두운 성직자 셔츠와 흰 칼라를 착용했고, 낡은 손전등 하나를 허리 앞에서 쥔다. 오른쪽 중간 상자의 인쇄된 표에는 판독 가능한 한글 항목들이 노출되어 있다.",
        "hard_violations": [
         "오른쪽 상자에 읽을 수 있는 한글 표기가 보여, 어디에도 읽을 수 있는 문자를 두지 말라는 명시적 조건을 위반한다."
        ],
        "physics": "손전등은 손잡이를 감싼 손에 확실히 지지된다. 앞으로 기울인 상체와 굽힌 팔꿈치, 통로를 가리키는 손은 급히 안내하는 동작으로 가능한 연결이다. 발은 프레임 밖이며 공중에 떠 있는 자세로 보이지 않는다. 상자는 서로 포개져 지지되고, 통로 안 판자는 아래쪽에 기대어 있다. 등불의 상부 걸이 부분은 화면 위로 이어지거나 잘려 있어 무지지 부유로 볼 근거가 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "중간 왼쪽 신부의 상체와 오른쪽 입구를 가리키는 손, 허리의 손전등은 지시대로지만, 석조·목재 공간과 판자로 가려진 입구는 장소 및 탈출 통로 설정에 덜 충실하다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "콘크리트 지하실과 열린 통로, 역동적인 상체 배치는 더 충실하지만, 오른쪽 상자에 읽을 수 있는 한글 표기가 노출되어 명시적인 문자 금지 조건을 위반한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "뻗은 검지는 오른쪽 배경의 목재 테두리 입구 안쪽 판자를 향한다. 입구라는 목적지는 분명하지만, 손끝 너머가 열린 이동 경로보다는 판자와 돌로 막힌 공간처럼 보인다. 신부의 눈은 통로가 아닌 카메라 쪽을 향한다. 손전등의 앞부분은 오른쪽으로 향하며 뚜렷한 조사광은 없다.",
        "built_space": "신부는 중간 왼쪽에 허리 부근까지 보이고, 오른쪽에는 목재 테두리 입구 하나가 있다. 켜진 등불은 신부 뒤와 오른쪽 가장자리에 하나씩, 총 두 개 보인다. 입구 안에는 비스듬한 판자 여러 개와 돌벽이 있으며 왼쪽 가장자리에는 목재 가구 일부가 보인다. 참고 장소의 거푸집 흔적이 있는 콘크리트 벽과 달리 거친 석재와 목재가 공간을 지배하고, 통로의 실제 접근 가능성도 불분명하다.",
        "entities": "등장인물은 신부 한 명뿐이다. 회색이 섞인 뒤로 넘긴 머리, 이마와 눈가 주름, 얼굴 윤곽은 참고 인물과 대체로 일치하는 한국인 60대 남성으로 보인다. 검은 성직자 옷과 흰 성직 칼라가 있으며, 아래쪽 손에는 낡은 금속 손전등 하나가 옆면으로 보인다. 추가 인물이나 읽을 수 있는 문자는 없다.",
        "hard_violations": [],
        "physics": "손전등 손잡이는 허리 앞의 손가락과 손바닥에 잡혀 있다. 반대쪽 팔은 어깨에서 자연스럽게 이어져 입구를 가리키며 상체 회전도 가능한 자세다. 하체는 프레임 밖이지만 몸이 떠 있다는 징후는 없다. 등불은 벽 쪽 부착부에 걸려 있고 판자는 벽과 입구에 기대어 있어 지지 없는 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "검지는 오른쪽 배경의 어두운 통로 개구부 안쪽을 향해, 손 너머의 탈출 방향을 명확하게 지정한다. 신부는 고개와 눈을 화면 왼쪽으로 돌려 통로 반대편을 보고 있다. 허리 앞 손전등은 오른쪽을 향하고 있으며 켜진 광선은 보이지 않는다.",
        "built_space": "신부는 중간 왼쪽에 허리까지 보이고, 오른쪽 배경에는 열린 입구 하나가 있다. 입구 오른쪽에는 접힌 주름 형태의 패널이 보이며 내부의 연속 공간은 어둠에 가려진다. 위쪽 왼편에는 등불 두 개, 뒤쪽 아래에는 낮은 목재 가구, 오른쪽에는 상자 세 개가 보인다. 벽의 콘크리트 질감과 이음선은 장소 참고에 A보다 가깝고, 신부와 입구의 위치도 요구한 배치에 맞는다.",
        "entities": "등장인물은 신부 한 명이다. 나이 든 한국인 남성의 얼굴, 희끗한 머리, 이마 주름과 수염 자국이 참고 인물에 대체로 부합한다. 어두운 성직자 셔츠와 흰 칼라를 착용했고, 낡은 손전등 하나를 허리 앞에서 쥔다. 오른쪽 중간 상자의 인쇄된 표에는 판독 가능한 한글 항목들이 노출되어 있다.",
        "hard_violations": [
         "오른쪽 상자에 읽을 수 있는 한글 표기가 보여, 어디에도 읽을 수 있는 문자를 두지 말라는 명시적 조건을 위반한다."
        ],
        "physics": "손전등은 손잡이를 감싼 손에 확실히 지지된다. 앞으로 기울인 상체와 굽힌 팔꿈치, 통로를 가리키는 손은 급히 안내하는 동작으로 가능한 연결이다. 발은 프레임 밖이며 공중에 떠 있는 자세로 보이지 않는다. 상자는 서로 포개져 지지되고, 통로 안 판자는 아래쪽에 기대어 있다. 등불의 상부 걸이 부분은 화면 위로 이어지거나 잘려 있어 무지지 부유로 볼 근거가 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.667,
    "B": 1.375
   },
   "adjusted": {
    "A": 1.417,
    "B": 1.125
   },
   "violations": {
    "B": [
     "[gemini-pro] 장소 고정(Location lock) 위반: 레퍼런스에 제시된 노출 콘크리트 벽 대신 거친 돌벽을 렌더링함."
    ],
    "A": [
     "[gpt-high] 오른쪽 상자에 읽을 수 있는 한글 표기가 보여, 어디에도 읽을 수 있는 문자를 두지 말라는 명시적 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1417,
   "B": 1125
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1417,
    "verdict_ko": "이전 샷의 콘크리트 벽면 질감을 정확히 유지하며 프롬프트의 지시대로 신부가 손전등을 쥔 채 통로를 가리키는 모습을 잘 구현했습니다.  ★위반: [gpt-high] 오른쪽 상자에 읽을 수 있는 한글 표기가 보여, 어디에도 읽을 수 있는 문자를 두지 말라는 명시적 조건을 위반한다."
   },
   {
    "label": "B",
    "score": 1125,
    "verdict_ko": "인물의 행동과 묘사는 좋으나, 이전 샷에서 확립된 콘크리트 벽면 대신 돌벽을 렌더링하여 장소 설정(Location Lock)을 위반했습니다.  ★위반: [gemini-pro] 장소 고정(Location lock) 위반: 레퍼런스에 제시된 노출 콘크리트 벽 대신 거친 돌벽을 렌더링함."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S32sh5_sel.png",
    "asset_id": "c0685468-30eb-4fe0-852d-d7aab65eb708",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 신부: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1402213>",
    "asset_id": "8696070d-ac09-4a5c-95f2-3daf1c015c32",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-5864-7b55-bf8b-69ae39603c13",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S32sh5"
  }
 },
 "S32sh8::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:24:38.853923+00:00",
  "fingerprint": "f38dcc274e73d909aee5079ded94332ce18b770ec209ac1b72826d24281d7456",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S32sh8_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S32sh8_sel.png",
  "source_sha256": "588a7904212afa9835fd77617f208b1ccd406cbdfa7613a82176ab71ab13181c",
  "file": "S32sh8_cine.png",
  "staged_sha256": "16ee4730979c8295c36c7736b427536ee2071845cb79e6e23e3e8b5c02cac538",
  "latency_ms": 9974
 },
 "S32sh9::signage": {
  "fp": "9d486fc277c0b7f6",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S32sh9": {
  "input_fingerprint": "fa1145a575d83b4f",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 통로 쪽으로 향하려는 신부의 팔을 다급하게 꽉 붙잡은 현우의 굳은 손 클로즈업.\n\nLOCATION (lock): Just before the hidden passage entrance inside the church basement, within the prayer room's lantern-lit area. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: insert close-up on a detail\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding neutral ambient illumination, using controlled contrast to make the gripping fingers and arrested arm clearly legible.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 신부 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the basement corner and the concealed passage entrance visible in the reference. Exclude the upstairs front door and outdoor searchlights.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The lit basement and its secret exit passage remain unchanged. Charlie still has the radio and wears the old coat and hat. 현우: He is at the departure point, with facial bruises, an injured leg and no outer garment. The contact card remains concealed in his shoe. 신부: He pauses on his way toward the secret passage, holding the flashlight and wearing his clerical collar.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 통로 쪽으로 향하려는 신부의 팔을 다급하게 꽉 붙잡은 현우의 굳은 손 클로즈업.\n\nLOCATION (lock): Just before the hidden passage entrance inside the church basement, within the prayer room's lantern-lit area. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: insert close-up on a detail\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding neutral ambient illumination, using controlled contrast to make the gripping fingers and arrested arm clearly legible.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 신부 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the basement corner and the concealed passage entrance visible in the reference. Exclude the upstairs front door and outdoor searchlights.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The lit basement and its secret exit passage remain unchanged. Charlie still has the radio and wears the old coat and hat. 현우: He is at the departure point, with facial bruises, an injured leg and no outer garment. The contact card remains concealed in his shoe. 신부: He pauses on his way toward the secret passage, holding the flashlight and wearing his clerical collar.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 통로 쪽으로 향하려는 신부의 팔을 다급하게 꽉 붙잡은 현우의 굳은 손 클로즈업.\n\nLOCATION (lock): Just before the hidden passage entrance inside the church basement, within the prayer room's lantern-lit area. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: insert close-up on a detail\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding neutral ambient illumination, using controlled contrast to make the gripping fingers and arrested arm clearly legible.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 신부 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the basement corner and the concealed passage entrance visible in the reference. Exclude the upstairs front door and outdoor searchlights.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The lit basement and its secret exit passage remain unchanged. Charlie still has the radio and wears the old coat and hat. 현우: He is at the departure point, with facial bruises, an injured leg and no outer garment. The contact card remains concealed in his shoe. 신부: He pauses on his way toward the secret passage, holding the flashlight and wearing his clerical collar.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "손이 팔이 아닌 신부의 가슴 및 목깃 부위를 향해 있음.",
    "built_space": "지하실 벽, 랜턴 1개, 통로 입구 및 일부 상자가 보임.",
    "entities": "신부의 셔츠와 칼라, 현우의 맨팔이 보이나 손의 형태가 온전하지 않음.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 해부학 (여러 개의 손과 손가락이 기괴하게 융합됨)"
    ],
    "physics": "팔과 손이 비정상적으로 엉켜 있어 어떻게 지탱되고 잡고 있는지 파악할 수 없음."
   },
   {
    "label": "B",
    "direction": "맨손이 통로 쪽으로 뻗은 어두운 셔츠의 팔을 향해 꽉 쥐고 있음.",
    "built_space": "콘크리트 벽면, 랜턴 2개, 우측의 종이상자 더미, 커튼이 쳐진 통로 입구가 레퍼런스와 일치하게 배치됨.",
    "entities": "신부의 어두운 셔츠 소매와 손전등(아래쪽), 현우의 상처 입은 맨손이 묘사됨.",
    "hard_violations": [],
    "physics": "맨손이 옷 입은 팔을 자연스럽게 쥐고 있으며, 구조와 무게감이 물리적으로 타당함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "레퍼런스의 배경과 인물 복장을 잘 유지하면서 텍스트가 요구한 '팔을 붙잡는' 클로즈업 연출을 충실하고 자연스럽게 구현함."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "요구된 '팔' 대신 가슴 부위를 잡고 있으며, 손의 해부학적 구조가 붕괴되는 심각한 오류가 발생함."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "맨손이 통로 쪽으로 뻗은 어두운 셔츠의 팔을 향해 꽉 쥐고 있음.",
        "built_space": "콘크리트 벽면, 랜턴 2개, 우측의 종이상자 더미, 커튼이 쳐진 통로 입구가 레퍼런스와 일치하게 배치됨.",
        "entities": "신부의 어두운 셔츠 소매와 손전등(아래쪽), 현우의 상처 입은 맨손이 묘사됨.",
        "hard_violations": [],
        "physics": "맨손이 옷 입은 팔을 자연스럽게 쥐고 있으며, 구조와 무게감이 물리적으로 타당함."
       },
       {
        "label": "A",
        "direction": "손이 팔이 아닌 신부의 가슴 및 목깃 부위를 향해 있음.",
        "built_space": "지하실 벽, 랜턴 1개, 통로 입구 및 일부 상자가 보임.",
        "entities": "신부의 셔츠와 칼라, 현우의 맨팔이 보이나 손의 형태가 온전하지 않음.",
        "hard_violations": [
         "물리적으로 불가능한 해부학 (여러 개의 손과 손가락이 기괴하게 융합됨)"
        ],
        "physics": "팔과 손이 비정상적으로 엉켜 있어 어떻게 지탱되고 잡고 있는지 파악할 수 없음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "레퍼런스의 배경과 인물 복장을 잘 유지하면서 텍스트가 요구한 '팔을 붙잡는' 클로즈업 연출을 충실하고 자연스럽게 구현함."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "요구된 '팔' 대신 가슴 부위를 잡고 있으며, 손의 해부학적 구조가 붕괴되는 심각한 오류가 발생함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "맨손이 통로 쪽으로 뻗은 어두운 셔츠의 팔을 향해 꽉 쥐고 있음.",
        "built_space": "콘크리트 벽면, 랜턴 2개, 우측의 종이상자 더미, 커튼이 쳐진 통로 입구가 레퍼런스와 일치하게 배치됨.",
        "entities": "신부의 어두운 셔츠 소매와 손전등(아래쪽), 현우의 상처 입은 맨손이 묘사됨.",
        "hard_violations": [],
        "physics": "맨손이 옷 입은 팔을 자연스럽게 쥐고 있으며, 구조와 무게감이 물리적으로 타당함."
       },
       {
        "label": "A",
        "direction": "손이 팔이 아닌 신부의 가슴 및 목깃 부위를 향해 있음.",
        "built_space": "지하실 벽, 랜턴 1개, 통로 입구 및 일부 상자가 보임.",
        "entities": "신부의 셔츠와 칼라, 현우의 맨팔이 보이나 손의 형태가 온전하지 않음.",
        "hard_violations": [
         "물리적으로 불가능한 해부학 (여러 개의 손과 손가락이 기괴하게 융합됨)"
        ],
        "physics": "팔과 손이 비정상적으로 엉켜 있어 어떻게 지탱되고 잡고 있는지 파악할 수 없음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "현우의 손이 통로 쪽으로 뻗은 신부의 팔을 붙잡는 인서트 클로즈업으로, 핵심 행동과 기존 공간·의상·손전등을 충실히 유지한다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "현우가 신부의 옷깃을 잡고 신부가 현우의 손목을 붙드는 장면으로 바뀌어, 요구된 팔 붙잡기 관계와 손 중심의 인서트 구도를 놓쳤다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "신부의 소매 입은 팔이 화면 왼쪽에서 오른쪽의 어두운 통로 입구 쪽으로 뻗어 있고, 현우의 손은 그 팔의 손목 바로 위를 붙잡는다. 아래쪽 신부의 다른 손에 든 손전등도 오른쪽을 향한다. 얼굴과 눈은 프레임 밖이므로 시선은 확인할 수 없다.",
        "built_space": "배경 오른쪽에 직사각형 통로 입구 하나와 문틀이 있고, 왼쪽 상단에는 등불 하나와 가장자리에 잘린 다른 등불 일부가 보인다. 왼쪽 아래의 목재 수납장과 오른쪽에 쌓인 상자 세 개의 전부 또는 일부, 낡은 벽이 이전 장면의 배치와 이어진다. 팔과 손이 전경을 차지하며 배경 시설은 그 뒤에 놓인다.",
        "entities": "현우는 젊고 비교적 매끈한 피부의 손만 보이며 손가락에 작은 상처가 있다. 신부는 이전 장면과 같은 짙은 회청색 긴소매와 나이 든 손으로 나타나고, 다른 손에는 낡은 금속 손전등이 있다. 두 사람의 얼굴·머리와 신부의 칼라는 이 인서트에서 제외되어 신원 세부를 직접 확인할 수 없다. 추가 인물은 없고, 상자의 인쇄 흔적은 흐려 읽히지 않는다.",
        "hard_violations": [],
        "physics": "현우의 손가락들이 신부의 소매를 감싸고 반대편 엄지가 맞물리며, 접촉 부위의 천이 주름진다. 현우의 손목 연결부는 신부의 팔 뒤로 가려져 있지만 손이 분리되거나 떠 있다고 볼 근거는 없다. 신부의 팔은 왼쪽 몸통으로 연결되고 손전등은 아래쪽 손이 실제로 움켜쥔다. 손과 팔을 제동하는 동작은 물리적으로 가능하다."
       },
       {
        "label": "B",
        "direction": "오른쪽에서 들어온 현우의 맨팔과 손은 신부의 팔이 아니라 목 아래 옷깃을 향해 잡아당긴다. 신부의 소매 입은 손은 반대로 현우의 손목을 붙든다. 신부의 턱은 오른쪽을 향하지만 눈은 잘려 있어 통로를 보는지는 알 수 없다. 요구된 붙잡는 주체와 대상의 관계가 다르다.",
        "built_space": "뒤쪽에 어두운 통로 입구 하나와 문틀이 있으며, 왼쪽 상단의 등불 하나 및 가장자리의 다른 등불 일부, 오른쪽의 상자 두 개, 왼쪽 아래 수납장 일부가 보인다. 낡은 벽과 배치는 이전 공간과 대체로 일치한다. 다만 신부의 어깨·가슴·턱과 현우의 긴 전완까지 크게 포함하여 손의 세부 인서트보다 상체 실랑이 구도에 가깝다.",
        "entities": "신부는 이전 장면의 짙은 회청색 성직자 셔츠와 흰 칼라, 나이 든 손 및 수염 난 턱으로 보인다. 현우는 짧은 남색 소매 끝과 맨팔, 젊은 손으로 나타나며 외투는 보이지 않는다. 얼굴이 충분히 나오지 않아 두 사람의 정확한 얼굴 정체성은 확인할 수 없다. 손전등과 하체 소지품은 프레임 밖이므로 미소지라고 판단할 수 없다. 추가 인물이나 판독 가능한 문구는 없다.",
        "hard_violations": [],
        "physics": "현우의 손은 실제로 셔츠 천을 움켜쥐어 목 아래에 당김 주름을 만들고, 신부의 손가락은 현우의 손목을 감싼다. 두 팔은 각각 소매와 화면 오른쪽으로 자연스럽게 이어지며 지지 없이 떠 있는 신체나 물건은 없다. 실랑이 자체는 가능하지만, 그 물리적 접촉이 프롬프트가 지정한 행동과 다르다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "현우의 손이 통로 쪽으로 뻗은 신부의 팔을 붙잡는 인서트 클로즈업으로, 핵심 행동과 기존 공간·의상·손전등을 충실히 유지한다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "현우가 신부의 옷깃을 잡고 신부가 현우의 손목을 붙드는 장면으로 바뀌어, 요구된 팔 붙잡기 관계와 손 중심의 인서트 구도를 놓쳤다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "신부의 소매 입은 팔이 화면 왼쪽에서 오른쪽의 어두운 통로 입구 쪽으로 뻗어 있고, 현우의 손은 그 팔의 손목 바로 위를 붙잡는다. 아래쪽 신부의 다른 손에 든 손전등도 오른쪽을 향한다. 얼굴과 눈은 프레임 밖이므로 시선은 확인할 수 없다.",
        "built_space": "배경 오른쪽에 직사각형 통로 입구 하나와 문틀이 있고, 왼쪽 상단에는 등불 하나와 가장자리에 잘린 다른 등불 일부가 보인다. 왼쪽 아래의 목재 수납장과 오른쪽에 쌓인 상자 세 개의 전부 또는 일부, 낡은 벽이 이전 장면의 배치와 이어진다. 팔과 손이 전경을 차지하며 배경 시설은 그 뒤에 놓인다.",
        "entities": "현우는 젊고 비교적 매끈한 피부의 손만 보이며 손가락에 작은 상처가 있다. 신부는 이전 장면과 같은 짙은 회청색 긴소매와 나이 든 손으로 나타나고, 다른 손에는 낡은 금속 손전등이 있다. 두 사람의 얼굴·머리와 신부의 칼라는 이 인서트에서 제외되어 신원 세부를 직접 확인할 수 없다. 추가 인물은 없고, 상자의 인쇄 흔적은 흐려 읽히지 않는다.",
        "hard_violations": [],
        "physics": "현우의 손가락들이 신부의 소매를 감싸고 반대편 엄지가 맞물리며, 접촉 부위의 천이 주름진다. 현우의 손목 연결부는 신부의 팔 뒤로 가려져 있지만 손이 분리되거나 떠 있다고 볼 근거는 없다. 신부의 팔은 왼쪽 몸통으로 연결되고 손전등은 아래쪽 손이 실제로 움켜쥔다. 손과 팔을 제동하는 동작은 물리적으로 가능하다."
       },
       {
        "label": "A",
        "direction": "오른쪽에서 들어온 현우의 맨팔과 손은 신부의 팔이 아니라 목 아래 옷깃을 향해 잡아당긴다. 신부의 소매 입은 손은 반대로 현우의 손목을 붙든다. 신부의 턱은 오른쪽을 향하지만 눈은 잘려 있어 통로를 보는지는 알 수 없다. 요구된 붙잡는 주체와 대상의 관계가 다르다.",
        "built_space": "뒤쪽에 어두운 통로 입구 하나와 문틀이 있으며, 왼쪽 상단의 등불 하나 및 가장자리의 다른 등불 일부, 오른쪽의 상자 두 개, 왼쪽 아래 수납장 일부가 보인다. 낡은 벽과 배치는 이전 공간과 대체로 일치한다. 다만 신부의 어깨·가슴·턱과 현우의 긴 전완까지 크게 포함하여 손의 세부 인서트보다 상체 실랑이 구도에 가깝다.",
        "entities": "신부는 이전 장면의 짙은 회청색 성직자 셔츠와 흰 칼라, 나이 든 손 및 수염 난 턱으로 보인다. 현우는 짧은 남색 소매 끝과 맨팔, 젊은 손으로 나타나며 외투는 보이지 않는다. 얼굴이 충분히 나오지 않아 두 사람의 정확한 얼굴 정체성은 확인할 수 없다. 손전등과 하체 소지품은 프레임 밖이므로 미소지라고 판단할 수 없다. 추가 인물이나 판독 가능한 문구는 없다.",
        "hard_violations": [],
        "physics": "현우의 손은 실제로 셔츠 천을 움켜쥐어 목 아래에 당김 주름을 만들고, 신부의 손가락은 현우의 손목을 감싼다. 두 팔은 각각 소매와 화면 오른쪽으로 자연스럽게 이어지며 지지 없이 떠 있는 신체나 물건은 없다. 실랑이 자체는 가능하지만, 그 물리적 접촉이 프롬프트가 지정한 행동과 다르다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.762,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.512,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 물리적으로 불가능한 해부학 (여러 개의 손과 손가락이 기괴하게 융합됨)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 512
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "레퍼런스의 배경과 인물 복장을 잘 유지하면서 텍스트가 요구한 '팔을 붙잡는' 클로즈업 연출을 충실하고 자연스럽게 구현함."
   },
   {
    "label": "A",
    "score": 512,
    "verdict_ko": "요구된 '팔' 대신 가슴 부위를 잡고 있으며, 손의 해부학적 구조가 붕괴되는 심각한 오류가 발생함.  ★위반: [gemini-pro] 물리적으로 불가능한 해부학 (여러 개의 손과 손가락이 기괴하게 융합됨)"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 신부 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S32sh8_sel.png",
    "asset_id": "4f9ee049-ec82-4d89-a9aa-0f12c6cddae3",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 신부: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1402213>",
    "asset_id": "8696070d-ac09-4a5c-95f2-3daf1c015c32",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-5a30-7ab2-a608-e218df2485b4",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S32sh8"
  }
 },
 "S32sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:25:39.779472+00:00",
  "fingerprint": "53c51aa3ab11715fb48c0170679bb07c8068ad374afa8860d6a25a8bccb45b82",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S32sh9_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S32sh9_sel.png",
  "source_sha256": "4b56f71e6b45fe49d28975839015664254d79be897d3a4b288bfa32ab651ec22",
  "file": "S32sh9_cine.png",
  "staged_sha256": "3cb2536234bc8c57b1a6784d05c2c512e4efda322fa7b0fdad5c19ded6af10a2",
  "latency_ms": 10316
 },
 "S33sh6::signage": {
  "fp": "8ecdc67edb389acf",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::27402fa9685326d1": {
  "subjects": [],
  "subject_text": "박철진이 탑승한 전투 헬기 내부\n조종석과 조수석, 통신 장비가 밀집한 비좁은 항공기 실내. 전면과 측면 창을 통해 아래쪽 지형이 내려다보인다.",
  "identity": "canonical",
  "scope_id": "L190",
  "scope_role": "location_interior",
  "scope_sha": "ef30841e56478db8"
 },
 "S33sh6::bgfirst_bg": {
  "input_fingerprint": "914dd9415c65b141",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 예배소 바닥의 나무 지하실 문을 손가락으로 가리키며 입을 크게 벌려 소리치는 민병대원의 역동적인 자세.\n\nLOCATION (lock): Inside the church worship hall, at the wooden basement hatch in the floor amid the nighttime search.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Wooden basement door indicated by the militia member in the lower-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Wooden basement door (Discovered in the chapel floor) — Its upper face is visible obliquely beneath the pointing hand; used as Receives the pointing gesture in the lower right while remaining smaller than the figure; Chapel floor (Being searched) — Seen from a downward oblique angle around the basement entrance; used as Connects the discoverer's stance to the entrance without introducing additional furnishings.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained interior illumination appropriate to the nighttime setting, preserving clear separation between the face, pointing hand, and floor entrance without specifying a source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 예배소 바닥의 나무 지하실 문을 손가락으로 가리키며 입을 크게 벌려 소리치는 민병대원의 역동적인 자세.\n\nLOCATION (lock): Inside the church worship hall, at the wooden basement hatch in the floor amid the nighttime search.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Wooden basement door indicated by the militia member in the lower-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Wooden basement door (Discovered in the chapel floor) — Its upper face is visible obliquely beneath the pointing hand; used as Receives the pointing gesture in the lower right while remaining smaller than the figure; Chapel floor (Being searched) — Seen from a downward oblique angle around the basement entrance; used as Connects the discoverer's stance to the entrance without introducing additional furnishings.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained interior illumination appropriate to the nighttime setting, preserving clear separation between the face, pointing hand, and floor entrance without specifying a source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S33sh6__bgfirst_bg.png",
  "asset_id": "d53a2385-242e-4c71-bb6c-2e769e64e6d2",
  "input_asset_ids": [
   "b50d98cd-e71c-4b90-928e-99153b8d142f",
   "c50aa81f-5e1c-4dd9-bbb0-46f0d8e60222"
  ]
 },
 "S33sh6": {
  "input_fingerprint": "212373fbfbae63ae",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 예배소 바닥의 나무 지하실 문을 손가락으로 가리키며 입을 크게 벌려 소리치는 민병대원의 역동적인 자세.\n\nLOCATION (lock): Inside the church worship hall, at the wooden basement hatch in the floor amid the nighttime search. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Wooden basement door indicated by the militia member in the lower-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Wooden basement door (Discovered in the chapel floor) — Its upper face is visible obliquely beneath the pointing hand; used as Receives the pointing gesture in the lower right while remaining smaller than the figure; Chapel floor (Being searched) — Seen from a downward oblique angle around the basement entrance; used as Connects the discoverer's stance to the entrance without introducing additional furnishings.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained interior illumination appropriate to the nighttime setting, preserving clear separation between the face, pointing hand, and floor entrance without specifying a source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A basement access door has been located in the church floor. The intercut seawall still has severe cracks and seepage, with a vehicle convoy travelling along the embankment road.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 예배소 바닥의 나무 지하실 문을 손가락으로 가리키며 입을 크게 벌려 소리치는 민병대원의 역동적인 자세.\n\nLOCATION (lock): Inside the church worship hall, at the wooden basement hatch in the floor amid the nighttime search. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Wooden basement door indicated by the militia member in the lower-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Wooden basement door (Discovered in the chapel floor) — Its upper face is visible obliquely beneath the pointing hand; used as Receives the pointing gesture in the lower right while remaining smaller than the figure; Chapel floor (Being searched) — Seen from a downward oblique angle around the basement entrance; used as Connects the discoverer's stance to the entrance without introducing additional furnishings.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained interior illumination appropriate to the nighttime setting, preserving clear separation between the face, pointing hand, and floor entrance without specifying a source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A basement access door has been located in the church floor. The intercut seawall still has severe cracks and seepage, with a vehicle convoy travelling along the embankment road.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 예배소 바닥의 나무 지하실 문을 손가락으로 가리키며 입을 크게 벌려 소리치는 민병대원의 역동적인 자세.\n\nLOCATION (lock): Inside the church worship hall, at the wooden basement hatch in the floor amid the nighttime search. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Wooden basement door indicated by the militia member in the lower-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Wooden basement door (Discovered in the chapel floor) — Its upper face is visible obliquely beneath the pointing hand; used as Receives the pointing gesture in the lower right while remaining smaller than the figure; Chapel floor (Being searched) — Seen from a downward oblique angle around the basement entrance; used as Connects the discoverer's stance to the entrance without introducing additional furnishings.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained interior illumination appropriate to the nighttime setting, preserving clear separation between the face, pointing hand, and floor entrance without specifying a source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A basement access door has been located in the church floor. The intercut seawall still has severe cracks and seepage, with a vehicle convoy travelling along the embankment road.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S33sh6__bgfirst_bg.png",
     "asset_id": "d53a2385-242e-4c71-bb6c-2e769e64e6d2",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S33sh6.png",
     "asset_id": "b50d98cd-e71c-4b90-928e-99153b8d142f",
     "role": "conti_light"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L190B02.png",
     "asset_id": "c50aa81f-5e1c-4dd9-bbb0-46f0d8e60222",
     "role": "location_plate"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "남성의 시선과 가리키는 손가락이 우측 하단의 나무 지하실 문을 정확히 향함.",
    "built_space": "레퍼런스와 일치하는 강대상, 나무 의자, 콘크리트 바닥, 지하실 문이 알맞은 위치에 배치됨.",
    "entities": "전술 장비를 착용한 한국인 남성 1명, 나무 지하실 문.",
    "hard_violations": [],
    "physics": "두 발을 바닥에 단단히 딛고 역동적으로 몸을 낮춘 안정적인 자세."
   },
   {
    "label": "B",
    "direction": "남성의 시선은 앞을 향하며, 손가락은 바닥의 지하실 문을 가리킴.",
    "built_space": "강대상과 의자가 보이나 레퍼런스의 정확한 공간 배치와는 차이가 있음.",
    "entities": "주 인물 1명 외에 무장을 한 병사 3명이 배경에 추가로 존재함.",
    "hard_violations": [
     "[gemini-pro] 프롬프트에 지시되지 않은 추가 인물(배경의 병사들) 등장",
     "[gpt-high] 샷에 명시된 민병대원 외에 배경 인물 최소 3명을 추가하여, 명시되지 않은 사람의 등장을 금지한 조건을 위반했습니다."
    ],
    "physics": "두 발로 바닥을 딛고 구부정하게 서 있는 자세."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "레퍼런스 공간을 정확히 재현하고, 불필요한 인물 없이 단일 인물의 역동적인 동작과 외침을 프롬프트대로 완벽히 구현함."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "프롬프트에 명시되지 않은 배경 인물들이 다수 등장하여 인물 제한 규칙을 심각하게 위반함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "남성의 시선과 가리키는 손가락이 우측 하단의 나무 지하실 문을 정확히 향함.",
        "built_space": "레퍼런스와 일치하는 강대상, 나무 의자, 콘크리트 바닥, 지하실 문이 알맞은 위치에 배치됨.",
        "entities": "전술 장비를 착용한 한국인 남성 1명, 나무 지하실 문.",
        "hard_violations": [],
        "physics": "두 발을 바닥에 단단히 딛고 역동적으로 몸을 낮춘 안정적인 자세."
       },
       {
        "label": "B",
        "direction": "남성의 시선은 앞을 향하며, 손가락은 바닥의 지하실 문을 가리킴.",
        "built_space": "강대상과 의자가 보이나 레퍼런스의 정확한 공간 배치와는 차이가 있음.",
        "entities": "주 인물 1명 외에 무장을 한 병사 3명이 배경에 추가로 존재함.",
        "hard_violations": [
         "프롬프트에 지시되지 않은 추가 인물(배경의 병사들) 등장"
        ],
        "physics": "두 발로 바닥을 딛고 구부정하게 서 있는 자세."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "레퍼런스 공간을 정확히 재현하고, 불필요한 인물 없이 단일 인물의 역동적인 동작과 외침을 프롬프트대로 완벽히 구현함."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "프롬프트에 명시되지 않은 배경 인물들이 다수 등장하여 인물 제한 규칙을 심각하게 위반함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "남성의 시선과 가리키는 손가락이 우측 하단의 나무 지하실 문을 정확히 향함.",
        "built_space": "레퍼런스와 일치하는 강대상, 나무 의자, 콘크리트 바닥, 지하실 문이 알맞은 위치에 배치됨.",
        "entities": "전술 장비를 착용한 한국인 남성 1명, 나무 지하실 문.",
        "hard_violations": [],
        "physics": "두 발을 바닥에 단단히 딛고 역동적으로 몸을 낮춘 안정적인 자세."
       },
       {
        "label": "B",
        "direction": "남성의 시선은 앞을 향하며, 손가락은 바닥의 지하실 문을 가리킴.",
        "built_space": "강대상과 의자가 보이나 레퍼런스의 정확한 공간 배치와는 차이가 있음.",
        "entities": "주 인물 1명 외에 무장을 한 병사 3명이 배경에 추가로 존재함.",
        "hard_violations": [
         "프롬프트에 지시되지 않은 추가 인물(배경의 병사들) 등장"
        ],
        "physics": "두 발로 바닥을 딛고 구부정하게 서 있는 자세."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "지하실 문을 가리키며 외치는 동작은 보이지만, 명시되지 않은 인물 최소 3명이 추가되어 실격이며 구도도 요구한 미디엄 숏보다 넓습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "한 명의 민병대원, 오른쪽 아래 나무문, 참고 장소와 외치는 동작을 충실히 구현했지만, 전신과 넓은 바닥을 담아 미디엄 숏 지시에는 실패했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앞쪽 남성의 검지는 오른쪽 아래로 뻗어 있으며 연장 방향이 지하실 나무문 왼쪽 부분에 닿습니다. 얼굴과 시선은 문보다 화면 오른쪽 바깥을 향해 누군가에게 외치는 모습입니다. 왼쪽 배경 인물이 든 총의 총구는 오른쪽 아래 바닥을 향하며, 특정 사람을 겨누지는 않습니다.",
        "built_space": "닫힌 바닥 나무문 1개와 왼쪽 경첩 2개, 뒤쪽 강단 1개와 긴 의자 1개가 보입니다. 주인물은 문 왼쪽 바닥에 웅크려 있습니다. 콘크리트 바닥은 유사하지만 참고 사진과 달리 뒤쪽 창문이 두드러지고 강단 주변 배치도 다릅니다. 나무문은 오른쪽 아래에 있고 인물보다 작지만, 인물의 하퇴까지 담은 구도로 미디엄 숏보다 넓습니다.",
        "entities": "주인물은 한국인 설정에 부합하는 외관의 성인 동아시아계 남성으로, 위장복과 전술 조끼를 착용하고 입을 크게 벌리고 있습니다. 바닥 출입문은 실제 목재 판재와 금속 경첩으로 표현되었습니다. 왼쪽, 중앙 뒤쪽, 오른쪽에 추가 군복 인물이 최소 3명 있어 한 명만 등장해야 하는 지시와 충돌합니다. 머리 조명과 무장도 추가되어 있습니다. 판독 가능한 글자는 보이지 않습니다.",
        "hard_violations": [
         "샷에 명시된 민병대원 외에 배경 인물 최소 3명을 추가하여, 명시되지 않은 사람의 등장을 금지한 조건을 위반했습니다."
        ],
        "physics": "주인물은 무릎을 굽히고 몸을 앞으로 기울였으며 뒤쪽 부츠가 바닥에 닿아 몸을 지탱합니다. 앞쪽 발은 화면 밖이므로 접지를 확인할 수 없지만 공중에 뜬 자세는 아닙니다. 가리키는 팔은 어깨에서 자연스럽게 이어지고, 조명은 머리띠에 고정되어 있으며 나무문은 바닥 틀에 받쳐져 있습니다. 배경 인물들도 바닥에 서 있습니다."
       },
       {
        "label": "B",
        "direction": "남성의 검지가 오른쪽 아래를 향하며 그 연장선은 나무문 윗면으로 이어집니다. 고개와 눈도 손끝과 문 쪽으로 내려가 있어 발견한 입구를 지목하는 관계가 명확합니다. 반대쪽 손은 균형을 잡듯 뒤로 벌어져 있습니다.",
        "built_space": "참고 사진과 같은 왼쪽 강단 1개, 그 뒤 작은 책장 1개, 뒤 벽의 긴 의자 1개, 오른쪽 벽 돌출부와 낡은 콘크리트 바닥이 보입니다. 바닥 나무문은 1개이고 왼쪽 경첩 2개가 보이며, 남성은 문 왼쪽의 빈 바닥에 서 있습니다. 문 윗면은 손 아래 오른쪽 하단에 비스듬히 보이고 인물보다 작습니다. 다만 두 발을 포함한 전신과 넓은 바닥을 보여 주어 지정된 미디엄 숏이 아니라 전신 숏에 가깝습니다.",
        "entities": "성인 동아시아계 남성 1명만 등장하며 한국인 민병대원 설정에 부합하는 외관입니다. 어두운 군용 복장과 전술 조끼를 착용하고 입을 크게 벌려 외칩니다. 나무 지하실 문, 금속 경첩, 예배소 바닥이 모두 식별되며 참고 장소의 재료와 색조도 잘 유지됩니다. 추가 인물이나 읽을 수 있는 글자는 없습니다. 프레임 밖의 방조제와 차량 행렬을 끌어들이지 않았습니다.",
        "hard_violations": [],
        "physics": "벌린 두 다리의 부츠가 모두 바닥에 닿고 무릎이 굽혀져 있어 앞으로 기울인 몸을 지탱합니다. 한쪽 팔을 뻗어 가리키고 다른 팔로 균형을 잡는 동작이 물리적으로 자연스럽습니다. 나무문은 바닥과 거의 같은 높이의 틀에 놓여 있으며 떠 있는 신체나 물체는 없습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "지하실 문을 가리키며 외치는 동작은 보이지만, 명시되지 않은 인물 최소 3명이 추가되어 실격이며 구도도 요구한 미디엄 숏보다 넓습니다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "한 명의 민병대원, 오른쪽 아래 나무문, 참고 장소와 외치는 동작을 충실히 구현했지만, 전신과 넓은 바닥을 담아 미디엄 숏 지시에는 실패했습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "앞쪽 남성의 검지는 오른쪽 아래로 뻗어 있으며 연장 방향이 지하실 나무문 왼쪽 부분에 닿습니다. 얼굴과 시선은 문보다 화면 오른쪽 바깥을 향해 누군가에게 외치는 모습입니다. 왼쪽 배경 인물이 든 총의 총구는 오른쪽 아래 바닥을 향하며, 특정 사람을 겨누지는 않습니다.",
        "built_space": "닫힌 바닥 나무문 1개와 왼쪽 경첩 2개, 뒤쪽 강단 1개와 긴 의자 1개가 보입니다. 주인물은 문 왼쪽 바닥에 웅크려 있습니다. 콘크리트 바닥은 유사하지만 참고 사진과 달리 뒤쪽 창문이 두드러지고 강단 주변 배치도 다릅니다. 나무문은 오른쪽 아래에 있고 인물보다 작지만, 인물의 하퇴까지 담은 구도로 미디엄 숏보다 넓습니다.",
        "entities": "주인물은 한국인 설정에 부합하는 외관의 성인 동아시아계 남성으로, 위장복과 전술 조끼를 착용하고 입을 크게 벌리고 있습니다. 바닥 출입문은 실제 목재 판재와 금속 경첩으로 표현되었습니다. 왼쪽, 중앙 뒤쪽, 오른쪽에 추가 군복 인물이 최소 3명 있어 한 명만 등장해야 하는 지시와 충돌합니다. 머리 조명과 무장도 추가되어 있습니다. 판독 가능한 글자는 보이지 않습니다.",
        "hard_violations": [
         "샷에 명시된 민병대원 외에 배경 인물 최소 3명을 추가하여, 명시되지 않은 사람의 등장을 금지한 조건을 위반했습니다."
        ],
        "physics": "주인물은 무릎을 굽히고 몸을 앞으로 기울였으며 뒤쪽 부츠가 바닥에 닿아 몸을 지탱합니다. 앞쪽 발은 화면 밖이므로 접지를 확인할 수 없지만 공중에 뜬 자세는 아닙니다. 가리키는 팔은 어깨에서 자연스럽게 이어지고, 조명은 머리띠에 고정되어 있으며 나무문은 바닥 틀에 받쳐져 있습니다. 배경 인물들도 바닥에 서 있습니다."
       },
       {
        "label": "A",
        "direction": "남성의 검지가 오른쪽 아래를 향하며 그 연장선은 나무문 윗면으로 이어집니다. 고개와 눈도 손끝과 문 쪽으로 내려가 있어 발견한 입구를 지목하는 관계가 명확합니다. 반대쪽 손은 균형을 잡듯 뒤로 벌어져 있습니다.",
        "built_space": "참고 사진과 같은 왼쪽 강단 1개, 그 뒤 작은 책장 1개, 뒤 벽의 긴 의자 1개, 오른쪽 벽 돌출부와 낡은 콘크리트 바닥이 보입니다. 바닥 나무문은 1개이고 왼쪽 경첩 2개가 보이며, 남성은 문 왼쪽의 빈 바닥에 서 있습니다. 문 윗면은 손 아래 오른쪽 하단에 비스듬히 보이고 인물보다 작습니다. 다만 두 발을 포함한 전신과 넓은 바닥을 보여 주어 지정된 미디엄 숏이 아니라 전신 숏에 가깝습니다.",
        "entities": "성인 동아시아계 남성 1명만 등장하며 한국인 민병대원 설정에 부합하는 외관입니다. 어두운 군용 복장과 전술 조끼를 착용하고 입을 크게 벌려 외칩니다. 나무 지하실 문, 금속 경첩, 예배소 바닥이 모두 식별되며 참고 장소의 재료와 색조도 잘 유지됩니다. 추가 인물이나 읽을 수 있는 글자는 없습니다. 프레임 밖의 방조제와 차량 행렬을 끌어들이지 않았습니다.",
        "hard_violations": [],
        "physics": "벌린 두 다리의 부츠가 모두 바닥에 닿고 무릎이 굽혀져 있어 앞으로 기울인 몸을 지탱합니다. 한쪽 팔을 뻗어 가리키고 다른 팔로 균형을 잡는 동작이 물리적으로 자연스럽습니다. 나무문은 바닥과 거의 같은 높이의 틀에 놓여 있으며 떠 있는 신체나 물체는 없습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.536
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.286
   },
   "violations": {
    "B": [
     "[gemini-pro] 프롬프트에 지시되지 않은 추가 인물(배경의 병사들) 등장",
     "[gpt-high] 샷에 명시된 민병대원 외에 배경 인물 최소 3명을 추가하여, 명시되지 않은 사람의 등장을 금지한 조건을 위반했습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 286
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "레퍼런스 공간을 정확히 재현하고, 불필요한 인물 없이 단일 인물의 역동적인 동작과 외침을 프롬프트대로 완벽히 구현함."
   },
   {
    "label": "B",
    "score": 286,
    "verdict_ko": "프롬프트에 명시되지 않은 배경 인물들이 다수 등장하여 인물 제한 규칙을 심각하게 위반함.  ★위반: [gemini-pro] 프롬프트에 지시되지 않은 추가 인물(배경의 병사들) 등장 / [gpt-high] 샷에 명시된 민병대원 외에 배경 인물 최소 3명을 추가하여, 명시되지 않은 사람의 등장을 금지한 조건을 위반했습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L190B02.png",
    "asset_id": "c50aa81f-5e1c-4dd9-bbb0-46f0d8e60222",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-5bfb-7df0-99b7-60c2fad36c08",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S33sh6__bgfirst_bg.png",
   "bg_asset_id": "d53a2385-242e-4c71-bb6c-2e769e64e6d2",
   "bg_record_key": "S33sh6::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S33sh6::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:50:36.778648+00:00",
  "fingerprint": "bed0e94769a6575c1c9ebe7b3230dd702845b5425b97e299b29ce1dca1b7ae1b",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S33sh6_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S33sh6_sel.png",
  "source_sha256": "6fc35f7e4a53547014f6a4b73d510b615407200c3fa1a4763d22faec77dbff14",
  "file": "S33sh6_cine.png",
  "staged_sha256": "8059d9853c2b3426f65f35dc932b019c99fa0e02c0aee8b4e1f3fbe856f795fd",
  "latency_ms": 12141
 },
 "S33sh11::signage": {
  "fp": "13e335e6bed67d92",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S33sh11": {
  "input_fingerprint": "38a1c48a1c0f8f63",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 거대한 인공제방 위를 달리는 차량을 향해 벌떼처럼 빽빽하게 날아드는 소형 드론들의 실루엣 전경.\n\nLOCATION (lock): In the open air over the seawall road, above a convoy of fleeing vehicles near the refugee settlement. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Approaching drone swarm in the lower-left of the frame, foreground, moves toward Convoy on the embankment road; Convoy on the embankment road in the upper-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Approaching small drones (Flying densely toward the convoy) — Rear and side profiles face the camera as the drones converge diagonally into depth; used as Create layered foreground movement, with each drone small enough to preserve the convoy's visibility; Convoy on the embankment road (Travelling in a single line) — Vehicle sides and rear quarters recede along the diagonal road; used as Provide the destination of the drone movement and establish the continuing travel axis; Artificial embankment (Intact before the breach) — The crest and inland face are visible obliquely; used as Separate foreground aerial space from the elevated convoy route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Hold a restrained nighttime tonal range, separating the dark drone silhouettes from the more legible road and convoy without adding an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall remains deeply cracked and leaking, with repair machinery still nearby. Drones are converging on the convoy and surrounding containers before the explosive breach.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 거대한 인공제방 위를 달리는 차량을 향해 벌떼처럼 빽빽하게 날아드는 소형 드론들의 실루엣 전경.\n\nLOCATION (lock): In the open air over the seawall road, above a convoy of fleeing vehicles near the refugee settlement. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Approaching drone swarm in the lower-left of the frame, foreground, moves toward Convoy on the embankment road; Convoy on the embankment road in the upper-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Approaching small drones (Flying densely toward the convoy) — Rear and side profiles face the camera as the drones converge diagonally into depth; used as Create layered foreground movement, with each drone small enough to preserve the convoy's visibility; Convoy on the embankment road (Travelling in a single line) — Vehicle sides and rear quarters recede along the diagonal road; used as Provide the destination of the drone movement and establish the continuing travel axis; Artificial embankment (Intact before the breach) — The crest and inland face are visible obliquely; used as Separate foreground aerial space from the elevated convoy route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Hold a restrained nighttime tonal range, separating the dark drone silhouettes from the more legible road and convoy without adding an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall remains deeply cracked and leaking, with repair machinery still nearby. Drones are converging on the convoy and surrounding containers before the explosive breach.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 거대한 인공제방 위를 달리는 차량을 향해 벌떼처럼 빽빽하게 날아드는 소형 드론들의 실루엣 전경.\n\nLOCATION (lock): In the open air over the seawall road, above a convoy of fleeing vehicles near the refugee settlement. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Approaching drone swarm in the lower-left of the frame, foreground, moves toward Convoy on the embankment road; Convoy on the embankment road in the upper-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Approaching small drones (Flying densely toward the convoy) — Rear and side profiles face the camera as the drones converge diagonally into depth; used as Create layered foreground movement, with each drone small enough to preserve the convoy's visibility; Convoy on the embankment road (Travelling in a single line) — Vehicle sides and rear quarters recede along the diagonal road; used as Provide the destination of the drone movement and establish the continuing travel axis; Artificial embankment (Intact before the breach) — The crest and inland face are visible obliquely; used as Separate foreground aerial space from the elevated convoy route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Hold a restrained nighttime tonal range, separating the dark drone silhouettes from the more legible road and convoy without adding an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall remains deeply cracked and leaking, with repair machinery still nearby. Drones are converging on the convoy and surrounding containers before the explosive breach.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "드론들이 좌측에서 우측으로 비행하며 도로 위의 호송대를 향하고 있음.",
    "built_space": "레퍼런스와 동일하게 우측에 바다와 테트라포드가 있는 도로 구도를 띠고 있으나, 요구된 내륙 측면의 단차와 공간 분리가 명확하지 않음.",
    "entities": "드론 떼와 차량 호송대는 존재하나, 프롬프트가 요구한 수리 장비가 보이지 않음.",
    "hard_violations": [
     "[gemini-pro] 지면과 난간에 파묻히거나 관통되어 있는 전경의 드론들 (물리적 불가능)"
    ],
    "physics": "공중에 떠 있어야 할 일부 드론들이 지지대나 그림자 없이 도로 바닥이나 구조물과 융합되어 있음."
   },
   {
    "label": "B",
    "direction": "좌측 하단의 드론 떼가 우측 상단의 고가 도로 위 호송대를 향해 상승하듯 대각선으로 날아가고 있음.",
    "built_space": "카메라가 제방의 내륙 측면을 비스듬히 바라보는 구도이며, 좌측 하단에 균열이 간 지면이 있고 우측 상단에 호송대가 달리는 고가 도로가 배치되어 공간이 분리됨.",
    "entities": "소형 드론 떼, 트럭과 승용차로 이루어진 호송대, 깊게 갈라진 콘크리트 제방, 균열 옆의 주황색 수리 장비가 모두 명확히 확인됨.",
    "hard_violations": [],
    "physics": "드론들은 공중에 안정적으로 떠 있으며, 차량과 수리 장비는 구조물 위에 올바르게 안착되어 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "지정된 프레임 레이아웃(좌측 하단 드론, 우측 상단 호송대)과 내륙 측면 앵글을 정확히 구현했으며, 균열과 수리 장비까지 모두 포함하여 지시사항을 훌륭하게 충족합니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "요구된 내륙 측면 구도와 수리 장비가 누락되었으며, 전경의 드론들이 지면을 관통하는 치명적인 물리적 오류가 있어 감점되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "드론들이 좌측에서 우측으로 비행하며 도로 위의 호송대를 향하고 있음.",
        "built_space": "레퍼런스와 동일하게 우측에 바다와 테트라포드가 있는 도로 구도를 띠고 있으나, 요구된 내륙 측면의 단차와 공간 분리가 명확하지 않음.",
        "entities": "드론 떼와 차량 호송대는 존재하나, 프롬프트가 요구한 수리 장비가 보이지 않음.",
        "hard_violations": [
         "지면과 난간에 파묻히거나 관통되어 있는 전경의 드론들 (물리적 불가능)"
        ],
        "physics": "공중에 떠 있어야 할 일부 드론들이 지지대나 그림자 없이 도로 바닥이나 구조물과 융합되어 있음."
       },
       {
        "label": "B",
        "direction": "좌측 하단의 드론 떼가 우측 상단의 고가 도로 위 호송대를 향해 상승하듯 대각선으로 날아가고 있음.",
        "built_space": "카메라가 제방의 내륙 측면을 비스듬히 바라보는 구도이며, 좌측 하단에 균열이 간 지면이 있고 우측 상단에 호송대가 달리는 고가 도로가 배치되어 공간이 분리됨.",
        "entities": "소형 드론 떼, 트럭과 승용차로 이루어진 호송대, 깊게 갈라진 콘크리트 제방, 균열 옆의 주황색 수리 장비가 모두 명확히 확인됨.",
        "hard_violations": [],
        "physics": "드론들은 공중에 안정적으로 떠 있으며, 차량과 수리 장비는 구조물 위에 올바르게 안착되어 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "지정된 프레임 레이아웃(좌측 하단 드론, 우측 상단 호송대)과 내륙 측면 앵글을 정확히 구현했으며, 균열과 수리 장비까지 모두 포함하여 지시사항을 훌륭하게 충족합니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "요구된 내륙 측면 구도와 수리 장비가 누락되었으며, 전경의 드론들이 지면을 관통하는 치명적인 물리적 오류가 있어 감점되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "드론들이 좌측에서 우측으로 비행하며 도로 위의 호송대를 향하고 있음.",
        "built_space": "레퍼런스와 동일하게 우측에 바다와 테트라포드가 있는 도로 구도를 띠고 있으나, 요구된 내륙 측면의 단차와 공간 분리가 명확하지 않음.",
        "entities": "드론 떼와 차량 호송대는 존재하나, 프롬프트가 요구한 수리 장비가 보이지 않음.",
        "hard_violations": [
         "지면과 난간에 파묻히거나 관통되어 있는 전경의 드론들 (물리적 불가능)"
        ],
        "physics": "공중에 떠 있어야 할 일부 드론들이 지지대나 그림자 없이 도로 바닥이나 구조물과 융합되어 있음."
       },
       {
        "label": "B",
        "direction": "좌측 하단의 드론 떼가 우측 상단의 고가 도로 위 호송대를 향해 상승하듯 대각선으로 날아가고 있음.",
        "built_space": "카메라가 제방의 내륙 측면을 비스듬히 바라보는 구도이며, 좌측 하단에 균열이 간 지면이 있고 우측 상단에 호송대가 달리는 고가 도로가 배치되어 공간이 분리됨.",
        "entities": "소형 드론 떼, 트럭과 승용차로 이루어진 호송대, 깊게 갈라진 콘크리트 제방, 균열 옆의 주황색 수리 장비가 모두 명확히 확인됨.",
        "hard_violations": [],
        "physics": "드론들은 공중에 안정적으로 떠 있으며, 차량과 수리 장비는 구조물 위에 올바르게 안착되어 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "좌하단 드론 전경과 우상단 차량 중경, 차량의 후측면 구도가 더 충실하지만, 여러 줄의 차량과 과도하게 벌어진 제방 손상은 지시와 다릅니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "드론과 차량 사이의 대각선 연결은 보이지만, 차량 정면이 카메라로 다가오고 드론이 도로 전반에 퍼져 지정된 후측면 행렬과 전경 구도에서 벗어납니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "드론 무리는 좌하단에서 중앙의 제방 아래쪽으로 좁아지며, 그 위 우상단 도로의 차량들이 접근 대상으로 읽힙니다. 다만 개별 기체의 전후 방향은 실루엣만으로 확정하기 어렵고, 무리의 수렴점도 차량보다는 제방 벽 아래에 가깝습니다. 차량은 후미와 측면을 보이며 도로를 따라 좌상단 원경으로 이어집니다.",
        "built_space": "높은 콘크리트 제방 하나와 상부 도로 하나, 도로 양쪽의 연속 난간이 보입니다. 경사 벽과 소파블록은 참고 장소의 주요 구조를 유지합니다. 하부에는 넓은 콘크리트 작업면과 장비 두 대가 보이며, 벽 중앙부터 작업면까지 큰 균열과 떨어져 나온 덩어리가 이어집니다. 상부 도로는 연결되어 있지만 하부 손상은 단순한 균열보다 붕괴에 가깝습니다. 차량은 한 줄이 아니라 여러 줄로 배치되어 있습니다.",
        "entities": "수십 대의 소형 다중회전익 드론, 화물차와 승용차 행렬, 콘크리트 제방, 난간, 소파블록, 물과 보수용으로 보이는 장비가 있습니다. 사람과 얼굴은 보이지 않습니다. 별도의 정착지 컨테이너는 식별되지 않으며 트럭 적재함만 명확합니다. 야간의 젖은 재질과 차량 불빛은 부합합니다. 읽을 수 있는 문구는 식별되지 않습니다.",
        "hard_violations": [],
        "physics": "드론에는 회전익과 회전 흐림이 보여 비행을 지탱하는 추진 장치가 있습니다. 차량 바퀴는 상부 도로에, 장비는 하부 작업면에 놓여 있습니다. 부서진 콘크리트도 하부 지면에 쌓여 있어 무지지 부유물은 보이지 않습니다. 물 고임과 젖은 흔적은 있지만 균열에서 실제로 새어 나오는 물줄기는 분명하지 않습니다."
       },
       {
        "label": "B",
        "direction": "드론은 좌하단과 도로 전경에서 우상단 차량 행렬 쪽으로 펼쳐져 있습니다. 일부 날개형 기체는 그 방향으로 향하지만 다중회전익 기체들의 전후 방향은 일정하게 판독되지 않습니다. 가장 눈에 띄는 화물차와 여러 승용차는 전조등과 정면을 카메라 쪽으로 향하고 있어, 후측면을 보이며 깊이 들어가는 지정 차량 행렬과 다릅니다.",
        "built_space": "연속된 콘크리트 제방 하나, 상부 도로 하나, 양쪽 난간과 바깥 경사면 아래의 소파블록이 보입니다. 참고 장소의 벽 재질과 난간 구조는 잘 유지됩니다. 도로와 벽에 균열이 있으나 큰 개구부는 없습니다. 차량은 여러 줄이며 일부는 서로 반대 방향을 향합니다. 보수 장비는 식별되지 않고, 내륙측 벽면보다는 소파블록이 있는 바다측 경사면이 주로 드러납니다.",
        "entities": "다중회전익 드론과 날개형 소형 무인기, 화물차와 승용차, 난간, 콘크리트 제방, 소파블록과 수면이 보입니다. 사람이나 얼굴은 없습니다. 정착지 컨테이너와 보수 장비는 확인되지 않습니다. 야간 분위기와 젖은 도로는 맞으며 읽을 수 있는 글자는 식별되지 않습니다.",
        "hard_violations": [],
        "physics": "다중회전익 드론은 로터로 지지되는 비행 상태이며, 날개형 기체도 날개를 가진 비행체로 보입니다. 정지 화면이라 추진 세부는 불명확하지만 무지지 물체로 단정할 근거는 없습니다. 차량은 바퀴로 도로에 지지되고, 난간과 소파블록도 구조물과 지면에 연결되어 있습니다. 균열은 보이지만 누수는 명확하지 않습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "좌하단 드론 전경과 우상단 차량 중경, 차량의 후측면 구도가 더 충실하지만, 여러 줄의 차량과 과도하게 벌어진 제방 손상은 지시와 다릅니다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "드론과 차량 사이의 대각선 연결은 보이지만, 차량 정면이 카메라로 다가오고 드론이 도로 전반에 퍼져 지정된 후측면 행렬과 전경 구도에서 벗어납니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "드론 무리는 좌하단에서 중앙의 제방 아래쪽으로 좁아지며, 그 위 우상단 도로의 차량들이 접근 대상으로 읽힙니다. 다만 개별 기체의 전후 방향은 실루엣만으로 확정하기 어렵고, 무리의 수렴점도 차량보다는 제방 벽 아래에 가깝습니다. 차량은 후미와 측면을 보이며 도로를 따라 좌상단 원경으로 이어집니다.",
        "built_space": "높은 콘크리트 제방 하나와 상부 도로 하나, 도로 양쪽의 연속 난간이 보입니다. 경사 벽과 소파블록은 참고 장소의 주요 구조를 유지합니다. 하부에는 넓은 콘크리트 작업면과 장비 두 대가 보이며, 벽 중앙부터 작업면까지 큰 균열과 떨어져 나온 덩어리가 이어집니다. 상부 도로는 연결되어 있지만 하부 손상은 단순한 균열보다 붕괴에 가깝습니다. 차량은 한 줄이 아니라 여러 줄로 배치되어 있습니다.",
        "entities": "수십 대의 소형 다중회전익 드론, 화물차와 승용차 행렬, 콘크리트 제방, 난간, 소파블록, 물과 보수용으로 보이는 장비가 있습니다. 사람과 얼굴은 보이지 않습니다. 별도의 정착지 컨테이너는 식별되지 않으며 트럭 적재함만 명확합니다. 야간의 젖은 재질과 차량 불빛은 부합합니다. 읽을 수 있는 문구는 식별되지 않습니다.",
        "hard_violations": [],
        "physics": "드론에는 회전익과 회전 흐림이 보여 비행을 지탱하는 추진 장치가 있습니다. 차량 바퀴는 상부 도로에, 장비는 하부 작업면에 놓여 있습니다. 부서진 콘크리트도 하부 지면에 쌓여 있어 무지지 부유물은 보이지 않습니다. 물 고임과 젖은 흔적은 있지만 균열에서 실제로 새어 나오는 물줄기는 분명하지 않습니다."
       },
       {
        "label": "A",
        "direction": "드론은 좌하단과 도로 전경에서 우상단 차량 행렬 쪽으로 펼쳐져 있습니다. 일부 날개형 기체는 그 방향으로 향하지만 다중회전익 기체들의 전후 방향은 일정하게 판독되지 않습니다. 가장 눈에 띄는 화물차와 여러 승용차는 전조등과 정면을 카메라 쪽으로 향하고 있어, 후측면을 보이며 깊이 들어가는 지정 차량 행렬과 다릅니다.",
        "built_space": "연속된 콘크리트 제방 하나, 상부 도로 하나, 양쪽 난간과 바깥 경사면 아래의 소파블록이 보입니다. 참고 장소의 벽 재질과 난간 구조는 잘 유지됩니다. 도로와 벽에 균열이 있으나 큰 개구부는 없습니다. 차량은 여러 줄이며 일부는 서로 반대 방향을 향합니다. 보수 장비는 식별되지 않고, 내륙측 벽면보다는 소파블록이 있는 바다측 경사면이 주로 드러납니다.",
        "entities": "다중회전익 드론과 날개형 소형 무인기, 화물차와 승용차, 난간, 콘크리트 제방, 소파블록과 수면이 보입니다. 사람이나 얼굴은 없습니다. 정착지 컨테이너와 보수 장비는 확인되지 않습니다. 야간 분위기와 젖은 도로는 맞으며 읽을 수 있는 글자는 식별되지 않습니다.",
        "hard_violations": [],
        "physics": "다중회전익 드론은 로터로 지지되는 비행 상태이며, 날개형 기체도 날개를 가진 비행체로 보입니다. 정지 화면이라 추진 세부는 불명확하지만 무지지 물체로 단정할 근거는 없습니다. 차량은 바퀴로 도로에 지지되고, 난간과 소파블록도 구조물과 지면에 연결되어 있습니다. 균열은 보이지만 누수는 명확하지 않습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.214,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.964,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 지면과 난간에 파묻히거나 관통되어 있는 전경의 드론들 (물리적 불가능)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 964
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "지정된 프레임 레이아웃(좌측 하단 드론, 우측 상단 호송대)과 내륙 측면 앵글을 정확히 구현했으며, 균열과 수리 장비까지 모두 포함하여 지시사항을 훌륭하게 충족합니다."
   },
   {
    "label": "A",
    "score": 964,
    "verdict_ko": "요구된 내륙 측면 구도와 수리 장비가 누락되었으며, 전경의 드론들이 지면을 관통하는 치명적인 물리적 오류가 있어 감점되었습니다.  ★위반: [gemini-pro] 지면과 난간에 파묻히거나 관통되어 있는 전경의 드론들 (물리적 불가능)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L190B01.png",
    "asset_id": "9dbf00e6-e0e7-42db-8305-7544a83a95ac",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-5f50-7f94-80f5-9946c03a3b75",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S33sh11::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:51:52.255857+00:00",
  "fingerprint": "cbeb9063194db515dd6fc26bcf870403d74a722d8bdcc7b1ec708dd95c090be6",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S33sh11_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S33sh11_sel.png",
  "source_sha256": "2637fa70cf1e02aefdb2cce477e1124332c670ed752636c5cdd3356913474ec1",
  "file": "S33sh11_cine.png",
  "staged_sha256": "fb7dc11330436265b14e320ca3185ac47889631581e726d4cbd1660b8f9a8b39",
  "latency_ms": 13828
 },
 "S33sh14::signage": {
  "fp": "fb42e2bd9ca127c6",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S33sh14": {
  "input_fingerprint": "245ab460d0a4f27e",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 폭발로 산산조각이 나며 무너져 내린 제방 틈새로 거대한 검은 바닷물이 폭포수처럼 쏟아져 들어오는 압도적인 광경.\n\nLOCATION (lock): At the breached upper section of the coastal seawall, where seawater pours into the refugee settlement at night. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Breached upper embankment (Broken by explosions and collapsing) — The inland face and broken edges of the upper opening are seen obliquely from below; used as Bracket the falling water and retain the established wall orientation during withdrawal; Seawater pouring through the breach (Rushing inward in a massive descending flow); used as Carry the principal movement from the elevated opening toward the bottom of the frame; Repair machinery and construction equipment (Being swept away by seawater) — Only partial forms remain visible below the breach near the lower frame edge; used as Supply scale without redirecting attention away from the breach.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain subdued nighttime exposure with controlled tonal separation between black seawater, broken embankment edges, and the remaining wall.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall's upper section has been blown apart, opening a breach through which seawater pours into the settlement. Nearby repair machinery and construction equipment are being swept away.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 폭발로 산산조각이 나며 무너져 내린 제방 틈새로 거대한 검은 바닷물이 폭포수처럼 쏟아져 들어오는 압도적인 광경.\n\nLOCATION (lock): At the breached upper section of the coastal seawall, where seawater pours into the refugee settlement at night. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Breached upper embankment (Broken by explosions and collapsing) — The inland face and broken edges of the upper opening are seen obliquely from below; used as Bracket the falling water and retain the established wall orientation during withdrawal; Seawater pouring through the breach (Rushing inward in a massive descending flow); used as Carry the principal movement from the elevated opening toward the bottom of the frame; Repair machinery and construction equipment (Being swept away by seawater) — Only partial forms remain visible below the breach near the lower frame edge; used as Supply scale without redirecting attention away from the breach.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain subdued nighttime exposure with controlled tonal separation between black seawater, broken embankment edges, and the remaining wall.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall's upper section has been blown apart, opening a breach through which seawater pours into the settlement. Nearby repair machinery and construction equipment are being swept away.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 폭발로 산산조각이 나며 무너져 내린 제방 틈새로 거대한 검은 바닷물이 폭포수처럼 쏟아져 들어오는 압도적인 광경.\n\nLOCATION (lock): At the breached upper section of the coastal seawall, where seawater pours into the refugee settlement at night. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Breached upper embankment (Broken by explosions and collapsing) — The inland face and broken edges of the upper opening are seen obliquely from below; used as Bracket the falling water and retain the established wall orientation during withdrawal; Seawater pouring through the breach (Rushing inward in a massive descending flow); used as Carry the principal movement from the elevated opening toward the bottom of the frame; Repair machinery and construction equipment (Being swept away by seawater) — Only partial forms remain visible below the breach near the lower frame edge; used as Supply scale without redirecting attention away from the breach.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain subdued nighttime exposure with controlled tonal separation between black seawater, broken embankment edges, and the remaining wall.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall's upper section has been blown apart, opening a breach through which seawater pours into the settlement. Nearby repair machinery and construction equipment are being swept away.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "거대한 바닷물이 무너진 제방 틈새에서 화면 하단을 향해 폭포수처럼 쏟아져 내리고 있음.",
    "built_space": "도로와 난간이 있는 거대한 콘크리트 방조제 중앙이 크게 붕괴되어 있으며, 화면 아래쪽 끄트머리에 건설 장비들이 위치함.",
    "entities": "검은 바닷물(일치), 무너진 상단 제방(일치), 화면 하단의 건설 장비 일부(일치).",
    "hard_violations": [
     "[gpt-high] 파열부 뒤에 정상 난간과 높이가 분리된 온전한 차량 도로를 추가하여, 기준 장소의 제방 정상 도로 배치를 별도의 도로층으로 바꿨습니다."
    ],
    "physics": "쏟아지는 물결은 중력에 의해 자연스럽게 낙하하며, 차량과 중장비들은 바닥이나 남은 구조물에 올바르게 지탱되어 있음."
   },
   {
    "label": "B",
    "direction": "방조제의 무너진 틈에서 바닷물이 화면 중앙에서 아래쪽을 향해 거세게 쏟아지고 있음.",
    "built_space": "무너진 방조제 왼편으로 판자촌 건물이 일부 보이며, 아래쪽에는 굴삭기, 트럭, 대형 철골 구조물들이 어지럽게 놓여 있음.",
    "entities": "바닷물(흰 거품이 두드러져 검은 바닷물 지시 불일치), 무너진 제방(일치), 건설 장비(프레임 하단에 부분적으로 보이라는 지시를 어기고 온전하게 전경을 차지함).",
    "hard_violations": [],
    "physics": "물이 중력에 따라 쏟아져 내리며, 굴삭기와 트럭 등은 지면에 제대로 지탱되어 서 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "프롬프트가 요구한 '검은 바닷물'의 톤을 잘 살렸으며, 화면 하단 가장자리에 건설 장비가 부분적으로만 보이도록 한 구도 지시를 정확히 따랐습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "바닷물에 흰 거품이 많아 '검은 바닷물' 지시와 거리가 멀고, 전경에 배치된 건설 장비와 철골이 너무 커서 시선을 분산시킵니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "거대한 바닷물이 무너진 제방 틈새에서 화면 하단을 향해 폭포수처럼 쏟아져 내리고 있음.",
        "built_space": "도로와 난간이 있는 거대한 콘크리트 방조제 중앙이 크게 붕괴되어 있으며, 화면 아래쪽 끄트머리에 건설 장비들이 위치함.",
        "entities": "검은 바닷물(일치), 무너진 상단 제방(일치), 화면 하단의 건설 장비 일부(일치).",
        "hard_violations": [],
        "physics": "쏟아지는 물결은 중력에 의해 자연스럽게 낙하하며, 차량과 중장비들은 바닥이나 남은 구조물에 올바르게 지탱되어 있음."
       },
       {
        "label": "B",
        "direction": "방조제의 무너진 틈에서 바닷물이 화면 중앙에서 아래쪽을 향해 거세게 쏟아지고 있음.",
        "built_space": "무너진 방조제 왼편으로 판자촌 건물이 일부 보이며, 아래쪽에는 굴삭기, 트럭, 대형 철골 구조물들이 어지럽게 놓여 있음.",
        "entities": "바닷물(흰 거품이 두드러져 검은 바닷물 지시 불일치), 무너진 제방(일치), 건설 장비(프레임 하단에 부분적으로 보이라는 지시를 어기고 온전하게 전경을 차지함).",
        "hard_violations": [],
        "physics": "물이 중력에 따라 쏟아져 내리며, 굴삭기와 트럭 등은 지면에 제대로 지탱되어 서 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "프롬프트가 요구한 '검은 바닷물'의 톤을 잘 살렸으며, 화면 하단 가장자리에 건설 장비가 부분적으로만 보이도록 한 구도 지시를 정확히 따랐습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "바닷물에 흰 거품이 많아 '검은 바닷물' 지시와 거리가 멀고, 전경에 배치된 건설 장비와 철골이 너무 커서 시선을 분산시킵니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "거대한 바닷물이 무너진 제방 틈새에서 화면 하단을 향해 폭포수처럼 쏟아져 내리고 있음.",
        "built_space": "도로와 난간이 있는 거대한 콘크리트 방조제 중앙이 크게 붕괴되어 있으며, 화면 아래쪽 끄트머리에 건설 장비들이 위치함.",
        "entities": "검은 바닷물(일치), 무너진 상단 제방(일치), 화면 하단의 건설 장비 일부(일치).",
        "hard_violations": [],
        "physics": "쏟아지는 물결은 중력에 의해 자연스럽게 낙하하며, 차량과 중장비들은 바닥이나 남은 구조물에 올바르게 지탱되어 있음."
       },
       {
        "label": "B",
        "direction": "방조제의 무너진 틈에서 바닷물이 화면 중앙에서 아래쪽을 향해 거세게 쏟아지고 있음.",
        "built_space": "무너진 방조제 왼편으로 판자촌 건물이 일부 보이며, 아래쪽에는 굴삭기, 트럭, 대형 철골 구조물들이 어지럽게 놓여 있음.",
        "entities": "바닷물(흰 거품이 두드러져 검은 바닷물 지시 불일치), 무너진 제방(일치), 건설 장비(프레임 하단에 부분적으로 보이라는 지시를 어기고 온전하게 전경을 차지함).",
        "hard_violations": [],
        "physics": "물이 중력에 따라 쏟아져 내리며, 굴삭기와 트럭 등은 지면에 제대로 지탱되어 서 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "파괴된 제방 사이에서 정착지로 낙하하는 검은 해수와 장소의 재질을 잘 구현했지만, 장비를 하단의 일부 형태로만 보여야 한다는 구도보다 노출이 많습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "거대한 낙수와 하단에 잘린 장비는 적절하지만, 정상 난간보다 낮은 틈새 뒤로 온전한 차량 도로를 추가해 기준 제방의 도로 배치를 어긋나게 했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "물은 중앙의 파괴된 개구부에서 카메라 쪽 내륙 공간으로 넘어와 화면 아래로 떨어집니다. 하단에서는 포말이 굴착기와 덤프트럭 주변으로 퍼집니다. 사람의 시선이나 무기는 없으며, 굴착기 버킷은 아래의 침수 구역을 향합니다.",
        "built_space": "하나의 콘크리트 제방이 중앙의 큰 파열부 양쪽으로 이어지고, 상단 난간도 그 양쪽에 남아 있습니다. 파열부 왼쪽에는 철근이 드러난 도로 슬래브 잔해가 보입니다. 왼쪽 제방 아래에는 낮은 정착지 건물들이, 하단에는 굴착기와 트럭 및 강재가 있습니다. 기준 사진의 경사진 콘크리트 벽체와 금속 난간은 유지되지만, 장비와 바닥을 넓게 내려다보여 개구부 아래에서 올려다보는 시점은 약합니다.",
        "entities": "검은 해수, 흰 포말, 부서진 콘크리트와 철근, 남은 제방, 수리·건설 장비가 보입니다. 왼쪽 굴착기와 중앙 덤프트럭은 대부분의 형태가 드러나며, 오른쪽 아래에는 다른 장비 일부가 잘려 있습니다. 인물이나 얼굴, 읽을 수 있는 문구는 보이지 않습니다. 야간의 낮은 노출과 젖은 콘크리트·금속의 재질은 요청에 부합합니다.",
        "hard_violations": [],
        "physics": "물은 높은 개구부에서 중력 방향으로 낙하하고 아래 수면에서 포말을 만듭니다. 차량 하부와 굴착기 궤도는 침수 구역에 잠겨 있으며, 굴착기 팔과 버킷은 기계 관절에 연결되어 있습니다. 강재는 잔해와 바닥에 걸쳐 있습니다. 지지 없이 공중에 뜬 물체는 보이지 않지만, 장비가 실제로 떠밀리는 순간보다는 물에 잠긴 상태가 더 명확합니다."
       },
       {
        "label": "B",
        "direction": "검은 물은 중앙 상부의 틈에서 전경 아래로 폭포처럼 떨어져 오른쪽과 하단 장비를 덮습니다. 굴착기 팔은 오른쪽 아래로 뻗어 있습니다. 사람의 시선이나 겨냥하는 무기는 없습니다. 틈새 뒤 차량들은 물의 낙하 방향을 가로지르는 도로에 놓여 있습니다.",
        "built_space": "좌우에 파손된 콘크리트 제방과 정상 난간이 있고, 오른쪽 벽 아래에는 정착지 건물들이 이어집니다. 그런데 두 정상 난간보다 낮은 위치에서 차량 여러 대가 선 온전한 도로가 파열부 뒤를 가로지릅니다. 이는 기준 사진의 제방 정상 도로와 별개의 낮은 도로층처럼 보여 장소의 고정 구조와 맞지 않습니다. 하단 장비는 부분적으로 잘렸지만, 전경 운전실과 오른쪽 굴착기 팔이 상당한 면적을 차지합니다.",
        "entities": "대규모 검은 낙수, 부서진 콘크리트 가장자리, 노출 철근, 제방 난간과 노란 건설 장비가 있습니다. 배경에는 승용차 네 대가 식별되며, 오른쪽에는 작은 건물들과 노란 상자형 설비가 보입니다. 인물이나 얼굴, 읽을 수 있는 글자는 없습니다. 야간 분위기와 물·콘크리트·금속의 물성은 대체로 적절합니다.",
        "hard_violations": [
         "파열부 뒤에 정상 난간과 높이가 분리된 온전한 차량 도로를 추가하여, 기준 장소의 제방 정상 도로 배치를 별도의 도로층으로 바꿨습니다."
        ],
        "physics": "낙수는 개구부에서 아래로 이어지고 장비 주변의 충돌 지점에 포말이 생깁니다. 전경 운전실은 일부 잠긴 차체에 연결되어 있고 오른쪽 굴착기 팔도 차체와 연결되어 있습니다. 기울어진 장비는 물에 휩쓸리는 상태로 읽힙니다. 차량은 배경 도로에 지지되어 있으며, 지지 없이 공중에 떠 있는 물체는 보이지 않습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "파괴된 제방 사이에서 정착지로 낙하하는 검은 해수와 장소의 재질을 잘 구현했지만, 장비를 하단의 일부 형태로만 보여야 한다는 구도보다 노출이 많습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "거대한 낙수와 하단에 잘린 장비는 적절하지만, 정상 난간보다 낮은 틈새 뒤로 온전한 차량 도로를 추가해 기준 제방의 도로 배치를 어긋나게 했습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "물은 중앙의 파괴된 개구부에서 카메라 쪽 내륙 공간으로 넘어와 화면 아래로 떨어집니다. 하단에서는 포말이 굴착기와 덤프트럭 주변으로 퍼집니다. 사람의 시선이나 무기는 없으며, 굴착기 버킷은 아래의 침수 구역을 향합니다.",
        "built_space": "하나의 콘크리트 제방이 중앙의 큰 파열부 양쪽으로 이어지고, 상단 난간도 그 양쪽에 남아 있습니다. 파열부 왼쪽에는 철근이 드러난 도로 슬래브 잔해가 보입니다. 왼쪽 제방 아래에는 낮은 정착지 건물들이, 하단에는 굴착기와 트럭 및 강재가 있습니다. 기준 사진의 경사진 콘크리트 벽체와 금속 난간은 유지되지만, 장비와 바닥을 넓게 내려다보여 개구부 아래에서 올려다보는 시점은 약합니다.",
        "entities": "검은 해수, 흰 포말, 부서진 콘크리트와 철근, 남은 제방, 수리·건설 장비가 보입니다. 왼쪽 굴착기와 중앙 덤프트럭은 대부분의 형태가 드러나며, 오른쪽 아래에는 다른 장비 일부가 잘려 있습니다. 인물이나 얼굴, 읽을 수 있는 문구는 보이지 않습니다. 야간의 낮은 노출과 젖은 콘크리트·금속의 재질은 요청에 부합합니다.",
        "hard_violations": [],
        "physics": "물은 높은 개구부에서 중력 방향으로 낙하하고 아래 수면에서 포말을 만듭니다. 차량 하부와 굴착기 궤도는 침수 구역에 잠겨 있으며, 굴착기 팔과 버킷은 기계 관절에 연결되어 있습니다. 강재는 잔해와 바닥에 걸쳐 있습니다. 지지 없이 공중에 뜬 물체는 보이지 않지만, 장비가 실제로 떠밀리는 순간보다는 물에 잠긴 상태가 더 명확합니다."
       },
       {
        "label": "A",
        "direction": "검은 물은 중앙 상부의 틈에서 전경 아래로 폭포처럼 떨어져 오른쪽과 하단 장비를 덮습니다. 굴착기 팔은 오른쪽 아래로 뻗어 있습니다. 사람의 시선이나 겨냥하는 무기는 없습니다. 틈새 뒤 차량들은 물의 낙하 방향을 가로지르는 도로에 놓여 있습니다.",
        "built_space": "좌우에 파손된 콘크리트 제방과 정상 난간이 있고, 오른쪽 벽 아래에는 정착지 건물들이 이어집니다. 그런데 두 정상 난간보다 낮은 위치에서 차량 여러 대가 선 온전한 도로가 파열부 뒤를 가로지릅니다. 이는 기준 사진의 제방 정상 도로와 별개의 낮은 도로층처럼 보여 장소의 고정 구조와 맞지 않습니다. 하단 장비는 부분적으로 잘렸지만, 전경 운전실과 오른쪽 굴착기 팔이 상당한 면적을 차지합니다.",
        "entities": "대규모 검은 낙수, 부서진 콘크리트 가장자리, 노출 철근, 제방 난간과 노란 건설 장비가 있습니다. 배경에는 승용차 네 대가 식별되며, 오른쪽에는 작은 건물들과 노란 상자형 설비가 보입니다. 인물이나 얼굴, 읽을 수 있는 글자는 없습니다. 야간 분위기와 물·콘크리트·금속의 물성은 대체로 적절합니다.",
        "hard_violations": [
         "파열부 뒤에 정상 난간과 높이가 분리된 온전한 차량 도로를 추가하여, 기준 장소의 제방 정상 도로 배치를 별도의 도로층으로 바꿨습니다."
        ],
        "physics": "낙수는 개구부에서 아래로 이어지고 장비 주변의 충돌 지점에 포말이 생깁니다. 전경 운전실은 일부 잠긴 차체에 연결되어 있고 오른쪽 굴착기 팔도 차체와 연결되어 있습니다. 기울어진 장비는 물에 휩쓸리는 상태로 읽힙니다. 차량은 배경 도로에 지지되어 있으며, 지지 없이 공중에 떠 있는 물체는 보이지 않습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.571,
    "B": 1.571
   },
   "adjusted": {
    "A": 1.321,
    "B": 1.571
   },
   "violations": {
    "A": [
     "[gpt-high] 파열부 뒤에 정상 난간과 높이가 분리된 온전한 차량 도로를 추가하여, 기준 장소의 제방 정상 도로 배치를 별도의 도로층으로 바꿨습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1321,
   "B": 1571
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1321,
    "verdict_ko": "프롬프트가 요구한 '검은 바닷물'의 톤을 잘 살렸으며, 화면 하단 가장자리에 건설 장비가 부분적으로만 보이도록 한 구도 지시를 정확히 따랐습니다.  ★위반: [gpt-high] 파열부 뒤에 정상 난간과 높이가 분리된 온전한 차량 도로를 추가하여, 기준 장소의 제방 정상 도로 배치를 별도의 도로층으로 바꿨습니다."
   },
   {
    "label": "B",
    "score": 1571,
    "verdict_ko": "바닷물에 흰 거품이 많아 '검은 바닷물' 지시와 거리가 멀고, 전경에 배치된 건설 장비와 철골이 너무 커서 시선을 분산시킵니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L190B01.png",
    "asset_id": "9dbf00e6-e0e7-42db-8305-7544a83a95ac",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-6115-74eb-9173-4bcc09c6aca4",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S33sh14::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:53:01.154664+00:00",
  "fingerprint": "18173c8537eba4d38bf859ec51d558d057bb7222ee07547151af661f2f374316",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S33sh14_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S33sh14_sel.png",
  "source_sha256": "d34fb74b7fd0b1b0c5756ad3781b444ccd9886d97573b32ef05d1813e2cec3b1",
  "file": "S33sh14_cine.png",
  "staged_sha256": "bfcc2d65ce57981b2ae5f0c76e9fbe841b684acddfedf79a5aedbf15dfb725f1",
  "latency_ms": 12394
 },
 "S34sh3::signage": {
  "fp": "a8efc31dab6ae9e0",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::698a95a010babfcd": {
  "subjects": [],
  "subject_text": "인천 난민촌 관리사무소 3층 복도\n여러 방의 출입문과 외부를 내다보는 창이 이어지는 복도. 구석에 소화기가 비치돼 있고 위층으로 향하는 계단과 연결된다.",
  "identity": "canonical",
  "scope_id": "L188",
  "scope_role": "location_interior",
  "scope_sha": "37994dc144aebf36"
 },
 "S34sh3::bgfirst_bg": {
  "input_fingerprint": "5194da41211331df",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 양손에 거머쥔 붉은 소화기가 단단한 문고리에 강하게 충돌한 순간의 역동적인 현우의 자세.\n\nLOCATION (lock): In the third-floor corridor of the illuminated refugee administration building, directly outside a locked detention-room door.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office door and handle (Door still resisting entry as the handle is struck) — The exterior face and projecting handle are visible at an oblique angle; used as Anchor the impact point at center-right and establish the threshold for the following entry; Red fire extinguisher (Held in both hands and striking the handle) — Its body crosses diagonally from 현우's grip toward the handle; used as Connect the two-handed effort to the contact point without obscuring either.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained nighttime interior illumination with enough local contrast to read the impact, retaining the extinguisher's supported red color without adding a colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 양손에 거머쥔 붉은 소화기가 단단한 문고리에 강하게 충돌한 순간의 역동적인 현우의 자세.\n\nLOCATION (lock): In the third-floor corridor of the illuminated refugee administration building, directly outside a locked detention-room door.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office door and handle (Door still resisting entry as the handle is struck) — The exterior face and projecting handle are visible at an oblique angle; used as Anchor the impact point at center-right and establish the threshold for the following entry; Red fire extinguisher (Held in both hands and striking the handle) — Its body crosses diagonally from 현우's grip toward the handle; used as Connect the two-handed effort to the contact point without obscuring either.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained nighttime interior illumination with enough local contrast to read the impact, retaining the extinguisher's supported red color without adding a colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S34sh3__bgfirst_bg.png",
  "asset_id": "cdae0e3e-8869-4d01-8715-a1ca5773fcb2",
  "input_asset_ids": [
   "3c73b6f6-486c-4b9f-93e3-d1822cce367e",
   "ed668b3a-a7db-491e-8f0e-bd093bd7b2fd"
  ]
 },
 "S34sh3": {
  "input_fingerprint": "da8293268c67a3cd",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 양손에 거머쥔 붉은 소화기가 단단한 문고리에 강하게 충돌한 순간의 역동적인 현우의 자세.\n\nLOCATION (lock): In the third-floor corridor of the illuminated refugee administration building, directly outside a locked detention-room door. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office door and handle (Door still resisting entry as the handle is struck) — The exterior face and projecting handle are visible at an oblique angle; used as Anchor the impact point at center-right and establish the threshold for the following entry; Red fire extinguisher (Held in both hands and striking the handle) — Its body crosses diagonally from 현우's grip toward the handle; used as Connect the two-handed effort to the contact point without obscuring either.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained nighttime interior illumination with enough local contrast to read the impact, retaining the extinguisher's supported red color without adding a colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The third-floor room door is still shut and resisting forced entry; its handle has not yet broken free. The management office remains lit. 현우: He holds the fire extinguisher for battering the door, with facial bruises, an injured leg and no outer garment. The contact card remains concealed in his shoe.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 양손에 거머쥔 붉은 소화기가 단단한 문고리에 강하게 충돌한 순간의 역동적인 현우의 자세.\n\nLOCATION (lock): In the third-floor corridor of the illuminated refugee administration building, directly outside a locked detention-room door. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office door and handle (Door still resisting entry as the handle is struck) — The exterior face and projecting handle are visible at an oblique angle; used as Anchor the impact point at center-right and establish the threshold for the following entry; Red fire extinguisher (Held in both hands and striking the handle) — Its body crosses diagonally from 현우's grip toward the handle; used as Connect the two-handed effort to the contact point without obscuring either.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained nighttime interior illumination with enough local contrast to read the impact, retaining the extinguisher's supported red color without adding a colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The third-floor room door is still shut and resisting forced entry; its handle has not yet broken free. The management office remains lit. 현우: He holds the fire extinguisher for battering the door, with facial bruises, an injured leg and no outer garment. The contact card remains concealed in his shoe.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 양손에 거머쥔 붉은 소화기가 단단한 문고리에 강하게 충돌한 순간의 역동적인 현우의 자세.\n\nLOCATION (lock): In the third-floor corridor of the illuminated refugee administration building, directly outside a locked detention-room door. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office door and handle (Door still resisting entry as the handle is struck) — The exterior face and projecting handle are visible at an oblique angle; used as Anchor the impact point at center-right and establish the threshold for the following entry; Red fire extinguisher (Held in both hands and striking the handle) — Its body crosses diagonally from 현우's grip toward the handle; used as Connect the two-handed effort to the contact point without obscuring either.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained nighttime interior illumination with enough local contrast to read the impact, retaining the extinguisher's supported red color without adding a colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The third-floor room door is still shut and resisting forced entry; its handle has not yet broken free. The management office remains lit. 현우: He holds the fire extinguisher for battering the door, with facial bruises, an injured leg and no outer garment. The contact card remains concealed in his shoe.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S34sh3__bgfirst_bg.png",
     "asset_id": "cdae0e3e-8869-4d01-8715-a1ca5773fcb2",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S34sh3.png",
     "asset_id": "3c73b6f6-486c-4b9f-93e3-d1822cce367e",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L188B02.png",
     "asset_id": "ed668b3a-a7db-491e-8f0e-bd093bd7b2fd",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우가 소화기를 문고리 쪽으로 밀어붙이고 있으며, 소화기의 타격점이 문고리 부근에 위치합니다.",
    "built_space": "복도 공간, 닫힌 문, 우측의 소화전 함 등 위치 레퍼런스의 주요 고정 요소들이 올바른 위치에 구현되어 있습니다.",
    "entities": "현우(상의 탈의, 멍든 얼굴, 헝클어진 머리)와 붉은 소화기(라벨 및 세부 묘사 생략됨)가 등장합니다.",
    "hard_violations": [
     "[gemini-pro] 소화기의 앞부분과 문의 손잡이가 금속 채로 기괴하게 융합되어 물리적으로 불가능한 형태를 띠고 있습니다."
    ],
    "physics": "양손으로 소화기를 쥐고 몸을 기울인 자세는 지탱되고 있으나, 타격 지점에서 물체가 서로 섞여버려 물리적 충돌의 묘사가 완전히 붕괴되었습니다."
   },
   {
    "label": "B",
    "direction": "현우의 강렬한 시선과 소화기의 방향이 타격 목표인 문고리를 정확히 향하고 있습니다.",
    "built_space": "위치 레퍼런스와 동일한 3층 복도입니다. 우측 벽의 붉은 소화전 함, 정면의 문과 문고리, 좌측의 창문 구조가 올바르게 배치되어 있습니다.",
    "entities": "현우(18세, 한국계 미국인, 헝클어진 머리, 멍든 얼굴, 상의 탈의)의 특징이 레퍼런스와 잘 일치하며, 붉은 소화기(라벨 포함)도 사실적으로 묘사되었습니다.",
    "hard_violations": [],
    "physics": "왼손으로 소화기 밑동을, 오른손으로 윗부분을 단단히 쥐고 문고리를 강하게 내리찍는 무게 중심과 자세가 매우 자연스럽고 역동적입니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "지정된 장소와 인물의 외양을 정확히 반영하였으며, 소화기로 문고리를 강하게 타격하는 역동적인 순간과 표정을 물리적 오류 없이 훌륭하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "인물과 공간의 설정은 대체로 따랐으나, 소화기의 앞부분과 문고리가 융합되는 치명적인 형태 왜곡이 발생하여 탈락입니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "현우의 강렬한 시선과 소화기의 방향이 타격 목표인 문고리를 정확히 향하고 있습니다.",
        "built_space": "위치 레퍼런스와 동일한 3층 복도입니다. 우측 벽의 붉은 소화전 함, 정면의 문과 문고리, 좌측의 창문 구조가 올바르게 배치되어 있습니다.",
        "entities": "현우(18세, 한국계 미국인, 헝클어진 머리, 멍든 얼굴, 상의 탈의)의 특징이 레퍼런스와 잘 일치하며, 붉은 소화기(라벨 포함)도 사실적으로 묘사되었습니다.",
        "hard_violations": [],
        "physics": "왼손으로 소화기 밑동을, 오른손으로 윗부분을 단단히 쥐고 문고리를 강하게 내리찍는 무게 중심과 자세가 매우 자연스럽고 역동적입니다."
       },
       {
        "label": "A",
        "direction": "현우가 소화기를 문고리 쪽으로 밀어붙이고 있으며, 소화기의 타격점이 문고리 부근에 위치합니다.",
        "built_space": "복도 공간, 닫힌 문, 우측의 소화전 함 등 위치 레퍼런스의 주요 고정 요소들이 올바른 위치에 구현되어 있습니다.",
        "entities": "현우(상의 탈의, 멍든 얼굴, 헝클어진 머리)와 붉은 소화기(라벨 및 세부 묘사 생략됨)가 등장합니다.",
        "hard_violations": [
         "소화기의 앞부분과 문의 손잡이가 금속 채로 기괴하게 융합되어 물리적으로 불가능한 형태를 띠고 있습니다."
        ],
        "physics": "양손으로 소화기를 쥐고 몸을 기울인 자세는 지탱되고 있으나, 타격 지점에서 물체가 서로 섞여버려 물리적 충돌의 묘사가 완전히 붕괴되었습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "지정된 장소와 인물의 외양을 정확히 반영하였으며, 소화기로 문고리를 강하게 타격하는 역동적인 순간과 표정을 물리적 오류 없이 훌륭하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "인물과 공간의 설정은 대체로 따랐으나, 소화기의 앞부분과 문고리가 융합되는 치명적인 형태 왜곡이 발생하여 탈락입니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 강렬한 시선과 소화기의 방향이 타격 목표인 문고리를 정확히 향하고 있습니다.",
        "built_space": "위치 레퍼런스와 동일한 3층 복도입니다. 우측 벽의 붉은 소화전 함, 정면의 문과 문고리, 좌측의 창문 구조가 올바르게 배치되어 있습니다.",
        "entities": "현우(18세, 한국계 미국인, 헝클어진 머리, 멍든 얼굴, 상의 탈의)의 특징이 레퍼런스와 잘 일치하며, 붉은 소화기(라벨 포함)도 사실적으로 묘사되었습니다.",
        "hard_violations": [],
        "physics": "왼손으로 소화기 밑동을, 오른손으로 윗부분을 단단히 쥐고 문고리를 강하게 내리찍는 무게 중심과 자세가 매우 자연스럽고 역동적입니다."
       },
       {
        "label": "A",
        "direction": "현우가 소화기를 문고리 쪽으로 밀어붙이고 있으며, 소화기의 타격점이 문고리 부근에 위치합니다.",
        "built_space": "복도 공간, 닫힌 문, 우측의 소화전 함 등 위치 레퍼런스의 주요 고정 요소들이 올바른 위치에 구현되어 있습니다.",
        "entities": "현우(상의 탈의, 멍든 얼굴, 헝클어진 머리)와 붉은 소화기(라벨 및 세부 묘사 생략됨)가 등장합니다.",
        "hard_violations": [
         "소화기의 앞부분과 문의 손잡이가 금속 채로 기괴하게 융합되어 물리적으로 불가능한 형태를 띠고 있습니다."
        ],
        "physics": "양손으로 소화기를 쥐고 몸을 기울인 자세는 지탱되고 있으나, 타격 지점에서 물체가 서로 섞여버려 물리적 충돌의 묘사가 완전히 붕괴되었습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "양손으로 소화기를 지지하며 닫힌 문의 손잡이를 정확히 타격하지만, 인물이 상대적으로 작고 복도 비중이 커 미디엄 숏 요구와 장소 재현에서 B보다 떨어진다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "상체 중심의 미디엄 숏에서 양손의 힘과 중앙 오른쪽 손잡이의 충돌을 연결하고 참조 복도를 더 충실히 재현하지만, 충격 잔해가 손잡이를 일부 가린다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 얼굴과 시선은 오른쪽 문손잡이 쪽을 향한다. 붉은 소화기는 왼쪽 아래의 받치는 손에서 오른쪽 위의 손잡이까지 대각선으로 놓이며, 밸브 쪽 끝이 실제 돌출 손잡이에 맞닿아 있다.",
        "built_space": "오른쪽에 닫힌 회색 금속문 하나, 부착된 돌출 레버 손잡이 하나, 그 옆 벽에 빈 붉은 매립함 하나가 보인다. 현우는 문 바깥 복도에 서 있다. 투톤 벽과 낡은 바닥, 상부 창, 천장 배관은 참조와 유사하지만, 반대편 벽의 큰 창과 가까운 원형 조명 때문에 참조보다 다른 복도처럼 보이는 부분이 있다. 손잡이와 문 외면은 비스듬히 보이며 불가능한 반사는 없다.",
        "entities": "젊은 동아시아계 남성 한 명으로, 헝클어진 검은 머리와 앳된 얼굴은 현우 참조에 대체로 부합한다. 한국계 미국인이라는 국적·배경은 영상만으로 확인할 수 없다. 얼굴에 멍이 있고 상체는 맨몸이며 어두운 바지를 입었다. 참조의 남색 티셔츠는 보이지 않는다. 붉은 소화기 한 개와 금속 손잡이는 식별되며 문은 닫혀 있다. 소화기 라벨에 인쇄 흔적은 있으나 확실히 읽히는 문구는 없다. 다리 부상과 신발 속 카드는 이 구도로 확인되지 않는다.",
        "hard_violations": [],
        "physics": "한 손은 소화기 상단을 잡고 다른 손은 밑면을 받쳐 무게를 지지한다. 벌어진 허벅지와 앞으로 기울인 몸통은 문을 향해 힘을 가하는 자세로 가능하다. 발은 화면 밖이므로 접지는 직접 보이지 않지만, 몸이 공중에 뜬 표현은 아니다. 접촉 부근의 작은 먼지와 파편은 충돌로 설명되며 손잡이는 문에 붙어 있다."
       },
       {
        "label": "B",
        "direction": "현우는 오른쪽 아래의 충돌 지점을 내려다본다. 소화기 몸통이 양손 사이에서 오른쪽 위로 뻗으며 상단 금속부가 문 가장자리의 손잡이 부위를 때린다. 타격 목표는 문판의 무관한 지점이 아니라 손잡이이며, 접촉점은 중앙 오른쪽에 놓인다.",
        "built_space": "닫힌 회녹색 금속문 하나와 손잡이 장치 하나, 바로 오른쪽의 빈 붉은 매립함 하나가 보인다. 왼쪽으로 이어지는 복도에는 투톤 거친 벽, 문 옆 상부 창과 뒤쪽의 작은 창, 천장 배관, 사각 조명 두 개, 야간 외부가 보이는 끝문이 배치되어 참조 장소와 잘 맞는다. 현우는 문 외부 복도에 있으며 문 외면과 문틀을 비스듬히 보는 시점도 성립한다. 충격 먼지와 소화기 상단이 손잡이 윤곽 일부를 가린다.",
        "entities": "헝클어진 검은 머리와 뺨의 멍이 있는 젊은 동아시아계 남성 한 명이다. 측면 얼굴은 현우의 연령대와 외형에 대체로 맞지만 정확한 얼굴 일치는 정면보다 판단하기 어렵다. 국적은 확인할 수 없다. 상체는 맨몸이고 회색 바지는 찢어지고 얼룩져 있으며, 참조의 남색 티셔츠는 없다. 붉은 소화기 한 개가 명확하고 읽을 수 있는 글자는 없다. 다리 쪽 손상 흔적은 있으나 부상 자체를 확정할 수 없고, 신발 속 카드는 화면 밖이다.",
        "hard_violations": [],
        "physics": "한 손이 소화기 상부 몸통을 감싸고 다른 손이 하단을 움켜쥐어 물체를 확실히 지지한다. 굽힌 팔과 앞으로 실린 몸통, 넓게 벌린 다리는 무거운 소화기를 밀어 타격하는 동작으로 가능하다. 발은 잘렸지만 부유나 비정상적인 신체 지지는 보이지 않는다. 접촉점에서 퍼지는 먼지와 작은 파편은 충격에 따른 움직임으로 설명되며, 손잡이 장치는 손상되어 보여도 문에서 완전히 떨어져 나가지는 않았다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "양손으로 소화기를 지지하며 닫힌 문의 손잡이를 정확히 타격하지만, 인물이 상대적으로 작고 복도 비중이 커 미디엄 숏 요구와 장소 재현에서 B보다 떨어진다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "상체 중심의 미디엄 숏에서 양손의 힘과 중앙 오른쪽 손잡이의 충돌을 연결하고 참조 복도를 더 충실히 재현하지만, 충격 잔해가 손잡이를 일부 가린다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 얼굴과 시선은 오른쪽 문손잡이 쪽을 향한다. 붉은 소화기는 왼쪽 아래의 받치는 손에서 오른쪽 위의 손잡이까지 대각선으로 놓이며, 밸브 쪽 끝이 실제 돌출 손잡이에 맞닿아 있다.",
        "built_space": "오른쪽에 닫힌 회색 금속문 하나, 부착된 돌출 레버 손잡이 하나, 그 옆 벽에 빈 붉은 매립함 하나가 보인다. 현우는 문 바깥 복도에 서 있다. 투톤 벽과 낡은 바닥, 상부 창, 천장 배관은 참조와 유사하지만, 반대편 벽의 큰 창과 가까운 원형 조명 때문에 참조보다 다른 복도처럼 보이는 부분이 있다. 손잡이와 문 외면은 비스듬히 보이며 불가능한 반사는 없다.",
        "entities": "젊은 동아시아계 남성 한 명으로, 헝클어진 검은 머리와 앳된 얼굴은 현우 참조에 대체로 부합한다. 한국계 미국인이라는 국적·배경은 영상만으로 확인할 수 없다. 얼굴에 멍이 있고 상체는 맨몸이며 어두운 바지를 입었다. 참조의 남색 티셔츠는 보이지 않는다. 붉은 소화기 한 개와 금속 손잡이는 식별되며 문은 닫혀 있다. 소화기 라벨에 인쇄 흔적은 있으나 확실히 읽히는 문구는 없다. 다리 부상과 신발 속 카드는 이 구도로 확인되지 않는다.",
        "hard_violations": [],
        "physics": "한 손은 소화기 상단을 잡고 다른 손은 밑면을 받쳐 무게를 지지한다. 벌어진 허벅지와 앞으로 기울인 몸통은 문을 향해 힘을 가하는 자세로 가능하다. 발은 화면 밖이므로 접지는 직접 보이지 않지만, 몸이 공중에 뜬 표현은 아니다. 접촉 부근의 작은 먼지와 파편은 충돌로 설명되며 손잡이는 문에 붙어 있다."
       },
       {
        "label": "A",
        "direction": "현우는 오른쪽 아래의 충돌 지점을 내려다본다. 소화기 몸통이 양손 사이에서 오른쪽 위로 뻗으며 상단 금속부가 문 가장자리의 손잡이 부위를 때린다. 타격 목표는 문판의 무관한 지점이 아니라 손잡이이며, 접촉점은 중앙 오른쪽에 놓인다.",
        "built_space": "닫힌 회녹색 금속문 하나와 손잡이 장치 하나, 바로 오른쪽의 빈 붉은 매립함 하나가 보인다. 왼쪽으로 이어지는 복도에는 투톤 거친 벽, 문 옆 상부 창과 뒤쪽의 작은 창, 천장 배관, 사각 조명 두 개, 야간 외부가 보이는 끝문이 배치되어 참조 장소와 잘 맞는다. 현우는 문 외부 복도에 있으며 문 외면과 문틀을 비스듬히 보는 시점도 성립한다. 충격 먼지와 소화기 상단이 손잡이 윤곽 일부를 가린다.",
        "entities": "헝클어진 검은 머리와 뺨의 멍이 있는 젊은 동아시아계 남성 한 명이다. 측면 얼굴은 현우의 연령대와 외형에 대체로 맞지만 정확한 얼굴 일치는 정면보다 판단하기 어렵다. 국적은 확인할 수 없다. 상체는 맨몸이고 회색 바지는 찢어지고 얼룩져 있으며, 참조의 남색 티셔츠는 없다. 붉은 소화기 한 개가 명확하고 읽을 수 있는 글자는 없다. 다리 쪽 손상 흔적은 있으나 부상 자체를 확정할 수 없고, 신발 속 카드는 화면 밖이다.",
        "hard_violations": [],
        "physics": "한 손이 소화기 상부 몸통을 감싸고 다른 손이 하단을 움켜쥐어 물체를 확실히 지지한다. 굽힌 팔과 앞으로 실린 몸통, 넓게 벌린 다리는 무거운 소화기를 밀어 타격하는 동작으로 가능하다. 발은 잘렸지만 부유나 비정상적인 신체 지지는 보이지 않는다. 접촉점에서 퍼지는 먼지와 작은 파편은 충격에 따른 움직임으로 설명되며, 손잡이 장치는 손상되어 보여도 문에서 완전히 떨어져 나가지는 않았다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.333,
    "B": 1.875
   },
   "adjusted": {
    "A": 1.083,
    "B": 1.875
   },
   "violations": {
    "A": [
     "[gemini-pro] 소화기의 앞부분과 문의 손잡이가 금속 채로 기괴하게 융합되어 물리적으로 불가능한 형태를 띠고 있습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1875,
   "A": 1083
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1875,
    "verdict_ko": "지정된 장소와 인물의 외양을 정확히 반영하였으며, 소화기로 문고리를 강하게 타격하는 역동적인 순간과 표정을 물리적 오류 없이 훌륭하게 구현했습니다."
   },
   {
    "label": "A",
    "score": 1083,
    "verdict_ko": "인물과 공간의 설정은 대체로 따랐으나, 소화기의 앞부분과 문고리가 융합되는 치명적인 형태 왜곡이 발생하여 탈락입니다.  ★위반: [gemini-pro] 소화기의 앞부분과 문의 손잡이가 금속 채로 기괴하게 융합되어 물리적으로 불가능한 형태를 띠고 있습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L188B02.png",
    "asset_id": "ed668b3a-a7db-491e-8f0e-bd093bd7b2fd",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-62c7-75d5-a861-fc05d813d1fc",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S34sh3__bgfirst_bg.png",
   "bg_asset_id": "cdae0e3e-8869-4d01-8715-a1ca5773fcb2",
   "bg_record_key": "S34sh3::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S34sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:54:44.520357+00:00",
  "fingerprint": "b4e9e849c83f8a8d5e57216fbaaad850869129cc738286026db279a2c5abb914",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S34sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S34sh3_sel.png",
  "source_sha256": "def3becda15c0dd55a4cc09124795c5add69e18254744ba65240c164ac8cdeff",
  "file": "S34sh3_cine.png",
  "staged_sha256": "2ef4d91592db8369b9b93fb2cdfacf508789bf53e82a07f9797602bc695b2f3e",
  "latency_ms": 11572
 },
 "S34sh7::signage": {
  "fp": "42b46010b0fc40a3",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S34sh7": {
  "input_fingerprint": "01d5e63eda701f13",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 미연과 현우가 바닥에 무릎을 꿇은 채 서로를 빈틈없이 꽉 끌어안은 굳은 자세.\n\nLOCATION (lock): Just inside the opened detention room on the administration building's third floor, under the building's nighttime lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office floor (Supporting the kneeling pair) — A narrow area is visible beneath their knees; used as Confirm their lowered posture while leaving the embrace as the primary subject.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the interior illumination restrained and gently modeled, allowing the contact between their faces, shoulders, and arms to remain readable without introducing a new source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The door handle has broken off and the third-floor room door is now open. The management office remains lit. 현우: He is inside the opened room in an embrace posture, with persistent facial bruises and an injured leg, without his outer garment. The contact card remains in his shoe. 미연: She is no longer shut behind the door and remains visibly injured, with a badly swollen face.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 미연과 현우가 바닥에 무릎을 꿇은 채 서로를 빈틈없이 꽉 끌어안은 굳은 자세.\n\nLOCATION (lock): Just inside the opened detention room on the administration building's third floor, under the building's nighttime lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office floor (Supporting the kneeling pair) — A narrow area is visible beneath their knees; used as Confirm their lowered posture while leaving the embrace as the primary subject.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the interior illumination restrained and gently modeled, allowing the contact between their faces, shoulders, and arms to remain readable without introducing a new source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The door handle has broken off and the third-floor room door is now open. The management office remains lit. 현우: He is inside the opened room in an embrace posture, with persistent facial bruises and an injured leg, without his outer garment. The contact card remains in his shoe. 미연: She is no longer shut behind the door and remains visibly injured, with a badly swollen face.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 미연과 현우가 바닥에 무릎을 꿇은 채 서로를 빈틈없이 꽉 끌어안은 굳은 자세.\n\nLOCATION (lock): Just inside the opened detention room on the administration building's third floor, under the building's nighttime lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office floor (Supporting the kneeling pair) — A narrow area is visible beneath their knees; used as Confirm their lowered posture while leaving the embrace as the primary subject.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the interior illumination restrained and gently modeled, allowing the contact between their faces, shoulders, and arms to remain readable without introducing a new source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The door handle has broken off and the third-floor room door is now open. The management office remains lit. 현우: He is inside the opened room in an embrace posture, with persistent facial bruises and an injured leg, without his outer garment. The contact card remains in his shoe. 미연: She is no longer shut behind the door and remains visibly injured, with a badly swollen face.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "두 인물이 방 안에서 무릎을 꿇고 서로 빈틈없이 껴안고 있음.",
    "built_space": "이전 샷의 투톤 벽과 나무 의자가 보이나, 열린 문의 손잡이가 파손되지 않고 멀쩡함.",
    "entities": "미연은 멍든 얼굴과 어두운 상의로 일치함. 현우는 어두운 상의와 반바지를 입고 운동화를 신음.",
    "hard_violations": [],
    "physics": "바닥에 무릎을 대고 체중을 지탱하며 물리적으로 문제없이 서로를 안고 있음."
   },
   {
    "label": "B",
    "direction": "두 인물이 문턱 근처에서 무릎을 꿇고 서로 단단히 껴안고 있음.",
    "built_space": "프레임 앞쪽 문틀 너머로 회색/흰색 벽과 탁자가 보임. 문에 둥근 손잡이가 남아 있음.",
    "entities": "미연은 멍든 얼굴과 어두운 상의로 레퍼런스와 일치함. 현우는 베이지색 상의를 입고 맨발임.",
    "hard_violations": [],
    "physics": "바닥에 무릎을 대고 안정적으로 지탱하며, 두 팔로 서로를 자연스럽게 감싸안음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "B는 지시된 미디엄 샷과 무릎 아래 좁은 공간 노출이라는 최우선 프레이밍 조건을 완벽히 충족하여 승리했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "A는 의상 색상과 신발 등 일부 디테일을 맞췄으나, 프레이밍이 풀 샷으로 넓어져 우선순위가 높은 샷 스케일 조건을 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 인물이 방 안에서 무릎을 꿇고 서로 빈틈없이 껴안고 있음.",
        "built_space": "이전 샷의 투톤 벽과 나무 의자가 보이나, 열린 문의 손잡이가 파손되지 않고 멀쩡함.",
        "entities": "미연은 멍든 얼굴과 어두운 상의로 일치함. 현우는 어두운 상의와 반바지를 입고 운동화를 신음.",
        "hard_violations": [],
        "physics": "바닥에 무릎을 대고 체중을 지탱하며 물리적으로 문제없이 서로를 안고 있음."
       },
       {
        "label": "B",
        "direction": "두 인물이 문턱 근처에서 무릎을 꿇고 서로 단단히 껴안고 있음.",
        "built_space": "프레임 앞쪽 문틀 너머로 회색/흰색 벽과 탁자가 보임. 문에 둥근 손잡이가 남아 있음.",
        "entities": "미연은 멍든 얼굴과 어두운 상의로 레퍼런스와 일치함. 현우는 베이지색 상의를 입고 맨발임.",
        "hard_violations": [],
        "physics": "바닥에 무릎을 대고 안정적으로 지탱하며, 두 팔로 서로를 자연스럽게 감싸안음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "B는 지시된 미디엄 샷과 무릎 아래 좁은 공간 노출이라는 최우선 프레이밍 조건을 완벽히 충족하여 승리했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "A는 의상 색상과 신발 등 일부 디테일을 맞췄으나, 프레이밍이 풀 샷으로 넓어져 우선순위가 높은 샷 스케일 조건을 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "두 인물이 방 안에서 무릎을 꿇고 서로 빈틈없이 껴안고 있음.",
        "built_space": "이전 샷의 투톤 벽과 나무 의자가 보이나, 열린 문의 손잡이가 파손되지 않고 멀쩡함.",
        "entities": "미연은 멍든 얼굴과 어두운 상의로 일치함. 현우는 어두운 상의와 반바지를 입고 운동화를 신음.",
        "hard_violations": [],
        "physics": "바닥에 무릎을 대고 체중을 지탱하며 물리적으로 문제없이 서로를 안고 있음."
       },
       {
        "label": "B",
        "direction": "두 인물이 문턱 근처에서 무릎을 꿇고 서로 단단히 껴안고 있음.",
        "built_space": "프레임 앞쪽 문틀 너머로 회색/흰색 벽과 탁자가 보임. 문에 둥근 손잡이가 남아 있음.",
        "entities": "미연은 멍든 얼굴과 어두운 상의로 레퍼런스와 일치함. 현우는 베이지색 상의를 입고 맨발임.",
        "hard_violations": [],
        "physics": "바닥에 무릎을 대고 안정적으로 지탱하며, 두 팔로 서로를 자연스럽게 감싸안음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "무릎 지지는 자연스럽지만, 미디엄 숏보다 넓은 문간 전신 구도이고 포옹에 틈이 있으며 현우의 상의·맨발과 남아 있는 문손잡이가 연속성에 어긋납니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "방 안에서 서로 밀착한 포옹과 기존 가구·현우의 남색 상의를 더 잘 유지하지만, 전신과 넓은 바닥을 보여 미디엄 숏 지시를 벗어나고 문손잡이도 남아 있습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 눈을 감고 미연 쪽으로 얼굴을 기울이며, 미연은 현우의 얼굴 쪽으로 시선을 낮춘다. 두 사람의 팔은 서로의 어깨와 등을 향하지만 허리와 몸통 사이에는 틈이 있어 빈틈없는 밀착은 약하다. 카메라를 보는 사람이나 방향을 확인할 휴대 소품은 없다.",
        "built_space": "카메라는 문밖에서 문간에 무릎 꿇은 두 사람을 바라본다. 왼쪽 앞에 열린 문짝 하나와 둥근 손잡이가 있고, 그 뒤로 별도의 문과 손잡이 하나가 더 보인다. 오른쪽 벽에는 스위치 하나가 있다. 회백색 투톤 벽은 참고와 유사하지만 참고의 책상과 목제 의자는 이 구도에서 확인되지 않는다. 열린 문 바로 안쪽보다는 문턱을 점유한 배치이며, 손잡이가 떨어졌다는 설정과 달리 앞문 손잡이가 붙어 있다. 반사는 없다.",
        "entities": "중년 동아시아계 여성 한 명과 앳된 동아시아계 남성 한 명만 보인다. 미연의 검은 단발, 남색 상의와 얼굴 부상은 참고에 대체로 부합한다. 현우의 헝클어진 검은 머리와 얼굴 멍은 맞지만 남색 참고 상의 대신 회갈색 티셔츠를 입고 있다. 긴 바지에는 부상 흔적으로 보이는 얼룩이 있고, 보이는 발은 맨발이어서 신발 속 연락 카드의 소지 상태를 뒷받침하지 못한다. 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "두 사람 모두 접힌 무릎과 하퇴가 바닥에 닿아 체중을 지지한다. 현우의 맨발과 미연의 구두도 바닥에 접촉한다. 서로의 어깨와 등을 잡은 손과 팔은 포옹으로 가능한 배치이며, 공중에 지지 없이 떠 있는 신체나 물체는 없다."
       },
       {
        "label": "B",
        "direction": "미연은 눈을 감고 현우의 어깨에 얼굴을 붙이고, 현우는 미연의 머리 뒤쪽으로 얼굴을 숙인다. 양쪽 팔이 상대의 등과 어깨를 감싸 몸통까지 밀착하므로 서로를 꽉 끌어안는 방향과 대상이 명확하다. 외부를 겨누거나 카메라를 응시하는 동작은 없다.",
        "built_space": "방 안에서 열린 출입문과 그 너머 복도를 보는 구도다. 왼쪽에 목제 의자 하나, 오른쪽에 금속 테두리 책상 하나, 뒤쪽에 열린 금속 문짝 하나와 벽 스위치 하나가 보인다. 참고의 투톤 벽과 낡은 의자·책상 재질을 이어 간다. 두 사람은 가구 사이의 방 안 바닥에 무릎을 꿇고 있다. 문손잡이 장착부는 파손되어 있지만 레버 자체는 여전히 붙어 있어 완전히 떨어졌다는 설정과 다르다. 복도의 작은 표지는 선명하게 판독하기 어렵고, 불가능한 반사는 없다.",
        "entities": "검은 단발의 중년 동아시아계 여성과 헝클어진 검은 머리의 젊은 동아시아계 남성 두 명만 있다. 미연의 남색 옷과 부은 눈가·볼의 멍은 이전 장면에 부합한다. 현우는 참고와 같은 남색 반팔 상의에 겉옷이 없고, 드러난 무릎과 정강이에 상처가 있다. 얼굴은 포옹에 가려 정확한 얼굴 일치와 멍의 정도를 충분히 확인할 수 없다. 운동화는 보이지만 내부의 연락 카드는 보이지 않으며, 이를 보여 주기 위해 구도를 바꿀 필요는 없다.",
        "hard_violations": [],
        "physics": "미연의 접힌 무릎과 하퇴, 현우의 무릎이 바닥에 닿아 두 사람을 지지한다. 뒤로 접힌 발과 신발도 바닥에 자연스럽게 놓인다. 손은 상대의 어깨와 등에 접촉하고, 상체를 서로 기대는 자세는 무릎 꿇은 정지 포옹으로 가능하다. 의자와 책상은 다리로 지지되며 떠 있는 물체는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "무릎 지지는 자연스럽지만, 미디엄 숏보다 넓은 문간 전신 구도이고 포옹에 틈이 있으며 현우의 상의·맨발과 남아 있는 문손잡이가 연속성에 어긋납니다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "방 안에서 서로 밀착한 포옹과 기존 가구·현우의 남색 상의를 더 잘 유지하지만, 전신과 넓은 바닥을 보여 미디엄 숏 지시를 벗어나고 문손잡이도 남아 있습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 눈을 감고 미연 쪽으로 얼굴을 기울이며, 미연은 현우의 얼굴 쪽으로 시선을 낮춘다. 두 사람의 팔은 서로의 어깨와 등을 향하지만 허리와 몸통 사이에는 틈이 있어 빈틈없는 밀착은 약하다. 카메라를 보는 사람이나 방향을 확인할 휴대 소품은 없다.",
        "built_space": "카메라는 문밖에서 문간에 무릎 꿇은 두 사람을 바라본다. 왼쪽 앞에 열린 문짝 하나와 둥근 손잡이가 있고, 그 뒤로 별도의 문과 손잡이 하나가 더 보인다. 오른쪽 벽에는 스위치 하나가 있다. 회백색 투톤 벽은 참고와 유사하지만 참고의 책상과 목제 의자는 이 구도에서 확인되지 않는다. 열린 문 바로 안쪽보다는 문턱을 점유한 배치이며, 손잡이가 떨어졌다는 설정과 달리 앞문 손잡이가 붙어 있다. 반사는 없다.",
        "entities": "중년 동아시아계 여성 한 명과 앳된 동아시아계 남성 한 명만 보인다. 미연의 검은 단발, 남색 상의와 얼굴 부상은 참고에 대체로 부합한다. 현우의 헝클어진 검은 머리와 얼굴 멍은 맞지만 남색 참고 상의 대신 회갈색 티셔츠를 입고 있다. 긴 바지에는 부상 흔적으로 보이는 얼룩이 있고, 보이는 발은 맨발이어서 신발 속 연락 카드의 소지 상태를 뒷받침하지 못한다. 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "두 사람 모두 접힌 무릎과 하퇴가 바닥에 닿아 체중을 지지한다. 현우의 맨발과 미연의 구두도 바닥에 접촉한다. 서로의 어깨와 등을 잡은 손과 팔은 포옹으로 가능한 배치이며, 공중에 지지 없이 떠 있는 신체나 물체는 없다."
       },
       {
        "label": "A",
        "direction": "미연은 눈을 감고 현우의 어깨에 얼굴을 붙이고, 현우는 미연의 머리 뒤쪽으로 얼굴을 숙인다. 양쪽 팔이 상대의 등과 어깨를 감싸 몸통까지 밀착하므로 서로를 꽉 끌어안는 방향과 대상이 명확하다. 외부를 겨누거나 카메라를 응시하는 동작은 없다.",
        "built_space": "방 안에서 열린 출입문과 그 너머 복도를 보는 구도다. 왼쪽에 목제 의자 하나, 오른쪽에 금속 테두리 책상 하나, 뒤쪽에 열린 금속 문짝 하나와 벽 스위치 하나가 보인다. 참고의 투톤 벽과 낡은 의자·책상 재질을 이어 간다. 두 사람은 가구 사이의 방 안 바닥에 무릎을 꿇고 있다. 문손잡이 장착부는 파손되어 있지만 레버 자체는 여전히 붙어 있어 완전히 떨어졌다는 설정과 다르다. 복도의 작은 표지는 선명하게 판독하기 어렵고, 불가능한 반사는 없다.",
        "entities": "검은 단발의 중년 동아시아계 여성과 헝클어진 검은 머리의 젊은 동아시아계 남성 두 명만 있다. 미연의 남색 옷과 부은 눈가·볼의 멍은 이전 장면에 부합한다. 현우는 참고와 같은 남색 반팔 상의에 겉옷이 없고, 드러난 무릎과 정강이에 상처가 있다. 얼굴은 포옹에 가려 정확한 얼굴 일치와 멍의 정도를 충분히 확인할 수 없다. 운동화는 보이지만 내부의 연락 카드는 보이지 않으며, 이를 보여 주기 위해 구도를 바꿀 필요는 없다.",
        "hard_violations": [],
        "physics": "미연의 접힌 무릎과 하퇴, 현우의 무릎이 바닥에 닿아 두 사람을 지지한다. 뒤로 접힌 발과 신발도 바닥에 자연스럽게 놓인다. 손은 상대의 어깨와 등에 접촉하고, 상체를 서로 기대는 자세는 무릎 꿇은 정지 포옹으로 가능하다. 의자와 책상은 다리로 지지되며 떠 있는 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.667,
    "B": 1.667
   },
   "adjusted": {
    "A": 1.667,
    "B": 1.667
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1667,
   "A": 1667
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1667,
    "verdict_ko": "B는 지시된 미디엄 샷과 무릎 아래 좁은 공간 노출이라는 최우선 프레이밍 조건을 완벽히 충족하여 승리했습니다."
   },
   {
    "label": "A",
    "score": 1667,
    "verdict_ko": "A는 의상 색상과 신발 등 일부 디테일을 맞췄으나, 프레이밍이 풀 샷으로 넓어져 우선순위가 높은 샷 스케일 조건을 위반했습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 미연 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S31sh9_sel.png",
    "asset_id": "837c4c98-8a22-4be5-9705-2ba9a25410f8",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1113064>",
    "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-661b-7c98-b2e7-70d47b76c917",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S31sh9"
  }
 },
 "S34sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:56:33.309771+00:00",
  "fingerprint": "1ae1b0913db202dfe7998ce69ca37e16c110f56922be87c1698b335134ebeeae",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S34sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S34sh7_sel.png",
  "source_sha256": "9d4c1a42036af9fcd057fc4f93b078ecfa351515af256e6e8ea6a8d0fd162026",
  "file": "S34sh7_cine.png",
  "staged_sha256": "6610dcad9186c9cb98ce501eedbc1612fdbcae966c93c834ae8350e7a3c8491b",
  "latency_ms": 12098
 },
 "S34sh10::signage": {
  "fp": "516c5a4e894f9913",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S34sh10": {
  "input_fingerprint": "ec9fa7392f61563b",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 강렬한 붉은 사이렌 불빛이 방 안을 비추는 찰나, 천장을 올려다보며 얼어붙은 두 사람의 굳은 상체.\n\nLOCATION (lock): Inside the third-floor detention room of the refugee administration building, washed by the red emergency light specified in the shot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office ceiling (Above the pair as they react to the alarm) — Its underside is partially visible above their raised faces; used as Provide a real spatial destination for the upward reaction without requiring a visible alarm fixture.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Capture the specified intense red alarm-light pulse across their faces and the room, preserving enough shadow detail to distinguish both reactions without showing an invented fixture.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the room surfaces and the opened doorway with its broken handle. Exclude the earlier unalarmed lighting state; allow the emergency red light to affect the room.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The third-floor room door remains open with its handle broken off. The management office remains lit; no additional alarm-light color is established. 현우: He remains in the opened room with facial bruises, an injured leg and no outer garment. The contact card remains concealed in his shoe. 미연: She remains in the opened room with her badly swollen face and beating injuries.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 강렬한 붉은 사이렌 불빛이 방 안을 비추는 찰나, 천장을 올려다보며 얼어붙은 두 사람의 굳은 상체.\n\nLOCATION (lock): Inside the third-floor detention room of the refugee administration building, washed by the red emergency light specified in the shot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office ceiling (Above the pair as they react to the alarm) — Its underside is partially visible above their raised faces; used as Provide a real spatial destination for the upward reaction without requiring a visible alarm fixture.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Capture the specified intense red alarm-light pulse across their faces and the room, preserving enough shadow detail to distinguish both reactions without showing an invented fixture.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the room surfaces and the opened doorway with its broken handle. Exclude the earlier unalarmed lighting state; allow the emergency red light to affect the room.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The third-floor room door remains open with its handle broken off. The management office remains lit; no additional alarm-light color is established. 현우: He remains in the opened room with facial bruises, an injured leg and no outer garment. The contact card remains concealed in his shoe. 미연: She remains in the opened room with her badly swollen face and beating injuries.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 강렬한 붉은 사이렌 불빛이 방 안을 비추는 찰나, 천장을 올려다보며 얼어붙은 두 사람의 굳은 상체.\n\nLOCATION (lock): Inside the third-floor detention room of the refugee administration building, washed by the red emergency light specified in the shot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office ceiling (Above the pair as they react to the alarm) — Its underside is partially visible above their raised faces; used as Provide a real spatial destination for the upward reaction without requiring a visible alarm fixture.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Capture the specified intense red alarm-light pulse across their faces and the room, preserving enough shadow detail to distinguish both reactions without showing an invented fixture.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the room surfaces and the opened doorway with its broken handle. Exclude the earlier unalarmed lighting state; allow the emergency red light to affect the room.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The third-floor room door remains open with its handle broken off. The management office remains lit; no additional alarm-light color is established. 현우: He remains in the opened room with facial bruises, an injured leg and no outer garment. The contact card remains concealed in his shoe. 미연: She remains in the opened room with her badly swollen face and beating injuries.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "두 인물의 시선과 고개가 모두 프레임 상단에 보이는 천장을 향해 고정되어 있음.",
    "built_space": "배경에 부서진 손잡이가 달린 열린 철문과 바깥 복도, 실내 천장이 레퍼런스와 동일한 구조로 배치됨. 붉은 조명이 공간 전체를 덮고 있음.",
    "entities": "현우(얼굴의 상처, 헝클어진 머리, 남색 상의)와 미연(멍든 얼굴, 머리스타일, 남색 상의) 모두 캐릭터 레퍼런스와 정확히 일치함.",
    "hard_violations": [],
    "physics": "두 사람 모두 중력에 맞게 상체를 꼿꼿이 세우고 있어 안정적으로 지지된 자세를 보여줌."
   },
   {
    "label": "B",
    "direction": "두 인물 모두 시선을 위로 향해 천장을 바라보고 있음.",
    "built_space": "천장의 구조물과 배경의 열린 문, 우측의 테이블이 보이나, 열린 문의 손잡이 디테일이 이전 샷 레퍼런스와 불일치함.",
    "entities": "현우와 미연의 얼굴 및 의상은 레퍼런스와 일치하지만, 프레임에 잡힌 현우의 오른손에 해부학적 오류가 있음.",
    "hard_violations": [
     "[gemini-pro] 현우의 오른손 손가락이 6개로 렌더링됨 (물리적으로 불가능한 해부학)."
    ],
    "physics": "허공에 들어올린 팔과 손짓은 어깨를 통해 지지되고 있으나, 신체 구조 자체에 근본적인 오류가 존재함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "붉은 조명 아래 천장을 올려다보는 미디엄 샷, 지정된 공간(부서진 문 손잡이 포함)과 두 인물의 외양 및 상처를 레퍼런스에 맞춰 매우 충실하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "샷의 연출과 조명 분위기는 양호하나, 인물의 손가락 구조에서 치명적인 해부학적 오류가 발생하여 규정을 크게 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 인물의 시선과 고개가 모두 프레임 상단에 보이는 천장을 향해 고정되어 있음.",
        "built_space": "배경에 부서진 손잡이가 달린 열린 철문과 바깥 복도, 실내 천장이 레퍼런스와 동일한 구조로 배치됨. 붉은 조명이 공간 전체를 덮고 있음.",
        "entities": "현우(얼굴의 상처, 헝클어진 머리, 남색 상의)와 미연(멍든 얼굴, 머리스타일, 남색 상의) 모두 캐릭터 레퍼런스와 정확히 일치함.",
        "hard_violations": [],
        "physics": "두 사람 모두 중력에 맞게 상체를 꼿꼿이 세우고 있어 안정적으로 지지된 자세를 보여줌."
       },
       {
        "label": "B",
        "direction": "두 인물 모두 시선을 위로 향해 천장을 바라보고 있음.",
        "built_space": "천장의 구조물과 배경의 열린 문, 우측의 테이블이 보이나, 열린 문의 손잡이 디테일이 이전 샷 레퍼런스와 불일치함.",
        "entities": "현우와 미연의 얼굴 및 의상은 레퍼런스와 일치하지만, 프레임에 잡힌 현우의 오른손에 해부학적 오류가 있음.",
        "hard_violations": [
         "현우의 오른손 손가락이 6개로 렌더링됨 (물리적으로 불가능한 해부학)."
        ],
        "physics": "허공에 들어올린 팔과 손짓은 어깨를 통해 지지되고 있으나, 신체 구조 자체에 근본적인 오류가 존재함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "붉은 조명 아래 천장을 올려다보는 미디엄 샷, 지정된 공간(부서진 문 손잡이 포함)과 두 인물의 외양 및 상처를 레퍼런스에 맞춰 매우 충실하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "샷의 연출과 조명 분위기는 양호하나, 인물의 손가락 구조에서 치명적인 해부학적 오류가 발생하여 규정을 크게 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "두 인물의 시선과 고개가 모두 프레임 상단에 보이는 천장을 향해 고정되어 있음.",
        "built_space": "배경에 부서진 손잡이가 달린 열린 철문과 바깥 복도, 실내 천장이 레퍼런스와 동일한 구조로 배치됨. 붉은 조명이 공간 전체를 덮고 있음.",
        "entities": "현우(얼굴의 상처, 헝클어진 머리, 남색 상의)와 미연(멍든 얼굴, 머리스타일, 남색 상의) 모두 캐릭터 레퍼런스와 정확히 일치함.",
        "hard_violations": [],
        "physics": "두 사람 모두 중력에 맞게 상체를 꼿꼿이 세우고 있어 안정적으로 지지된 자세를 보여줌."
       },
       {
        "label": "B",
        "direction": "두 인물 모두 시선을 위로 향해 천장을 바라보고 있음.",
        "built_space": "천장의 구조물과 배경의 열린 문, 우측의 테이블이 보이나, 열린 문의 손잡이 디테일이 이전 샷 레퍼런스와 불일치함.",
        "entities": "현우와 미연의 얼굴 및 의상은 레퍼런스와 일치하지만, 프레임에 잡힌 현우의 오른손에 해부학적 오류가 있음.",
        "hard_violations": [
         "현우의 오른손 손가락이 6개로 렌더링됨 (물리적으로 불가능한 해부학)."
        ],
        "physics": "허공에 들어올린 팔과 손짓은 어깨를 통해 지지되고 있으나, 신체 구조 자체에 근본적인 오류가 존재함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "올려다보는 시선과 강렬한 붉은빛은 맞지만, 천장과 벌린 팔의 비중이 커 굳은 상체의 순간이 덜 집중되고 미연의 짧은 소매가 이전 숏과 다릅니다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "두 사람의 굳은 상체와 천장을 향한 시선을 집중적으로 담고 열린 파손 문과 붉은 조명을 잘 유지하지만, 문손잡이 레버는 완전히 떨어져 있지 않습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 현우와 오른쪽 미연 모두 턱을 들고 카메라 위쪽의 천장 방향을 바라봅니다. 두 사람의 시선은 렌즈가 아니라 머리 위 공간으로 향합니다. 현우의 펼친 손은 화면 왼쪽으로 나와 있지만 특정 대상을 가리키지는 않습니다.",
        "built_space": "뒤쪽에 열린 금속문 하나와 문틀 하나, 투톤 벽, 오른쪽 책상 일부, 두 사람 사이 뒤쪽에 의자 등받이 하나가 보입니다. 위쪽 천장에는 길쭉한 매입 설비 두 개와 사각 점검구 하나가 크게 드러납니다. 문 너머 복도의 천장 설비도 보입니다. 방의 재질은 대체로 이어지지만, 천장이 화면 상단 약 삼분의 일을 차지하고 의자 위치는 이전 숏과 달라 보입니다. 손잡이 파손부는 인물과 문 가장자리에 가려 확인하기 어렵습니다.",
        "entities": "인물은 두 명뿐입니다. 현우는 앳된 동아시아계 남성으로 헝클어진 검은 머리, 남색 반소매 티셔츠, 얼굴의 상처가 보입니다. 미연은 중년 동아시아계 여성으로 검은 단발머리와 얼굴의 멍이 보이지만, 드러난 짧은 소매는 이전 숏의 긴 소매와 다릅니다. 얼굴 정체성은 대체로 참고 이미지에 부합하나 미연의 심한 부기는 뚜렷하지 않습니다. 붉은빛이 얼굴과 벽을 강하게 물들이며 별도 경보등이나 읽을 수 있는 글자는 없습니다. 신발 속 카드와 다리 부상은 프레임 밖입니다.",
        "hard_violations": [],
        "physics": "두 사람의 하체와 바닥 접점은 프레임 밖이라 서 있는지 무릎을 꿇었는지는 확정할 수 없습니다. 상체는 자연스럽게 아래쪽 몸통으로 이어지고 공중에 뜬 정황은 없습니다. 현우의 든 팔은 어깨와 굽힌 팔꿈치로 지탱되며, 미연의 손은 현우의 어깨와 가슴 경계에 닿아 있습니다. 열린 문은 문틀에 연결되어 있고 책상은 보이는 다리로 지지됩니다."
       },
       {
        "label": "B",
        "direction": "왼쪽 현우와 오른쪽 미연 모두 목을 젖히고 눈을 위로 향해 방 천장을 바라봅니다. 머리 위 천장 아랫면이 함께 보여 반응의 공간적 대상이 분명합니다. 두 사람 모두 카메라를 바라보지 않습니다.",
        "built_space": "왼쪽으로 열린 망입유리 금속문 하나, 그 뒤 문틀 하나, 흰색과 회색의 투톤 벽, 상단의 사무실 천장 일부가 보입니다. 복도에는 길쭉한 천장 설비 두 개와 켜진 사각 조명 하나가 식별됩니다. 두 사람은 열린 출입구 앞 방 안에 밀착해 있으며, 낮은 시점에서 상체를 중심으로 담았습니다. 파손된 손잡이 장착부는 보이지만 레버 일부는 붙어 있습니다. 의자와 책상은 이 구도 밖이므로 누락으로 볼 필요가 없습니다.",
        "entities": "현우와 미연 두 명만 등장합니다. 현우의 어린 남성 얼굴, 헝클어진 검은 머리, 남색 반소매와 얼굴 타박상은 참고 이미지에 잘 부합합니다. 미연도 중년 여성의 얼굴 윤곽, 검은 단발, 어두운 남색 상의와 볼의 부상이 이어집니다. 다만 심한 얼굴 부기는 명확하지 않습니다. 국적은 외형만으로 확인할 수 없습니다. 강한 붉은 조명이 두 얼굴과 방에 닿으면서 눈과 표정의 음영은 남아 있습니다. 별도 경보등, 추가 인물, 읽을 수 있는 글자는 없고 다리와 신발 속 카드는 프레임 밖입니다.",
        "hard_violations": [],
        "physics": "두 사람의 하체 접점은 잘려 있지만 몸통과 목의 연결이 자연스럽고, 뒤로 젖힌 머리는 목과 상체로 지탱됩니다. 서로 붙어 멈춘 자세에 비현실적인 부유나 관절 꺾임은 없습니다. 손에 든 물체는 없으며 열린 금속문도 문틀에 연결된 상태입니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "올려다보는 시선과 강렬한 붉은빛은 맞지만, 천장과 벌린 팔의 비중이 커 굳은 상체의 순간이 덜 집중되고 미연의 짧은 소매가 이전 숏과 다릅니다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "두 사람의 굳은 상체와 천장을 향한 시선을 집중적으로 담고 열린 파손 문과 붉은 조명을 잘 유지하지만, 문손잡이 레버는 완전히 떨어져 있지 않습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽 현우와 오른쪽 미연 모두 턱을 들고 카메라 위쪽의 천장 방향을 바라봅니다. 두 사람의 시선은 렌즈가 아니라 머리 위 공간으로 향합니다. 현우의 펼친 손은 화면 왼쪽으로 나와 있지만 특정 대상을 가리키지는 않습니다.",
        "built_space": "뒤쪽에 열린 금속문 하나와 문틀 하나, 투톤 벽, 오른쪽 책상 일부, 두 사람 사이 뒤쪽에 의자 등받이 하나가 보입니다. 위쪽 천장에는 길쭉한 매입 설비 두 개와 사각 점검구 하나가 크게 드러납니다. 문 너머 복도의 천장 설비도 보입니다. 방의 재질은 대체로 이어지지만, 천장이 화면 상단 약 삼분의 일을 차지하고 의자 위치는 이전 숏과 달라 보입니다. 손잡이 파손부는 인물과 문 가장자리에 가려 확인하기 어렵습니다.",
        "entities": "인물은 두 명뿐입니다. 현우는 앳된 동아시아계 남성으로 헝클어진 검은 머리, 남색 반소매 티셔츠, 얼굴의 상처가 보입니다. 미연은 중년 동아시아계 여성으로 검은 단발머리와 얼굴의 멍이 보이지만, 드러난 짧은 소매는 이전 숏의 긴 소매와 다릅니다. 얼굴 정체성은 대체로 참고 이미지에 부합하나 미연의 심한 부기는 뚜렷하지 않습니다. 붉은빛이 얼굴과 벽을 강하게 물들이며 별도 경보등이나 읽을 수 있는 글자는 없습니다. 신발 속 카드와 다리 부상은 프레임 밖입니다.",
        "hard_violations": [],
        "physics": "두 사람의 하체와 바닥 접점은 프레임 밖이라 서 있는지 무릎을 꿇었는지는 확정할 수 없습니다. 상체는 자연스럽게 아래쪽 몸통으로 이어지고 공중에 뜬 정황은 없습니다. 현우의 든 팔은 어깨와 굽힌 팔꿈치로 지탱되며, 미연의 손은 현우의 어깨와 가슴 경계에 닿아 있습니다. 열린 문은 문틀에 연결되어 있고 책상은 보이는 다리로 지지됩니다."
       },
       {
        "label": "A",
        "direction": "왼쪽 현우와 오른쪽 미연 모두 목을 젖히고 눈을 위로 향해 방 천장을 바라봅니다. 머리 위 천장 아랫면이 함께 보여 반응의 공간적 대상이 분명합니다. 두 사람 모두 카메라를 바라보지 않습니다.",
        "built_space": "왼쪽으로 열린 망입유리 금속문 하나, 그 뒤 문틀 하나, 흰색과 회색의 투톤 벽, 상단의 사무실 천장 일부가 보입니다. 복도에는 길쭉한 천장 설비 두 개와 켜진 사각 조명 하나가 식별됩니다. 두 사람은 열린 출입구 앞 방 안에 밀착해 있으며, 낮은 시점에서 상체를 중심으로 담았습니다. 파손된 손잡이 장착부는 보이지만 레버 일부는 붙어 있습니다. 의자와 책상은 이 구도 밖이므로 누락으로 볼 필요가 없습니다.",
        "entities": "현우와 미연 두 명만 등장합니다. 현우의 어린 남성 얼굴, 헝클어진 검은 머리, 남색 반소매와 얼굴 타박상은 참고 이미지에 잘 부합합니다. 미연도 중년 여성의 얼굴 윤곽, 검은 단발, 어두운 남색 상의와 볼의 부상이 이어집니다. 다만 심한 얼굴 부기는 명확하지 않습니다. 국적은 외형만으로 확인할 수 없습니다. 강한 붉은 조명이 두 얼굴과 방에 닿으면서 눈과 표정의 음영은 남아 있습니다. 별도 경보등, 추가 인물, 읽을 수 있는 글자는 없고 다리와 신발 속 카드는 프레임 밖입니다.",
        "hard_violations": [],
        "physics": "두 사람의 하체 접점은 잘려 있지만 몸통과 목의 연결이 자연스럽고, 뒤로 젖힌 머리는 목과 상체로 지탱됩니다. 서로 붙어 멈춘 자세에 비현실적인 부유나 관절 꺾임은 없습니다. 손에 든 물체는 없으며 열린 금속문도 문틀에 연결된 상태입니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.153
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.903
   },
   "violations": {
    "B": [
     "[gemini-pro] 현우의 오른손 손가락이 6개로 렌더링됨 (물리적으로 불가능한 해부학)."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 903
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "붉은 조명 아래 천장을 올려다보는 미디엄 샷, 지정된 공간(부서진 문 손잡이 포함)과 두 인물의 외양 및 상처를 레퍼런스에 맞춰 매우 충실하게 구현했습니다."
   },
   {
    "label": "B",
    "score": 903,
    "verdict_ko": "샷의 연출과 조명 분위기는 양호하나, 인물의 손가락 구조에서 치명적인 해부학적 오류가 발생하여 규정을 크게 위반했습니다.  ★위반: [gemini-pro] 현우의 오른손 손가락이 6개로 렌더링됨 (물리적으로 불가능한 해부학)."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 미연, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S34sh7_sel.png",
    "asset_id": "84cd4095-cf56-4dc3-851e-09494116a0b0",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1113064>",
    "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-67d8-72ba-a114-eb88e4dae87e",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S34sh7"
  }
 },
 "S34sh10::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:57:34.723238+00:00",
  "fingerprint": "b9ccfe2a6bc3772a8fdf5700142cdcc96fe9ea46d703d6da9ed77d9f00adbbf9",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S34sh10_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S34sh10_sel.png",
  "source_sha256": "59c2c86ceaac20394e4566f217c11faf47a2e1bfb52ce9afc5378217f2a05a0d",
  "file": "S34sh10_cine.png",
  "staged_sha256": "48ca7a6918f8851c5763c0dc1cfe7bc93e388cddb46d537e0f8b2da81a8e5f0d",
  "latency_ms": 11331
 },
 "S35sh1::signage": {
  "fp": "516d3e716cf35998",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::aaf3aac37588d4aa": {
  "subjects": [],
  "subject_text": "해일이 덮치는 인천 난민촌 판잣집 구역\n거대한 콘크리트 방벽 아래 낮은 판잣집과 낡은 임시 주거 시설이 조밀하게 모여 있는 구역. 건물 사이로 좁은 길이 이어진다.",
  "identity": "canonical",
  "scope_id": "L192",
  "scope_role": "location_exterior",
  "scope_sha": "f93d49212003ab29"
 },
 "groupbg::camp_inundation_street": {
  "input_fingerprint": "252c428074a19472",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "camp_inundation_street",
    "tags": [
     "S35sh1",
     "S35sh3"
    ]
   },
   "context_sig": "d17b8ea2372b19fa"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On an outdoor lane among the refugee settlement's makeshift homes, under emergency warning lights as the ground trembles.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n해일이 덮치는 인천 난민촌 판잣집 구역: 거대한 바닷물 장벽이 허술한 주거지구 구조물들을 쓸어버리는 재난 순간. (특징: 집어삼킬 듯 밀려오는 검푸른 바닷물; 부서지며 떠오르는 판잣집과 컨테이너 파편들)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 바닷물이 밀려들고, 맨몸으로 대피하는 사람들! 빠르게 덮치는 쓰나미.\n\nTIME OF DAY (lock): night.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On an outdoor lane among the refugee settlement's makeshift homes, under emergency warning lights as the ground trembles.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n해일이 덮치는 인천 난민촌 판잣집 구역: 거대한 바닷물 장벽이 허술한 주거지구 구조물들을 쓸어버리는 재난 순간. (특징: 집어삼킬 듯 밀려오는 검푸른 바닷물; 부서지며 떠오르는 판잣집과 컨테이너 파편들)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 바닷물이 밀려들고, 맨몸으로 대피하는 사람들! 빠르게 덮치는 쓰나미.\n\nTIME OF DAY (lock): night.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_camp_inundation_street_5076fe.png",
  "asset_id": "19f46218-6a77-42f8-9095-640b2dac045d",
  "input_asset_ids": [
   "d7134b9e-61cf-401a-932e-ef0971017649"
  ],
  "origin_tag": "S35sh1",
  "place_text": "On an outdoor lane among the refugee settlement's makeshift homes, under emergency warning lights as the ground trembles.",
  "origin_inputs": {
   "place_text": "On an outdoor lane among the refugee settlement's makeshift homes, under emergency warning lights as the ground trembles.",
   "time_of_day_en": "night",
   "conti_asset_id": "d7134b9e-61cf-401a-932e-ef0971017649"
  }
 },
 "S35sh1::bgfirst_bg": {
  "input_fingerprint": "eb9e5b46b5acf54d",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 요란한 비상 사이렌 불빛 아래, 흔들리는 바닥을 디딘 채 공포에 질린 표정으로 허공을 올려다보는 난민촌 사람들의 굳은 전경.\n\nLOCATION (lock): On an outdoor lane among the refugee settlement's makeshift homes, under emergency warning lights as the ground trembles.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Refugee-camp ground (Beginning to shake beneath the residents) — Visible beneath their unevenly braced feet and into the foreground; used as Connect bodily instability to the space along which the camera will retreat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the specified emergency-alarm illumination register intermittently within the nighttime exposure, without assigning an unsupported color or adding visible fixtures.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 요란한 비상 사이렌 불빛 아래, 흔들리는 바닥을 디딘 채 공포에 질린 표정으로 허공을 올려다보는 난민촌 사람들의 굳은 전경.\n\nLOCATION (lock): On an outdoor lane among the refugee settlement's makeshift homes, under emergency warning lights as the ground trembles.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Refugee-camp ground (Beginning to shake beneath the residents) — Visible beneath their unevenly braced feet and into the foreground; used as Connect bodily instability to the space along which the camera will retreat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the specified emergency-alarm illumination register intermittently within the nighttime exposure, without assigning an unsupported color or adding visible fixtures.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S35sh1__bgfirst_bg.png",
  "asset_id": "c4a1bc33-f30d-4b6f-ae75-66d7dbdc2b52",
  "input_asset_ids": [
   "d7134b9e-61cf-401a-932e-ef0971017649",
   "19f46218-6a77-42f8-9095-640b2dac045d"
  ]
 },
 "S35sh1": {
  "input_fingerprint": "71cf5c9fa68d83cf",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 요란한 비상 사이렌 불빛 아래, 흔들리는 바닥을 디딘 채 공포에 질린 표정으로 허공을 올려다보는 난민촌 사람들의 굳은 전경.\n\nLOCATION (lock): On an outdoor lane among the refugee settlement's makeshift homes, under emergency warning lights as the ground trembles. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Refugee-camp ground (Beginning to shake beneath the residents) — Visible beneath their unevenly braced feet and into the foreground; used as Connect bodily instability to the space along which the camera will retreat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the specified emergency-alarm illumination register intermittently within the nighttime exposure, without assigning an unsupported color or adding visible fixtures.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall breach remains open and seawater is advancing into the nighttime settlement. The ground is beginning to shake.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 요란한 비상 사이렌 불빛 아래, 흔들리는 바닥을 디딘 채 공포에 질린 표정으로 허공을 올려다보는 난민촌 사람들의 굳은 전경.\n\nLOCATION (lock): On an outdoor lane among the refugee settlement's makeshift homes, under emergency warning lights as the ground trembles. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Refugee-camp ground (Beginning to shake beneath the residents) — Visible beneath their unevenly braced feet and into the foreground; used as Connect bodily instability to the space along which the camera will retreat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the specified emergency-alarm illumination register intermittently within the nighttime exposure, without assigning an unsupported color or adding visible fixtures.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall breach remains open and seawater is advancing into the nighttime settlement. The ground is beginning to shake.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 요란한 비상 사이렌 불빛 아래, 흔들리는 바닥을 디딘 채 공포에 질린 표정으로 허공을 올려다보는 난민촌 사람들의 굳은 전경.\n\nLOCATION (lock): On an outdoor lane among the refugee settlement's makeshift homes, under emergency warning lights as the ground trembles. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Refugee-camp ground (Beginning to shake beneath the residents) — Visible beneath their unevenly braced feet and into the foreground; used as Connect bodily instability to the space along which the camera will retreat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the specified emergency-alarm illumination register intermittently within the nighttime exposure, without assigning an unsupported color or adding visible fixtures.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The seawall breach remains open and seawater is advancing into the nighttime settlement. The ground is beginning to shake.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S35sh1__bgfirst_bg.png",
     "asset_id": "c4a1bc33-f30d-4b6f-ae75-66d7dbdc2b52",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S35sh1.png",
     "asset_id": "d7134b9e-61cf-401a-932e-ef0971017649",
     "role": "conti_light"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_camp_inundation_street_5076fe.png",
     "asset_id": "19f46218-6a77-42f8-9095-640b2dac045d",
     "role": "bgfirst_group_bg"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "인물들의 시선이 모두 허공과 파도를 향해 위로 향해 있습니다.",
    "built_space": "판자촌과 갈라진 바닥은 구현되었으나, 전경 인물들의 크기가 주변 컨테이너 건물에 비해 과도하게 큽니다.",
    "entities": "다민족 설정에 비해 인물들의 외형과 복장이 단조롭고 작위적입니다.",
    "hard_violations": [
     "[gemini-pro] 주변 구조물 대비 전경 인물들의 물리적 축척 및 비율 오류",
     "[gemini-pro] 우측 아기를 안은 여성의 해부학적으로 불가능하게 늘어난 다리 구조"
    ],
    "physics": "발은 땅에 있으나 자세가 부자연스럽고, 현장에 없는 강한 백색광이 인물들을 비추고 있습니다."
   },
   {
    "label": "B",
    "direction": "인물들이 위쪽 하늘과 밀려오는 파도를 향해 시선을 고정하고 있습니다.",
    "built_space": "참고 이미지의 판자촌 골목과 비상등을 정확히 재현했으며, 인물들이 구조물과 올바른 비율로 배치되었습니다.",
    "entities": "지정된 대로 다양한 인종과 연령대의 사람들이 다채로운 복장을 입고 공포에 질린 표정을 짓고 있습니다.",
    "hard_violations": [],
    "physics": "인물들이 바닥을 단단히 딛거나 벽에 기대어 몸을 지탱하고 있으며, 붉은 조명이 물리적으로 자연스럽게 반사됩니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "다양한 인종의 난민들이 공포에 질려 허공을 바라보는 모습을 지정된 환경 조명과 정확한 공간 비율 내에서 사실적으로 잘 구현했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "배경 구조물 대비 인물의 크기가 비정상적으로 크며, 공간에 존재하지 않는 인위적인 조명과 해부학적 오류가 있어 사실성이 떨어집니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "인물들의 시선이 모두 허공과 파도를 향해 위로 향해 있습니다.",
        "built_space": "판자촌과 갈라진 바닥은 구현되었으나, 전경 인물들의 크기가 주변 컨테이너 건물에 비해 과도하게 큽니다.",
        "entities": "다민족 설정에 비해 인물들의 외형과 복장이 단조롭고 작위적입니다.",
        "hard_violations": [
         "주변 구조물 대비 전경 인물들의 물리적 축척 및 비율 오류",
         "우측 아기를 안은 여성의 해부학적으로 불가능하게 늘어난 다리 구조"
        ],
        "physics": "발은 땅에 있으나 자세가 부자연스럽고, 현장에 없는 강한 백색광이 인물들을 비추고 있습니다."
       },
       {
        "label": "B",
        "direction": "인물들이 위쪽 하늘과 밀려오는 파도를 향해 시선을 고정하고 있습니다.",
        "built_space": "참고 이미지의 판자촌 골목과 비상등을 정확히 재현했으며, 인물들이 구조물과 올바른 비율로 배치되었습니다.",
        "entities": "지정된 대로 다양한 인종과 연령대의 사람들이 다채로운 복장을 입고 공포에 질린 표정을 짓고 있습니다.",
        "hard_violations": [],
        "physics": "인물들이 바닥을 단단히 딛거나 벽에 기대어 몸을 지탱하고 있으며, 붉은 조명이 물리적으로 자연스럽게 반사됩니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "다양한 인종의 난민들이 공포에 질려 허공을 바라보는 모습을 지정된 환경 조명과 정확한 공간 비율 내에서 사실적으로 잘 구현했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "배경 구조물 대비 인물의 크기가 비정상적으로 크며, 공간에 존재하지 않는 인위적인 조명과 해부학적 오류가 있어 사실성이 떨어집니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "인물들의 시선이 모두 허공과 파도를 향해 위로 향해 있습니다.",
        "built_space": "판자촌과 갈라진 바닥은 구현되었으나, 전경 인물들의 크기가 주변 컨테이너 건물에 비해 과도하게 큽니다.",
        "entities": "다민족 설정에 비해 인물들의 외형과 복장이 단조롭고 작위적입니다.",
        "hard_violations": [
         "주변 구조물 대비 전경 인물들의 물리적 축척 및 비율 오류",
         "우측 아기를 안은 여성의 해부학적으로 불가능하게 늘어난 다리 구조"
        ],
        "physics": "발은 땅에 있으나 자세가 부자연스럽고, 현장에 없는 강한 백색광이 인물들을 비추고 있습니다."
       },
       {
        "label": "B",
        "direction": "인물들이 위쪽 하늘과 밀려오는 파도를 향해 시선을 고정하고 있습니다.",
        "built_space": "참고 이미지의 판자촌 골목과 비상등을 정확히 재현했으며, 인물들이 구조물과 올바른 비율로 배치되었습니다.",
        "entities": "지정된 대로 다양한 인종과 연령대의 사람들이 다채로운 복장을 입고 공포에 질린 표정을 짓고 있습니다.",
        "hard_violations": [],
        "physics": "인물들이 바닥을 단단히 딛거나 벽에 기대어 몸을 지탱하고 있으며, 붉은 조명이 물리적으로 자연스럽게 반사됩니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "야간의 공포와 서로 의지하는 주민들은 잘 표현했지만, 사람들이 길 양옆에 정렬되어 흔들리는 바닥과의 관계가 약하고 전경 전봇대 등 장소 배치도 참조와 차이가 난다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "참조 장소의 구조를 유지하면서 넓은 전경 바닥, 불균형하게 버티는 발, 허공을 향한 시선을 한 와이드 숏에 연결해 핵심 순간을 더 충실히 구현한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 앞의 두 남성과 여성 무리, 오른쪽 앞의 가족은 고개를 들어 화면 위쪽의 보이지 않는 대상을 바라본다. 다만 중앙 후방의 일부 주민은 정면이나 주변 사람을 보는 듯하여 상향 시선이 군중 전체에 일관되지는 않는다. 무기나 특정 대상을 겨누는 소품은 없다.",
        "built_space": "골목 양쪽에 골판금·목재·방수포 가옥이 있고, 중앙 통로가 깊숙이 열린다. 붉은 경고등은 일곱 개가 식별되며 전선과 전봇대에 연결되어 있다. 왼쪽 전경의 굵은 전봇대에 두 남성이 기대고, 나머지 주민 대부분은 길 양옆에 모여 있다. 참조의 재료와 방벽 배경은 유지되지만 전경 전봇대와 가까운 가옥의 배치가 달라 정확한 장소 일치도는 떨어진다. 젖은 바닥의 붉은 반사는 가능한 배치다.",
        "entities": "성인 남녀와 어린이로 구성된 다인종 주민 무리가 보이며, 동아시아계로 보이는 인물이 다수이고 오른쪽에는 흑인으로 보이는 남성도 있다. 개별 국적은 확인할 수 없다. 낡은 일상복, 임시 가옥, 전선, 경고등, 갈라진 젖은 도로, 방벽 너머의 큰 파도가 보인다. 눈과 얼굴은 정상적인 인간 형태다. 전봇대에 작은 표식이 있으나 읽을 수 있는 문구는 식별되지 않는다. 방벽의 열린 파손부와 바닷물의 전진 자체는 명확히 드러나지 않는다.",
        "hard_violations": [],
        "physics": "앞쪽 주민들의 발은 노면이나 돌무더기에 닿아 있고, 왼쪽 남성들은 손으로 전봇대 또는 벽을 짚는다. 여성들은 서로 팔을 붙잡고 오른쪽 가족도 아이를 감싸며 지지한다. 뒤쪽 주민 몇 명은 무릎을 굽히고 몸을 낮춘다. 지지 없이 떠 있는 몸이나 물체는 보이지 않는다. 다만 앞쪽 여러 사람의 비교적 곧은 자세는 지면 진동에 대응하는 불균형을 약하게 전달한다."
       },
       {
        "label": "B",
        "direction": "왼쪽 앞 남성과 그 손을 잡은 아이는 화면 위쪽을 올려다본다. 중앙 남성은 얼굴 앞에 팔을 들고 위를 보며, 아기를 안은 여성과 오른쪽 남성도 허공을 향해 고개를 든다. 뒤쪽의 웅크린 주민들은 시선이 덜 분명하지만, 주요 인물들의 시선 대상은 일관되게 화면 밖 상공이다. 팔은 균형을 잡거나 얼굴을 가리는 방향으로 뻗어 있다.",
        "built_space": "양쪽 임시 가옥 사이의 넓은 골목, 왼쪽 앞 파란 드럼통 한 개, 오른쪽 수레 한 대, 머리 위 전선, 오른쪽 후방 방벽이 참조와 잘 대응한다. 붉은 경고등은 일곱 개가 보이며 전경부터 원경까지 거리감에 맞춰 작아진다. 주민들은 길 위 여러 깊이에 분산되어 있고, 발밑에서 화면 하단까지 갈라지고 젖은 노면이 넓게 이어진다. 낮은 와이드 시점에서 사람과 바닥의 관계가 명확하며 물웅덩이의 반사도 광원 위치와 양립한다.",
        "entities": "성인 다섯 명과 어린이·아기 네 명으로 보이는 주민 아홉 명이 있다. 여러 피부색의 남녀가 보이며 국적이나 언어는 이미지로 확정할 수 없다. 낡은 셔츠와 바지, 치마, 품에 안긴 아기, 임시 가옥, 경고등, 균열과 물웅덩이, 방벽 너머 파도가 확인된다. 사람들은 구체적인 얼굴과 정상적인 눈을 가진 실사 인물로 표현되어 있다. 읽을 수 있는 글자는 보이지 않는다. 열린 방벽 파손부와 유입수의 이동 방향은 명확히 확인되지 않는다.",
        "hard_violations": [],
        "physics": "왼쪽 남성은 양발을 넓게 디디고 아이의 손을 잡으며, 아이도 기울어진 몸을 지면에 닿은 발과 잡은 손으로 지탱한다. 중앙 남성은 무릎을 굽히고 발을 벌려 버티며 뒤쪽 인물들도 웅크리거나 중심을 낮춘다. 여성은 두 발로 서서 팔과 몸통으로 아기를 받친다. 오른쪽 남성 역시 벌린 발과 뻗은 팔로 균형을 잡는다. 지지 없는 부유나 불가능한 자세는 보이지 않으며, 흔들리기 시작한 지면에 반응하는 순간으로 읽힌다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "야간의 공포와 서로 의지하는 주민들은 잘 표현했지만, 사람들이 길 양옆에 정렬되어 흔들리는 바닥과의 관계가 약하고 전경 전봇대 등 장소 배치도 참조와 차이가 난다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "참조 장소의 구조를 유지하면서 넓은 전경 바닥, 불균형하게 버티는 발, 허공을 향한 시선을 한 와이드 숏에 연결해 핵심 순간을 더 충실히 구현한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽 앞의 두 남성과 여성 무리, 오른쪽 앞의 가족은 고개를 들어 화면 위쪽의 보이지 않는 대상을 바라본다. 다만 중앙 후방의 일부 주민은 정면이나 주변 사람을 보는 듯하여 상향 시선이 군중 전체에 일관되지는 않는다. 무기나 특정 대상을 겨누는 소품은 없다.",
        "built_space": "골목 양쪽에 골판금·목재·방수포 가옥이 있고, 중앙 통로가 깊숙이 열린다. 붉은 경고등은 일곱 개가 식별되며 전선과 전봇대에 연결되어 있다. 왼쪽 전경의 굵은 전봇대에 두 남성이 기대고, 나머지 주민 대부분은 길 양옆에 모여 있다. 참조의 재료와 방벽 배경은 유지되지만 전경 전봇대와 가까운 가옥의 배치가 달라 정확한 장소 일치도는 떨어진다. 젖은 바닥의 붉은 반사는 가능한 배치다.",
        "entities": "성인 남녀와 어린이로 구성된 다인종 주민 무리가 보이며, 동아시아계로 보이는 인물이 다수이고 오른쪽에는 흑인으로 보이는 남성도 있다. 개별 국적은 확인할 수 없다. 낡은 일상복, 임시 가옥, 전선, 경고등, 갈라진 젖은 도로, 방벽 너머의 큰 파도가 보인다. 눈과 얼굴은 정상적인 인간 형태다. 전봇대에 작은 표식이 있으나 읽을 수 있는 문구는 식별되지 않는다. 방벽의 열린 파손부와 바닷물의 전진 자체는 명확히 드러나지 않는다.",
        "hard_violations": [],
        "physics": "앞쪽 주민들의 발은 노면이나 돌무더기에 닿아 있고, 왼쪽 남성들은 손으로 전봇대 또는 벽을 짚는다. 여성들은 서로 팔을 붙잡고 오른쪽 가족도 아이를 감싸며 지지한다. 뒤쪽 주민 몇 명은 무릎을 굽히고 몸을 낮춘다. 지지 없이 떠 있는 몸이나 물체는 보이지 않는다. 다만 앞쪽 여러 사람의 비교적 곧은 자세는 지면 진동에 대응하는 불균형을 약하게 전달한다."
       },
       {
        "label": "A",
        "direction": "왼쪽 앞 남성과 그 손을 잡은 아이는 화면 위쪽을 올려다본다. 중앙 남성은 얼굴 앞에 팔을 들고 위를 보며, 아기를 안은 여성과 오른쪽 남성도 허공을 향해 고개를 든다. 뒤쪽의 웅크린 주민들은 시선이 덜 분명하지만, 주요 인물들의 시선 대상은 일관되게 화면 밖 상공이다. 팔은 균형을 잡거나 얼굴을 가리는 방향으로 뻗어 있다.",
        "built_space": "양쪽 임시 가옥 사이의 넓은 골목, 왼쪽 앞 파란 드럼통 한 개, 오른쪽 수레 한 대, 머리 위 전선, 오른쪽 후방 방벽이 참조와 잘 대응한다. 붉은 경고등은 일곱 개가 보이며 전경부터 원경까지 거리감에 맞춰 작아진다. 주민들은 길 위 여러 깊이에 분산되어 있고, 발밑에서 화면 하단까지 갈라지고 젖은 노면이 넓게 이어진다. 낮은 와이드 시점에서 사람과 바닥의 관계가 명확하며 물웅덩이의 반사도 광원 위치와 양립한다.",
        "entities": "성인 다섯 명과 어린이·아기 네 명으로 보이는 주민 아홉 명이 있다. 여러 피부색의 남녀가 보이며 국적이나 언어는 이미지로 확정할 수 없다. 낡은 셔츠와 바지, 치마, 품에 안긴 아기, 임시 가옥, 경고등, 균열과 물웅덩이, 방벽 너머 파도가 확인된다. 사람들은 구체적인 얼굴과 정상적인 눈을 가진 실사 인물로 표현되어 있다. 읽을 수 있는 글자는 보이지 않는다. 열린 방벽 파손부와 유입수의 이동 방향은 명확히 확인되지 않는다.",
        "hard_violations": [],
        "physics": "왼쪽 남성은 양발을 넓게 디디고 아이의 손을 잡으며, 아이도 기울어진 몸을 지면에 닿은 발과 잡은 손으로 지탱한다. 중앙 남성은 무릎을 굽히고 발을 벌려 버티며 뒤쪽 인물들도 웅크리거나 중심을 낮춘다. 여성은 두 발로 서서 팔과 몸통으로 아기를 받친다. 오른쪽 남성 역시 벌린 발과 뻗은 팔로 균형을 잡는다. 지지 없는 부유나 불가능한 자세는 보이지 않으며, 흔들리기 시작한 지면에 반응하는 순간으로 읽힌다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.571,
    "B": 1.778
   },
   "adjusted": {
    "A": 1.321,
    "B": 1.778
   },
   "violations": {
    "A": [
     "[gemini-pro] 주변 구조물 대비 전경 인물들의 물리적 축척 및 비율 오류",
     "[gemini-pro] 우측 아기를 안은 여성의 해부학적으로 불가능하게 늘어난 다리 구조"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1778,
   "A": 1321
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1778,
    "verdict_ko": "다양한 인종의 난민들이 공포에 질려 허공을 바라보는 모습을 지정된 환경 조명과 정확한 공간 비율 내에서 사실적으로 잘 구현했습니다."
   },
   {
    "label": "A",
    "score": 1321,
    "verdict_ko": "배경 구조물 대비 인물의 크기가 비정상적으로 크며, 공간에 존재하지 않는 인위적인 조명과 해부학적 오류가 있어 사실성이 떨어집니다.  ★위반: [gemini-pro] 주변 구조물 대비 전경 인물들의 물리적 축척 및 비율 오류 / [gemini-pro] 우측 아기를 안은 여성의 해부학적으로 불가능하게 늘어난 다리 구조"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_camp_inundation_street_5076fe.png",
    "asset_id": "19f46218-6a77-42f8-9095-640b2dac045d",
    "role": "bgfirst_group_bg"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-698f-7787-9423-de0ee35a4c16",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S35sh1__bgfirst_bg.png",
   "bg_asset_id": "c4a1bc33-f30d-4b6f-ae75-66d7dbdc2b52",
   "bg_record_key": "S35sh1::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "camp_inundation_street",
   "groupbg_asset_id": "19f46218-6a77-42f8-9095-640b2dac045d"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S35sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T05:59:42.512754+00:00",
  "fingerprint": "7664191713986e8cf0e08f7faa7bb861399605fe0c829d5ff5864dc92baa0dbd",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S35sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S35sh1_sel.png",
  "source_sha256": "f872bc579aec4c4894567c2071326fbabddaeb2eeac92a6f5654c4d5dbca5cfc",
  "file": "S35sh1_cine.png",
  "staged_sha256": "dda543c4253d6a91851566eceae3ec1f26d182b924d3d798d1ae27d06223df09",
  "latency_ms": 13800
 },
 "S35sh3::signage": {
  "fp": "aa08dc17650491ac",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S35sh3": {
  "input_fingerprint": "1a43543febe4bbc9",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 판잣집 지붕들 위로 솟구쳐 오른 거대한 검은 바닷물이 하늘을 완전히 뒤덮은 압도적인 찰나.\n\nLOCATION (lock): Above the roofs of the refugee settlement's shacks, where a towering nighttime surge rises over the dwellings. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Shack roofline retained as scale reference in the lower-center of the frame, midground; Towering seawater above the roofs in the upper-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Shack roofline (Below the rising tsunami) — Roof edges and partial sloping faces are seen sharply from below; used as Remain along the bottom edge as the essential scale reference; Towering seawater (Rising above the roofs and covering the sky); used as Occupy the field above the roofline as a continuous environmental threat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the black seawater within a restrained nighttime tonal range, keeping its advancing contours distinguishable from the roof silhouettes without introducing a source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rapidly advancing tsunami is engulfing the settlement following the seawall breach. The flooding has not subsided.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 판잣집 지붕들 위로 솟구쳐 오른 거대한 검은 바닷물이 하늘을 완전히 뒤덮은 압도적인 찰나.\n\nLOCATION (lock): Above the roofs of the refugee settlement's shacks, where a towering nighttime surge rises over the dwellings. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Shack roofline retained as scale reference in the lower-center of the frame, midground; Towering seawater above the roofs in the upper-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Shack roofline (Below the rising tsunami) — Roof edges and partial sloping faces are seen sharply from below; used as Remain along the bottom edge as the essential scale reference; Towering seawater (Rising above the roofs and covering the sky); used as Occupy the field above the roofline as a continuous environmental threat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the black seawater within a restrained nighttime tonal range, keeping its advancing contours distinguishable from the roof silhouettes without introducing a source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rapidly advancing tsunami is engulfing the settlement following the seawall breach. The flooding has not subsided.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 판잣집 지붕들 위로 솟구쳐 오른 거대한 검은 바닷물이 하늘을 완전히 뒤덮은 압도적인 찰나.\n\nLOCATION (lock): Above the roofs of the refugee settlement's shacks, where a towering nighttime surge rises over the dwellings. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Shack roofline retained as scale reference in the lower-center of the frame, midground; Towering seawater above the roofs in the upper-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Shack roofline (Below the rising tsunami) — Roof edges and partial sloping faces are seen sharply from below; used as Remain along the bottom edge as the essential scale reference; Towering seawater (Rising above the roofs and covering the sky); used as Occupy the field above the roofline as a continuous environmental threat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the black seawater within a restrained nighttime tonal range, keeping its advancing contours distinguishable from the roof silhouettes without introducing a source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rapidly advancing tsunami is engulfing the settlement following the seawall breach. The flooding has not subsided.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "해당 사항 없음.",
    "built_space": "판잣촌 골목길 시점으로 양옆에 판잣집들이 있고 중앙 하단은 지붕이 아닌 골목길 바닥이 보임.",
    "entities": "거대한 검은 파도, 판잣집 외벽, 붉은 조명. 지시된 대로 인물은 등장하지 않음.",
    "hard_violations": [
     "[gemini-pro] 명시된 카메라 위치('지붕들 위')와 구도('프레임 하단 중앙에 지붕선 배치')를 어기고 골목길 시점으로 렌더링함."
    ],
    "physics": "거대한 파도가 골목길 끝에서 솟구쳐 오르는 모습."
   },
   {
    "label": "B",
    "direction": "해당 사항 없음 (인물 및 조준 방향 없음).",
    "built_space": "카메라가 판잣집 지붕들 위로 설정되어 프레임 하단을 지붕과 방수포들이 채우고 있으며, 레퍼런스와 일치하는 붉은 조명들이 배치됨.",
    "entities": "프레임 상단을 뒤덮은 거대한 검은 파도, 하단의 판잣집 지붕들. 지시된 대로 인물은 등장하지 않음.",
    "hard_violations": [
     "[gpt-high] 장소 참조에 없는 위성접시 여러 개를 지붕 위에 새로 배치했다."
    ],
    "physics": "거대한 파도가 지붕들 너머로 솟구쳐 오르는 모습이 물리적으로 타당하게 묘사됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "'지붕들 위'라는 카메라 위치와 프레임 하단에 지붕선을 배치하라는 구도 지시를 완벽하게 수행했으며, 지시된 대로 사람이 없는 상태의 위압적인 파도를 잘 묘사했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "사람을 제거하라는 지시는 따랐으나, 명시된 '지붕들 위' 시점과 하단 지붕선 구도를 무시하고 레퍼런스 이미지의 골목길 바닥 시점을 그대로 답습했습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "해당 사항 없음 (인물 및 조준 방향 없음).",
        "built_space": "카메라가 판잣집 지붕들 위로 설정되어 프레임 하단을 지붕과 방수포들이 채우고 있으며, 레퍼런스와 일치하는 붉은 조명들이 배치됨.",
        "entities": "프레임 상단을 뒤덮은 거대한 검은 파도, 하단의 판잣집 지붕들. 지시된 대로 인물은 등장하지 않음.",
        "hard_violations": [],
        "physics": "거대한 파도가 지붕들 너머로 솟구쳐 오르는 모습이 물리적으로 타당하게 묘사됨."
       },
       {
        "label": "A",
        "direction": "해당 사항 없음.",
        "built_space": "판잣촌 골목길 시점으로 양옆에 판잣집들이 있고 중앙 하단은 지붕이 아닌 골목길 바닥이 보임.",
        "entities": "거대한 검은 파도, 판잣집 외벽, 붉은 조명. 지시된 대로 인물은 등장하지 않음.",
        "hard_violations": [
         "명시된 카메라 위치('지붕들 위')와 구도('프레임 하단 중앙에 지붕선 배치')를 어기고 골목길 시점으로 렌더링함."
        ],
        "physics": "거대한 파도가 골목길 끝에서 솟구쳐 오르는 모습."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "'지붕들 위'라는 카메라 위치와 프레임 하단에 지붕선을 배치하라는 구도 지시를 완벽하게 수행했으며, 지시된 대로 사람이 없는 상태의 위압적인 파도를 잘 묘사했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "사람을 제거하라는 지시는 따랐으나, 명시된 '지붕들 위' 시점과 하단 지붕선 구도를 무시하고 레퍼런스 이미지의 골목길 바닥 시점을 그대로 답습했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "해당 사항 없음 (인물 및 조준 방향 없음).",
        "built_space": "카메라가 판잣집 지붕들 위로 설정되어 프레임 하단을 지붕과 방수포들이 채우고 있으며, 레퍼런스와 일치하는 붉은 조명들이 배치됨.",
        "entities": "프레임 상단을 뒤덮은 거대한 검은 파도, 하단의 판잣집 지붕들. 지시된 대로 인물은 등장하지 않음.",
        "hard_violations": [],
        "physics": "거대한 파도가 지붕들 너머로 솟구쳐 오르는 모습이 물리적으로 타당하게 묘사됨."
       },
       {
        "label": "A",
        "direction": "해당 사항 없음.",
        "built_space": "판잣촌 골목길 시점으로 양옆에 판잣집들이 있고 중앙 하단은 지붕이 아닌 골목길 바닥이 보임.",
        "entities": "거대한 검은 파도, 판잣집 외벽, 붉은 조명. 지시된 대로 인물은 등장하지 않음.",
        "hard_violations": [
         "명시된 카메라 위치('지붕들 위')와 구도('프레임 하단 중앙에 지붕선 배치')를 어기고 골목길 시점으로 렌더링함."
        ],
        "physics": "거대한 파도가 골목길 끝에서 솟구쳐 오르는 모습."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "검은 해일의 규모는 전달하지만, 지붕을 위에서 넓게 내려다보는 구도로 지정된 아래쪽 지붕선 배치를 벗어나며 참조에 없는 위성접시들을 추가했다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "아래에서 올려다보는 지붕 처마와 참조의 골목 재질·조명을 더 충실히 유지하지만, 양옆 건물이 과도하게 높이 들어오고 해일 위에 하늘이 남는다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "거대한 파면이 뒤쪽에서 전면의 판잣집 지붕들을 향해 밀려오는 방향으로 보인다. 사람의 시선이나 무기는 없다. 물결의 마루 위로 어두운 하늘이 남아 있어 바닷물이 하늘을 완전히 덮은 순간은 아니다.",
        "built_space": "중앙의 좁은 통로 양쪽에 골함석 판잣집들이 있고, 지붕 윗면이 화면 아래 약 40%를 차지한다. 카메라는 처마 아래가 아니라 지붕 높이 이상에서 내려다본다. 붉은 경고등 네 개와 여러 전선·기둥이 보인다. 참조의 함석과 방수포 재질은 이어지지만, 지붕을 하단 중경의 규모 기준으로만 남기는 배치와 다르다.",
        "entities": "검은 바닷물, 젖은 골함석 지붕, 방수포, 붉은 경고등은 요구된 환경과 대체로 맞는다. 사람이나 얼굴, 판독 가능한 글자는 보이지 않는다. 참조에서 확인되지 않는 위성접시 여러 개가 지붕 위의 뚜렷한 소품으로 추가되어 있다.",
        "hard_violations": [
         "장소 참조에 없는 위성접시 여러 개를 지붕 위에 새로 배치했다."
        ],
        "physics": "지붕은 판잣집 벽체에, 접시들은 지붕의 거치대에, 전선은 기둥에 지지되어 있다. 방수포와 무게추는 지붕 표면에 놓여 있다. 해일은 화면 아래로 이어지는 연속된 수괴이며 마루에서 포말이 날린다. 지지 없이 떠 있는 독립 물체나 신체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "뒤쪽의 거대한 물벽이 골목과 양옆 판잣집을 향해 전진하는 것으로 읽힌다. 카메라는 처마 너머 물벽을 올려다본다. 사람의 시선이나 조준 대상은 없다. 특히 왼쪽 위에 하늘이 드러나므로 하늘을 완전히 덮었다는 조건에는 미달한다.",
        "built_space": "골목 양쪽에 함석 벽체와 방수포를 두른 판잣집들이 늘어서고, 가까운 처마가 좌우에서 안쪽으로 돌출된다. 중앙 왼쪽의 큰 전신주 한 개와 골목을 따라 이어지는 작은 기둥들, 여러 붉은 경고등과 따뜻한 실내등이 보인다. 참조의 골목 구조와 낡은 재질을 더 잘 유지하며 처마를 아래에서 보는 시점도 맞는다. 다만 양옆 지붕과 벽이 화면 높이의 상당 부분을 차지해 하단 중앙에 지붕선만 남기는 구성과는 차이가 있다.",
        "entities": "거대한 검푸른 해일, 골함석 판잣집, 방수포, 전신주와 전선, 붉은 경고등이 보인다. 밤의 색조와 기존 조명의 성격이 참조에 가깝다. 식별 가능한 사람·얼굴과 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "처마는 벽체와 지붕 구조에 연결되고 방수포는 건물에 걸려 있다. 전선은 기둥 사이에 연결되어 있다. 물벽은 하부가 건물 뒤에 가려진 연속된 수괴로 보이며 포말도 파면과 이어진다. 지지 없이 공중에 멈춘 물체나 신체는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "검은 해일의 규모는 전달하지만, 지붕을 위에서 넓게 내려다보는 구도로 지정된 아래쪽 지붕선 배치를 벗어나며 참조에 없는 위성접시들을 추가했다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "아래에서 올려다보는 지붕 처마와 참조의 골목 재질·조명을 더 충실히 유지하지만, 양옆 건물이 과도하게 높이 들어오고 해일 위에 하늘이 남는다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "거대한 파면이 뒤쪽에서 전면의 판잣집 지붕들을 향해 밀려오는 방향으로 보인다. 사람의 시선이나 무기는 없다. 물결의 마루 위로 어두운 하늘이 남아 있어 바닷물이 하늘을 완전히 덮은 순간은 아니다.",
        "built_space": "중앙의 좁은 통로 양쪽에 골함석 판잣집들이 있고, 지붕 윗면이 화면 아래 약 40%를 차지한다. 카메라는 처마 아래가 아니라 지붕 높이 이상에서 내려다본다. 붉은 경고등 네 개와 여러 전선·기둥이 보인다. 참조의 함석과 방수포 재질은 이어지지만, 지붕을 하단 중경의 규모 기준으로만 남기는 배치와 다르다.",
        "entities": "검은 바닷물, 젖은 골함석 지붕, 방수포, 붉은 경고등은 요구된 환경과 대체로 맞는다. 사람이나 얼굴, 판독 가능한 글자는 보이지 않는다. 참조에서 확인되지 않는 위성접시 여러 개가 지붕 위의 뚜렷한 소품으로 추가되어 있다.",
        "hard_violations": [
         "장소 참조에 없는 위성접시 여러 개를 지붕 위에 새로 배치했다."
        ],
        "physics": "지붕은 판잣집 벽체에, 접시들은 지붕의 거치대에, 전선은 기둥에 지지되어 있다. 방수포와 무게추는 지붕 표면에 놓여 있다. 해일은 화면 아래로 이어지는 연속된 수괴이며 마루에서 포말이 날린다. 지지 없이 떠 있는 독립 물체나 신체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "뒤쪽의 거대한 물벽이 골목과 양옆 판잣집을 향해 전진하는 것으로 읽힌다. 카메라는 처마 너머 물벽을 올려다본다. 사람의 시선이나 조준 대상은 없다. 특히 왼쪽 위에 하늘이 드러나므로 하늘을 완전히 덮었다는 조건에는 미달한다.",
        "built_space": "골목 양쪽에 함석 벽체와 방수포를 두른 판잣집들이 늘어서고, 가까운 처마가 좌우에서 안쪽으로 돌출된다. 중앙 왼쪽의 큰 전신주 한 개와 골목을 따라 이어지는 작은 기둥들, 여러 붉은 경고등과 따뜻한 실내등이 보인다. 참조의 골목 구조와 낡은 재질을 더 잘 유지하며 처마를 아래에서 보는 시점도 맞는다. 다만 양옆 지붕과 벽이 화면 높이의 상당 부분을 차지해 하단 중앙에 지붕선만 남기는 구성과는 차이가 있다.",
        "entities": "거대한 검푸른 해일, 골함석 판잣집, 방수포, 전신주와 전선, 붉은 경고등이 보인다. 밤의 색조와 기존 조명의 성격이 참조에 가깝다. 식별 가능한 사람·얼굴과 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "처마는 벽체와 지붕 구조에 연결되고 방수포는 건물에 걸려 있다. 전선은 기둥 사이에 연결되어 있다. 물벽은 하부가 건물 뒤에 가려진 연속된 수괴로 보이며 포말도 파면과 이어진다. 지지 없이 공중에 멈춘 물체나 신체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.5,
    "B": 1.571
   },
   "adjusted": {
    "A": 1.25,
    "B": 1.321
   },
   "violations": {
    "A": [
     "[gemini-pro] 명시된 카메라 위치('지붕들 위')와 구도('프레임 하단 중앙에 지붕선 배치')를 어기고 골목길 시점으로 렌더링함."
    ],
    "B": [
     "[gpt-high] 장소 참조에 없는 위성접시 여러 개를 지붕 위에 새로 배치했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1321,
   "A": 1250
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1321,
    "verdict_ko": "'지붕들 위'라는 카메라 위치와 프레임 하단에 지붕선을 배치하라는 구도 지시를 완벽하게 수행했으며, 지시된 대로 사람이 없는 상태의 위압적인 파도를 잘 묘사했습니다.  ★위반: [gpt-high] 장소 참조에 없는 위성접시 여러 개를 지붕 위에 새로 배치했다."
   },
   {
    "label": "A",
    "score": 1250,
    "verdict_ko": "사람을 제거하라는 지시는 따랐으나, 명시된 '지붕들 위' 시점과 하단 지붕선 구도를 무시하고 레퍼런스 이미지의 골목길 바닥 시점을 그대로 답습했습니다.  ★위반: [gemini-pro] 명시된 카메라 위치('지붕들 위')와 구도('프레임 하단 중앙에 지붕선 배치')를 어기고 골목길 시점으로 렌더링함."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S35sh1_sel.png",
    "asset_id": "5e49bfe7-5f64-4f29-81c7-b25afbfd5938",
    "role": "prev_still"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-6e81-768b-ba15-5b779143c64a",
  "ref_mode": "prev만 (배경 전용·공유 계획)",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S35sh1"
  },
  "lane_policy": "ab_select_bypass:bg_only:share_plan_prev_bgonly"
 },
 "S35sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:00:39.330668+00:00",
  "fingerprint": "876a27f13ac0d9cabef4a866d5a179859b0c7a9a35443e17978253bcbc70e507",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S35sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S35sh3_sel.png",
  "source_sha256": "aafbb6d42b1474c5b617ac63045189bf2767fbe6833d98b37ce4c64004181de9",
  "file": "S35sh3_cine.png",
  "staged_sha256": "5720db55e95852943285ddf22836668f78f2a4c8e1391b89f524a537b29c188b",
  "latency_ms": 11811
 },
 "S36sh4::signage": {
  "fp": "77b6793b97b3620d",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::43286f194755c581": {
  "subjects": [],
  "subject_text": "바닷물에 침수된 인천 난민촌 지하 하수도\n맨홀과 연결된 좁고 어두운 지하 배수 통로. 길게 이어지는 벽과 천장 아래로 바닷물이 가득 차 통로의 윤곽만 드러난다.",
  "identity": "canonical",
  "scope_id": "L176",
  "scope_role": "location_interior",
  "scope_sha": "3095aa32a1f7998e"
 },
 "S36sh4::bgfirst_bg": {
  "input_fingerprint": "3a741e19ed84fdf6",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 구도환의 등 뒤쪽 터널 끝에서 거대한 바닷물이 무서운 기세로 터져 나오는 찰나.\n\nLOCATION (lock): Inside the dark underground sewer beneath the refugee settlement, at the tunnel section where seawater bursts toward the fleeing group.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Unobstructed tunnel and erupting seawater behind 구도환 in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Sewer passage behind 구도환 (Being invaded by seawater) — Its length recedes obliquely behind the figure rather than being blocked by his torso; used as Provide an unobstructed central corridor for the flood reveal; Erupting seawater (Bursting into the tunnel behind 구도환); used as Advance through the central and right background toward the foreground escape route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued illumination appropriate to the sewer interior, preserving readable separation between 구도환, the passage, and the erupting water without specifying an unsupported source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 구도환의 등 뒤쪽 터널 끝에서 거대한 바닷물이 무서운 기세로 터져 나오는 찰나.\n\nLOCATION (lock): Inside the dark underground sewer beneath the refugee settlement, at the tunnel section where seawater bursts toward the fleeing group.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Unobstructed tunnel and erupting seawater behind 구도환 in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Sewer passage behind 구도환 (Being invaded by seawater) — Its length recedes obliquely behind the figure rather than being blocked by his torso; used as Provide an unobstructed central corridor for the flood reveal; Erupting seawater (Bursting into the tunnel behind 구도환); used as Advance through the central and right background toward the foreground escape route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued illumination appropriate to the sewer interior, preserving readable separation between 구도환, the passage, and the erupting water without specifying an unsupported source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S36sh4__bgfirst_bg.png",
  "asset_id": "9ca46561-613d-4a43-8f4a-8931c01b07eb",
  "input_asset_ids": [
   "6c12b889-9578-4c38-95ec-c46ab7596741",
   "ebb4956c-b06f-42c3-8ed3-c2e4d3794c4b"
  ]
 },
 "S36sh4": {
  "input_fingerprint": "00aa43afaf101f65",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 구도환의 등 뒤쪽 터널 끝에서 거대한 바닷물이 무서운 기세로 터져 나오는 찰나.\n\nLOCATION (lock): Inside the dark underground sewer beneath the refugee settlement, at the tunnel section where seawater bursts toward the fleeing group. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Unobstructed tunnel and erupting seawater behind 구도환 in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Sewer passage behind 구도환 (Being invaded by seawater) — Its length recedes obliquely behind the figure rather than being blocked by his torso; used as Provide an unobstructed central corridor for the flood reveal; Erupting seawater (Bursting into the tunnel behind 구도환); used as Advance through the central and right background toward the foreground escape route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued illumination appropriate to the sewer interior, preserving readable separation between 구도환, the passage, and the erupting water without specifying an unsupported source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Seawater bursts into the underground sewer passages and begins filling them. Charlie retains his old coat, hat and radio as the flood reaches the escape route. 구도환: He is running through the sewer escape route.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 구도환의 등 뒤쪽 터널 끝에서 거대한 바닷물이 무서운 기세로 터져 나오는 찰나.\n\nLOCATION (lock): Inside the dark underground sewer beneath the refugee settlement, at the tunnel section where seawater bursts toward the fleeing group. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Unobstructed tunnel and erupting seawater behind 구도환 in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Sewer passage behind 구도환 (Being invaded by seawater) — Its length recedes obliquely behind the figure rather than being blocked by his torso; used as Provide an unobstructed central corridor for the flood reveal; Erupting seawater (Bursting into the tunnel behind 구도환); used as Advance through the central and right background toward the foreground escape route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued illumination appropriate to the sewer interior, preserving readable separation between 구도환, the passage, and the erupting water without specifying an unsupported source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Seawater bursts into the underground sewer passages and begins filling them. Charlie retains his old coat, hat and radio as the flood reaches the escape route. 구도환: He is running through the sewer escape route.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 구도환의 등 뒤쪽 터널 끝에서 거대한 바닷물이 무서운 기세로 터져 나오는 찰나.\n\nLOCATION (lock): Inside the dark underground sewer beneath the refugee settlement, at the tunnel section where seawater bursts toward the fleeing group. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Unobstructed tunnel and erupting seawater behind 구도환 in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Sewer passage behind 구도환 (Being invaded by seawater) — Its length recedes obliquely behind the figure rather than being blocked by his torso; used as Provide an unobstructed central corridor for the flood reveal; Erupting seawater (Bursting into the tunnel behind 구도환); used as Advance through the central and right background toward the foreground escape route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued illumination appropriate to the sewer interior, preserving readable separation between 구도환, the passage, and the erupting water without specifying an unsupported source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Seawater bursts into the underground sewer passages and begins filling them. Charlie retains his old coat, hat and radio as the flood reaches the escape route. 구도환: He is running through the sewer escape route.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S36sh4__bgfirst_bg.png",
     "asset_id": "9ca46561-613d-4a43-8f4a-8931c01b07eb",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S36sh4.png",
     "asset_id": "6c12b889-9578-4c38-95ec-c46ab7596741",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1311816>",
     "asset_id": "623f0421-67dc-4592-a009-148d8e957276",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L176B01.png",
     "asset_id": "ebb4956c-b06f-42c3-8ed3-c2e4d3794c4b",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1311816>",
     "asset_id": "623f0421-67dc-4592-a009-148d8e957276",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "인물이 화면 앞쪽을 향해 달리고 있으며, 거대한 물기둥은 지문대로 인물의 등 뒤(터널 끝)에서 분출하고 있음.",
    "built_space": "참조 사진의 하수구 구조(사다리, 천장 구멍, 양측 통로)가 올바른 위치에 정확히 묘사됨.",
    "entities": "얼굴 생김새, 나이대, 머리 스타일이 참조 이미지의 구도환과 완벽히 일치함.",
    "hard_violations": [],
    "physics": "거센 물살을 헤치며 달리는 하체의 체중 이동과 무전기를 쥔 손의 형태가 자연스러움."
   },
   {
    "label": "B",
    "direction": "인물이 카메라를 등지고 물기둥이 터지는 터널 안쪽을 향해 달려감(지문의 '등 뒤쪽' 방향과 반대).",
    "built_space": "터널의 물리적 구조는 맞으나, 좌측 통로에 지문에 지시되지 않은 다수의 군중이 서 있음.",
    "entities": "인물의 뒷모습만 보여 신원 확인이 불가하며, 좌측에 등장한 사람들은 지문에 없는 추가 인물들임.",
    "hard_violations": [
     "[gemini-pro] extra bodies (좌측 통로에 지문에 없는 인물들 등장)",
     "[gemini-pro] staging violation (물이 등 뒤가 아닌 정면에서 덮침)",
     "[gpt-high] 구도환만 허용된 장면에 왼쪽 측면 개구부의 추가 인물이 최소 두 명 등장한다."
    ],
    "physics": "물속을 뛰는 하체의 지지는 안정적이나 전체적인 상황과 맞지 않는 방향의 움직임임."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "등 뒤에서 물이 터져 나오는 찰나에 앞을 향해 달리는 구도를 훌륭히 소화했으며, 구도환의 얼굴 특징을 참조와 정확하게 일치시킴."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "지문에 없는 추가 인물들이 화면 좌측에 등장하며, 인물이 물을 등지지 않고 물을 향해 뛰어가는 치명적인 연출 오류가 있음."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "인물이 화면 앞쪽을 향해 달리고 있으며, 거대한 물기둥은 지문대로 인물의 등 뒤(터널 끝)에서 분출하고 있음.",
        "built_space": "참조 사진의 하수구 구조(사다리, 천장 구멍, 양측 통로)가 올바른 위치에 정확히 묘사됨.",
        "entities": "얼굴 생김새, 나이대, 머리 스타일이 참조 이미지의 구도환과 완벽히 일치함.",
        "hard_violations": [],
        "physics": "거센 물살을 헤치며 달리는 하체의 체중 이동과 무전기를 쥔 손의 형태가 자연스러움."
       },
       {
        "label": "B",
        "direction": "인물이 카메라를 등지고 물기둥이 터지는 터널 안쪽을 향해 달려감(지문의 '등 뒤쪽' 방향과 반대).",
        "built_space": "터널의 물리적 구조는 맞으나, 좌측 통로에 지문에 지시되지 않은 다수의 군중이 서 있음.",
        "entities": "인물의 뒷모습만 보여 신원 확인이 불가하며, 좌측에 등장한 사람들은 지문에 없는 추가 인물들임.",
        "hard_violations": [
         "extra bodies (좌측 통로에 지문에 없는 인물들 등장)",
         "staging violation (물이 등 뒤가 아닌 정면에서 덮침)"
        ],
        "physics": "물속을 뛰는 하체의 지지는 안정적이나 전체적인 상황과 맞지 않는 방향의 움직임임."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "등 뒤에서 물이 터져 나오는 찰나에 앞을 향해 달리는 구도를 훌륭히 소화했으며, 구도환의 얼굴 특징을 참조와 정확하게 일치시킴."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "지문에 없는 추가 인물들이 화면 좌측에 등장하며, 인물이 물을 등지지 않고 물을 향해 뛰어가는 치명적인 연출 오류가 있음."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "인물이 화면 앞쪽을 향해 달리고 있으며, 거대한 물기둥은 지문대로 인물의 등 뒤(터널 끝)에서 분출하고 있음.",
        "built_space": "참조 사진의 하수구 구조(사다리, 천장 구멍, 양측 통로)가 올바른 위치에 정확히 묘사됨.",
        "entities": "얼굴 생김새, 나이대, 머리 스타일이 참조 이미지의 구도환과 완벽히 일치함.",
        "hard_violations": [],
        "physics": "거센 물살을 헤치며 달리는 하체의 체중 이동과 무전기를 쥔 손의 형태가 자연스러움."
       },
       {
        "label": "B",
        "direction": "인물이 카메라를 등지고 물기둥이 터지는 터널 안쪽을 향해 달려감(지문의 '등 뒤쪽' 방향과 반대).",
        "built_space": "터널의 물리적 구조는 맞으나, 좌측 통로에 지문에 지시되지 않은 다수의 군중이 서 있음.",
        "entities": "인물의 뒷모습만 보여 신원 확인이 불가하며, 좌측에 등장한 사람들은 지문에 없는 추가 인물들임.",
        "hard_violations": [
         "extra bodies (좌측 통로에 지문에 없는 인물들 등장)",
         "staging violation (물이 등 뒤가 아닌 정면에서 덮침)"
        ],
        "physics": "물속을 뛰는 하체의 지지는 안정적이나 전체적인 상황과 맞지 않는 방향의 움직임임."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "통로와 폭발하는 해수는 잘 드러나지만, 금지된 추가 인물들이 등장하고 구도환이 물을 등지고 도망치지 않고 물 쪽으로 달린다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "해수를 등지고 전경으로 도망치는 방향과 열린 통로는 맞지만, 인물이 요구된 배경 배치보다 크고 가까우며 복장이 구도환 참조와 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "오른쪽 구도환은 등을 카메라에 보이고 머리와 몸을 터널 안쪽으로 향한 채 달린다. 따라서 진행 목표는 해수가 터지는 터널 끝으로 읽히며, 해수를 등지고 전경 탈출로로 달려야 하는 관계와 반대다. 해수 자체는 중앙 깊숙한 곳에서 전경으로 밀려온다. 왼쪽 추가 인물들은 측면 개구부 안쪽을 향한다.",
        "built_space": "낮고 평평한 콘크리트 천장, 중앙 직사각형 통로, 좌우 측면 개구부 각 하나, 오른쪽 벽 사다리 하나와 그 위 원형 천장 구멍 하나가 보인다. 주요 구조와 젖고 오염된 재질은 장소 참조와 일치한다. 오른쪽 인물 때문에 중앙 통로가 막히지는 않으며, 통로는 인물의 왼쪽 뒤로 이어진다. 왼쪽 개구부에는 허용되지 않은 사람들이 배치되어 있다.",
        "entities": "주요 인물 한 명 외에 왼쪽 개구부에서 최소 두 명의 추가 인물이 보인다. 주요 인물의 검은 머리와 짙은 남색 정장, 체격은 구도환 참조에 대체로 맞지만, 뒤통수만 보여 얼굴과 정확한 연령은 확인할 수 없다. 거대한 포말성 물과 하수도는 분명히 구현되어 있다. 읽을 수 있는 글자나 화면 표식은 없다.",
        "hard_violations": [
         "구도환만 허용된 장면에 왼쪽 측면 개구부의 추가 인물이 최소 두 명 등장한다."
        ],
        "physics": "구도환의 한쪽 다리는 수면 아래 바닥으로 내려가고 다른 발은 뒤로 들려 있어 달리는 보행 주기로 읽힌다. 발바닥 접촉은 물에 가려져 있지만 몸이 근거 없이 떠 있는 모습은 아니다. 추가 인물들도 개구부 바닥에 서 있는 자세다. 해수의 큰 물마루와 비말은 뒤에서 밀려오는 연속된 수류에 연결되어 있어 물리적으로 가능한 급류다."
       },
       {
        "label": "B",
        "direction": "인물은 터널 깊숙한 곳의 해수를 등지고 카메라 쪽, 약간 화면 오른쪽의 탈출 방향으로 달린다. 얼굴과 시선도 화면 오른쪽 앞의 화면 밖 공간을 향한다. 해수는 인물 뒤 중앙 통로에서 전경으로 밀려오므로 추격하는 물과 도주하는 사람의 방향 관계가 맞는다.",
        "built_space": "낮은 콘크리트 천장과 반복되는 직사각형 통로, 좌우 측면 개구부 각 하나, 오른쪽 벽 사다리 하나, 그 위 원형 천장 구멍 하나가 보인다. 오른쪽 개구부는 인물에게 일부 가려지지만 중앙의 해수 분출 구간은 열려 있다. 구조와 표면 재질은 장소 참조에 충실하다. 다만 인물이 화면 중간 오른쪽 배경이 아니라 오른쪽 중경에서 전경에 가까운 크기로 배치되어 있다.",
        "entities": "중년 동아시아계 남성 한 명만 보이며 검은 머리 일부와 정상적인 얼굴이 드러난다. 구도환 참조와 얼굴이 정확히 같은지는 확실하지 않다. 참조의 남색 정장과 흰 셔츠 대신 회갈색 긴 외투, 모자, 어두운 셔츠를 착용하고 무전기와 가방을 지닌다. 이는 구도환보다 본문에서 찰리에게 지정한 소지 상태를 적용한 모습이다. 거대한 해수와 지하 하수도는 구현되어 있으며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "앞다리는 물속 바닥을 향해 뻗고 반대 다리는 뒤에서 굽혀져 있어 급류를 헤치며 달리는 자세로 읽힌다. 실제 발 접촉은 물보라에 가려졌으나 지지 없는 공중 부양은 아니다. 무전기는 가슴 앞 손에 잡혀 있고 가방은 어깨끈으로 지지된다. 외투 자락의 들림은 달리는 동작에 부합하며, 뒤의 물보라는 바닥의 연속된 급류에서 솟는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "통로와 폭발하는 해수는 잘 드러나지만, 금지된 추가 인물들이 등장하고 구도환이 물을 등지고 도망치지 않고 물 쪽으로 달린다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "해수를 등지고 전경으로 도망치는 방향과 열린 통로는 맞지만, 인물이 요구된 배경 배치보다 크고 가까우며 복장이 구도환 참조와 다르다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "오른쪽 구도환은 등을 카메라에 보이고 머리와 몸을 터널 안쪽으로 향한 채 달린다. 따라서 진행 목표는 해수가 터지는 터널 끝으로 읽히며, 해수를 등지고 전경 탈출로로 달려야 하는 관계와 반대다. 해수 자체는 중앙 깊숙한 곳에서 전경으로 밀려온다. 왼쪽 추가 인물들은 측면 개구부 안쪽을 향한다.",
        "built_space": "낮고 평평한 콘크리트 천장, 중앙 직사각형 통로, 좌우 측면 개구부 각 하나, 오른쪽 벽 사다리 하나와 그 위 원형 천장 구멍 하나가 보인다. 주요 구조와 젖고 오염된 재질은 장소 참조와 일치한다. 오른쪽 인물 때문에 중앙 통로가 막히지는 않으며, 통로는 인물의 왼쪽 뒤로 이어진다. 왼쪽 개구부에는 허용되지 않은 사람들이 배치되어 있다.",
        "entities": "주요 인물 한 명 외에 왼쪽 개구부에서 최소 두 명의 추가 인물이 보인다. 주요 인물의 검은 머리와 짙은 남색 정장, 체격은 구도환 참조에 대체로 맞지만, 뒤통수만 보여 얼굴과 정확한 연령은 확인할 수 없다. 거대한 포말성 물과 하수도는 분명히 구현되어 있다. 읽을 수 있는 글자나 화면 표식은 없다.",
        "hard_violations": [
         "구도환만 허용된 장면에 왼쪽 측면 개구부의 추가 인물이 최소 두 명 등장한다."
        ],
        "physics": "구도환의 한쪽 다리는 수면 아래 바닥으로 내려가고 다른 발은 뒤로 들려 있어 달리는 보행 주기로 읽힌다. 발바닥 접촉은 물에 가려져 있지만 몸이 근거 없이 떠 있는 모습은 아니다. 추가 인물들도 개구부 바닥에 서 있는 자세다. 해수의 큰 물마루와 비말은 뒤에서 밀려오는 연속된 수류에 연결되어 있어 물리적으로 가능한 급류다."
       },
       {
        "label": "A",
        "direction": "인물은 터널 깊숙한 곳의 해수를 등지고 카메라 쪽, 약간 화면 오른쪽의 탈출 방향으로 달린다. 얼굴과 시선도 화면 오른쪽 앞의 화면 밖 공간을 향한다. 해수는 인물 뒤 중앙 통로에서 전경으로 밀려오므로 추격하는 물과 도주하는 사람의 방향 관계가 맞는다.",
        "built_space": "낮은 콘크리트 천장과 반복되는 직사각형 통로, 좌우 측면 개구부 각 하나, 오른쪽 벽 사다리 하나, 그 위 원형 천장 구멍 하나가 보인다. 오른쪽 개구부는 인물에게 일부 가려지지만 중앙의 해수 분출 구간은 열려 있다. 구조와 표면 재질은 장소 참조에 충실하다. 다만 인물이 화면 중간 오른쪽 배경이 아니라 오른쪽 중경에서 전경에 가까운 크기로 배치되어 있다.",
        "entities": "중년 동아시아계 남성 한 명만 보이며 검은 머리 일부와 정상적인 얼굴이 드러난다. 구도환 참조와 얼굴이 정확히 같은지는 확실하지 않다. 참조의 남색 정장과 흰 셔츠 대신 회갈색 긴 외투, 모자, 어두운 셔츠를 착용하고 무전기와 가방을 지닌다. 이는 구도환보다 본문에서 찰리에게 지정한 소지 상태를 적용한 모습이다. 거대한 해수와 지하 하수도는 구현되어 있으며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "앞다리는 물속 바닥을 향해 뻗고 반대 다리는 뒤에서 굽혀져 있어 급류를 헤치며 달리는 자세로 읽힌다. 실제 발 접촉은 물보라에 가려졌으나 지지 없는 공중 부양은 아니다. 무전기는 가슴 앞 손에 잡혀 있고 가방은 어깨끈으로 지지된다. 외투 자락의 들림은 달리는 동작에 부합하며, 뒤의 물보라는 바닥의 연속된 급류에서 솟는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.714
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.464
   },
   "violations": {
    "B": [
     "[gemini-pro] extra bodies (좌측 통로에 지문에 없는 인물들 등장)",
     "[gemini-pro] staging violation (물이 등 뒤가 아닌 정면에서 덮침)",
     "[gpt-high] 구도환만 허용된 장면에 왼쪽 측면 개구부의 추가 인물이 최소 두 명 등장한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 464
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "등 뒤에서 물이 터져 나오는 찰나에 앞을 향해 달리는 구도를 훌륭히 소화했으며, 구도환의 얼굴 특징을 참조와 정확하게 일치시킴."
   },
   {
    "label": "B",
    "score": 464,
    "verdict_ko": "지문에 없는 추가 인물들이 화면 좌측에 등장하며, 인물이 물을 등지지 않고 물을 향해 뛰어가는 치명적인 연출 오류가 있음.  ★위반: [gemini-pro] extra bodies (좌측 통로에 지문에 없는 인물들 등장) / [gemini-pro] staging violation (물이 등 뒤가 아닌 정면에서 덮침) / [gpt-high] 구도환만 허용된 장면에 왼쪽 측면 개구부의 추가 인물이 최소 두 명 등장한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L176B01.png",
    "asset_id": "ebb4956c-b06f-42c3-8ed3-c2e4d3794c4b",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1311816>",
    "asset_id": "623f0421-67dc-4592-a009-148d8e957276",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-7042-740b-b8aa-9ecf208123e8",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S36sh4__bgfirst_bg.png",
   "bg_asset_id": "9ca46561-613d-4a43-8f4a-8931c01b07eb",
   "bg_record_key": "S36sh4::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S36sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:02:23.342107+00:00",
  "fingerprint": "44bcdc77ad2dfb24fdaec85b0bcd847e7ec2b7c05626aaa3e0ff0e8f54b10b48",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S36sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S36sh4_sel.png",
  "source_sha256": "fc0915345b96ae95c146fa665f78c633963585aeee6385449a97ebef3b4b6e59",
  "file": "S36sh4_cine.png",
  "staged_sha256": "816a674307217b3fe9eb1f2aa6c2c690f0962df47836b54978a2895321e9687c",
  "latency_ms": 13660
 },
 "S36sh7::signage": {
  "fp": "965f03999b11a1e1",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S36sh7::bgfirst_bg": {
  "input_fingerprint": "824ac1b7a3bd52bc",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 수중에서 앰버의 작은 손목을 꽉 움켜쥔 찰리의 커다란 금속 손 클로즈업.\n\nLOCATION (lock): Under the floodwater outside the sewer outlet, in the inundated refugee settlement.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Seawater outside the sewer (Surrounding the submerged pair as they swim) — The hands and partial bodies are viewed directly through the surrounding water; used as Maintain underwater spatial continuity while leaving the joined hands clearly legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued underwater illumination with controlled contrast on the skin-to-metal contact, adding no unsupported color, bubbles, or visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 수중에서 앰버의 작은 손목을 꽉 움켜쥔 찰리의 커다란 금속 손 클로즈업.\n\nLOCATION (lock): Under the floodwater outside the sewer outlet, in the inundated refugee settlement.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Seawater outside the sewer (Surrounding the submerged pair as they swim) — The hands and partial bodies are viewed directly through the surrounding water; used as Maintain underwater spatial continuity while leaving the joined hands clearly legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued underwater illumination with controlled contrast on the skin-to-metal contact, adding no unsupported color, bubbles, or visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S36sh7__bgfirst_bg.png",
  "asset_id": "8484c4a2-ba85-42fe-b5be-7c3cd2e4cc63",
  "input_asset_ids": [
   "e8a72824-fe1d-4ef0-b1b3-784cdbd25fdc",
   "ebb4956c-b06f-42c3-8ed3-c2e4d3794c4b"
  ]
 },
 "S36sh7": {
  "input_fingerprint": "904666a23f8e0003",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 수중에서 앰버의 작은 손목을 꽉 움켜쥔 찰리의 커다란 금속 손 클로즈업.\n\nLOCATION (lock): Under the floodwater outside the sewer outlet, in the inundated refugee settlement. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Seawater outside the sewer (Surrounding the submerged pair as they swim) — The hands and partial bodies are viewed directly through the surrounding water; used as Maintain underwater spatial continuity while leaving the joined hands clearly legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued underwater illumination with controlled contrast on the skin-to-metal contact, adding no unsupported color, bubbles, or visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The sewer is inundated, and the flood has forced the escape into open water. Charlie is submerged and swimming, with his worn metal body and old disguise now soaked. 앰버: She has been swept out of the sewer and is submerged, struggling in the seawater with soaked clothing.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 수중에서 앰버의 작은 손목을 꽉 움켜쥔 찰리의 커다란 금속 손 클로즈업.\n\nLOCATION (lock): Under the floodwater outside the sewer outlet, in the inundated refugee settlement. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Seawater outside the sewer (Surrounding the submerged pair as they swim) — The hands and partial bodies are viewed directly through the surrounding water; used as Maintain underwater spatial continuity while leaving the joined hands clearly legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued underwater illumination with controlled contrast on the skin-to-metal contact, adding no unsupported color, bubbles, or visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The sewer is inundated, and the flood has forced the escape into open water. Charlie is submerged and swimming, with his worn metal body and old disguise now soaked. 앰버: She has been swept out of the sewer and is submerged, struggling in the seawater with soaked clothing.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 수중에서 앰버의 작은 손목을 꽉 움켜쥔 찰리의 커다란 금속 손 클로즈업.\n\nLOCATION (lock): Under the floodwater outside the sewer outlet, in the inundated refugee settlement. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Seawater outside the sewer (Surrounding the submerged pair as they swim) — The hands and partial bodies are viewed directly through the surrounding water; used as Maintain underwater spatial continuity while leaving the joined hands clearly legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued underwater illumination with controlled contrast on the skin-to-metal contact, adding no unsupported color, bubbles, or visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The sewer is inundated, and the flood has forced the escape into open water. Charlie is submerged and swimming, with his worn metal body and old disguise now soaked. 앰버: She has been swept out of the sewer and is submerged, struggling in the seawater with soaked clothing.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S36sh7__bgfirst_bg.png",
     "asset_id": "8484c4a2-ba85-42fe-b5be-7c3cd2e4cc63",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S36sh7.png",
     "asset_id": "e8a72824-fe1d-4ef0-b1b3-784cdbd25fdc",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L176B01.png",
     "asset_id": "ebb4956c-b06f-42c3-8ed3-c2e4d3794c4b",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "B",
    "direction": "찰리의 팔은 왼쪽 아래에서 중앙으로 뻗고 금속 손가락은 앰버의 손 쪽으로 닫혀 있다. 앰버의 팔은 오른쪽 위에서 접촉부로 내려온다. 접촉 대상은 맞지만, 드러난 손목보다 손바닥과 손가락 부분을 감싸 잡은 것으로 보인다. 얼굴과 눈은 잘려 시선 방향은 확인할 수 없다.",
    "built_space": "배경은 어둡고 흐린 물이며 식별 가능한 벽, 사다리, 출입구 등의 고정 시설은 없다. 따라서 정확한 장소를 입증하지는 못하지만, 출구 밖 물속이라는 설정과 충돌하는 구조물도 없다. 금속 손과 팔이 전경을 크게 차지하고 앰버의 상체 일부만 뒤에 보여 손 중심 클로즈업에 가깝다.",
    "entities": "큰 손과 전완에는 찰리 참고의 샌드 베이지 장갑판, 검은 관절, 마모된 금속 표면이 보인다. 앰버는 작은 체구와 밝은 피부의 팔, 금발, 짙은 반소매 상의로 나타난다. 얼굴이 제외되어 정확한 얼굴 정체성과 혼혈 외모는 판별할 수 없다. 추가 인물이나 읽을 수 있는 글자는 없다. 물속에 작은 밝은 입자가 다수 보여 기포를 추가하지 말라는 조건에는 다소 불확실성이 남는다.",
    "hard_violations": [],
    "physics": "금속 손은 전완에 연결되어 있고 앰버의 손과 실제로 접촉한다. 앰버의 상체와 머리카락은 주변 물속에 잠겨 있으며, 수중 부력과 손의 연결로 설명되는 자세다. 공중에 지지 없이 떠 있는 신체나 분리된 물체는 없다. 수영의 추진 동작은 화면 밖이라 확인할 수 없다."
   },
   {
    "label": "A",
    "direction": "찰리의 팔이 오른쪽 위에서 왼쪽 아래로 뻗어 앰버의 작은 손목을 감싼다. 앰버의 주먹은 금속 손 너머 오른쪽으로 나오므로 손목을 잡는 관계가 명확하다. 앰버의 고개는 오른쪽 아래의 잡힌 손 방향으로 숙여져 있으나 눈 자체는 보이지 않는다.",
    "built_space": "왼쪽 콘크리트 벽의 사각 개구부 한 개, 중앙 뒤편의 어두운 사각 통로 한 개, 오른쪽 사다리 한 개와 상단의 원형 개구부 한 개가 보인다. 양쪽 벽과 통로가 인물을 둘러싸 하수도 내부로 읽히며, 출구 밖 개방된 바닷물이라는 위치 지정과 다르다. 앰버의 머리와 어깨, 찰리의 상완과 몸통까지 크게 포함하여 손의 접촉부보다 넓은 장면을 보여준다.",
    "entities": "앰버의 금발, 작은 팔, 짙은 청색 반소매 옷은 참고와 부합한다. 옆얼굴 일부만 보여 정확한 얼굴 일치 여부는 제한적으로만 판단할 수 있다. 찰리의 큰 기계 손과 팔에는 베이지 장갑판, 검은 관절과 금속 구동부가 보이며 참고의 재질과 체격에 맞는다. 추가 인물이나 읽을 수 있는 글자는 없다.",
    "hard_violations": [
     "장면을 하수도 출구 밖 침수 정착지의 바닷물이 아니라 벽·사다리·상부 개구부가 둘러싼 하수도 내부에 배치했다."
    ],
    "physics": "찰리의 굽힌 금속 손가락이 앰버의 손목 양쪽을 감싸며, 두 팔은 각각 화면에 보이는 신체에 연결된다. 수중에서 뻗은 팔과 떠오른 머리카락은 부력과 물의 움직임으로 설명 가능하다. 손목 접촉이나 관절에서 명백한 물리적 불가능은 없고, 전체 수영 자세와 추진 동작은 프레임 밖이다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": null,
     "normalized": null,
     "ok": false
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "큰 금속 손 중심의 수중 클로즈업과 개방된 물속 배경은 부합하지만, 손목을 꽉 움켜쥐기보다는 앰버의 손을 감싸 잡은 모습에 가깝다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "작은 손목을 감싼 금속 손의 동작은 정확하지만, 출구 밖 바닷물이 아니라 하수도 내부에 배치했고 주변 인물과 구조물을 지나치게 넓게 보여준다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 팔은 왼쪽 아래에서 중앙으로 뻗고 금속 손가락은 앰버의 손 쪽으로 닫혀 있다. 앰버의 팔은 오른쪽 위에서 접촉부로 내려온다. 접촉 대상은 맞지만, 드러난 손목보다 손바닥과 손가락 부분을 감싸 잡은 것으로 보인다. 얼굴과 눈은 잘려 시선 방향은 확인할 수 없다.",
        "built_space": "배경은 어둡고 흐린 물이며 식별 가능한 벽, 사다리, 출입구 등의 고정 시설은 없다. 따라서 정확한 장소를 입증하지는 못하지만, 출구 밖 물속이라는 설정과 충돌하는 구조물도 없다. 금속 손과 팔이 전경을 크게 차지하고 앰버의 상체 일부만 뒤에 보여 손 중심 클로즈업에 가깝다.",
        "entities": "큰 손과 전완에는 찰리 참고의 샌드 베이지 장갑판, 검은 관절, 마모된 금속 표면이 보인다. 앰버는 작은 체구와 밝은 피부의 팔, 금발, 짙은 반소매 상의로 나타난다. 얼굴이 제외되어 정확한 얼굴 정체성과 혼혈 외모는 판별할 수 없다. 추가 인물이나 읽을 수 있는 글자는 없다. 물속에 작은 밝은 입자가 다수 보여 기포를 추가하지 말라는 조건에는 다소 불확실성이 남는다.",
        "hard_violations": [],
        "physics": "금속 손은 전완에 연결되어 있고 앰버의 손과 실제로 접촉한다. 앰버의 상체와 머리카락은 주변 물속에 잠겨 있으며, 수중 부력과 손의 연결로 설명되는 자세다. 공중에 지지 없이 떠 있는 신체나 분리된 물체는 없다. 수영의 추진 동작은 화면 밖이라 확인할 수 없다."
       },
       {
        "label": "B",
        "direction": "찰리의 팔이 오른쪽 위에서 왼쪽 아래로 뻗어 앰버의 작은 손목을 감싼다. 앰버의 주먹은 금속 손 너머 오른쪽으로 나오므로 손목을 잡는 관계가 명확하다. 앰버의 고개는 오른쪽 아래의 잡힌 손 방향으로 숙여져 있으나 눈 자체는 보이지 않는다.",
        "built_space": "왼쪽 콘크리트 벽의 사각 개구부 한 개, 중앙 뒤편의 어두운 사각 통로 한 개, 오른쪽 사다리 한 개와 상단의 원형 개구부 한 개가 보인다. 양쪽 벽과 통로가 인물을 둘러싸 하수도 내부로 읽히며, 출구 밖 개방된 바닷물이라는 위치 지정과 다르다. 앰버의 머리와 어깨, 찰리의 상완과 몸통까지 크게 포함하여 손의 접촉부보다 넓은 장면을 보여준다.",
        "entities": "앰버의 금발, 작은 팔, 짙은 청색 반소매 옷은 참고와 부합한다. 옆얼굴 일부만 보여 정확한 얼굴 일치 여부는 제한적으로만 판단할 수 있다. 찰리의 큰 기계 손과 팔에는 베이지 장갑판, 검은 관절과 금속 구동부가 보이며 참고의 재질과 체격에 맞는다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "장면을 하수도 출구 밖 침수 정착지의 바닷물이 아니라 벽·사다리·상부 개구부가 둘러싼 하수도 내부에 배치했다."
        ],
        "physics": "찰리의 굽힌 금속 손가락이 앰버의 손목 양쪽을 감싸며, 두 팔은 각각 화면에 보이는 신체에 연결된다. 수중에서 뻗은 팔과 떠오른 머리카락은 부력과 물의 움직임으로 설명 가능하다. 손목 접촉이나 관절에서 명백한 물리적 불가능은 없고, 전체 수영 자세와 추진 동작은 프레임 밖이다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "큰 금속 손 중심의 수중 클로즈업과 개방된 물속 배경은 부합하지만, 손목을 꽉 움켜쥐기보다는 앰버의 손을 감싸 잡은 모습에 가깝다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "작은 손목을 감싼 금속 손의 동작은 정확하지만, 출구 밖 바닷물이 아니라 하수도 내부에 배치했고 주변 인물과 구조물을 지나치게 넓게 보여준다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 팔은 왼쪽 아래에서 중앙으로 뻗고 금속 손가락은 앰버의 손 쪽으로 닫혀 있다. 앰버의 팔은 오른쪽 위에서 접촉부로 내려온다. 접촉 대상은 맞지만, 드러난 손목보다 손바닥과 손가락 부분을 감싸 잡은 것으로 보인다. 얼굴과 눈은 잘려 시선 방향은 확인할 수 없다.",
        "built_space": "배경은 어둡고 흐린 물이며 식별 가능한 벽, 사다리, 출입구 등의 고정 시설은 없다. 따라서 정확한 장소를 입증하지는 못하지만, 출구 밖 물속이라는 설정과 충돌하는 구조물도 없다. 금속 손과 팔이 전경을 크게 차지하고 앰버의 상체 일부만 뒤에 보여 손 중심 클로즈업에 가깝다.",
        "entities": "큰 손과 전완에는 찰리 참고의 샌드 베이지 장갑판, 검은 관절, 마모된 금속 표면이 보인다. 앰버는 작은 체구와 밝은 피부의 팔, 금발, 짙은 반소매 상의로 나타난다. 얼굴이 제외되어 정확한 얼굴 정체성과 혼혈 외모는 판별할 수 없다. 추가 인물이나 읽을 수 있는 글자는 없다. 물속에 작은 밝은 입자가 다수 보여 기포를 추가하지 말라는 조건에는 다소 불확실성이 남는다.",
        "hard_violations": [],
        "physics": "금속 손은 전완에 연결되어 있고 앰버의 손과 실제로 접촉한다. 앰버의 상체와 머리카락은 주변 물속에 잠겨 있으며, 수중 부력과 손의 연결로 설명되는 자세다. 공중에 지지 없이 떠 있는 신체나 분리된 물체는 없다. 수영의 추진 동작은 화면 밖이라 확인할 수 없다."
       },
       {
        "label": "A",
        "direction": "찰리의 팔이 오른쪽 위에서 왼쪽 아래로 뻗어 앰버의 작은 손목을 감싼다. 앰버의 주먹은 금속 손 너머 오른쪽으로 나오므로 손목을 잡는 관계가 명확하다. 앰버의 고개는 오른쪽 아래의 잡힌 손 방향으로 숙여져 있으나 눈 자체는 보이지 않는다.",
        "built_space": "왼쪽 콘크리트 벽의 사각 개구부 한 개, 중앙 뒤편의 어두운 사각 통로 한 개, 오른쪽 사다리 한 개와 상단의 원형 개구부 한 개가 보인다. 양쪽 벽과 통로가 인물을 둘러싸 하수도 내부로 읽히며, 출구 밖 개방된 바닷물이라는 위치 지정과 다르다. 앰버의 머리와 어깨, 찰리의 상완과 몸통까지 크게 포함하여 손의 접촉부보다 넓은 장면을 보여준다.",
        "entities": "앰버의 금발, 작은 팔, 짙은 청색 반소매 옷은 참고와 부합한다. 옆얼굴 일부만 보여 정확한 얼굴 일치 여부는 제한적으로만 판단할 수 있다. 찰리의 큰 기계 손과 팔에는 베이지 장갑판, 검은 관절과 금속 구동부가 보이며 참고의 재질과 체격에 맞는다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "장면을 하수도 출구 밖 침수 정착지의 바닷물이 아니라 벽·사다리·상부 개구부가 둘러싼 하수도 내부에 배치했다."
        ],
        "physics": "찰리의 굽힌 금속 손가락이 앰버의 손목 양쪽을 감싸며, 두 팔은 각각 화면에 보이는 신체에 연결된다. 수중에서 뻗은 팔과 떠오른 머리카락은 부력과 물의 움직임으로 설명 가능하다. 손목 접촉이나 관절에서 명백한 물리적 불가능은 없고, 전체 수영 자세와 추진 동작은 프레임 밖이다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gemini-pro"
   ],
   "route": "single_reverse"
  },
  "totals": {
   "B": 7,
   "A": 3
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 7,
    "verdict_ko": "큰 금속 손 중심의 수중 클로즈업과 개방된 물속 배경은 부합하지만, 손목을 꽉 움켜쥐기보다는 앰버의 손을 감싸 잡은 모습에 가깝다."
   },
   {
    "label": "A",
    "score": 3,
    "verdict_ko": "작은 손목을 감싼 금속 손의 동작은 정확하지만, 출구 밖 바닷물이 아니라 하수도 내부에 배치했고 주변 인물과 구조물을 지나치게 넓게 보여준다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L176B01.png",
    "asset_id": "ebb4956c-b06f-42c3-8ed3-c2e4d3794c4b",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-73ab-7d1a-bf3f-e4fc80c2c1bf",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S36sh7__bgfirst_bg.png",
   "bg_asset_id": "8484c4a2-ba85-42fe-b5be-7c3cd2e4cc63",
   "bg_record_key": "S36sh7::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S36sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:04:03.602266+00:00",
  "fingerprint": "868b8c308b32daaae3bad69438b8b8cc4690fde838bc2a3b21c040ebf934d115",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S36sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S36sh7_sel.png",
  "source_sha256": "fbdae060dfebb6a155a412b806f134a49bf4b21e82347746894a78752d4fd325",
  "file": "S36sh7_cine.png",
  "staged_sha256": "bb91fdca79d614a49e9487a0a053fc709c4aa5f21866967a8f95c06e675f3624",
  "latency_ms": 12689
 },
 "S37sh4::signage": {
  "fp": "88e9bb51cd7b6166",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S37sh4::bgfirst_bg": {
  "input_fingerprint": "92fbdfb31bc64ace",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 두 눈을 질끈 감은 채 현우의 머리를 자신의 가슴팍에 빈틈없이 끌어안은 미연의 절박한 자세.\n\nLOCATION (lock): Beside corridor windows on an upper floor of the refugee administration building, with the dark incoming surge visible outside.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Upper-floor corridor (The pair have stopped here and embraced) — The corridor interior appears obliquely around the pair; used as Narrow peripheral context keeps the embrace physically grounded without competing with the faces and arms.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained nighttime ambient illumination preserves detail in the closed eyes and clasping arms without introducing a visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 두 눈을 질끈 감은 채 현우의 머리를 자신의 가슴팍에 빈틈없이 끌어안은 미연의 절박한 자세.\n\nLOCATION (lock): Beside corridor windows on an upper floor of the refugee administration building, with the dark incoming surge visible outside.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Upper-floor corridor (The pair have stopped here and embraced) — The corridor interior appears obliquely around the pair; used as Narrow peripheral context keeps the embrace physically grounded without competing with the faces and arms.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained nighttime ambient illumination preserves detail in the closed eyes and clasping arms without introducing a visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S37sh4__bgfirst_bg.png",
  "asset_id": "6b563ad8-6d4e-4a21-9db4-4cdc611ddc68",
  "input_asset_ids": [
   "eda28974-c427-4e36-a410-4567430cdb20",
   "09d33278-e815-4961-9ee0-fa14a5350447"
  ]
 },
 "S37sh4": {
  "input_fingerprint": "155461718edf0c11",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 두 눈을 질끈 감은 채 현우의 머리를 자신의 가슴팍에 빈틈없이 끌어안은 미연의 절박한 자세.\n\nLOCATION (lock): Beside corridor windows on an upper floor of the refugee administration building, with the dark incoming surge visible outside. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Upper-floor corridor (The pair have stopped here and embraced) — The corridor interior appears obliquely around the pair; used as Narrow peripheral context keeps the embrace physically grounded without competing with the faces and arms.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained nighttime ambient illumination preserves detail in the closed eyes and clasping arms without introducing a visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The management-office corridor windows are still intact as the approaching tidal wave fills the view outside. 미연: Her face remains badly swollen from the beating. She stands with both arms closed tightly against her chest in a protective embrace. 현우: He retains facial bruising and the untreated dog-bite wound on his leg, and his outer garment remains removed. Yoon's contact card is concealed inside his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 두 눈을 질끈 감은 채 현우의 머리를 자신의 가슴팍에 빈틈없이 끌어안은 미연의 절박한 자세.\n\nLOCATION (lock): Beside corridor windows on an upper floor of the refugee administration building, with the dark incoming surge visible outside. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Upper-floor corridor (The pair have stopped here and embraced) — The corridor interior appears obliquely around the pair; used as Narrow peripheral context keeps the embrace physically grounded without competing with the faces and arms.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained nighttime ambient illumination preserves detail in the closed eyes and clasping arms without introducing a visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The management-office corridor windows are still intact as the approaching tidal wave fills the view outside. 미연: Her face remains badly swollen from the beating. She stands with both arms closed tightly against her chest in a protective embrace. 현우: He retains facial bruising and the untreated dog-bite wound on his leg, and his outer garment remains removed. Yoon's contact card is concealed inside his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 두 눈을 질끈 감은 채 현우의 머리를 자신의 가슴팍에 빈틈없이 끌어안은 미연의 절박한 자세.\n\nLOCATION (lock): Beside corridor windows on an upper floor of the refugee administration building, with the dark incoming surge visible outside. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Upper-floor corridor (The pair have stopped here and embraced) — The corridor interior appears obliquely around the pair; used as Narrow peripheral context keeps the embrace physically grounded without competing with the faces and arms.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained nighttime ambient illumination preserves detail in the closed eyes and clasping arms without introducing a visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The management-office corridor windows are still intact as the approaching tidal wave fills the view outside. 미연: Her face remains badly swollen from the beating. She stands with both arms closed tightly against her chest in a protective embrace. 현우: He retains facial bruising and the untreated dog-bite wound on his leg, and his outer garment remains removed. Yoon's contact card is concealed inside his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S37sh4__bgfirst_bg.png",
     "asset_id": "6b563ad8-6d4e-4a21-9db4-4cdc611ddc68",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S37sh4.png",
     "asset_id": "eda28974-c427-4e36-a410-4567430cdb20",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1113064>",
     "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L188B01.png",
     "asset_id": "09d33278-e815-4961-9ee0-fa14a5350447",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1113064>",
     "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "미연은 눈을 감은 채 고개를 숙이고, 현우의 머리는 미연의 가슴에 완전히 파묻혀 있음.",
    "built_space": "기준 사진과 완벽히 일치하는 복도(소화기, 부서진 문, 계단 위치). 좌측 창밖에는 해일 대신 도시의 야경이 보임.",
    "entities": "미연은 기준 이미지의 남색 옷과 인상이 일치하며 얼굴에 상처와 부기가 있음. 현우는 뒷머리와 상의 일부만 확인됨.",
    "hard_violations": [],
    "physics": "미연의 양팔이 현우의 상체를 자연스럽게 끌어안아 물리적으로 안정되게 지탱하고 있음."
   },
   {
    "label": "B",
    "direction": "미연은 눈을 질끈 감고 있으며, 현우는 고개를 숙인 채 얼굴을 드러내고 있음.",
    "built_space": "일반적인 복도 형태이나 기준 사진에 있는 소화기와 특정 문이 생략됨. 창밖으로는 거대한 파도가 밀려오는 것이 보임.",
    "entities": "미연이 기준 이미지의 남색 상의가 아닌 베이지색 셔츠를 입고 있음. 현우는 얼굴의 멍 자국과 외모 특징이 확인됨.",
    "hard_violations": [
     "[gemini-pro] physically impossible anatomy (미연의 오른팔과 손이 여러 개로 겹쳐 묘사되어 신체 구조가 불가능함)",
     "[gpt-high] 정확한 장소로 지정된 참조의 오른쪽 출입문 벽 구간에 별도의 큰 창문을 만들어 고정 시설 구성을 바꿨다."
    ],
    "physics": "인물을 끌어안는 팔과 손의 개수가 비정상적으로 많아 물리적 지탱 구조가 성립하지 않음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "창밖의 해일 묘사가 누락되었으나, 지정된 복도 배경과 인물의 의상을 정확히 재현하고 밀착된 포옹 자세를 훌륭히 구현함."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "창밖의 해일은 표현되었으나, 기준 의상과 배경 디테일을 무시했으며 신체 구조(여분의 팔)에 치명적인 오류가 있음."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "미연은 눈을 감은 채 고개를 숙이고, 현우의 머리는 미연의 가슴에 완전히 파묻혀 있음.",
        "built_space": "기준 사진과 완벽히 일치하는 복도(소화기, 부서진 문, 계단 위치). 좌측 창밖에는 해일 대신 도시의 야경이 보임.",
        "entities": "미연은 기준 이미지의 남색 옷과 인상이 일치하며 얼굴에 상처와 부기가 있음. 현우는 뒷머리와 상의 일부만 확인됨.",
        "hard_violations": [],
        "physics": "미연의 양팔이 현우의 상체를 자연스럽게 끌어안아 물리적으로 안정되게 지탱하고 있음."
       },
       {
        "label": "B",
        "direction": "미연은 눈을 질끈 감고 있으며, 현우는 고개를 숙인 채 얼굴을 드러내고 있음.",
        "built_space": "일반적인 복도 형태이나 기준 사진에 있는 소화기와 특정 문이 생략됨. 창밖으로는 거대한 파도가 밀려오는 것이 보임.",
        "entities": "미연이 기준 이미지의 남색 상의가 아닌 베이지색 셔츠를 입고 있음. 현우는 얼굴의 멍 자국과 외모 특징이 확인됨.",
        "hard_violations": [
         "physically impossible anatomy (미연의 오른팔과 손이 여러 개로 겹쳐 묘사되어 신체 구조가 불가능함)"
        ],
        "physics": "인물을 끌어안는 팔과 손의 개수가 비정상적으로 많아 물리적 지탱 구조가 성립하지 않음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "창밖의 해일 묘사가 누락되었으나, 지정된 복도 배경과 인물의 의상을 정확히 재현하고 밀착된 포옹 자세를 훌륭히 구현함."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "창밖의 해일은 표현되었으나, 기준 의상과 배경 디테일을 무시했으며 신체 구조(여분의 팔)에 치명적인 오류가 있음."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "미연은 눈을 감은 채 고개를 숙이고, 현우의 머리는 미연의 가슴에 완전히 파묻혀 있음.",
        "built_space": "기준 사진과 완벽히 일치하는 복도(소화기, 부서진 문, 계단 위치). 좌측 창밖에는 해일 대신 도시의 야경이 보임.",
        "entities": "미연은 기준 이미지의 남색 옷과 인상이 일치하며 얼굴에 상처와 부기가 있음. 현우는 뒷머리와 상의 일부만 확인됨.",
        "hard_violations": [],
        "physics": "미연의 양팔이 현우의 상체를 자연스럽게 끌어안아 물리적으로 안정되게 지탱하고 있음."
       },
       {
        "label": "B",
        "direction": "미연은 눈을 질끈 감고 있으며, 현우는 고개를 숙인 채 얼굴을 드러내고 있음.",
        "built_space": "일반적인 복도 형태이나 기준 사진에 있는 소화기와 특정 문이 생략됨. 창밖으로는 거대한 파도가 밀려오는 것이 보임.",
        "entities": "미연이 기준 이미지의 남색 상의가 아닌 베이지색 셔츠를 입고 있음. 현우는 얼굴의 멍 자국과 외모 특징이 확인됨.",
        "hard_violations": [
         "physically impossible anatomy (미연의 오른팔과 손이 여러 개로 겹쳐 묘사되어 신체 구조가 불가능함)"
        ],
        "physics": "인물을 끌어안는 팔과 손의 개수가 비정상적으로 많아 물리적 지탱 구조가 성립하지 않음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "질끈 감은 눈과 밀착 포옹, 창밖 해일은 잘 보이지만, 참조의 오른쪽 출입문 벽을 창문으로 바꾼 공간과 두 사람의 의상 변경이 충실도를 떨어뜨린다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "현우의 머리를 가슴에 묻고 감싸는 클로즈업과 미연의 정체성, 참조 복도 구조가 더 충실하지만, 창밖을 채워야 할 해일이 없고 눈을 질끈 감는 절박함이 약하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "미연은 두 눈을 강하게 감고 얼굴을 현우 쪽으로 기울인다. 현우는 아래쪽을 향하며 얼굴 옆면을 미연의 윗가슴에 밀착한다. 미연의 한 손은 정수리를 누르고 다른 팔은 목 뒤를 감싸므로 포옹의 대상과 방향은 맞는다.",
        "built_space": "왼쪽에는 가까운 창 한 조와 뒤쪽 창 한 조, 오른쪽에는 큰 창 한 조가 보인다. 천장 배관과 켜진 천장등 하나, 오른쪽 아래 붉은 매립함 일부가 있다. 낡은 투톤 벽은 유사하지만 참조에서 금속 출입문과 벽이 있는 오른쪽에 창문을 추가했다. 두 사람은 창 옆 복도에 있으며, 배경이 양옆으로 상당히 드러나 얼굴과 팔 중심의 좁은 주변 맥락보다는 조금 넓다. 왼쪽의 온전한 유리 밖에는 어두운 물결과 포말이 보인다.",
        "entities": "검은 머리의 중년 한국인 여성과 앳된 한국계 남성으로 보이는 두 사람만 있다. 현우의 헝클어진 머리와 얼굴 멍은 맞고, 미연도 볼에 멍과 부기가 있으나 심하게 부어오른 정도는 약하다. 미연의 머리는 참조보다 짧거나 뒤로 모인 모습이며, 두 사람 모두 참조의 남색 상의 대신 밝은 회색 계열의 칼라 있는 상의를 입었다. 현우의 겉옷 제거 상태는 이 의상만으로 명확하지 않다. 다리 상처와 신발 속 카드는 화면 밖이므로 판단하지 않는다. 읽을 수 있는 글자는 없다. 켜진 천장등이 보여 가시 광원을 배제한 조명 지시는 어긴다.",
        "hard_violations": [
         "정확한 장소로 지정된 참조의 오른쪽 출입문 벽 구간에 별도의 큰 창문을 만들어 고정 시설 구성을 바꿨다."
        ],
        "physics": "미연의 손바닥과 팔이 현우의 머리와 목에 실제로 닿아 받치고, 현우의 팔도 미연의 어깨와 몸통을 감싼다. 머리를 낮추고 상체를 기울인 포옹은 가능한 자세다. 하체와 발은 잘렸지만 공중에 떠 있다는 징후는 없으며, 보이는 손과 팔의 연결에도 명백한 불가능성은 없다."
       },
       {
        "label": "B",
        "direction": "미연은 눈을 감고 고개를 현우의 정수리 쪽으로 숙인다. 현우는 얼굴을 미연의 가슴 안쪽으로 돌려 묻고 있어 시선은 보이지 않는다. 미연의 손은 뒤통수를 잡고 반대 팔은 현우를 안쪽으로 감싸며, 머리를 자신의 가슴으로 끌어안는 방향이 명확하다. 다만 눈꺼풀과 미간의 수축은 질끈 감은 표정보다는 차분하게 감은 표정에 가깝다.",
        "built_space": "왼쪽 전경에 온전한 창 한 조, 복도 끝에 창 한 조와 위로 올라가는 계단 한 구간이 보인다. 오른쪽에는 열린 금속문 하나, 붉은 경보함 하나, 매립 소방함 하나, 벽에 걸린 소화기 하나가 있다. 천장 배관과 천장등 하나도 보여 참조의 재료와 시설 배치를 대체로 유지한다. 두 사람은 왼쪽 창 바로 옆에 있고 복도가 뒤로 비스듬히 이어진다. 얼굴과 포옹하는 팔이 화면을 크게 차지한다. 창밖 건물과 산 능선은 높은 층의 야경을 뒷받침하지만 접근하는 해일은 보이지 않는다.",
        "entities": "보이는 인물은 미연과 현우 두 명뿐이다. 미연은 참조에 가까운 중년 얼굴 윤곽, 검은 단발, 남색 상의를 갖췄고 볼에 타박상이 보이지만 심한 얼굴 부종은 충분하지 않다. 현우는 젊은 남성의 체격과 헝클어진 검은 머리를 보이며 얼굴 대부분은 포옹에 가려져 정확한 얼굴 일치와 멍을 확인할 수 없다. 보이는 상의는 참조의 남색과 다른 베이지색이고, 별도의 외투는 뚜렷하지 않다. 다리와 신발은 화면 밖이다. 소화기 표식은 읽을 수 없으며 별도의 자막이나 문자는 없다. 천장등과 경보등은 가시 광원을 넣지 말라는 조건과 다르다.",
        "hard_violations": [],
        "physics": "현우의 머리는 미연의 가슴에 닿아 있고 미연의 손이 뒤통수를 받친다. 반대 팔은 현우의 몸통을 둘러 포옹을 유지한다. 현우가 상체와 목을 앞으로 굽힌 자세는 자연스럽고, 손과 팔의 접촉 및 옷의 눌림도 그 동작에 맞는다. 발은 화면 밖이므로 접지 자체는 확인할 수 없지만 부유하거나 지지 없이 매달린 몸은 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "질끈 감은 눈과 밀착 포옹, 창밖 해일은 잘 보이지만, 참조의 오른쪽 출입문 벽을 창문으로 바꾼 공간과 두 사람의 의상 변경이 충실도를 떨어뜨린다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "현우의 머리를 가슴에 묻고 감싸는 클로즈업과 미연의 정체성, 참조 복도 구조가 더 충실하지만, 창밖을 채워야 할 해일이 없고 눈을 질끈 감는 절박함이 약하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "미연은 두 눈을 강하게 감고 얼굴을 현우 쪽으로 기울인다. 현우는 아래쪽을 향하며 얼굴 옆면을 미연의 윗가슴에 밀착한다. 미연의 한 손은 정수리를 누르고 다른 팔은 목 뒤를 감싸므로 포옹의 대상과 방향은 맞는다.",
        "built_space": "왼쪽에는 가까운 창 한 조와 뒤쪽 창 한 조, 오른쪽에는 큰 창 한 조가 보인다. 천장 배관과 켜진 천장등 하나, 오른쪽 아래 붉은 매립함 일부가 있다. 낡은 투톤 벽은 유사하지만 참조에서 금속 출입문과 벽이 있는 오른쪽에 창문을 추가했다. 두 사람은 창 옆 복도에 있으며, 배경이 양옆으로 상당히 드러나 얼굴과 팔 중심의 좁은 주변 맥락보다는 조금 넓다. 왼쪽의 온전한 유리 밖에는 어두운 물결과 포말이 보인다.",
        "entities": "검은 머리의 중년 한국인 여성과 앳된 한국계 남성으로 보이는 두 사람만 있다. 현우의 헝클어진 머리와 얼굴 멍은 맞고, 미연도 볼에 멍과 부기가 있으나 심하게 부어오른 정도는 약하다. 미연의 머리는 참조보다 짧거나 뒤로 모인 모습이며, 두 사람 모두 참조의 남색 상의 대신 밝은 회색 계열의 칼라 있는 상의를 입었다. 현우의 겉옷 제거 상태는 이 의상만으로 명확하지 않다. 다리 상처와 신발 속 카드는 화면 밖이므로 판단하지 않는다. 읽을 수 있는 글자는 없다. 켜진 천장등이 보여 가시 광원을 배제한 조명 지시는 어긴다.",
        "hard_violations": [
         "정확한 장소로 지정된 참조의 오른쪽 출입문 벽 구간에 별도의 큰 창문을 만들어 고정 시설 구성을 바꿨다."
        ],
        "physics": "미연의 손바닥과 팔이 현우의 머리와 목에 실제로 닿아 받치고, 현우의 팔도 미연의 어깨와 몸통을 감싼다. 머리를 낮추고 상체를 기울인 포옹은 가능한 자세다. 하체와 발은 잘렸지만 공중에 떠 있다는 징후는 없으며, 보이는 손과 팔의 연결에도 명백한 불가능성은 없다."
       },
       {
        "label": "A",
        "direction": "미연은 눈을 감고 고개를 현우의 정수리 쪽으로 숙인다. 현우는 얼굴을 미연의 가슴 안쪽으로 돌려 묻고 있어 시선은 보이지 않는다. 미연의 손은 뒤통수를 잡고 반대 팔은 현우를 안쪽으로 감싸며, 머리를 자신의 가슴으로 끌어안는 방향이 명확하다. 다만 눈꺼풀과 미간의 수축은 질끈 감은 표정보다는 차분하게 감은 표정에 가깝다.",
        "built_space": "왼쪽 전경에 온전한 창 한 조, 복도 끝에 창 한 조와 위로 올라가는 계단 한 구간이 보인다. 오른쪽에는 열린 금속문 하나, 붉은 경보함 하나, 매립 소방함 하나, 벽에 걸린 소화기 하나가 있다. 천장 배관과 천장등 하나도 보여 참조의 재료와 시설 배치를 대체로 유지한다. 두 사람은 왼쪽 창 바로 옆에 있고 복도가 뒤로 비스듬히 이어진다. 얼굴과 포옹하는 팔이 화면을 크게 차지한다. 창밖 건물과 산 능선은 높은 층의 야경을 뒷받침하지만 접근하는 해일은 보이지 않는다.",
        "entities": "보이는 인물은 미연과 현우 두 명뿐이다. 미연은 참조에 가까운 중년 얼굴 윤곽, 검은 단발, 남색 상의를 갖췄고 볼에 타박상이 보이지만 심한 얼굴 부종은 충분하지 않다. 현우는 젊은 남성의 체격과 헝클어진 검은 머리를 보이며 얼굴 대부분은 포옹에 가려져 정확한 얼굴 일치와 멍을 확인할 수 없다. 보이는 상의는 참조의 남색과 다른 베이지색이고, 별도의 외투는 뚜렷하지 않다. 다리와 신발은 화면 밖이다. 소화기 표식은 읽을 수 없으며 별도의 자막이나 문자는 없다. 천장등과 경보등은 가시 광원을 넣지 말라는 조건과 다르다.",
        "hard_violations": [],
        "physics": "현우의 머리는 미연의 가슴에 닿아 있고 미연의 손이 뒤통수를 받친다. 반대 팔은 현우의 몸통을 둘러 포옹을 유지한다. 현우가 상체와 목을 앞으로 굽힌 자세는 자연스럽고, 손과 팔의 접촉 및 옷의 눌림도 그 동작에 맞는다. 발은 화면 밖이므로 접지 자체는 확인할 수 없지만 부유하거나 지지 없이 매달린 몸은 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.857
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.607
   },
   "violations": {
    "B": [
     "[gemini-pro] physically impossible anatomy (미연의 오른팔과 손이 여러 개로 겹쳐 묘사되어 신체 구조가 불가능함)",
     "[gpt-high] 정확한 장소로 지정된 참조의 오른쪽 출입문 벽 구간에 별도의 큰 창문을 만들어 고정 시설 구성을 바꿨다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 607
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "창밖의 해일 묘사가 누락되었으나, 지정된 복도 배경과 인물의 의상을 정확히 재현하고 밀착된 포옹 자세를 훌륭히 구현함."
   },
   {
    "label": "B",
    "score": 607,
    "verdict_ko": "창밖의 해일은 표현되었으나, 기준 의상과 배경 디테일을 무시했으며 신체 구조(여분의 팔)에 치명적인 오류가 있음.  ★위반: [gemini-pro] physically impossible anatomy (미연의 오른팔과 손이 여러 개로 겹쳐 묘사되어 신체 구조가 불가능함) / [gpt-high] 정확한 장소로 지정된 참조의 오른쪽 출입문 벽 구간에 별도의 큰 창문을 만들어 고정 시설 구성을 바꿨다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L188B01.png",
    "asset_id": "09d33278-e815-4961-9ee0-fa14a5350447",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1113064>",
    "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-7712-71bc-81f5-495220efa0a6",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S37sh4__bgfirst_bg.png",
   "bg_asset_id": "6b563ad8-6d4e-4a21-9db4-4cdc611ddc68",
   "bg_record_key": "S37sh4::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S37sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:05:44.469587+00:00",
  "fingerprint": "6d9cd00642d770a203ca83debf59c42105ac8ff1b68160bbc8ff903d3a7dfbf8",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S37sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S37sh4_sel.png",
  "source_sha256": "b4f318696123c6255a017ddc1381e74612e5e7d09ddcaee96d869a3d8d251802",
  "file": "S37sh4_cine.png",
  "staged_sha256": "fd493b60a9f85f633b2dfe11ee235e8f1790f009a3272263ec5c5ee23c1da4a4",
  "latency_ms": 12513
 },
 "S37sh5::signage": {
  "fp": "1942506ba6b1cd01",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S37sh5": {
  "input_fingerprint": "920b0fba36ce7da9",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 산산조각 나 흩어지는 복도 창문 파편들과 함께 거대한 해일이 두 사람을 덮치는 폭발적인 찰나.\n\nLOCATION (lock): Inside the upper-floor corridor of the refugee administration building, at the windows bursting inward under the nighttime wave. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Breaking corridor window in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Corridor window (Breaking into scattered fragments as the wave enters) — Viewed obliquely from inside the corridor, with the approaching water visible through the breaking opening; used as Upper-right impact boundary above the embracing pair; Giant wave (Crashing through the window and engulfing the pair) — The advancing water enters from the window side toward the corridor interior; used as Connects the window break to the human figures without eliminating spatial context; Upper-floor corridor (Being overtaken by the wave) — The interior extends around and behind the pair; used as Provides scale and defines the camera's retreat away from the window.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding nighttime tonal treatment through the impact, preserving readable separation between the embracing figures, water, and breaking glass.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The corridor windows shatter, scattering glass as seawater surges into the management office. 미연: Her face remains swollen from the beating, and her arms remain tightly closed in the protective embrace as the water strikes. 현우: He still has facial bruises and the untreated leg wound, with his outer garment removed. Yoon's contact card remains concealed inside his shoe through the flood.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 산산조각 나 흩어지는 복도 창문 파편들과 함께 거대한 해일이 두 사람을 덮치는 폭발적인 찰나.\n\nLOCATION (lock): Inside the upper-floor corridor of the refugee administration building, at the windows bursting inward under the nighttime wave. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Breaking corridor window in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Corridor window (Breaking into scattered fragments as the wave enters) — Viewed obliquely from inside the corridor, with the approaching water visible through the breaking opening; used as Upper-right impact boundary above the embracing pair; Giant wave (Crashing through the window and engulfing the pair) — The advancing water enters from the window side toward the corridor interior; used as Connects the window break to the human figures without eliminating spatial context; Upper-floor corridor (Being overtaken by the wave) — The interior extends around and behind the pair; used as Provides scale and defines the camera's retreat away from the window.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding nighttime tonal treatment through the impact, preserving readable separation between the embracing figures, water, and breaking glass.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The corridor windows shatter, scattering glass as seawater surges into the management office. 미연: Her face remains swollen from the beating, and her arms remain tightly closed in the protective embrace as the water strikes. 현우: He still has facial bruises and the untreated leg wound, with his outer garment removed. Yoon's contact card remains concealed inside his shoe through the flood.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 산산조각 나 흩어지는 복도 창문 파편들과 함께 거대한 해일이 두 사람을 덮치는 폭발적인 찰나.\n\nLOCATION (lock): Inside the upper-floor corridor of the refugee administration building, at the windows bursting inward under the nighttime wave. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Breaking corridor window in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Corridor window (Breaking into scattered fragments as the wave enters) — Viewed obliquely from inside the corridor, with the approaching water visible through the breaking opening; used as Upper-right impact boundary above the embracing pair; Giant wave (Crashing through the window and engulfing the pair) — The advancing water enters from the window side toward the corridor interior; used as Connects the window break to the human figures without eliminating spatial context; Upper-floor corridor (Being overtaken by the wave) — The interior extends around and behind the pair; used as Provides scale and defines the camera's retreat away from the window.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding nighttime tonal treatment through the impact, preserving readable separation between the embracing figures, water, and breaking glass.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The corridor windows shatter, scattering glass as seawater surges into the management office. 미연: Her face remains swollen from the beating, and her arms remain tightly closed in the protective embrace as the water strikes. 현우: He still has facial bruises and the untreated leg wound, with his outer garment removed. Yoon's contact card remains concealed inside his shoe through the flood.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "우측 창문에서 쏟아지는 해일이 현우의 등을 향하고 있으며, 미연은 창문(위협) 쪽을 바라보고 있음.",
    "built_space": "우측에 창문, 좌측에 문과 화재경보기가 있는 복도 구조. 카메라가 복도 안쪽을 비추고 있음.",
    "entities": "미연과 현우. 현우의 긴팔 셔츠는 레퍼런스와 일치하나, 두 사람의 위치와 보호 관계가 완전히 뒤바뀜.",
    "hard_violations": [
     "[gemini-pro] 설정 및 연출 위반: 미연이 창문을 등지고 현우를 보호해야 하는 프롬프트 지시를 어기고, 현우가 창문 쪽에 서서 해일에 노출되도록 배치됨."
    ],
    "physics": "바닥에 물이 차오르고 공중에 유리 파편이 떠 있으나 운동감이 다소 부자연스러움. 두 사람은 서서 서로를 지탱함."
   },
   {
    "label": "B",
    "direction": "우측 창문에서 부서져 들어오는 해일이 현우를 감싸 안은 미연의 등을 정확히 타격하고 있음.",
    "built_space": "우측의 연속된 창문, 좌측의 문과 소화기함이 배치된 상층부 복도. 프롬프트가 요구한 배경의 깊이감이 잘 표현됨.",
    "entities": "미연과 현우. 현우가 레퍼런스와 달리 반팔을 입고 있으며, 포즈 반전으로 인해 미연의 멍이 왼쪽 뺨에 위치함.",
    "hard_violations": [],
    "physics": "창문을 산산조각 내며 쏟아지는 물보라와 바닥에 고이는 파도의 물리적 타격감이 사실적이며, 인물들의 지탱하는 자세도 안정적임."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "현우의 의상(반팔)과 미연의 멍 위치가 반전된 오류가 있으나, 미연이 창문을 등지고 해일을 막아내는 핵심 보호 자세(Staging)와 지정된 프레임 구성을 훌륭하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "미연이 현우를 보호해야 하는 연출(protective embrace)을 심각하게 위반하여, 오히려 현우의 등이 해일에 노출되는 치명적인 오류를 범했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "우측 창문에서 쏟아지는 해일이 현우의 등을 향하고 있으며, 미연은 창문(위협) 쪽을 바라보고 있음.",
        "built_space": "우측에 창문, 좌측에 문과 화재경보기가 있는 복도 구조. 카메라가 복도 안쪽을 비추고 있음.",
        "entities": "미연과 현우. 현우의 긴팔 셔츠는 레퍼런스와 일치하나, 두 사람의 위치와 보호 관계가 완전히 뒤바뀜.",
        "hard_violations": [
         "설정 및 연출 위반: 미연이 창문을 등지고 현우를 보호해야 하는 프롬프트 지시를 어기고, 현우가 창문 쪽에 서서 해일에 노출되도록 배치됨."
        ],
        "physics": "바닥에 물이 차오르고 공중에 유리 파편이 떠 있으나 운동감이 다소 부자연스러움. 두 사람은 서서 서로를 지탱함."
       },
       {
        "label": "B",
        "direction": "우측 창문에서 부서져 들어오는 해일이 현우를 감싸 안은 미연의 등을 정확히 타격하고 있음.",
        "built_space": "우측의 연속된 창문, 좌측의 문과 소화기함이 배치된 상층부 복도. 프롬프트가 요구한 배경의 깊이감이 잘 표현됨.",
        "entities": "미연과 현우. 현우가 레퍼런스와 달리 반팔을 입고 있으며, 포즈 반전으로 인해 미연의 멍이 왼쪽 뺨에 위치함.",
        "hard_violations": [],
        "physics": "창문을 산산조각 내며 쏟아지는 물보라와 바닥에 고이는 파도의 물리적 타격감이 사실적이며, 인물들의 지탱하는 자세도 안정적임."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "현우의 의상(반팔)과 미연의 멍 위치가 반전된 오류가 있으나, 미연이 창문을 등지고 해일을 막아내는 핵심 보호 자세(Staging)와 지정된 프레임 구성을 훌륭하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "미연이 현우를 보호해야 하는 연출(protective embrace)을 심각하게 위반하여, 오히려 현우의 등이 해일에 노출되는 치명적인 오류를 범했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "우측 창문에서 쏟아지는 해일이 현우의 등을 향하고 있으며, 미연은 창문(위협) 쪽을 바라보고 있음.",
        "built_space": "우측에 창문, 좌측에 문과 화재경보기가 있는 복도 구조. 카메라가 복도 안쪽을 비추고 있음.",
        "entities": "미연과 현우. 현우의 긴팔 셔츠는 레퍼런스와 일치하나, 두 사람의 위치와 보호 관계가 완전히 뒤바뀜.",
        "hard_violations": [
         "설정 및 연출 위반: 미연이 창문을 등지고 현우를 보호해야 하는 프롬프트 지시를 어기고, 현우가 창문 쪽에 서서 해일에 노출되도록 배치됨."
        ],
        "physics": "바닥에 물이 차오르고 공중에 유리 파편이 떠 있으나 운동감이 다소 부자연스러움. 두 사람은 서서 서로를 지탱함."
       },
       {
        "label": "B",
        "direction": "우측 창문에서 부서져 들어오는 해일이 현우를 감싸 안은 미연의 등을 정확히 타격하고 있음.",
        "built_space": "우측의 연속된 창문, 좌측의 문과 소화기함이 배치된 상층부 복도. 프롬프트가 요구한 배경의 깊이감이 잘 표현됨.",
        "entities": "미연과 현우. 현우가 레퍼런스와 달리 반팔을 입고 있으며, 포즈 반전으로 인해 미연의 멍이 왼쪽 뺨에 위치함.",
        "hard_violations": [],
        "physics": "창문을 산산조각 내며 쏟아지는 물보라와 바닥에 고이는 파도의 물리적 타격감이 사실적이며, 인물들의 지탱하는 자세도 안정적임."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "우측 상단의 파손 창에서 밀려온 물이 포옹한 두 사람의 어깨와 등을 직접 덮쳐, 요구한 충돌 순간을 B보다 정확히 구현한다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "미연의 외모와 보호하는 포옹은 잘 유지하지만, 물이 주로 두 사람 옆과 하체로 쏟아져 해일이 두 사람을 덮치는 순간이 A보다 약하고 현우의 긴소매도 의상 연속성이 덜 명확하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "미연은 눈을 감고 현우의 머리 쪽으로 얼굴을 숙이며, 현우도 미연의 가슴 쪽으로 고개를 묻는다. 물과 유리 파편은 우측 상단의 깨진 창에서 좌측 아래 복도 내부로 밀려오며, 물줄기가 두 사람의 어깨와 등에 실제로 닿는다. 카메라를 바라보는 인물이나 무기는 없다.",
        "built_space": "오른쪽에는 크게 깨진 전경 창 구획과 뒤로 이어지는 창열이 있고, 왼쪽에는 가까운 출입구 하나와 뒤쪽 출입구들이 보인다. 붉은 매립함 하나, 벽걸이 소화기 하나, 작은 붉은 벽 장치 하나, 천장등 세 개와 노출 배관이 보이며 끝에는 계단과 난간이 있다. 두 사람은 창 바로 안쪽 복도에 서 있다. 낡은 이색 도장 벽과 금속 창틀은 참조 장소와 잘 이어진다. 공간은 넓게 보이지만 인물은 하체가 잘린 중경에 가까워 명시된 와이드 숏보다 다소 타이트하다.",
        "entities": "인물은 검은 단발머리의 중년 동아시아계 여성 한 명과 헝클어진 검은 머리의 젊은 동아시아계 남성 한 명뿐이다. 미연의 남색 상의, 현우의 베이지색 상의와 벗겨진 겉옷 상태가 이전 장면에 대체로 부합한다. 미연의 뺨과 현우의 얼굴에 타박 흔적이 보인다. 현우의 정확한 얼굴 일치는 숙인 자세 때문에 제한적으로만 확인된다. 다리 상처와 신발 속 카드는 화면 밖이라 판단할 수 없다. 파손 유리, 바닷물, 밤의 어두운 외부가 있으며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 사람은 몸을 세운 채 서로를 팔로 감싸고 있으며 공중에 뜬 자세가 아니다. 발은 화면 밖이지만 몸통의 자세는 바닥에 선 상태와 양립한다. 미연의 손은 현우의 머리와 등에 닿아 보호한다. 물과 유리의 비산은 창을 깨고 들어온 수압으로 설명되고, 물은 벽과 인물에 부딪혀 바닥으로 떨어진다. 지지나 운동 원인이 없는 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "미연은 눈을 내리감고 현우의 머리를 향해 얼굴을 숙이며, 현우는 미연의 가슴 쪽으로 얼굴을 숨긴다. 물과 유리는 오른쪽의 파손 창에서 왼쪽 아래 실내로 진행한다. 다만 큰 물덩어리는 두 사람의 오른쪽과 하체 앞쪽에 집중되고, 상체를 덮치는 접촉은 A보다 약하다.",
        "built_space": "오른쪽에는 크게 파손된 가까운 창과 뒤쪽 창열이 있고, 깨진 개구부 안으로 금속 창틀 일부가 기울어져 있다. 왼쪽에는 가까운 출입구 하나와 뒤쪽 문틀들, 붉은 경보 장치 하나가 보이며 천장등 두 개와 노출 배관이 확인된다. 복도 끝에는 작은 창이 보이고 하부는 인물에 가려진다. 참조의 소화기와 매립함은 이 화면에서 확인되지 않는다. 낡은 벽과 창틀, 야간 조명은 참조와 유사하다. 인물은 왼쪽에 놓여 창과 분리되지만 하체가 잘려 완전한 와이드 숏보다는 다소 타이트하다.",
        "entities": "검은 단발머리의 중년 동아시아계 여성과 검은 머리의 젊은 동아시아계 남성, 두 명만 보인다. 미연의 얼굴 윤곽과 남색 상의, 뺨의 멍은 참조에 가깝다. 현우는 베이지색 긴소매 옷을 입고 있으며 얼굴이 거의 가려져 얼굴의 멍과 정확한 동일성은 확인하기 어렵다. 이전 장면에서 보인 베이지색 상의 색은 유지하지만 긴소매 의상의 형태와 겉옷 제거 상태는 덜 명확하다. 다리 상처와 숨긴 카드는 가려져 판단 대상이 아니다. 물, 유리 파편, 밤의 도시 불빛이 보이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "미연의 손과 팔이 현우의 머리와 등을 감싸고 현우의 팔도 미연의 허리를 두른다. 두 몸은 수직으로 서 있으며 떠 있는 징후는 없고, 발 접촉은 화면 아래 물과 크롭에 가려진다. 물은 창턱을 넘어 아래로 쏟아지고 유리 파편은 수압에 밀려 실내로 비산한다. 기울어진 창틀은 파손 창 주변에 걸쳐 있어 충격으로 꺾이는 순간으로 해석 가능하다. 명백하게 지지 원인이 없는 신체나 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "우측 상단의 파손 창에서 밀려온 물이 포옹한 두 사람의 어깨와 등을 직접 덮쳐, 요구한 충돌 순간을 B보다 정확히 구현한다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "미연의 외모와 보호하는 포옹은 잘 유지하지만, 물이 주로 두 사람 옆과 하체로 쏟아져 해일이 두 사람을 덮치는 순간이 A보다 약하고 현우의 긴소매도 의상 연속성이 덜 명확하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "미연은 눈을 감고 현우의 머리 쪽으로 얼굴을 숙이며, 현우도 미연의 가슴 쪽으로 고개를 묻는다. 물과 유리 파편은 우측 상단의 깨진 창에서 좌측 아래 복도 내부로 밀려오며, 물줄기가 두 사람의 어깨와 등에 실제로 닿는다. 카메라를 바라보는 인물이나 무기는 없다.",
        "built_space": "오른쪽에는 크게 깨진 전경 창 구획과 뒤로 이어지는 창열이 있고, 왼쪽에는 가까운 출입구 하나와 뒤쪽 출입구들이 보인다. 붉은 매립함 하나, 벽걸이 소화기 하나, 작은 붉은 벽 장치 하나, 천장등 세 개와 노출 배관이 보이며 끝에는 계단과 난간이 있다. 두 사람은 창 바로 안쪽 복도에 서 있다. 낡은 이색 도장 벽과 금속 창틀은 참조 장소와 잘 이어진다. 공간은 넓게 보이지만 인물은 하체가 잘린 중경에 가까워 명시된 와이드 숏보다 다소 타이트하다.",
        "entities": "인물은 검은 단발머리의 중년 동아시아계 여성 한 명과 헝클어진 검은 머리의 젊은 동아시아계 남성 한 명뿐이다. 미연의 남색 상의, 현우의 베이지색 상의와 벗겨진 겉옷 상태가 이전 장면에 대체로 부합한다. 미연의 뺨과 현우의 얼굴에 타박 흔적이 보인다. 현우의 정확한 얼굴 일치는 숙인 자세 때문에 제한적으로만 확인된다. 다리 상처와 신발 속 카드는 화면 밖이라 판단할 수 없다. 파손 유리, 바닷물, 밤의 어두운 외부가 있으며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 사람은 몸을 세운 채 서로를 팔로 감싸고 있으며 공중에 뜬 자세가 아니다. 발은 화면 밖이지만 몸통의 자세는 바닥에 선 상태와 양립한다. 미연의 손은 현우의 머리와 등에 닿아 보호한다. 물과 유리의 비산은 창을 깨고 들어온 수압으로 설명되고, 물은 벽과 인물에 부딪혀 바닥으로 떨어진다. 지지나 운동 원인이 없는 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "미연은 눈을 내리감고 현우의 머리를 향해 얼굴을 숙이며, 현우는 미연의 가슴 쪽으로 얼굴을 숨긴다. 물과 유리는 오른쪽의 파손 창에서 왼쪽 아래 실내로 진행한다. 다만 큰 물덩어리는 두 사람의 오른쪽과 하체 앞쪽에 집중되고, 상체를 덮치는 접촉은 A보다 약하다.",
        "built_space": "오른쪽에는 크게 파손된 가까운 창과 뒤쪽 창열이 있고, 깨진 개구부 안으로 금속 창틀 일부가 기울어져 있다. 왼쪽에는 가까운 출입구 하나와 뒤쪽 문틀들, 붉은 경보 장치 하나가 보이며 천장등 두 개와 노출 배관이 확인된다. 복도 끝에는 작은 창이 보이고 하부는 인물에 가려진다. 참조의 소화기와 매립함은 이 화면에서 확인되지 않는다. 낡은 벽과 창틀, 야간 조명은 참조와 유사하다. 인물은 왼쪽에 놓여 창과 분리되지만 하체가 잘려 완전한 와이드 숏보다는 다소 타이트하다.",
        "entities": "검은 단발머리의 중년 동아시아계 여성과 검은 머리의 젊은 동아시아계 남성, 두 명만 보인다. 미연의 얼굴 윤곽과 남색 상의, 뺨의 멍은 참조에 가깝다. 현우는 베이지색 긴소매 옷을 입고 있으며 얼굴이 거의 가려져 얼굴의 멍과 정확한 동일성은 확인하기 어렵다. 이전 장면에서 보인 베이지색 상의 색은 유지하지만 긴소매 의상의 형태와 겉옷 제거 상태는 덜 명확하다. 다리 상처와 숨긴 카드는 가려져 판단 대상이 아니다. 물, 유리 파편, 밤의 도시 불빛이 보이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "미연의 손과 팔이 현우의 머리와 등을 감싸고 현우의 팔도 미연의 허리를 두른다. 두 몸은 수직으로 서 있으며 떠 있는 징후는 없고, 발 접촉은 화면 아래 물과 크롭에 가려진다. 물은 창턱을 넘어 아래로 쏟아지고 유리 파편은 수압에 밀려 실내로 비산한다. 기울어진 창틀은 파손 창 주변에 걸쳐 있어 충격으로 꺾이는 순간으로 해석 가능하다. 명백하게 지지 원인이 없는 신체나 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.304,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.054,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 설정 및 연출 위반: 미연이 창문을 등지고 현우를 보호해야 하는 프롬프트 지시를 어기고, 현우가 창문 쪽에 서서 해일에 노출되도록 배치됨."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1054
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "현우의 의상(반팔)과 미연의 멍 위치가 반전된 오류가 있으나, 미연이 창문을 등지고 해일을 막아내는 핵심 보호 자세(Staging)와 지정된 프레임 구성을 훌륭하게 구현했습니다."
   },
   {
    "label": "A",
    "score": 1054,
    "verdict_ko": "미연이 현우를 보호해야 하는 연출(protective embrace)을 심각하게 위반하여, 오히려 현우의 등이 해일에 노출되는 치명적인 오류를 범했습니다.  ★위반: [gemini-pro] 설정 및 연출 위반: 미연이 창문을 등지고 현우를 보호해야 하는 프롬프트 지시를 어기고, 현우가 창문 쪽에 서서 해일에 노출되도록 배치됨."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 미연, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S37sh4_sel.png",
    "asset_id": "d51cfc30-c44d-48f0-8280-8a7957e7c97c",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1113064>",
    "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-7a80-7c52-ab73-affd2cefc16f",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S37sh4"
  }
 },
 "S37sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:08:17.147893+00:00",
  "fingerprint": "e6ad963ff5e67622ba15490c6709c09142e0675ccae073ce0014c5ffd8df1e3a",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S37sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S37sh5_sel.png",
  "source_sha256": "7d31da70697c5bcedeee1ef943ec907152f36051a171321c913e04bf9085a056",
  "file": "S37sh5_cine.png",
  "staged_sha256": "2edf9403fd6f16ecabec729c59bfa763f9edf0c88c6545eccd30b4b21d82db1f",
  "latency_ms": 16282
 },
 "S38sh2::signage": {
  "fp": "160a517829789c7a",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::5c8660f181826b50": {
  "subjects": [],
  "subject_text": "수몰된 인천 난민촌 수면과 떠다니는 컨테이너\n넓게 펼쳐진 바닷물 위에 철제 컨테이너와 판자가 떠 있는 공간. 담벼락 일부가 수면 위로 드러나고 희미한 새벽빛이 물에 비친다.",
  "identity": "canonical",
  "scope_id": "L193",
  "scope_role": "location_exterior",
  "scope_sha": "a3674d85c9b2a6df"
 },
 "groupbg::flood_raft_water": {
  "input_fingerprint": "dba0c01f7f8f09a9",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "flood_raft_water",
    "tags": [
     "S38sh13",
     "S38sh2",
     "S38sh7",
     "S63sh16"
    ]
   },
   "context_sig": "70ee00269da35a94"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On a floating container panel amid the submerged refugee settlement, visible across open floodwater at dawn.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n수몰된 인천 난민촌 수면과 떠다니는 컨테이너: 마을 전체가 물에 잠겨 바다처럼 변해버린 아침 풍경. (특징: 잔잔해진 광활한 수면; 뗏목처럼 부유하는 직육면체 컨테이너 박스들; 수면을 비추는 헬기의 서치라이트 빔; 물에 떠 있는 파편과 희생자들) / 현우와 수빈이 사용하는 낡은 트럭 운전석: 지붕 덮개가 떨어져 나가 하늘이 개방된 오래된 트럭의 앞좌석. (특징: 지붕 패널이 없는 오픈형 트럭 실내; 낡은 운전대와 대시보드; 마주 보는 운전자와 조수석 인물의 상반신)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 저 멀리 수면 위에 컨테이너 판자 위에 누워있는 미연.\n- / INSERT(회상F.B): 앰버를... 하면서 숨을 거두기 직전의 엄마 모습 /\n\nTIME OF DAY (lock): dawn.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On a floating container panel amid the submerged refugee settlement, visible across open floodwater at dawn.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n수몰된 인천 난민촌 수면과 떠다니는 컨테이너: 마을 전체가 물에 잠겨 바다처럼 변해버린 아침 풍경. (특징: 잔잔해진 광활한 수면; 뗏목처럼 부유하는 직육면체 컨테이너 박스들; 수면을 비추는 헬기의 서치라이트 빔; 물에 떠 있는 파편과 희생자들) / 현우와 수빈이 사용하는 낡은 트럭 운전석: 지붕 덮개가 떨어져 나가 하늘이 개방된 오래된 트럭의 앞좌석. (특징: 지붕 패널이 없는 오픈형 트럭 실내; 낡은 운전대와 대시보드; 마주 보는 운전자와 조수석 인물의 상반신)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 저 멀리 수면 위에 컨테이너 판자 위에 누워있는 미연.\n- / INSERT(회상F.B): 앰버를... 하면서 숨을 거두기 직전의 엄마 모습 /\n\nTIME OF DAY (lock): dawn.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_flood_raft_water_456400.png",
  "asset_id": "0797d468-af7c-4934-b8f4-036bf4cc5448",
  "input_asset_ids": [
   "316d049e-c00b-4b04-9191-f52cdd862974"
  ],
  "origin_tag": "S38sh2",
  "place_text": "On a floating container panel amid the submerged refugee settlement, visible across open floodwater at dawn.",
  "origin_inputs": {
   "place_text": "On a floating container panel amid the submerged refugee settlement, visible across open floodwater at dawn.",
   "time_of_day_en": "dawn",
   "conti_asset_id": "316d049e-c00b-4b04-9191-f52cdd862974"
  }
 },
 "S38sh2::bgfirst_bg": {
  "input_fingerprint": "dc101a1a466114c0",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 멀리 떨어진 판자 위에 축 늘어진 채 누워있는 미연의 핏기 없는 전신을 바라보는 현우의 시점 쇼트.\n\nLOCATION (lock): On a floating container panel amid the submerged refugee settlement, visible across open floodwater at dawn.\n\nTIME OF DAY (lock): dawn.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 미연's container board (Floating on the water with 미연 lying on it) — Its supporting face and a narrow edge are visible from the low, oblique viewpoint; used as Supports the distant full-body silhouette and remains small within the surrounding water; Intervening sea (Separating 현우's viewing position from 미연's board); used as A broad, uninterrupted interval makes the distance and helplessness readable.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Early-dawn ambient light holds a subdued tonal range, leaving the distant body legible without isolating it with an invented source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 멀리 떨어진 판자 위에 축 늘어진 채 누워있는 미연의 핏기 없는 전신을 바라보는 현우의 시점 쇼트.\n\nLOCATION (lock): On a floating container panel amid the submerged refugee settlement, visible across open floodwater at dawn.\n\nTIME OF DAY (lock): dawn.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 미연's container board (Floating on the water with 미연 lying on it) — Its supporting face and a narrow edge are visible from the low, oblique viewpoint; used as Supports the distant full-body silhouette and remains small within the surrounding water; Intervening sea (Separating 현우's viewing position from 미연's board); used as A broad, uninterrupted interval makes the distance and helplessness readable.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Early-dawn ambient light holds a subdued tonal range, leaving the distant body legible without isolating it with an invented source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S38sh2__bgfirst_bg.png",
  "asset_id": "6a473823-e079-4c39-92bf-e7c364af89d6",
  "input_asset_ids": [
   "316d049e-c00b-4b04-9191-f52cdd862974",
   "0797d468-af7c-4934-b8f4-036bf4cc5448"
  ]
 },
 "S38sh2": {
  "input_fingerprint": "dc2c06d3fc199ef5",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 멀리 떨어진 판자 위에 축 늘어진 채 누워있는 미연의 핏기 없는 전신을 바라보는 현우의 시점 쇼트.\n\nLOCATION (lock): On a floating container panel amid the submerged refugee settlement, visible across open floodwater at dawn. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 미연's container board (Floating on the water with 미연 lying on it) — Its supporting face and a narrow edge are visible from the low, oblique viewpoint; used as Supports the distant full-body silhouette and remains small within the surrounding water; Intervening sea (Separating 현우's viewing position from 미연's board); used as A broad, uninterrupted interval makes the distance and helplessness readable.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Early-dawn ambient light holds a subdued tonal range, leaving the distant body legible without isolating it with an invented source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Miyeon is stretched out on a floating container panel, her body supported by the panel and her abdomen bleeding from a puncture wound. The source does not specify whether she rests on her back or side, her head's direction, or the arrangement of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): At dawn, separate container panels float on the sea amid the flood wreckage. Storage boxes remain inside the container that has served as overnight shelter. 미연: She lies soaked and motionless on a floating container panel, with her face still swollen. Her abdominal puncture wound is already bleeding, although it is discovered only upon closer approach.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 멀리 떨어진 판자 위에 축 늘어진 채 누워있는 미연의 핏기 없는 전신을 바라보는 현우의 시점 쇼트.\n\nLOCATION (lock): On a floating container panel amid the submerged refugee settlement, visible across open floodwater at dawn. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 미연's container board (Floating on the water with 미연 lying on it) — Its supporting face and a narrow edge are visible from the low, oblique viewpoint; used as Supports the distant full-body silhouette and remains small within the surrounding water; Intervening sea (Separating 현우's viewing position from 미연's board); used as A broad, uninterrupted interval makes the distance and helplessness readable.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Early-dawn ambient light holds a subdued tonal range, leaving the distant body legible without isolating it with an invented source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Miyeon is stretched out on a floating container panel, her body supported by the panel and her abdomen bleeding from a puncture wound. The source does not specify whether she rests on her back or side, her head's direction, or the arrangement of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): At dawn, separate container panels float on the sea amid the flood wreckage. Storage boxes remain inside the container that has served as overnight shelter. 미연: She lies soaked and motionless on a floating container panel, with her face still swollen. Her abdominal puncture wound is already bleeding, although it is discovered only upon closer approach.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 멀리 떨어진 판자 위에 축 늘어진 채 누워있는 미연의 핏기 없는 전신을 바라보는 현우의 시점 쇼트.\n\nLOCATION (lock): On a floating container panel amid the submerged refugee settlement, visible across open floodwater at dawn. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 미연's container board (Floating on the water with 미연 lying on it) — Its supporting face and a narrow edge are visible from the low, oblique viewpoint; used as Supports the distant full-body silhouette and remains small within the surrounding water; Intervening sea (Separating 현우's viewing position from 미연's board); used as A broad, uninterrupted interval makes the distance and helplessness readable.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Early-dawn ambient light holds a subdued tonal range, leaving the distant body legible without isolating it with an invented source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Miyeon is stretched out on a floating container panel, her body supported by the panel and her abdomen bleeding from a puncture wound. The source does not specify whether she rests on her back or side, her head's direction, or the arrangement of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): At dawn, separate container panels float on the sea amid the flood wreckage. Storage boxes remain inside the container that has served as overnight shelter. 미연: She lies soaked and motionless on a floating container panel, with her face still swollen. Her abdominal puncture wound is already bleeding, although it is discovered only upon closer approach.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S38sh2__bgfirst_bg.png",
     "asset_id": "6a473823-e079-4c39-92bf-e7c364af89d6",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S38sh2.png",
     "asset_id": "316d049e-c00b-4b04-9191-f52cdd862974",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1113064>",
     "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_flood_raft_water_456400.png",
     "asset_id": "0797d468-af7c-4934-b8f4-036bf4cc5448",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1113064>",
     "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시점에서 수면 건너편 멀리 떠 있는 컨테이너 판자와 그 위의 인물을 똑바로 향하고 있습니다.",
    "built_space": "레퍼런스와 동일한 수몰된 도시의 새벽 수면 배경이며, 중앙 우측에 인물이 눕기 위한 컨테이너 판자가 적절한 크기로 배치되어 있습니다.",
    "entities": "판자 위에 40대 여성(미연)이 누워 있으나, 프롬프트가 요구한 복부의 찔린 상처와 출혈 자국이 전혀 보이지 않습니다.",
    "hard_violations": [],
    "physics": "인물의 전신이 판자 표면에 밀착되어 중력에 맞게 완전히 지탱되고 있습니다."
   },
   {
    "label": "B",
    "direction": "현우의 시점을 대변하여 멀리 떨어진 컨테이너 판자와 그 위의 인물을 정확히 바라보고 있습니다.",
    "built_space": "위치 레퍼런스의 배경과 조명을 그대로 유지한 채, 요구된 컨테이너 판자가 올바른 비례와 원근감으로 배치되었습니다.",
    "entities": "지정된 인물(미연)이 누워 있으며, 복부 부근에 어둡고 붉은 출혈 자국이 묘사되어 캐논 설정을 충족합니다.",
    "hard_violations": [],
    "physics": "신체 전체가 부력과 중력에 순응하여 판자 위에 축 늘어진 채 안정적으로 지탱되어 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지정된 와이드 쇼트 화각과 배경 묘사를 완벽히 구현했으며, 캐논 설정인 복부의 출혈 자국까지 충실하게 시각화했습니다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "요구된 위치와 카메라 시점은 잘 반영되었으나, 필수 설정인 복부의 찔린 상처와 출혈 묘사가 누락되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시점에서 수면 건너편 멀리 떠 있는 컨테이너 판자와 그 위의 인물을 똑바로 향하고 있습니다.",
        "built_space": "레퍼런스와 동일한 수몰된 도시의 새벽 수면 배경이며, 중앙 우측에 인물이 눕기 위한 컨테이너 판자가 적절한 크기로 배치되어 있습니다.",
        "entities": "판자 위에 40대 여성(미연)이 누워 있으나, 프롬프트가 요구한 복부의 찔린 상처와 출혈 자국이 전혀 보이지 않습니다.",
        "hard_violations": [],
        "physics": "인물의 전신이 판자 표면에 밀착되어 중력에 맞게 완전히 지탱되고 있습니다."
       },
       {
        "label": "B",
        "direction": "현우의 시점을 대변하여 멀리 떨어진 컨테이너 판자와 그 위의 인물을 정확히 바라보고 있습니다.",
        "built_space": "위치 레퍼런스의 배경과 조명을 그대로 유지한 채, 요구된 컨테이너 판자가 올바른 비례와 원근감으로 배치되었습니다.",
        "entities": "지정된 인물(미연)이 누워 있으며, 복부 부근에 어둡고 붉은 출혈 자국이 묘사되어 캐논 설정을 충족합니다.",
        "hard_violations": [],
        "physics": "신체 전체가 부력과 중력에 순응하여 판자 위에 축 늘어진 채 안정적으로 지탱되어 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지정된 와이드 쇼트 화각과 배경 묘사를 완벽히 구현했으며, 캐논 설정인 복부의 출혈 자국까지 충실하게 시각화했습니다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "요구된 위치와 카메라 시점은 잘 반영되었으나, 필수 설정인 복부의 찔린 상처와 출혈 묘사가 누락되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시점에서 수면 건너편 멀리 떠 있는 컨테이너 판자와 그 위의 인물을 똑바로 향하고 있습니다.",
        "built_space": "레퍼런스와 동일한 수몰된 도시의 새벽 수면 배경이며, 중앙 우측에 인물이 눕기 위한 컨테이너 판자가 적절한 크기로 배치되어 있습니다.",
        "entities": "판자 위에 40대 여성(미연)이 누워 있으나, 프롬프트가 요구한 복부의 찔린 상처와 출혈 자국이 전혀 보이지 않습니다.",
        "hard_violations": [],
        "physics": "인물의 전신이 판자 표면에 밀착되어 중력에 맞게 완전히 지탱되고 있습니다."
       },
       {
        "label": "B",
        "direction": "현우의 시점을 대변하여 멀리 떨어진 컨테이너 판자와 그 위의 인물을 정확히 바라보고 있습니다.",
        "built_space": "위치 레퍼런스의 배경과 조명을 그대로 유지한 채, 요구된 컨테이너 판자가 올바른 비례와 원근감으로 배치되었습니다.",
        "entities": "지정된 인물(미연)이 누워 있으며, 복부 부근에 어둡고 붉은 출혈 자국이 묘사되어 캐논 설정을 충족합니다.",
        "hard_violations": [],
        "physics": "신체 전체가 부력과 중력에 순응하여 판자 위에 축 늘어진 채 안정적으로 지탱되어 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "작게 보이는 미연의 전신과 넓게 열린 중간 수면, 절제된 새벽빛이 멀리 떨어진 미연을 바라보는 시점 쇼트에 더 충실하다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "전신과 지지 관계는 적절하지만 같은 판자에서 미연이 더 크게 보여 거리감이 약하고, 밝은 회색 상의와 강한 주황빛도 지정된 의상·분위기에서 멀어진다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 넓은 수면 너머 오른쪽 판자에 누운 미연을 바라본다. 미연은 머리가 왼쪽, 발이 오른쪽이며 얼굴은 위쪽을 향한다. 눈의 시선은 이 거리에서 판독되지 않는다. 현우의 신체가 나오지 않아 그의 시점이라는 설정과 맞는다.",
        "built_space": "미연을 받치는 직사각형 컨테이너 패널 하나가 오른쪽 중경에 있고, 윗면과 물에 닿는 측면이 보인다. 왼쪽 아래에는 별도의 전경 패널 하나가 있다. 왼쪽의 크게 기운 컨테이너 한 동, 중앙 왼쪽의 반쯤 잠긴 잔해, 오른쪽의 기울어진 컨테이너와 골조, 먼 도시 윤곽이 장소 참조와 대응한다. 전경과 미연 사이 중앙 수면은 넓게 열려 있다. 햇빛 반사는 왼쪽의 낮은 태양 아래 수면에 이어져 자연스럽다.",
        "entities": "보이는 사람은 검은 머리의 여성 한 명뿐이며, 작은 크기에서도 머리부터 발까지 사람의 형태가 구별된다. 드러난 얼굴과 발은 창백하게 보인다. 한국인 중년 여성이라는 설정과 뚜렷한 충돌은 없지만 정확한 얼굴 일치와 부기는 확인하기 어렵다. 상의는 어두운 회색·남색 계열로 보이며 참조의 남색 상의에 비교적 가깝다. 복부의 어두운 부분만으로 출혈이나 관통상을 확정할 수는 없다. 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "미연의 머리, 몸통, 골반과 다리는 패널 윗면에 놓여 있고, 보이는 팔과 손도 몸 옆 지지면에 내려앉아 있다. 발끝의 위쪽 방향은 뒤꿈치가 놓인 누운 자세로 설명된다. 근육으로 신체를 들어 올리거나 허공에 떠 있는 부분은 확인되지 않는다. 패널은 수면에 잠긴 하단으로 부력을 받는 것으로 읽힌다."
       },
       {
        "label": "B",
        "direction": "카메라는 수면 건너 오른쪽 패널 위 미연을 향한다. 미연은 왼쪽에 머리, 오른쪽에 발을 두고 얼굴을 하늘 쪽으로 향한 채 누워 있다. 눈의 구체적인 시선은 식별되지 않으며 카메라를 응시하는 모습은 아니다. 현우는 화면에 나타나지 않는다.",
        "built_space": "오른쪽 중경의 지지 패널 하나와 왼쪽 아래 전경 패널 하나가 보인다. 지지 패널의 윗면과 측면, 왼쪽의 기운 컨테이너, 중앙 왼쪽의 침수 잔해, 오른쪽 컨테이너와 골조가 참조 장소의 배치를 따른다. 사이의 수면도 넓게 확보되어 있다. 다만 같은 크기로 보이는 지지 패널 위에서 미연이 A보다 훨씬 긴 면적을 차지해 더 가까운 인물처럼 읽힌다. 낮은 태양과 수면의 반사 방향은 서로 맞는다.",
        "entities": "검은 머리의 여성 한 명이 전신으로 보이며 다른 사람은 없다. 창백한 얼굴은 보이지만 정확한 나이, 참조 인물의 얼굴 윤곽과 부기를 확정할 정도의 크기는 아니다. 밝은 회색 상의는 참조의 남색 상의와 다르다. 복부에 놓인 손과 어두운 부분이 보이지만 관통상 출혈은 명확히 식별되지 않는다. 패널과 주변 잔해는 실제 금속·목재로 읽히며 읽을 수 있는 문자는 없다.",
        "hard_violations": [],
        "physics": "머리와 몸통, 다리는 패널에 받쳐져 있다. 한쪽 팔은 몸 옆에 내려놓았고 다른 손은 복부 위에 얹혀 있어 각각 패널과 몸의 지지를 받는다. 뒤꿈치가 패널에 놓인 상태에서 발끝이 위로 향하는 자세도 가능하다. 지지 없이 떠 있는 신체나 물체는 보이지 않으며 패널은 수면에서 부유한다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "작게 보이는 미연의 전신과 넓게 열린 중간 수면, 절제된 새벽빛이 멀리 떨어진 미연을 바라보는 시점 쇼트에 더 충실하다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "전신과 지지 관계는 적절하지만 같은 판자에서 미연이 더 크게 보여 거리감이 약하고, 밝은 회색 상의와 강한 주황빛도 지정된 의상·분위기에서 멀어진다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "카메라는 넓은 수면 너머 오른쪽 판자에 누운 미연을 바라본다. 미연은 머리가 왼쪽, 발이 오른쪽이며 얼굴은 위쪽을 향한다. 눈의 시선은 이 거리에서 판독되지 않는다. 현우의 신체가 나오지 않아 그의 시점이라는 설정과 맞는다.",
        "built_space": "미연을 받치는 직사각형 컨테이너 패널 하나가 오른쪽 중경에 있고, 윗면과 물에 닿는 측면이 보인다. 왼쪽 아래에는 별도의 전경 패널 하나가 있다. 왼쪽의 크게 기운 컨테이너 한 동, 중앙 왼쪽의 반쯤 잠긴 잔해, 오른쪽의 기울어진 컨테이너와 골조, 먼 도시 윤곽이 장소 참조와 대응한다. 전경과 미연 사이 중앙 수면은 넓게 열려 있다. 햇빛 반사는 왼쪽의 낮은 태양 아래 수면에 이어져 자연스럽다.",
        "entities": "보이는 사람은 검은 머리의 여성 한 명뿐이며, 작은 크기에서도 머리부터 발까지 사람의 형태가 구별된다. 드러난 얼굴과 발은 창백하게 보인다. 한국인 중년 여성이라는 설정과 뚜렷한 충돌은 없지만 정확한 얼굴 일치와 부기는 확인하기 어렵다. 상의는 어두운 회색·남색 계열로 보이며 참조의 남색 상의에 비교적 가깝다. 복부의 어두운 부분만으로 출혈이나 관통상을 확정할 수는 없다. 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "미연의 머리, 몸통, 골반과 다리는 패널 윗면에 놓여 있고, 보이는 팔과 손도 몸 옆 지지면에 내려앉아 있다. 발끝의 위쪽 방향은 뒤꿈치가 놓인 누운 자세로 설명된다. 근육으로 신체를 들어 올리거나 허공에 떠 있는 부분은 확인되지 않는다. 패널은 수면에 잠긴 하단으로 부력을 받는 것으로 읽힌다."
       },
       {
        "label": "A",
        "direction": "카메라는 수면 건너 오른쪽 패널 위 미연을 향한다. 미연은 왼쪽에 머리, 오른쪽에 발을 두고 얼굴을 하늘 쪽으로 향한 채 누워 있다. 눈의 구체적인 시선은 식별되지 않으며 카메라를 응시하는 모습은 아니다. 현우는 화면에 나타나지 않는다.",
        "built_space": "오른쪽 중경의 지지 패널 하나와 왼쪽 아래 전경 패널 하나가 보인다. 지지 패널의 윗면과 측면, 왼쪽의 기운 컨테이너, 중앙 왼쪽의 침수 잔해, 오른쪽 컨테이너와 골조가 참조 장소의 배치를 따른다. 사이의 수면도 넓게 확보되어 있다. 다만 같은 크기로 보이는 지지 패널 위에서 미연이 A보다 훨씬 긴 면적을 차지해 더 가까운 인물처럼 읽힌다. 낮은 태양과 수면의 반사 방향은 서로 맞는다.",
        "entities": "검은 머리의 여성 한 명이 전신으로 보이며 다른 사람은 없다. 창백한 얼굴은 보이지만 정확한 나이, 참조 인물의 얼굴 윤곽과 부기를 확정할 정도의 크기는 아니다. 밝은 회색 상의는 참조의 남색 상의와 다르다. 복부에 놓인 손과 어두운 부분이 보이지만 관통상 출혈은 명확히 식별되지 않는다. 패널과 주변 잔해는 실제 금속·목재로 읽히며 읽을 수 있는 문자는 없다.",
        "hard_violations": [],
        "physics": "머리와 몸통, 다리는 패널에 받쳐져 있다. 한쪽 팔은 몸 옆에 내려놓았고 다른 손은 복부 위에 얹혀 있어 각각 패널과 몸의 지지를 받는다. 뒤꿈치가 패널에 놓인 상태에서 발끝이 위로 향하는 자세도 가능하다. 지지 없이 떠 있는 신체나 물체는 보이지 않으며 패널은 수면에서 부유한다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.492,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.492,
    "B": 2.0
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1492
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "지정된 와이드 쇼트 화각과 배경 묘사를 완벽히 구현했으며, 캐논 설정인 복부의 출혈 자국까지 충실하게 시각화했습니다."
   },
   {
    "label": "A",
    "score": 1492,
    "verdict_ko": "요구된 위치와 카메라 시점은 잘 반영되었으나, 필수 설정인 복부의 찔린 상처와 출혈 묘사가 누락되었습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_flood_raft_water_456400.png",
    "asset_id": "0797d468-af7c-4934-b8f4-036bf4cc5448",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1113064>",
    "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-7c42-7c5e-a07b-91605bee761d",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S38sh2__bgfirst_bg.png",
   "bg_asset_id": "6a473823-e079-4c39-92bf-e7c364af89d6",
   "bg_record_key": "S38sh2::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "flood_raft_water",
   "groupbg_asset_id": "0797d468-af7c-4934-b8f4-036bf4cc5448"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S38sh2::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:10:11.152678+00:00",
  "fingerprint": "e0b1a36c85ff60d4ca9d96afec9efaf3f112b80899529d3880a7022106310662",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S38sh2_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S38sh2_sel.png",
  "source_sha256": "bf3a60103bba2903490fecd61d80722f6fd15e933c08fe82cbf271a454808a71",
  "file": "S38sh2_cine.png",
  "staged_sha256": "b7d93f36821316a55ae4a5e2772fad2420d99121e13e7789b3cf9778b43c9b9a",
  "latency_ms": 13930
 },
 "S38sh7::signage": {
  "fp": "7ea4e02b84278fc1",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S38sh7": {
  "input_fingerprint": "16aa1e855c1d0a73",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 숨을 거둔 미연의 몸을 끌어안은 채 새벽하늘을 향해 입을 크게 벌리고 절규하는 현우의 얼굴.\n\nLOCATION (lock): On the floating container panel supporting the injured mother, surrounded by the flooded settlement at dawn. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container board (Floating beneath the pair) — Only a small portion of its supporting face is visible at the lower edge; used as Maintains the physical basis of 현우's seated embrace; Sea (Surrounding the floating board); used as A peripheral strip preserves the exposed setting without distracting from the faces; Dawn sky (Visible above the sea at daybreak); used as Provides uncluttered space above 현우's upward cry.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Soft early-dawn ambient illumination preserves facial detail and restrained contrast, with no change in light treatment to signal 미연's death.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the floating container panel, surrounding floodwater, and dawn coloration. Exclude unrelated floating platforms and do not import the reference's human figure as part of the setting.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Miyeon's motionless body is held against Hyunwoo in his embrace on the floating container panel after her death. The source does not specify how far he lifts her torso, how her head rests, or the positions of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The container panels and surviving containers remain afloat on the dawn sea. 미연: She is now dead, with closed eyes, a swollen face and a bleeding abdominal puncture wound. Her body remains on the floating container panel. 현우: He is on the floating panel, crying with his arms closed in an embrace. His facial bruises and leg wound remain, and the contact card is still hidden inside his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 숨을 거둔 미연의 몸을 끌어안은 채 새벽하늘을 향해 입을 크게 벌리고 절규하는 현우의 얼굴.\n\nLOCATION (lock): On the floating container panel supporting the injured mother, surrounded by the flooded settlement at dawn. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container board (Floating beneath the pair) — Only a small portion of its supporting face is visible at the lower edge; used as Maintains the physical basis of 현우's seated embrace; Sea (Surrounding the floating board); used as A peripheral strip preserves the exposed setting without distracting from the faces; Dawn sky (Visible above the sea at daybreak); used as Provides uncluttered space above 현우's upward cry.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Soft early-dawn ambient illumination preserves facial detail and restrained contrast, with no change in light treatment to signal 미연's death.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the floating container panel, surrounding floodwater, and dawn coloration. Exclude unrelated floating platforms and do not import the reference's human figure as part of the setting.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Miyeon's motionless body is held against Hyunwoo in his embrace on the floating container panel after her death. The source does not specify how far he lifts her torso, how her head rests, or the positions of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The container panels and surviving containers remain afloat on the dawn sea. 미연: She is now dead, with closed eyes, a swollen face and a bleeding abdominal puncture wound. Her body remains on the floating container panel. 현우: He is on the floating panel, crying with his arms closed in an embrace. His facial bruises and leg wound remain, and the contact card is still hidden inside his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 숨을 거둔 미연의 몸을 끌어안은 채 새벽하늘을 향해 입을 크게 벌리고 절규하는 현우의 얼굴.\n\nLOCATION (lock): On the floating container panel supporting the injured mother, surrounded by the flooded settlement at dawn. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container board (Floating beneath the pair) — Only a small portion of its supporting face is visible at the lower edge; used as Maintains the physical basis of 현우's seated embrace; Sea (Surrounding the floating board); used as A peripheral strip preserves the exposed setting without distracting from the faces; Dawn sky (Visible above the sea at daybreak); used as Provides uncluttered space above 현우's upward cry.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Soft early-dawn ambient illumination preserves facial detail and restrained contrast, with no change in light treatment to signal 미연's death.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the floating container panel, surrounding floodwater, and dawn coloration. Exclude unrelated floating platforms and do not import the reference's human figure as part of the setting.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Miyeon's motionless body is held against Hyunwoo in his embrace on the floating container panel after her death. The source does not specify how far he lifts her torso, how her head rests, or the positions of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The container panels and surviving containers remain afloat on the dawn sea. 미연: She is now dead, with closed eyes, a swollen face and a bleeding abdominal puncture wound. Her body remains on the floating container panel. 현우: He is on the floating panel, crying with his arms closed in an embrace. His facial bruises and leg wound remain, and the contact card is still hidden inside his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선과 입 방향이 새벽 하늘을 향하고 있음.",
    "built_space": "레퍼런스와 동일한 형태의 컨테이너 패널 모서리와 침수된 도시 스카이라인이 정확한 위치에 배치됨.",
    "entities": "현우(젊은 외모, 상처)와 미연(복부 상처, 감은 눈) 모두 레퍼런스 특징을 잘 반영함.",
    "hard_violations": [],
    "physics": "현우가 컨테이너 위에 앉아 미연을 안고 있으며, 미연의 몸은 중력에 따라 자연스럽게 늘어져 지탱됨."
   },
   {
    "label": "B",
    "direction": "현우가 위쪽 하늘을 바라보며 절규함.",
    "built_space": "바다와 컨테이너 일부만 보이며, 레퍼런스의 특징적인 배경 스카이라인 구조물들이 부재함.",
    "entities": "현우의 얼굴이 지나치게 나이 들어 보이며, 미연이 지정된 의상(짙은 파란색)이 아닌 갈색 옷을 입고 있음.",
    "hard_violations": [],
    "physics": "현우의 팔이 미연의 몸을 감싸 안아 물리적으로 지탱하고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "지정된 배경(스카이라인, 컨테이너 모서리)과 인물의 외모를 정확히 재현했으며, 요구된 포즈와 감정선이 사실적으로 표현되었습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "필수적인 배경 디테일이 누락되었고, 미연의 의상 설정 위반 및 현우의 연령대 묘사에 오류가 있습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선과 입 방향이 새벽 하늘을 향하고 있음.",
        "built_space": "레퍼런스와 동일한 형태의 컨테이너 패널 모서리와 침수된 도시 스카이라인이 정확한 위치에 배치됨.",
        "entities": "현우(젊은 외모, 상처)와 미연(복부 상처, 감은 눈) 모두 레퍼런스 특징을 잘 반영함.",
        "hard_violations": [],
        "physics": "현우가 컨테이너 위에 앉아 미연을 안고 있으며, 미연의 몸은 중력에 따라 자연스럽게 늘어져 지탱됨."
       },
       {
        "label": "B",
        "direction": "현우가 위쪽 하늘을 바라보며 절규함.",
        "built_space": "바다와 컨테이너 일부만 보이며, 레퍼런스의 특징적인 배경 스카이라인 구조물들이 부재함.",
        "entities": "현우의 얼굴이 지나치게 나이 들어 보이며, 미연이 지정된 의상(짙은 파란색)이 아닌 갈색 옷을 입고 있음.",
        "hard_violations": [],
        "physics": "현우의 팔이 미연의 몸을 감싸 안아 물리적으로 지탱하고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "지정된 배경(스카이라인, 컨테이너 모서리)과 인물의 외모를 정확히 재현했으며, 요구된 포즈와 감정선이 사실적으로 표현되었습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "필수적인 배경 디테일이 누락되었고, 미연의 의상 설정 위반 및 현우의 연령대 묘사에 오류가 있습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선과 입 방향이 새벽 하늘을 향하고 있음.",
        "built_space": "레퍼런스와 동일한 형태의 컨테이너 패널 모서리와 침수된 도시 스카이라인이 정확한 위치에 배치됨.",
        "entities": "현우(젊은 외모, 상처)와 미연(복부 상처, 감은 눈) 모두 레퍼런스 특징을 잘 반영함.",
        "hard_violations": [],
        "physics": "현우가 컨테이너 위에 앉아 미연을 안고 있으며, 미연의 몸은 중력에 따라 자연스럽게 늘어져 지탱됨."
       },
       {
        "label": "B",
        "direction": "현우가 위쪽 하늘을 바라보며 절규함.",
        "built_space": "바다와 컨테이너 일부만 보이며, 레퍼런스의 특징적인 배경 스카이라인 구조물들이 부재함.",
        "entities": "현우의 얼굴이 지나치게 나이 들어 보이며, 미연이 지정된 의상(짙은 파란색)이 아닌 갈색 옷을 입고 있음.",
        "hard_violations": [],
        "physics": "현우의 팔이 미연의 몸을 감싸 안아 물리적으로 지탱하고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "상향 절규와 지지되는 포옹은 맞지만 얼굴 클로즈업보다 넓고, 미연의 회갈색 상의와 현우의 불분명한 얼굴 멍, 평범한 바다 배경이 연속성을 약화한다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "클로즈업보다 넓고 바다 노출이 과하지만, 하늘을 향한 절규와 얼굴 상처, 미연의 어두운 상의, 침수 정착지와 금속 패널의 연속성이 A보다 충실하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 눈을 감고 턱과 얼굴을 위로 들어 입을 크게 벌린다. 절규의 방향은 머리 위 새벽하늘을 향한다. 미연은 눈을 감고 얼굴을 화면 왼쪽으로 기울여 현우의 목과 가슴에 기대고 있다. 겨누거나 조작하는 물건은 없다.",
        "built_space": "두 사람 아래에 녹슨 청록색 금속 지지체 하나가 있고, 뒤쪽과 양옆의 솟은 테두리가 보여 평평한 컨테이너 패널보다는 얕은 상자 같은 인상이 강하다. 현우는 그 안에 앉아 미연을 무릎과 가슴 쪽으로 안는다. 바다는 하단 주변 띠로, 하늘은 넓게 보인다. 얼굴뿐 아니라 팔 전체와 복부, 무릎 일부까지 들어와 요청한 얼굴 클로즈업보다 넓다. 이전 장면의 침수 건물이나 잔존 컨테이너는 보이지 않는다.",
        "entities": "인물은 젊은 동아시아계 남성 한 명과 중년 동아시아계 여성 한 명뿐이다. 현우의 검은 머리와 남색 반소매는 참조와 대체로 맞지만 얼굴의 멍은 뚜렷하지 않다. 미연의 검은 머리, 감긴 눈, 중년 얼굴은 맞으나 회갈색 상의는 참조의 어두운 남색 상의와 다르다. 복부에는 혈흔이 있지만 천자상 자체와 얼굴 부종은 명확하지 않다. 다리 상처와 신발 속 카드는 프레임 밖이므로 확인 대상이 아니다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우의 앉은 하체는 하단 금속면에 놓여 있고, 두 팔과 손이 미연의 어깨 및 몸통에 접촉해 끌어안는다. 미연의 머리는 현우의 목과 가슴에 기대며 몸통은 팔과 무릎 쪽에 받쳐져 있다. 보이는 미연의 팔은 아래로 내려가 있고 스스로 치켜든 부위는 없다. 금속 지지체는 주변 수면에 떠 있는 것으로 읽히며, 지지 없이 공중에 뜬 신체나 물건은 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "현우는 얼굴과 눈을 화면 오른쪽 위의 새벽하늘로 향하고 입을 크게 벌려 절규한다. 시선과 얼굴 방향 모두 요청한 하늘을 향한다. 미연은 눈을 감은 채 얼굴이 화면 왼쪽으로 기울어 있으며 능동적으로 바라보는 대상은 없다. 방향을 검사할 별도 도구는 없다.",
        "built_space": "하단에 녹슬고 젖은 컨테이너 패널 하나가 보이며, 왼쪽 아래에는 구멍이 있는 모서리 철물 하나와 보강선들이 있다. 현우는 패널 위에 앉아 미연을 자기 몸 앞에 안는다. 뒤에는 수면, 먼 건물 윤곽, 기울어진 컨테이너와 구조물들이 작게 배치되어 이전 침수 정착지와 연결된다. 별도의 사람이 탄 플랫폼은 없다. 다만 패널 지지면과 바다가 상당한 면적을 차지하고 복부까지 보여, 얼굴 클로즈업과 주변의 좁은 바다 띠라는 구도 지시는 충족하지 못한다.",
        "entities": "젊은 동아시아계 남성과 중년 동아시아계 여성, 두 사람만 보인다. 현우는 헝클어진 검은 머리와 앳된 얼굴을 갖고 있으며 볼의 멍과 찰과상이 선명하다. 상의는 참조보다 회색 기운이 강하다. 미연의 검은 단발과 어두운 남색 상의는 참조에 더 가깝고, 감긴 눈과 힘 빠진 표정, 입가 혈흔이 보인다. 복부에는 출혈하는 상처가 드러나지만 얼굴 부종은 강하게 표현되지 않았다. 패널, 바다, 새벽하늘이 모두 있으며 읽을 수 있는 글자는 없다. 프레임 밖 다리 상처나 숨겨진 카드는 판단하지 않는다.",
        "hard_violations": [],
        "physics": "현우의 하체는 패널 위에 놓이고, 한 팔은 미연의 윗가슴을 감싸 반대쪽 어깨를 잡으며 다른 손은 복부 옆에 닿아 몸통을 받친다. 미연의 기울어진 머리와 목은 현우의 어깨·윗가슴에 기대고, 몸통은 그의 팔과 무릎 쪽에 실려 있다. 보이는 미연의 팔은 아래로 늘어져 프레임 밖으로 이어진다. 시신이 스스로 자세를 유지하거나 지지 없이 떠 있는 명백한 부위는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "상향 절규와 지지되는 포옹은 맞지만 얼굴 클로즈업보다 넓고, 미연의 회갈색 상의와 현우의 불분명한 얼굴 멍, 평범한 바다 배경이 연속성을 약화한다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "클로즈업보다 넓고 바다 노출이 과하지만, 하늘을 향한 절규와 얼굴 상처, 미연의 어두운 상의, 침수 정착지와 금속 패널의 연속성이 A보다 충실하다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 눈을 감고 턱과 얼굴을 위로 들어 입을 크게 벌린다. 절규의 방향은 머리 위 새벽하늘을 향한다. 미연은 눈을 감고 얼굴을 화면 왼쪽으로 기울여 현우의 목과 가슴에 기대고 있다. 겨누거나 조작하는 물건은 없다.",
        "built_space": "두 사람 아래에 녹슨 청록색 금속 지지체 하나가 있고, 뒤쪽과 양옆의 솟은 테두리가 보여 평평한 컨테이너 패널보다는 얕은 상자 같은 인상이 강하다. 현우는 그 안에 앉아 미연을 무릎과 가슴 쪽으로 안는다. 바다는 하단 주변 띠로, 하늘은 넓게 보인다. 얼굴뿐 아니라 팔 전체와 복부, 무릎 일부까지 들어와 요청한 얼굴 클로즈업보다 넓다. 이전 장면의 침수 건물이나 잔존 컨테이너는 보이지 않는다.",
        "entities": "인물은 젊은 동아시아계 남성 한 명과 중년 동아시아계 여성 한 명뿐이다. 현우의 검은 머리와 남색 반소매는 참조와 대체로 맞지만 얼굴의 멍은 뚜렷하지 않다. 미연의 검은 머리, 감긴 눈, 중년 얼굴은 맞으나 회갈색 상의는 참조의 어두운 남색 상의와 다르다. 복부에는 혈흔이 있지만 천자상 자체와 얼굴 부종은 명확하지 않다. 다리 상처와 신발 속 카드는 프레임 밖이므로 확인 대상이 아니다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우의 앉은 하체는 하단 금속면에 놓여 있고, 두 팔과 손이 미연의 어깨 및 몸통에 접촉해 끌어안는다. 미연의 머리는 현우의 목과 가슴에 기대며 몸통은 팔과 무릎 쪽에 받쳐져 있다. 보이는 미연의 팔은 아래로 내려가 있고 스스로 치켜든 부위는 없다. 금속 지지체는 주변 수면에 떠 있는 것으로 읽히며, 지지 없이 공중에 뜬 신체나 물건은 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "현우는 얼굴과 눈을 화면 오른쪽 위의 새벽하늘로 향하고 입을 크게 벌려 절규한다. 시선과 얼굴 방향 모두 요청한 하늘을 향한다. 미연은 눈을 감은 채 얼굴이 화면 왼쪽으로 기울어 있으며 능동적으로 바라보는 대상은 없다. 방향을 검사할 별도 도구는 없다.",
        "built_space": "하단에 녹슬고 젖은 컨테이너 패널 하나가 보이며, 왼쪽 아래에는 구멍이 있는 모서리 철물 하나와 보강선들이 있다. 현우는 패널 위에 앉아 미연을 자기 몸 앞에 안는다. 뒤에는 수면, 먼 건물 윤곽, 기울어진 컨테이너와 구조물들이 작게 배치되어 이전 침수 정착지와 연결된다. 별도의 사람이 탄 플랫폼은 없다. 다만 패널 지지면과 바다가 상당한 면적을 차지하고 복부까지 보여, 얼굴 클로즈업과 주변의 좁은 바다 띠라는 구도 지시는 충족하지 못한다.",
        "entities": "젊은 동아시아계 남성과 중년 동아시아계 여성, 두 사람만 보인다. 현우는 헝클어진 검은 머리와 앳된 얼굴을 갖고 있으며 볼의 멍과 찰과상이 선명하다. 상의는 참조보다 회색 기운이 강하다. 미연의 검은 단발과 어두운 남색 상의는 참조에 더 가깝고, 감긴 눈과 힘 빠진 표정, 입가 혈흔이 보인다. 복부에는 출혈하는 상처가 드러나지만 얼굴 부종은 강하게 표현되지 않았다. 패널, 바다, 새벽하늘이 모두 있으며 읽을 수 있는 글자는 없다. 프레임 밖 다리 상처나 숨겨진 카드는 판단하지 않는다.",
        "hard_violations": [],
        "physics": "현우의 하체는 패널 위에 놓이고, 한 팔은 미연의 윗가슴을 감싸 반대쪽 어깨를 잡으며 다른 손은 복부 옆에 닿아 몸통을 받친다. 미연의 기울어진 머리와 목은 현우의 어깨·윗가슴에 기대고, 몸통은 그의 팔과 무릎 쪽에 실려 있다. 보이는 미연의 팔은 아래로 늘어져 프레임 밖으로 이어진다. 시신이 스스로 자세를 유지하거나 지지 없이 떠 있는 명백한 부위는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.089
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.089
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1089
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지정된 배경(스카이라인, 컨테이너 모서리)과 인물의 외모를 정확히 재현했으며, 요구된 포즈와 감정선이 사실적으로 표현되었습니다."
   },
   {
    "label": "B",
    "score": 1089,
    "verdict_ko": "필수적인 배경 디테일이 누락되었고, 미연의 의상 설정 위반 및 현우의 연령대 묘사에 오류가 있습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 미연 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S38sh2_sel.png",
    "asset_id": "554835af-6c75-4628-96f5-e1c8780e94a0",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1113064>",
    "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-810a-7383-af79-e62baa6b0fc8",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S38sh2"
  }
 },
 "S38sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:11:58.099452+00:00",
  "fingerprint": "32709d6d856a3afe6ea78d979ef8bdeb29919a19e8f10c5ef99d8dd54976961f",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S38sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S38sh7_sel.png",
  "source_sha256": "4c1a97930929bb72985592fa92cd822ca63d3f5f659e42bdb2ceffa6bd9a7dc0",
  "file": "S38sh7_cine.png",
  "staged_sha256": "5eb5310e308797199627e5fef848c83085a2e213b0092db8cdbf546c4808fe46",
  "latency_ms": 12195
 },
 "S38sh13::signage": {
  "fp": "f4fe3ea44c2a1244",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S38sh13": {
  "input_fingerprint": "1f531ce66dddd499",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 오열하는 앰버를 굽어보며 어깨를 축 늘어뜨린 채 슬픈 표정을 짓는 찰리의 낡은 금속 상체.\n\nLOCATION (lock): On floating container wreckage beside the family's raft-like panel, above the inundated refugee settlement at dawn. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Container top (Supporting 찰리 and 앰버 after their escape) — A limited portion of the upper supporting face is visible beneath the figures; used as Connects their different frame heights within the same physical space; Sea (Surrounding the containers); used as Peripheral background preserves the aftermath setting.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the early-dawn illumination subdued and continuous, with the final fade treated as an editorial darkening rather than a change of physical light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same floating panel, immediately surrounding floodwater, and dawn light. Exclude the separate enclosed container interior and its storage boxes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The floating containers and container panel have drawn close together on the dawn sea. 찰리: He stands on the container with a sorrowful expression, retaining his worn metal body and old coat-and-hat disguise.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 오열하는 앰버를 굽어보며 어깨를 축 늘어뜨린 채 슬픈 표정을 짓는 찰리의 낡은 금속 상체.\n\nLOCATION (lock): On floating container wreckage beside the family's raft-like panel, above the inundated refugee settlement at dawn. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Container top (Supporting 찰리 and 앰버 after their escape) — A limited portion of the upper supporting face is visible beneath the figures; used as Connects their different frame heights within the same physical space; Sea (Surrounding the containers); used as Peripheral background preserves the aftermath setting.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the early-dawn illumination subdued and continuous, with the final fade treated as an editorial darkening rather than a change of physical light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same floating panel, immediately surrounding floodwater, and dawn light. Exclude the separate enclosed container interior and its storage boxes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The floating containers and container panel have drawn close together on the dawn sea. 찰리: He stands on the container with a sorrowful expression, retaining his worn metal body and old coat-and-hat disguise.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 오열하는 앰버를 굽어보며 어깨를 축 늘어뜨린 채 슬픈 표정을 짓는 찰리의 낡은 금속 상체.\n\nLOCATION (lock): On floating container wreckage beside the family's raft-like panel, above the inundated refugee settlement at dawn. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Container top (Supporting 찰리 and 앰버 after their escape) — A limited portion of the upper supporting face is visible beneath the figures; used as Connects their different frame heights within the same physical space; Sea (Surrounding the containers); used as Peripheral background preserves the aftermath setting.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the early-dawn illumination subdued and continuous, with the final fade treated as an editorial darkening rather than a change of physical light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same floating panel, immediately surrounding floodwater, and dawn light. Exclude the separate enclosed container interior and its storage boxes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The floating containers and container panel have drawn close together on the dawn sea. 찰리: He stands on the container with a sorrowful expression, retaining his worn metal body and old coat-and-hat disguise.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "찰리가 패널 위에 쓰러진 앰버를 내려다보고 있음.",
    "built_space": "바다 위 부유하는 패널 구조물. 찰리가 패널 공간을 벗어나 있음.",
    "entities": "찰리의 금속 몸체는 묘사되었으나 코트와 모자 변장이 누락됨. 앰버가 오열하고 있음.",
    "hard_violations": [
     "[gemini-pro] 찰리가 컨테이너 위에서 지탱되지 않고 물속에 서 있어 프롬프트의 스테이징 지시를 위반함"
    ],
    "physics": "찰리의 하반신이 물에 잠겨 있으며, 컨테이너 위에서 체중을 지탱하지 못함. 앰버는 패널 위에 누워 있음."
   },
   {
    "label": "B",
    "direction": "찰리의 시선과 자세가 쓰러져 오열하는 앰버를 향해 아래로 굽어보고 있음.",
    "built_space": "바다 위에 떠 있는 컨테이너 상단 표면. 배경에 다른 컨테이너들이 보임.",
    "entities": "찰리는 낡은 금속 몸체에 지시된 코트와 모자를 착용함. 앰버는 슬퍼하는 표정으로 묘사됨.",
    "hard_violations": [
     "[gpt-high] 기존 장소에 없고 지시에도 없는 대형 목재 등받이형 구조물을 만들어 여성을 기대어 눕혔다. 이는 사소한 배경 차이가 아니라 인물의 지지와 배치를 바꾼 새 구조물이다."
    ],
    "physics": "찰리와 앰버 모두 컨테이너 상단과 나무 판자에 안정적으로 지탱되어 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "찰리의 코트와 모자 변장을 정확히 반영했으며, 두 인물 모두 컨테이너 위에서 지탱되는 스테이징을 올바르게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "찰리의 코트와 모자 변장을 누락했고, 컨테이너 위가 아닌 물속에 서 있어 치명적인 스테이징 위반이 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 시선과 자세가 쓰러져 오열하는 앰버를 향해 아래로 굽어보고 있음.",
        "built_space": "바다 위에 떠 있는 컨테이너 상단 표면. 배경에 다른 컨테이너들이 보임.",
        "entities": "찰리는 낡은 금속 몸체에 지시된 코트와 모자를 착용함. 앰버는 슬퍼하는 표정으로 묘사됨.",
        "hard_violations": [],
        "physics": "찰리와 앰버 모두 컨테이너 상단과 나무 판자에 안정적으로 지탱되어 있음."
       },
       {
        "label": "A",
        "direction": "찰리가 패널 위에 쓰러진 앰버를 내려다보고 있음.",
        "built_space": "바다 위 부유하는 패널 구조물. 찰리가 패널 공간을 벗어나 있음.",
        "entities": "찰리의 금속 몸체는 묘사되었으나 코트와 모자 변장이 누락됨. 앰버가 오열하고 있음.",
        "hard_violations": [
         "찰리가 컨테이너 위에서 지탱되지 않고 물속에 서 있어 프롬프트의 스테이징 지시를 위반함"
        ],
        "physics": "찰리의 하반신이 물에 잠겨 있으며, 컨테이너 위에서 체중을 지탱하지 못함. 앰버는 패널 위에 누워 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "찰리의 코트와 모자 변장을 정확히 반영했으며, 두 인물 모두 컨테이너 위에서 지탱되는 스테이징을 올바르게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "찰리의 코트와 모자 변장을 누락했고, 컨테이너 위가 아닌 물속에 서 있어 치명적인 스테이징 위반이 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 시선과 자세가 쓰러져 오열하는 앰버를 향해 아래로 굽어보고 있음.",
        "built_space": "바다 위에 떠 있는 컨테이너 상단 표면. 배경에 다른 컨테이너들이 보임.",
        "entities": "찰리는 낡은 금속 몸체에 지시된 코트와 모자를 착용함. 앰버는 슬퍼하는 표정으로 묘사됨.",
        "hard_violations": [],
        "physics": "찰리와 앰버 모두 컨테이너 상단과 나무 판자에 안정적으로 지탱되어 있음."
       },
       {
        "label": "A",
        "direction": "찰리가 패널 위에 쓰러진 앰버를 내려다보고 있음.",
        "built_space": "바다 위 부유하는 패널 구조물. 찰리가 패널 공간을 벗어나 있음.",
        "entities": "찰리의 금속 몸체는 묘사되었으나 코트와 모자 변장이 누락됨. 앰버가 오열하고 있음.",
        "hard_violations": [
         "찰리가 컨테이너 위에서 지탱되지 않고 물속에 서 있어 프롬프트의 스테이징 지시를 위반함"
        ],
        "physics": "찰리의 하반신이 물에 잠겨 있으며, 컨테이너 위에서 체중을 지탱하지 못함. 앰버는 패널 위에 누워 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "낡은 코트와 모자, 아래를 보는 찰리는 맞지만, 커다란 목재 받침을 새로 만들고 인물의 하체와 컨테이너를 넓게 보여 주어 지정된 장소와 상체 중심 구도를 훼손했다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "찰리가 낮은 위치의 여성을 내려다보는 관계와 기존 패널·새벽 바다의 연속성이 더 정확하지만, 필수 코트와 모자가 없고 상체 중심 미디엄 숏보다 범위가 넓다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 얼굴은 화면 왼쪽 아래로 숙여져 있으며, 바로 아래 목재에 기대 우는 여성을 향한다. 여성은 눈을 감고 얼굴을 아래로 기울여 찰리와 눈을 맞추지 않는다. 무기나 방향을 확인할 휴대 도구는 없다.",
        "built_space": "화면 아래에 넓은 골판 금속 컨테이너 지붕이 있고, 여성 뒤에는 여러 장의 큰 목재 판자가 비스듬히 세워져 있다. 찰리는 그 뒤쪽, 여성은 앞쪽 낮은 위치에 있다. 주변에는 파란색과 붉은색 컨테이너 여러 개가 상당히 크게 보인다. 참조의 녹슨 평판과 구멍 난 모서리 철물 대신 골판 지붕과 목재 받침이 주요 공간을 차지하며, 지지면도 요구한 제한적 노출보다 넓다.",
        "entities": "찰리 한 체와 성인 여성 한 명이 보인다. 찰리의 육중한 베이지 장갑, 긴 팔, 흰 각진 마스크 얼굴은 참조와 대체로 일치하고 낡은 코트와 모자도 있다. 여성은 검은 머리, 짙은 반소매 상의, 밝은 바지 차림으로 이전 사진 여성의 외양·의상과 매우 유사하다. 앰버의 독립적인 인물 참조가 없어 그 정체성은 확인할 수 없다. 바다와 컨테이너는 있지만 기존 가족 패널의 특징은 약하다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "기존 장소에 없고 지시에도 없는 대형 목재 등받이형 구조물을 만들어 여성을 기대어 눕혔다. 이는 사소한 배경 차이가 아니라 인물의 지지와 배치를 바꾼 새 구조물이다."
        ],
        "physics": "여성의 몸통은 비스듬한 목재에 기대고 골반과 다리는 아래쪽 판재 및 금속 지지면에 놓여 있어 공중에 뜬 자세는 아니다. 찰리의 하체는 지붕 위로 이어지지만 발 접촉은 화면 밖이라 직접 확인되지 않는다. 팔은 아래로 처지고 코트는 중력 방향으로 늘어진다. 목재 받침의 하단은 지붕에 닿지만 고정 방식은 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "찰리는 고개를 숙여 화면 아래 전경의 여성을 내려다본다. 여성은 눈을 감고 얼굴을 바닥 쪽으로 돌린 채 몸을 웅크리고 있어 시선 교환은 없다. 찰리의 응시 방향은 아래에 있는 앰버라는 장면 관계에 부합한다.",
        "built_space": "하단 왼쪽에 녹슨 평판과 원형 구멍 하나가 있는 모서리 철물이 보이며 이전 사진의 패널 특징을 유지한다. 찰리는 뒤쪽 높은 위치, 여성은 앞쪽 낮은 위치에 있다. 주변 바다와 멀리 기울어진 잔해, 도시 윤곽도 참조와 가깝다. 불필요한 실내나 보관 상자는 없다. 다만 찰리의 허벅지와 여성의 몸통까지 포함해 상체 중심 촬영보다 넓다.",
        "entities": "찰리 한 체와 성인 여성 한 명이 보인다. 찰리의 베이지 장갑판, 흰 마스크 얼굴, 긴 기계 팔과 육중한 체격은 참조에 가깝지만, 유지해야 할 낡은 코트와 모자는 모두 없다. 여성의 검은 머리와 짙은 상의·밝은 바지는 이전 사진 여성과 유사해 인물 외양을 가져오지 말라는 조건에 대한 우려가 있다. 독립적인 앰버 참조가 없어 동일인 여부는 단정할 수 없다. 새벽 바다와 녹슨 패널은 일치하며 읽을 수 있는 문자는 없다.",
        "hard_violations": [],
        "physics": "여성은 하단 패널 위에 옆으로 몸을 낮추고 있으며 머리 가까이 놓인 팔과 아래쪽 몸통이 지지면으로 이어진다. 찰리의 다리는 여성 뒤로 가려져 발 접촉이 직접 보이지 않지만, 패널 위에 서 있는 배치와 모순되지 않는다. 양팔은 어깨와 팔꿈치 관절에서 자연스럽게 아래로 이어진다. 지지 없이 공중에 떠 있는 몸이나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "낡은 코트와 모자, 아래를 보는 찰리는 맞지만, 커다란 목재 받침을 새로 만들고 인물의 하체와 컨테이너를 넓게 보여 주어 지정된 장소와 상체 중심 구도를 훼손했다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "찰리가 낮은 위치의 여성을 내려다보는 관계와 기존 패널·새벽 바다의 연속성이 더 정확하지만, 필수 코트와 모자가 없고 상체 중심 미디엄 숏보다 범위가 넓다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 얼굴은 화면 왼쪽 아래로 숙여져 있으며, 바로 아래 목재에 기대 우는 여성을 향한다. 여성은 눈을 감고 얼굴을 아래로 기울여 찰리와 눈을 맞추지 않는다. 무기나 방향을 확인할 휴대 도구는 없다.",
        "built_space": "화면 아래에 넓은 골판 금속 컨테이너 지붕이 있고, 여성 뒤에는 여러 장의 큰 목재 판자가 비스듬히 세워져 있다. 찰리는 그 뒤쪽, 여성은 앞쪽 낮은 위치에 있다. 주변에는 파란색과 붉은색 컨테이너 여러 개가 상당히 크게 보인다. 참조의 녹슨 평판과 구멍 난 모서리 철물 대신 골판 지붕과 목재 받침이 주요 공간을 차지하며, 지지면도 요구한 제한적 노출보다 넓다.",
        "entities": "찰리 한 체와 성인 여성 한 명이 보인다. 찰리의 육중한 베이지 장갑, 긴 팔, 흰 각진 마스크 얼굴은 참조와 대체로 일치하고 낡은 코트와 모자도 있다. 여성은 검은 머리, 짙은 반소매 상의, 밝은 바지 차림으로 이전 사진 여성의 외양·의상과 매우 유사하다. 앰버의 독립적인 인물 참조가 없어 그 정체성은 확인할 수 없다. 바다와 컨테이너는 있지만 기존 가족 패널의 특징은 약하다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "기존 장소에 없고 지시에도 없는 대형 목재 등받이형 구조물을 만들어 여성을 기대어 눕혔다. 이는 사소한 배경 차이가 아니라 인물의 지지와 배치를 바꾼 새 구조물이다."
        ],
        "physics": "여성의 몸통은 비스듬한 목재에 기대고 골반과 다리는 아래쪽 판재 및 금속 지지면에 놓여 있어 공중에 뜬 자세는 아니다. 찰리의 하체는 지붕 위로 이어지지만 발 접촉은 화면 밖이라 직접 확인되지 않는다. 팔은 아래로 처지고 코트는 중력 방향으로 늘어진다. 목재 받침의 하단은 지붕에 닿지만 고정 방식은 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "찰리는 고개를 숙여 화면 아래 전경의 여성을 내려다본다. 여성은 눈을 감고 얼굴을 바닥 쪽으로 돌린 채 몸을 웅크리고 있어 시선 교환은 없다. 찰리의 응시 방향은 아래에 있는 앰버라는 장면 관계에 부합한다.",
        "built_space": "하단 왼쪽에 녹슨 평판과 원형 구멍 하나가 있는 모서리 철물이 보이며 이전 사진의 패널 특징을 유지한다. 찰리는 뒤쪽 높은 위치, 여성은 앞쪽 낮은 위치에 있다. 주변 바다와 멀리 기울어진 잔해, 도시 윤곽도 참조와 가깝다. 불필요한 실내나 보관 상자는 없다. 다만 찰리의 허벅지와 여성의 몸통까지 포함해 상체 중심 촬영보다 넓다.",
        "entities": "찰리 한 체와 성인 여성 한 명이 보인다. 찰리의 베이지 장갑판, 흰 마스크 얼굴, 긴 기계 팔과 육중한 체격은 참조에 가깝지만, 유지해야 할 낡은 코트와 모자는 모두 없다. 여성의 검은 머리와 짙은 상의·밝은 바지는 이전 사진 여성과 유사해 인물 외양을 가져오지 말라는 조건에 대한 우려가 있다. 독립적인 앰버 참조가 없어 동일인 여부는 단정할 수 없다. 새벽 바다와 녹슨 패널은 일치하며 읽을 수 있는 문자는 없다.",
        "hard_violations": [],
        "physics": "여성은 하단 패널 위에 옆으로 몸을 낮추고 있으며 머리 가까이 놓인 팔과 아래쪽 몸통이 지지면으로 이어진다. 찰리의 다리는 여성 뒤로 가려져 발 접촉이 직접 보이지 않지만, 패널 위에 서 있는 배치와 모순되지 않는다. 양팔은 어깨와 팔꿈치 관절에서 자연스럽게 아래로 이어진다. 지지 없이 공중에 떠 있는 몸이나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.429,
    "B": 1.5
   },
   "adjusted": {
    "A": 1.179,
    "B": 1.25
   },
   "violations": {
    "A": [
     "[gemini-pro] 찰리가 컨테이너 위에서 지탱되지 않고 물속에 서 있어 프롬프트의 스테이징 지시를 위반함"
    ],
    "B": [
     "[gpt-high] 기존 장소에 없고 지시에도 없는 대형 목재 등받이형 구조물을 만들어 여성을 기대어 눕혔다. 이는 사소한 배경 차이가 아니라 인물의 지지와 배치를 바꾼 새 구조물이다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1250,
   "A": 1179
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1250,
    "verdict_ko": "찰리의 코트와 모자 변장을 정확히 반영했으며, 두 인물 모두 컨테이너 위에서 지탱되는 스테이징을 올바르게 구현했습니다.  ★위반: [gpt-high] 기존 장소에 없고 지시에도 없는 대형 목재 등받이형 구조물을 만들어 여성을 기대어 눕혔다. 이는 사소한 배경 차이가 아니라 인물의 지지와 배치를 바꾼 새 구조물이다."
   },
   {
    "label": "A",
    "score": 1179,
    "verdict_ko": "찰리의 코트와 모자 변장을 누락했고, 컨테이너 위가 아닌 물속에 서 있어 치명적인 스테이징 위반이 발생했습니다.  ★위반: [gemini-pro] 찰리가 컨테이너 위에서 지탱되지 않고 물속에 서 있어 프롬프트의 스테이징 지시를 위반함"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S38sh7_sel.png",
    "asset_id": "f4087254-1c5a-4a73-b72d-40eea4ee6c67",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-82c6-725e-bdad-41418c972e54",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S38sh7"
  }
 },
 "S38sh13::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:12:58.748063+00:00",
  "fingerprint": "89304417cea17d1c275bc8d879df146a3f6b7e4eb3b9daa1f80d2469e7b692e4",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S38sh13_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S38sh13_sel.png",
  "source_sha256": "596769fa7a117ac68b2be479a238b9c5c245bc5df15f5799e2b15116f14eee57",
  "file": "S38sh13_cine.png",
  "staged_sha256": "9d1d471df0b4918ba4e4e1d40eb360a6915ec6e4a8272ae4b6c76e4c8060ec37",
  "latency_ms": 15394
 },
 "S39sh2::signage": {
  "fp": "1dda73e67ecab9ed",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "groupbg::flood_search_water": {
  "input_fingerprint": "299a7dbb0b706ee5",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "flood_search_water",
    "tags": [
     "S39sh2",
     "S39sh6"
    ]
   },
   "context_sig": "a97b192acf54dd12"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At the muddy water surface over the submerged refugee settlement, among floating boards and debris in daylight.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n수몰된 인천 난민촌 수면과 떠다니는 컨테이너: 마을 전체가 물에 잠겨 바다처럼 변해버린 아침 풍경. (특징: 잔잔해진 광활한 수면; 뗏목처럼 부유하는 직육면체 컨테이너 박스들; 수면을 비추는 헬기의 서치라이트 빔; 물에 떠 있는 파편과 희생자들)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 수몰된 난민촌의 풍경. 페드로와 구도환 등도 둥둥 떠다니고\n- 고무보트를 타고 수몰된 난민촌을 가로지르는 민병대원들.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At the muddy water surface over the submerged refugee settlement, among floating boards and debris in daylight.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n수몰된 인천 난민촌 수면과 떠다니는 컨테이너: 마을 전체가 물에 잠겨 바다처럼 변해버린 아침 풍경. (특징: 잔잔해진 광활한 수면; 뗏목처럼 부유하는 직육면체 컨테이너 박스들; 수면을 비추는 헬기의 서치라이트 빔; 물에 떠 있는 파편과 희생자들)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 수몰된 난민촌의 풍경. 페드로와 구도환 등도 둥둥 떠다니고\n- 고무보트를 타고 수몰된 난민촌을 가로지르는 민병대원들.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_flood_search_water_8a2584.png",
  "asset_id": "5bd9a9bb-1d8b-43aa-95b1-ba9803c3d1da",
  "input_asset_ids": [
   "019a6799-a901-4c35-8fb2-1592f569dde6"
  ],
  "origin_tag": "S39sh2",
  "place_text": "At the muddy water surface over the submerged refugee settlement, among floating boards and debris in daylight.",
  "origin_inputs": {
   "place_text": "At the muddy water surface over the submerged refugee settlement, among floating boards and debris in daylight.",
   "time_of_day_en": "day",
   "conti_asset_id": "019a6799-a901-4c35-8fb2-1592f569dde6"
  }
 },
 "S39sh2::bgfirst_bg": {
  "input_fingerprint": "ac709eafdaaff59f",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 탁한 흙탕물 속에서 널빤지를 꽉 움켜쥔 채 버티고 있는 페드로와 구도환의 지친 상체.\n\nLOCATION (lock): At the muddy water surface over the submerged refugee settlement, among floating boards and debris in daylight.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Floating plank (Gripped tightly by both survivors) — Its upper face and near edge run obliquely across the lower middle; used as Links the two different gripping postures without occupying more than two-fifths of the frame; Floodwater (Opaque and muddy around the floating survivors); used as Surrounds the visible upper bodies and establishes their lack of secure footing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination and controlled contrast keep the hands and tired faces readable against the muddy water.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 탁한 흙탕물 속에서 널빤지를 꽉 움켜쥔 채 버티고 있는 페드로와 구도환의 지친 상체.\n\nLOCATION (lock): At the muddy water surface over the submerged refugee settlement, among floating boards and debris in daylight.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Floating plank (Gripped tightly by both survivors) — Its upper face and near edge run obliquely across the lower middle; used as Links the two different gripping postures without occupying more than two-fifths of the frame; Floodwater (Opaque and muddy around the floating survivors); used as Surrounds the visible upper bodies and establishes their lack of secure footing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination and controlled contrast keep the hands and tired faces readable against the muddy water.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S39sh2__bgfirst_bg.png",
  "asset_id": "22172d82-3681-4cc3-9212-7d3b5566ad92",
  "input_asset_ids": [
   "019a6799-a901-4c35-8fb2-1592f569dde6",
   "5bd9a9bb-1d8b-43aa-95b1-ba9803c3d1da"
  ]
 },
 "S39sh2": {
  "input_fingerprint": "add3bb9beac3a9bd",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 탁한 흙탕물 속에서 널빤지를 꽉 움켜쥔 채 버티고 있는 페드로와 구도환의 지친 상체.\n\nLOCATION (lock): At the muddy water surface over the submerged refugee settlement, among floating boards and debris in daylight. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Floating plank (Gripped tightly by both survivors) — Its upper face and near edge run obliquely across the lower middle; used as Links the two different gripping postures without occupying more than two-fifths of the frame; Floodwater (Opaque and muddy around the floating survivors); used as Surrounds the visible upper bodies and establishes their lack of secure footing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination and controlled contrast keep the hands and tired faces readable against the muddy water.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The refugee settlement is submerged in daylight, with containers and other wreckage afloat. 페드로: He is afloat in the flooded settlement, wet from immersion. 구도환: He is afloat in the flooded settlement, wet from immersion; his condition is not yet established.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 페드로와 구도환 right now, so 페드로와 구도환's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 페드로와 구도환: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼); 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 탁한 흙탕물 속에서 널빤지를 꽉 움켜쥔 채 버티고 있는 페드로와 구도환의 지친 상체.\n\nLOCATION (lock): At the muddy water surface over the submerged refugee settlement, among floating boards and debris in daylight. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Floating plank (Gripped tightly by both survivors) — Its upper face and near edge run obliquely across the lower middle; used as Links the two different gripping postures without occupying more than two-fifths of the frame; Floodwater (Opaque and muddy around the floating survivors); used as Surrounds the visible upper bodies and establishes their lack of secure footing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination and controlled contrast keep the hands and tired faces readable against the muddy water.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The refugee settlement is submerged in daylight, with containers and other wreckage afloat. 페드로: He is afloat in the flooded settlement, wet from immersion. 구도환: He is afloat in the flooded settlement, wet from immersion; his condition is not yet established.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 페드로와 구도환 right now, so 페드로와 구도환's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 페드로와 구도환: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼); 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 탁한 흙탕물 속에서 널빤지를 꽉 움켜쥔 채 버티고 있는 페드로와 구도환의 지친 상체.\n\nLOCATION (lock): At the muddy water surface over the submerged refugee settlement, among floating boards and debris in daylight. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Floating plank (Gripped tightly by both survivors) — Its upper face and near edge run obliquely across the lower middle; used as Links the two different gripping postures without occupying more than two-fifths of the frame; Floodwater (Opaque and muddy around the floating survivors); used as Surrounds the visible upper bodies and establishes their lack of secure footing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination and controlled contrast keep the hands and tired faces readable against the muddy water.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The refugee settlement is submerged in daylight, with containers and other wreckage afloat. 페드로: He is afloat in the flooded settlement, wet from immersion. 구도환: He is afloat in the flooded settlement, wet from immersion; his condition is not yet established.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 페드로와 구도환 right now, so 페드로와 구도환's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 페드로와 구도환: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 페드로 (18세 남성, 라틴계 혼혈, 앳된 얼굴, 짙은 머리칼); 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S39sh2__bgfirst_bg.png",
     "asset_id": "22172d82-3681-4cc3-9212-7d3b5566ad92",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S39sh2.png",
     "asset_id": "019a6799-a901-4c35-8fb2-1592f569dde6",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 페드로: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1278830>",
     "asset_id": "b09df655-64d4-4db4-a1b3-2f0bb5d29c95",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1311816>",
     "asset_id": "623f0421-67dc-4592-a009-148d8e957276",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_flood_search_water_8a2584.png",
     "asset_id": "5bd9a9bb-1d8b-43aa-95b1-ba9803c3d1da",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 페드로: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1278830>",
     "asset_id": "b09df655-64d4-4db4-a1b3-2f0bb5d29c95",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1311816>",
     "asset_id": "623f0421-67dc-4592-a009-148d8e957276",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "두 인물 모두 시선을 아래 널빤지와 수면 쪽으로 향한 채 지쳐 있는 모습이다.",
    "built_space": "레퍼런스와 동일한 위치로, 왼쪽에 침수된 컨테이너와 오른쪽에 물에 잠긴 건물이 배치되어 있다.",
    "entities": "페드로(왼쪽)와 구도환(오른쪽)의 외모가 레퍼런스와 일치하며, 흙탕물 속에서 널빤지에 기대어 있다.",
    "hard_violations": [
     "[gemini-pro] Physically impossible anatomy: 페드로의 왼팔이 널빤지 위에 올려져 있으나 손이 없고 나무 질감과 기괴하게 융합되어 절단된 것처럼 보임"
    ],
    "physics": "물의 부력과 널빤지에 기대어 상체를 지탱하고 있으나, 페드로의 왼팔 지지 형태가 물리적/해부학적으로 불가능하게 렌더링되었다."
   },
   {
    "label": "B",
    "direction": "두 인물 모두 시선을 정면 쪽으로 향하며 널빤지를 꽉 움켜쥐고 버티고 있다.",
    "built_space": "레퍼런스 사진과 정확히 일치하는 침수된 정착지(왼쪽의 컨테이너, 중앙의 부유하는 드럼통, 오른쪽의 잠긴 건물)가 사실적으로 구현되었다.",
    "entities": "페드로(왼쪽)와 구도환(오른쪽)의 얼굴, 머리, 연령대가 캐릭터 레퍼런스와 매우 정확히 일치하며 탁한 흙탕물 속에서 널빤지를 잡고 있다.",
    "hard_violations": [],
    "physics": "부력에 의해 상체가 물에 떠 있으며, 인물들이 손으로 널빤지의 가까운 가장자리를 꽉 움켜쥐고 체중을 지탱하는 자연스러운 상태를 보여준다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "지친 상태로 널빤지를 움켜쥔 두 인물의 모습, 정확한 캐릭터 외모 구현, 그리고 침수된 정착지의 배경이 프롬프트와 완벽하게 일치합니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "인물들의 지친 분위기는 잘 표현되었으나, 널빤지 위에 올려진 페드로의 왼팔에 손이 없고 나무와 뭉개지는 치명적인 신체 묘사 오류가 있습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 인물 모두 시선을 아래 널빤지와 수면 쪽으로 향한 채 지쳐 있는 모습이다.",
        "built_space": "레퍼런스와 동일한 위치로, 왼쪽에 침수된 컨테이너와 오른쪽에 물에 잠긴 건물이 배치되어 있다.",
        "entities": "페드로(왼쪽)와 구도환(오른쪽)의 외모가 레퍼런스와 일치하며, 흙탕물 속에서 널빤지에 기대어 있다.",
        "hard_violations": [
         "Physically impossible anatomy: 페드로의 왼팔이 널빤지 위에 올려져 있으나 손이 없고 나무 질감과 기괴하게 융합되어 절단된 것처럼 보임"
        ],
        "physics": "물의 부력과 널빤지에 기대어 상체를 지탱하고 있으나, 페드로의 왼팔 지지 형태가 물리적/해부학적으로 불가능하게 렌더링되었다."
       },
       {
        "label": "B",
        "direction": "두 인물 모두 시선을 정면 쪽으로 향하며 널빤지를 꽉 움켜쥐고 버티고 있다.",
        "built_space": "레퍼런스 사진과 정확히 일치하는 침수된 정착지(왼쪽의 컨테이너, 중앙의 부유하는 드럼통, 오른쪽의 잠긴 건물)가 사실적으로 구현되었다.",
        "entities": "페드로(왼쪽)와 구도환(오른쪽)의 얼굴, 머리, 연령대가 캐릭터 레퍼런스와 매우 정확히 일치하며 탁한 흙탕물 속에서 널빤지를 잡고 있다.",
        "hard_violations": [],
        "physics": "부력에 의해 상체가 물에 떠 있으며, 인물들이 손으로 널빤지의 가까운 가장자리를 꽉 움켜쥐고 체중을 지탱하는 자연스러운 상태를 보여준다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "지친 상태로 널빤지를 움켜쥔 두 인물의 모습, 정확한 캐릭터 외모 구현, 그리고 침수된 정착지의 배경이 프롬프트와 완벽하게 일치합니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "인물들의 지친 분위기는 잘 표현되었으나, 널빤지 위에 올려진 페드로의 왼팔에 손이 없고 나무와 뭉개지는 치명적인 신체 묘사 오류가 있습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "두 인물 모두 시선을 아래 널빤지와 수면 쪽으로 향한 채 지쳐 있는 모습이다.",
        "built_space": "레퍼런스와 동일한 위치로, 왼쪽에 침수된 컨테이너와 오른쪽에 물에 잠긴 건물이 배치되어 있다.",
        "entities": "페드로(왼쪽)와 구도환(오른쪽)의 외모가 레퍼런스와 일치하며, 흙탕물 속에서 널빤지에 기대어 있다.",
        "hard_violations": [
         "Physically impossible anatomy: 페드로의 왼팔이 널빤지 위에 올려져 있으나 손이 없고 나무 질감과 기괴하게 융합되어 절단된 것처럼 보임"
        ],
        "physics": "물의 부력과 널빤지에 기대어 상체를 지탱하고 있으나, 페드로의 왼팔 지지 형태가 물리적/해부학적으로 불가능하게 렌더링되었다."
       },
       {
        "label": "B",
        "direction": "두 인물 모두 시선을 정면 쪽으로 향하며 널빤지를 꽉 움켜쥐고 버티고 있다.",
        "built_space": "레퍼런스 사진과 정확히 일치하는 침수된 정착지(왼쪽의 컨테이너, 중앙의 부유하는 드럼통, 오른쪽의 잠긴 건물)가 사실적으로 구현되었다.",
        "entities": "페드로(왼쪽)와 구도환(오른쪽)의 얼굴, 머리, 연령대가 캐릭터 레퍼런스와 매우 정확히 일치하며 탁한 흙탕물 속에서 널빤지를 잡고 있다.",
        "hard_violations": [],
        "physics": "부력에 의해 상체가 물에 떠 있으며, 인물들이 손으로 널빤지의 가까운 가장자리를 꽉 움켜쥐고 체중을 지탱하는 자연스러운 상태를 보여준다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "흙탕물에서 같은 널빤지를 붙드는 상체 미디엄 숏은 충족하지만, 정면 위주의 표정은 탈진감이 약하고 두 사람의 의상이 인물 참조와 다르다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "서로 다른 자세로 대각선 널빤지를 움켜쥔 지친 상체와 젖은 얼굴을 더 충실히 구현하며, 구도환의 셔츠 차림이 참조의 정장과 다른 점은 남는다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 페드로는 화면 앞쪽의 약간 오른편을 보고, 오른쪽 구도환은 거의 카메라 쪽을 본다. 두 사람 모두 같은 널빤지에 손을 뻗어 가장자리를 붙들고 있다. 지정된 시선 표적은 없지만, 고개를 든 모습은 탈진보다는 주변을 살피는 순간에 가깝다.",
        "built_space": "왼쪽 뒤에 큰 침수 컨테이너 하나, 중앙 뒤에 기울어진 드럼통 하나, 오른쪽 끝에 침수 건물 하나와 전신주가 보인다. 그 사이에 작은 컨테이너들과 잔해가 흩어져 있어 장소 참조의 배치를 따른다. 두 사람은 구조물 위가 아니라 수면에 있으며, 공유하는 긴 널빤지 하나가 화면 하단 중앙을 왼쪽 아래에서 오른쪽 위로 가로지른다. 널빤지의 윗면과 가까운 모서리가 보이고 면적은 화면의 5분의 2보다 작다.",
        "entities": "사람은 두 명뿐이다. 페드로는 앳된 얼굴과 짙은 젖은 머리의 청년 남성으로, 구도환은 검은 머리의 한국인 중년 남성으로 표현되어 대체로 해당 인물 설정에 맞는다. 다만 페드로의 탁한 색 긴소매와 구도환의 흙빛 셔츠는 각각 참조의 남색 반소매와 남색 재킷·흰 셔츠와 다르다. 불투명한 흙탕물, 젖은 목재, 떠 있는 잔해와 낮빛이 보이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 사람의 몸통은 가슴 부근까지 물에 잠겨 부력을 받고 있다. 페드로의 보이는 한 손과 구도환의 두 손이 널빤지 가장자리를 감싸고, 팔은 수면과 판자에 기대어 있다. 판자 자체도 물에 떠 있으므로 몸과 물체의 지지가 설명된다. 보이지 않는 발의 지면 접촉을 가정할 필요가 없으며, 명백한 무지지 부유나 불가능한 관절은 없다."
       },
       {
        "label": "B",
        "direction": "페드로는 앞쪽 아래의 널빤지와 수면 쪽으로 시선을 떨구고, 구도환은 자신의 손과 널빤지를 내려다본다. 페드로는 한 손을 앞으로 길게 내밀어 가까운 모서리를 잡고, 구도환은 두 손으로 판자를 붙든다. 서로 다른 팔 자세와 아래로 처진 시선이 지쳐 버티는 행동에 맞는다.",
        "built_space": "왼쪽 배경에 큰 침수 컨테이너 하나, 중앙에 기울어진 드럼통 하나, 오른쪽에 지붕이 드러난 침수 건물 하나와 전신주 하나가 보인다. 먼 컨테이너들과 부유 잔해는 작게 유지되어 참조 장소의 규모와 재료를 따른다. 페드로는 왼쪽 가까이, 구도환은 오른쪽 조금 뒤의 물속에 있다. 두 사람을 연결하는 널빤지 하나의 윗면과 가까운 모서리가 하단 중앙을 비스듬히 지나며 화면의 5분의 2를 넘지 않는다.",
        "entities": "추가 인물 없이 청년 남성 페드로와 한국인 중년 남성 구도환만 보인다. 페드로의 앳된 얼굴, 짙은 머리와 남색 반소매는 참조에 가깝다. 구도환의 중년 얼굴과 검은 머리도 참조에 부합하지만, 파란 겉셔츠와 흰 티셔츠는 참조의 정장 재킷과 흰 칼라 셔츠를 대신하고 있다. 두 사람의 젖은 피부와 옷, 탁한 갈색 물, 실제 목재 질감과 낮의 조명이 분명하며 읽을 수 있는 문자는 없다.",
        "hard_violations": [],
        "physics": "두 사람은 가슴 아래가 물에 잠겨 있고, 떠 있는 판자에 손과 팔을 걸쳐 상체를 지탱한다. 페드로의 앞쪽 손은 판자의 가까운 모서리를 감싸며 다른 팔은 수면에 잠기고, 구도환의 두 손도 판자에 확실히 접촉한다. 수중 몸통의 부력과 판자의 부력이 함께 지지를 제공하므로 발판 없이 버티는 자세가 가능하다. 지지 없는 신체나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "흙탕물에서 같은 널빤지를 붙드는 상체 미디엄 숏은 충족하지만, 정면 위주의 표정은 탈진감이 약하고 두 사람의 의상이 인물 참조와 다르다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "서로 다른 자세로 대각선 널빤지를 움켜쥔 지친 상체와 젖은 얼굴을 더 충실히 구현하며, 구도환의 셔츠 차림이 참조의 정장과 다른 점은 남는다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽 페드로는 화면 앞쪽의 약간 오른편을 보고, 오른쪽 구도환은 거의 카메라 쪽을 본다. 두 사람 모두 같은 널빤지에 손을 뻗어 가장자리를 붙들고 있다. 지정된 시선 표적은 없지만, 고개를 든 모습은 탈진보다는 주변을 살피는 순간에 가깝다.",
        "built_space": "왼쪽 뒤에 큰 침수 컨테이너 하나, 중앙 뒤에 기울어진 드럼통 하나, 오른쪽 끝에 침수 건물 하나와 전신주가 보인다. 그 사이에 작은 컨테이너들과 잔해가 흩어져 있어 장소 참조의 배치를 따른다. 두 사람은 구조물 위가 아니라 수면에 있으며, 공유하는 긴 널빤지 하나가 화면 하단 중앙을 왼쪽 아래에서 오른쪽 위로 가로지른다. 널빤지의 윗면과 가까운 모서리가 보이고 면적은 화면의 5분의 2보다 작다.",
        "entities": "사람은 두 명뿐이다. 페드로는 앳된 얼굴과 짙은 젖은 머리의 청년 남성으로, 구도환은 검은 머리의 한국인 중년 남성으로 표현되어 대체로 해당 인물 설정에 맞는다. 다만 페드로의 탁한 색 긴소매와 구도환의 흙빛 셔츠는 각각 참조의 남색 반소매와 남색 재킷·흰 셔츠와 다르다. 불투명한 흙탕물, 젖은 목재, 떠 있는 잔해와 낮빛이 보이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 사람의 몸통은 가슴 부근까지 물에 잠겨 부력을 받고 있다. 페드로의 보이는 한 손과 구도환의 두 손이 널빤지 가장자리를 감싸고, 팔은 수면과 판자에 기대어 있다. 판자 자체도 물에 떠 있으므로 몸과 물체의 지지가 설명된다. 보이지 않는 발의 지면 접촉을 가정할 필요가 없으며, 명백한 무지지 부유나 불가능한 관절은 없다."
       },
       {
        "label": "A",
        "direction": "페드로는 앞쪽 아래의 널빤지와 수면 쪽으로 시선을 떨구고, 구도환은 자신의 손과 널빤지를 내려다본다. 페드로는 한 손을 앞으로 길게 내밀어 가까운 모서리를 잡고, 구도환은 두 손으로 판자를 붙든다. 서로 다른 팔 자세와 아래로 처진 시선이 지쳐 버티는 행동에 맞는다.",
        "built_space": "왼쪽 배경에 큰 침수 컨테이너 하나, 중앙에 기울어진 드럼통 하나, 오른쪽에 지붕이 드러난 침수 건물 하나와 전신주 하나가 보인다. 먼 컨테이너들과 부유 잔해는 작게 유지되어 참조 장소의 규모와 재료를 따른다. 페드로는 왼쪽 가까이, 구도환은 오른쪽 조금 뒤의 물속에 있다. 두 사람을 연결하는 널빤지 하나의 윗면과 가까운 모서리가 하단 중앙을 비스듬히 지나며 화면의 5분의 2를 넘지 않는다.",
        "entities": "추가 인물 없이 청년 남성 페드로와 한국인 중년 남성 구도환만 보인다. 페드로의 앳된 얼굴, 짙은 머리와 남색 반소매는 참조에 가깝다. 구도환의 중년 얼굴과 검은 머리도 참조에 부합하지만, 파란 겉셔츠와 흰 티셔츠는 참조의 정장 재킷과 흰 칼라 셔츠를 대신하고 있다. 두 사람의 젖은 피부와 옷, 탁한 갈색 물, 실제 목재 질감과 낮의 조명이 분명하며 읽을 수 있는 문자는 없다.",
        "hard_violations": [],
        "physics": "두 사람은 가슴 아래가 물에 잠겨 있고, 떠 있는 판자에 손과 팔을 걸쳐 상체를 지탱한다. 페드로의 앞쪽 손은 판자의 가까운 모서리를 감싸며 다른 팔은 수면에 잠기고, 구도환의 두 손도 판자에 확실히 접촉한다. 수중 몸통의 부력과 판자의 부력이 함께 지지를 제공하므로 발판 없이 버티는 자세가 가능하다. 지지 없는 신체나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.444,
    "B": 1.778
   },
   "adjusted": {
    "A": 1.194,
    "B": 1.778
   },
   "violations": {
    "A": [
     "[gemini-pro] Physically impossible anatomy: 페드로의 왼팔이 널빤지 위에 올려져 있으나 손이 없고 나무 질감과 기괴하게 융합되어 절단된 것처럼 보임"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1778,
   "A": 1194
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1778,
    "verdict_ko": "지친 상태로 널빤지를 움켜쥔 두 인물의 모습, 정확한 캐릭터 외모 구현, 그리고 침수된 정착지의 배경이 프롬프트와 완벽하게 일치합니다."
   },
   {
    "label": "A",
    "score": 1194,
    "verdict_ko": "인물들의 지친 분위기는 잘 표현되었으나, 널빤지 위에 올려진 페드로의 왼팔에 손이 없고 나무와 뭉개지는 치명적인 신체 묘사 오류가 있습니다.  ★위반: [gemini-pro] Physically impossible anatomy: 페드로의 왼팔이 널빤지 위에 올려져 있으나 손이 없고 나무 질감과 기괴하게 융합되어 절단된 것처럼 보임"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_flood_search_water_8a2584.png",
    "asset_id": "5bd9a9bb-1d8b-43aa-95b1-ba9803c3d1da",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 페드로: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1278830>",
    "asset_id": "b09df655-64d4-4db4-a1b3-2f0bb5d29c95",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1311816>",
    "asset_id": "623f0421-67dc-4592-a009-148d8e957276",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-8480-71f7-8262-ee9cac1de86a",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S39sh2__bgfirst_bg.png",
   "bg_asset_id": "22172d82-3681-4cc3-9212-7d3b5566ad92",
   "bg_record_key": "S39sh2::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "flood_search_water",
   "groupbg_asset_id": "5bd9a9bb-1d8b-43aa-95b1-ba9803c3d1da"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S39sh2::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:15:33.417999+00:00",
  "fingerprint": "f4c22f016bfb61b9241e70d936dc4070a4c73501b9cb152286d96bd9ad21a29a",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S39sh2_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S39sh2_sel.png",
  "source_sha256": "e5366c3de5c7cba329da8a75c2bd706d018c954e4d44b78aca90bb1aca7ee387",
  "file": "S39sh2_cine.png",
  "staged_sha256": "6315a246fb4e269354f2557f77de729aceb4e523650afaaa292b8bd60bb06989",
  "latency_ms": 10709
 },
 "S39sh6::signage": {
  "fp": "aebc244dc47b0747",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S39sh6": {
  "input_fingerprint": "27f4467f64be14f6",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 보트 끝에 비스듬히 서서, 긴 막대기 끝으로 수면 위의 시체 어깨를 꾹 누르며 힘을 가하고 있는 민병대원의 거친 mid-action 자세.\n\nLOCATION (lock): At the open bow of an inflatable patrol boat moving through the flooded refugee settlement. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Rubber boat (Carrying the militia member across the flooded settlement) — The end and adjacent outer side are visible from alongside; used as Provides a limited right-side support and scale reference for the braced posture; Long pole (Its tip presses into the floating body's shoulder) — Runs diagonally from the militia member's hands toward the lower-left contact point; used as Creates the visible line of force between the standing figure and the body; Floodwater (Surrounding the boat and floating body); used as Separates the corpse from the boat while keeping both within one continuous space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral daytime ambient light maintains clear physical contact and restrained tonal contrast without extending the helicopter searchlight into this shot.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): The corpse floats at the water's surface while a militia member presses a pole against its shoulder to inspect it. The source does not establish whether the body faces upward or downward, or how its head and limbs are arranged.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Rubber boats move through the submerged settlement among floating containers and wreckage.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 민병대원 right now, so 민병대원's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 민병대원: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 보트 끝에 비스듬히 서서, 긴 막대기 끝으로 수면 위의 시체 어깨를 꾹 누르며 힘을 가하고 있는 민병대원의 거친 mid-action 자세.\n\nLOCATION (lock): At the open bow of an inflatable patrol boat moving through the flooded refugee settlement. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Rubber boat (Carrying the militia member across the flooded settlement) — The end and adjacent outer side are visible from alongside; used as Provides a limited right-side support and scale reference for the braced posture; Long pole (Its tip presses into the floating body's shoulder) — Runs diagonally from the militia member's hands toward the lower-left contact point; used as Creates the visible line of force between the standing figure and the body; Floodwater (Surrounding the boat and floating body); used as Separates the corpse from the boat while keeping both within one continuous space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral daytime ambient light maintains clear physical contact and restrained tonal contrast without extending the helicopter searchlight into this shot.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): The corpse floats at the water's surface while a militia member presses a pole against its shoulder to inspect it. The source does not establish whether the body faces upward or downward, or how its head and limbs are arranged.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Rubber boats move through the submerged settlement among floating containers and wreckage.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 민병대원 right now, so 민병대원's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 민병대원: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 보트 끝에 비스듬히 서서, 긴 막대기 끝으로 수면 위의 시체 어깨를 꾹 누르며 힘을 가하고 있는 민병대원의 거친 mid-action 자세.\n\nLOCATION (lock): At the open bow of an inflatable patrol boat moving through the flooded refugee settlement. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Rubber boat (Carrying the militia member across the flooded settlement) — The end and adjacent outer side are visible from alongside; used as Provides a limited right-side support and scale reference for the braced posture; Long pole (Its tip presses into the floating body's shoulder) — Runs diagonally from the militia member's hands toward the lower-left contact point; used as Creates the visible line of force between the standing figure and the body; Floodwater (Surrounding the boat and floating body); used as Separates the corpse from the boat while keeping both within one continuous space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral daytime ambient light maintains clear physical contact and restrained tonal contrast without extending the helicopter searchlight into this shot.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): The corpse floats at the water's surface while a militia member presses a pole against its shoulder to inspect it. The source does not establish whether the body faces upward or downward, or how its head and limbs are arranged.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Rubber boats move through the submerged settlement among floating containers and wreckage.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 민병대원 right now, so 민병대원's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 민병대원: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "민병대원의 시선은 물에 뜬 시체를 향하며, 막대기의 끝은 시체의 가슴(흉부) 중앙을 찌르고 있습니다.",
    "built_space": "검은색 고무보트가 물 위에 있으며, 배경의 흩어진 컨테이너와 먼 해안선 등은 레퍼런스 이미지의 수역과 일치합니다. 민병대원은 보트 내부에 서 있습니다.",
    "entities": "모자와 전술조끼를 착용한 새로운 얼굴의 민병대원, 긴 막대기, 하늘을 보고 떠 있는 시체가 존재합니다. 레퍼런스의 인물은 포함되지 않았습니다.",
    "hard_violations": [],
    "physics": "민병대원은 보트 바닥을 두 발로 딛고 서서 두 손으로 막대기를 지지하고 있으며, 시체는 물의 부력으로 떠 있습니다."
   },
   {
    "label": "B",
    "direction": "민병대원의 시선이 시체를 향하고, 막대기의 끝은 엎드려 있는 시체의 어깨 뒷부분을 꾹 누르고 있습니다.",
    "built_space": "고무보트가 떠 있으나, 배경에 레퍼런스에는 없는 빽빽한 수상 판자촌 구조물들이 넓게 펼쳐져 있어 장소 설정이 어긋납니다.",
    "entities": "긴 막대기, 엎드린 시체, 전술복을 입은 민병대원이 있습니다. 민병대원의 얼굴이 레퍼런스 이미지 우측의 나이 든 남성과 정확히 일치하여 지시를 위반했습니다.",
    "hard_violations": [],
    "physics": "민병대원은 보트의 튜브 부분에 한쪽 발을 올리고 두 손으로 막대기를 잡고 있으며, 시체는 물에 의해 지지되어 떠 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "레퍼런스의 장소와 '이전 인물 배제' 지시를 충실히 따랐으나, 막대기가 시체의 어깨가 아닌 가슴을 누르고 있는 점이 아쉽습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "막대기가 어깨를 누르는 묘사는 정확하나, 레퍼런스의 인물 얼굴을 그대로 복사하고 배경에 없는 밀집 판자촌을 생성하여 주요 지시사항을 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "민병대원의 시선은 물에 뜬 시체를 향하며, 막대기의 끝은 시체의 가슴(흉부) 중앙을 찌르고 있습니다.",
        "built_space": "검은색 고무보트가 물 위에 있으며, 배경의 흩어진 컨테이너와 먼 해안선 등은 레퍼런스 이미지의 수역과 일치합니다. 민병대원은 보트 내부에 서 있습니다.",
        "entities": "모자와 전술조끼를 착용한 새로운 얼굴의 민병대원, 긴 막대기, 하늘을 보고 떠 있는 시체가 존재합니다. 레퍼런스의 인물은 포함되지 않았습니다.",
        "hard_violations": [],
        "physics": "민병대원은 보트 바닥을 두 발로 딛고 서서 두 손으로 막대기를 지지하고 있으며, 시체는 물의 부력으로 떠 있습니다."
       },
       {
        "label": "B",
        "direction": "민병대원의 시선이 시체를 향하고, 막대기의 끝은 엎드려 있는 시체의 어깨 뒷부분을 꾹 누르고 있습니다.",
        "built_space": "고무보트가 떠 있으나, 배경에 레퍼런스에는 없는 빽빽한 수상 판자촌 구조물들이 넓게 펼쳐져 있어 장소 설정이 어긋납니다.",
        "entities": "긴 막대기, 엎드린 시체, 전술복을 입은 민병대원이 있습니다. 민병대원의 얼굴이 레퍼런스 이미지 우측의 나이 든 남성과 정확히 일치하여 지시를 위반했습니다.",
        "hard_violations": [],
        "physics": "민병대원은 보트의 튜브 부분에 한쪽 발을 올리고 두 손으로 막대기를 잡고 있으며, 시체는 물에 의해 지지되어 떠 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "레퍼런스의 장소와 '이전 인물 배제' 지시를 충실히 따랐으나, 막대기가 시체의 어깨가 아닌 가슴을 누르고 있는 점이 아쉽습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "막대기가 어깨를 누르는 묘사는 정확하나, 레퍼런스의 인물 얼굴을 그대로 복사하고 배경에 없는 밀집 판자촌을 생성하여 주요 지시사항을 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "민병대원의 시선은 물에 뜬 시체를 향하며, 막대기의 끝은 시체의 가슴(흉부) 중앙을 찌르고 있습니다.",
        "built_space": "검은색 고무보트가 물 위에 있으며, 배경의 흩어진 컨테이너와 먼 해안선 등은 레퍼런스 이미지의 수역과 일치합니다. 민병대원은 보트 내부에 서 있습니다.",
        "entities": "모자와 전술조끼를 착용한 새로운 얼굴의 민병대원, 긴 막대기, 하늘을 보고 떠 있는 시체가 존재합니다. 레퍼런스의 인물은 포함되지 않았습니다.",
        "hard_violations": [],
        "physics": "민병대원은 보트 바닥을 두 발로 딛고 서서 두 손으로 막대기를 지지하고 있으며, 시체는 물의 부력으로 떠 있습니다."
       },
       {
        "label": "B",
        "direction": "민병대원의 시선이 시체를 향하고, 막대기의 끝은 엎드려 있는 시체의 어깨 뒷부분을 꾹 누르고 있습니다.",
        "built_space": "고무보트가 떠 있으나, 배경에 레퍼런스에는 없는 빽빽한 수상 판자촌 구조물들이 넓게 펼쳐져 있어 장소 설정이 어긋납니다.",
        "entities": "긴 막대기, 엎드린 시체, 전술복을 입은 민병대원이 있습니다. 민병대원의 얼굴이 레퍼런스 이미지 우측의 나이 든 남성과 정확히 일치하여 지시를 위반했습니다.",
        "hard_violations": [],
        "physics": "민병대원은 보트의 튜브 부분에 한쪽 발을 올리고 두 손으로 막대기를 잡고 있으며, 시체는 물에 의해 지지되어 떠 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "오른쪽 보트에서 몸을 비스듬히 버티고 왼쪽 아래 시체의 어깨를 누르는 핵심 구도와 동작이 정확하지만, 밀집한 판잣집 배경은 참조의 개방된 침수 공간과 다릅니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "참조 장소의 컨테이너와 잔해는 잘 이어지지만, 보트와 접촉점의 좌우 배치가 지시와 반대이고 막대 끝도 어깨보다 가슴 옆에 닿습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "민병대원은 왼쪽 아래의 시체와 막대 접촉점을 내려다봅니다. 두 손에서 뻗은 막대는 왼쪽 아래로 내려가 엎드린 시체의 어깨 뒤쪽에 닿아, 지정된 힘의 방향과 목표를 충족합니다. 보트가 움직이는 방향은 뚜렷하게 드러나지 않습니다.",
        "built_space": "전경에는 고무보트 한 척의 둥근 선수와 인접한 바깥 측면이 오른쪽에 보입니다. 민병대원은 선수 안쪽에서 다리를 벌리고 있으며, 튜브가 하체 일부를 가립니다. 배경에는 작은 고무보트 한 척과 여러 판잣집·천막·컨테이너가 있습니다. 물과 부유 잔해는 참조와 통하지만, 참조의 넓게 열린 수면보다 주거 구조물이 훨씬 밀집해 장소 연속성이 약합니다.",
        "entities": "전경에서 명확히 식별되는 사람은 성인 동아시아계 남성 민병대원 한 명과 성인 남성 시체 한 구입니다. 민병대원은 녹색 군용 작업복과 조끼를 착용하며, 시체는 어두운 상의와 바지를 입었습니다. 별도의 인물 참조나 구체적인 연령·복장 지정은 없어 세부 신원 일치 여부는 확정할 수 없습니다. 긴 목재 막대, 고무보트, 탁한 홍수 물과 부유 잔해가 보이며 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "민병대원은 무릎을 굽히고 상체를 앞으로 기울인 채 양손으로 막대를 잡아 누릅니다. 발은 튜브에 가려져 있지만 하체는 보트 내부로 이어져 바닥에 체중을 싣는 자세로 읽힙니다. 시체의 얼굴과 몸통은 수면에 잠기거나 걸쳐 있고 팔다리는 물에 지지되어 있으며, 능동적으로 들어 올린 자세는 아닙니다. 막대 끝의 어깨 접촉과 주변 물결이 압박 동작을 뒷받침합니다."
       },
       {
        "label": "B",
        "direction": "민병대원은 오른쪽 아래의 시체 쪽을 바라봅니다. 막대는 두 손에서 오른쪽 아래로 뻗어 시체의 가슴 옆 또는 겨드랑이 아래에 닿습니다. 시체를 누르는 관계는 보이지만, 지정된 왼쪽 아래 접촉점 및 어깨 목표와는 다릅니다. 보트의 이동 방향은 분명하지 않습니다.",
        "built_space": "고무보트 한 척의 선수와 바깥 측면이 화면 왼쪽을 차지하고, 민병대원은 선수 안쪽에 서 있습니다. 시체는 보트 오른쪽 수면에 떨어져 있습니다. 따라서 보트가 오른쪽의 제한된 지지물이어야 한다는 배치는 뒤집혔습니다. 배경의 왼쪽 대형 컨테이너, 중앙의 기울어진 녹슨 통, 넓은 수면과 목재 잔해는 참조 장소와 잘 이어집니다.",
        "entities": "성인 동아시아계 남성 민병대원 한 명과 성인 남성 시체 한 구가 보입니다. 민병대원은 모자, 얼룩무늬 군복과 전술 조끼를 착용하고, 시체는 어두운 상의와 바지를 입었습니다. 구체적인 인물 신원이나 복장을 고정할 별도 참조는 없습니다. 긴 목재 막대와 고무보트, 홍수 물, 컨테이너 및 부유 잔해가 확인되며 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "민병대원은 보트 안에서 다리를 벌리고 몸을 기울여 양손으로 막대를 밀고 있습니다. 발은 튜브 뒤에 가려져 있으나 보트 바닥에 서 있는 것으로 자연스럽게 연결됩니다. 시체는 등을 아래로 하여 물에 떠 있고 머리와 팔도 수면에 놓여 있어 부력으로 지지됩니다. 막대 역시 양손과 몸통 접촉점으로 지지되지만, 압박 위치가 요청된 어깨에서 벗어납니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "오른쪽 보트에서 몸을 비스듬히 버티고 왼쪽 아래 시체의 어깨를 누르는 핵심 구도와 동작이 정확하지만, 밀집한 판잣집 배경은 참조의 개방된 침수 공간과 다릅니다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "참조 장소의 컨테이너와 잔해는 잘 이어지지만, 보트와 접촉점의 좌우 배치가 지시와 반대이고 막대 끝도 어깨보다 가슴 옆에 닿습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "민병대원은 왼쪽 아래의 시체와 막대 접촉점을 내려다봅니다. 두 손에서 뻗은 막대는 왼쪽 아래로 내려가 엎드린 시체의 어깨 뒤쪽에 닿아, 지정된 힘의 방향과 목표를 충족합니다. 보트가 움직이는 방향은 뚜렷하게 드러나지 않습니다.",
        "built_space": "전경에는 고무보트 한 척의 둥근 선수와 인접한 바깥 측면이 오른쪽에 보입니다. 민병대원은 선수 안쪽에서 다리를 벌리고 있으며, 튜브가 하체 일부를 가립니다. 배경에는 작은 고무보트 한 척과 여러 판잣집·천막·컨테이너가 있습니다. 물과 부유 잔해는 참조와 통하지만, 참조의 넓게 열린 수면보다 주거 구조물이 훨씬 밀집해 장소 연속성이 약합니다.",
        "entities": "전경에서 명확히 식별되는 사람은 성인 동아시아계 남성 민병대원 한 명과 성인 남성 시체 한 구입니다. 민병대원은 녹색 군용 작업복과 조끼를 착용하며, 시체는 어두운 상의와 바지를 입었습니다. 별도의 인물 참조나 구체적인 연령·복장 지정은 없어 세부 신원 일치 여부는 확정할 수 없습니다. 긴 목재 막대, 고무보트, 탁한 홍수 물과 부유 잔해가 보이며 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "민병대원은 무릎을 굽히고 상체를 앞으로 기울인 채 양손으로 막대를 잡아 누릅니다. 발은 튜브에 가려져 있지만 하체는 보트 내부로 이어져 바닥에 체중을 싣는 자세로 읽힙니다. 시체의 얼굴과 몸통은 수면에 잠기거나 걸쳐 있고 팔다리는 물에 지지되어 있으며, 능동적으로 들어 올린 자세는 아닙니다. 막대 끝의 어깨 접촉과 주변 물결이 압박 동작을 뒷받침합니다."
       },
       {
        "label": "A",
        "direction": "민병대원은 오른쪽 아래의 시체 쪽을 바라봅니다. 막대는 두 손에서 오른쪽 아래로 뻗어 시체의 가슴 옆 또는 겨드랑이 아래에 닿습니다. 시체를 누르는 관계는 보이지만, 지정된 왼쪽 아래 접촉점 및 어깨 목표와는 다릅니다. 보트의 이동 방향은 분명하지 않습니다.",
        "built_space": "고무보트 한 척의 선수와 바깥 측면이 화면 왼쪽을 차지하고, 민병대원은 선수 안쪽에 서 있습니다. 시체는 보트 오른쪽 수면에 떨어져 있습니다. 따라서 보트가 오른쪽의 제한된 지지물이어야 한다는 배치는 뒤집혔습니다. 배경의 왼쪽 대형 컨테이너, 중앙의 기울어진 녹슨 통, 넓은 수면과 목재 잔해는 참조 장소와 잘 이어집니다.",
        "entities": "성인 동아시아계 남성 민병대원 한 명과 성인 남성 시체 한 구가 보입니다. 민병대원은 모자, 얼룩무늬 군복과 전술 조끼를 착용하고, 시체는 어두운 상의와 바지를 입었습니다. 구체적인 인물 신원이나 복장을 고정할 별도 참조는 없습니다. 긴 목재 막대와 고무보트, 홍수 물, 컨테이너 및 부유 잔해가 확인되며 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "민병대원은 보트 안에서 다리를 벌리고 몸을 기울여 양손으로 막대를 밀고 있습니다. 발은 튜브 뒤에 가려져 있으나 보트 바닥에 서 있는 것으로 자연스럽게 연결됩니다. 시체는 등을 아래로 하여 물에 떠 있고 머리와 팔도 수면에 놓여 있어 부력으로 지지됩니다. 막대 역시 양손과 몸통 접촉점으로 지지되지만, 압박 위치가 요청된 어깨에서 벗어납니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.75,
    "B": 1.571
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.571
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1750,
   "B": 1571
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "레퍼런스의 장소와 '이전 인물 배제' 지시를 충실히 따랐으나, 막대기가 시체의 어깨가 아닌 가슴을 누르고 있는 점이 아쉽습니다."
   },
   {
    "label": "B",
    "score": 1571,
    "verdict_ko": "막대기가 어깨를 누르는 묘사는 정확하나, 레퍼런스의 인물 얼굴을 그대로 복사하고 배경에 없는 밀집 판자촌을 생성하여 주요 지시사항을 위반했습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S39sh2_sel.png",
    "asset_id": "79a32a77-08d3-473a-b8c0-e701d92267e0",
    "role": "prev_still"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-895d-780e-b157-6a94605a4fd5",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S39sh2"
  }
 },
 "S39sh6::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:16:50.010769+00:00",
  "fingerprint": "39a8d992d5cf742805d50cbeb363d51449a397e738cf727c17a29123d4769541",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S39sh6_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S39sh6_sel.png",
  "source_sha256": "ec162e167de698dae0b9022ee50418c3a9cf6062b1e2508cd67a6dede5e8e747",
  "file": "S39sh6_cine.png",
  "staged_sha256": "0de89f98f5975e53069eb88400b58ad9700ca478294e93be62bef581117cf340",
  "latency_ms": 11762
 },
 "S40sh1::signage": {
  "fp": "2ce0de0d9ed5006c",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S40sh1::bgfirst_bg": {
  "input_fingerprint": "991a18db2d387744",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 텅 빈 잿빛 아스팔트 도로 위, 땀과 눈물로 엉망이 된 얼굴로 멍하니 서 있는 앰버의 지친 상체.\n\nLOCATION (lock): On an empty asphalt road at the outer edge of the refugee settlement, along the escape route toward the city.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Asphalt road (Empty in the visible portion around 앰버) — The road surface recedes obliquely behind her; used as An open right-side wedge isolates her stalled movement without introducing other street activity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination reveals the tears and perspiration without adding a dramatic source or altering the street's restrained gray tonality.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 텅 빈 잿빛 아스팔트 도로 위, 땀과 눈물로 엉망이 된 얼굴로 멍하니 서 있는 앰버의 지친 상체.\n\nLOCATION (lock): On an empty asphalt road at the outer edge of the refugee settlement, along the escape route toward the city.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Asphalt road (Empty in the visible portion around 앰버) — The road surface recedes obliquely behind her; used as An open right-side wedge isolates her stalled movement without introducing other street activity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination reveals the tears and perspiration without adding a dramatic source or altering the street's restrained gray tonality.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S40sh1__bgfirst_bg.png",
  "asset_id": "3f8ff6ab-9364-4e5b-a480-4d9ff441d727",
  "input_asset_ids": [
   "e04d942a-da09-45cf-bfc4-228283629bc9",
   "2c19fbb8-19d6-4259-a297-df4959b88a58"
  ]
 },
 "S40sh1": {
  "input_fingerprint": "d59927c500e38652",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 텅 빈 잿빛 아스팔트 도로 위, 땀과 눈물로 엉망이 된 얼굴로 멍하니 서 있는 앰버의 지친 상체.\n\nLOCATION (lock): On an empty asphalt road at the outer edge of the refugee settlement, along the escape route toward the city. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Asphalt road (Empty in the visible portion around 앰버) — The road surface recedes obliquely behind her; used as An open right-side wedge isolates her stalled movement without introducing other street activity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination reveals the tears and perspiration without adding a dramatic source or altering the street's restrained gray tonality.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): It is daylight on the outskirts of the flooded refugee settlement. 앰버: She has stopped from exhaustion, with tears covering her face and her attention repeatedly turning backward.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 텅 빈 잿빛 아스팔트 도로 위, 땀과 눈물로 엉망이 된 얼굴로 멍하니 서 있는 앰버의 지친 상체.\n\nLOCATION (lock): On an empty asphalt road at the outer edge of the refugee settlement, along the escape route toward the city. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Asphalt road (Empty in the visible portion around 앰버) — The road surface recedes obliquely behind her; used as An open right-side wedge isolates her stalled movement without introducing other street activity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination reveals the tears and perspiration without adding a dramatic source or altering the street's restrained gray tonality.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): It is daylight on the outskirts of the flooded refugee settlement. 앰버: She has stopped from exhaustion, with tears covering her face and her attention repeatedly turning backward.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 텅 빈 잿빛 아스팔트 도로 위, 땀과 눈물로 엉망이 된 얼굴로 멍하니 서 있는 앰버의 지친 상체.\n\nLOCATION (lock): On an empty asphalt road at the outer edge of the refugee settlement, along the escape route toward the city. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Asphalt road (Empty in the visible portion around 앰버) — The road surface recedes obliquely behind her; used as An open right-side wedge isolates her stalled movement without introducing other street activity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination reveals the tears and perspiration without adding a dramatic source or altering the street's restrained gray tonality.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): It is daylight on the outskirts of the flooded refugee settlement. 앰버: She has stopped from exhaustion, with tears covering her face and her attention repeatedly turning backward.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S40sh1__bgfirst_bg.png",
     "asset_id": "3f8ff6ab-9364-4e5b-a480-4d9ff441d727",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S40sh1.png",
     "asset_id": "e04d942a-da09-45cf-bfc4-228283629bc9",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L194B02.png",
     "asset_id": "2c19fbb8-19d6-4259-a297-df4959b88a58",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "시선과 주의가 프레임 왼쪽 뒤편을 향하고 있음.",
    "built_space": "아스팔트 도로 위에 서 있으며, 왼쪽에 레퍼런스 이미지와 동일한 낡은 상점 건물이 있고 오른쪽으로 도로가 사선으로 뻗어 있음.",
    "entities": "앰버의 이목구비, 금발이 레퍼런스와 일치하며, 땀과 눈물, 얼룩으로 얼굴이 엉망이 된 상태임.",
    "hard_violations": [],
    "physics": "땅에 발을 딛고 상체를 살짝 숙인 채 지친 자세를 스스로 지탱하고 있음."
   },
   {
    "label": "B",
    "direction": "시선은 프레임 왼쪽 밖을 멍하니 바라보고 있음.",
    "built_space": "아스팔트 도로 위에 서 있으며 도로가 오른쪽으로 뻗어 있으나, 배경은 텐트들만 보이고 지정된 상점 건물이 없음.",
    "entities": "앰버의 외모는 레퍼런스와 일치하며 얼굴에 땀과 눈물이 있으나 비교적 깨끗한 상태임.",
    "hard_violations": [],
    "physics": "땅에 발을 딛고 안정적으로 서 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 로케이션의 상점 건물을 배경에 정확히 배치하고, 사선으로 이어지는 도로 구도와 땀과 눈물로 엉망이 된 앰버의 지친 상태를 훌륭하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "요구된 화면 구도와 인물의 기본 외형은 일치하나, 로케이션 레퍼런스의 핵심 건축물이 누락되었고 지친 묘사가 다소 밋밋합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선과 주의가 프레임 왼쪽 뒤편을 향하고 있음.",
        "built_space": "아스팔트 도로 위에 서 있으며, 왼쪽에 레퍼런스 이미지와 동일한 낡은 상점 건물이 있고 오른쪽으로 도로가 사선으로 뻗어 있음.",
        "entities": "앰버의 이목구비, 금발이 레퍼런스와 일치하며, 땀과 눈물, 얼룩으로 얼굴이 엉망이 된 상태임.",
        "hard_violations": [],
        "physics": "땅에 발을 딛고 상체를 살짝 숙인 채 지친 자세를 스스로 지탱하고 있음."
       },
       {
        "label": "B",
        "direction": "시선은 프레임 왼쪽 밖을 멍하니 바라보고 있음.",
        "built_space": "아스팔트 도로 위에 서 있으며 도로가 오른쪽으로 뻗어 있으나, 배경은 텐트들만 보이고 지정된 상점 건물이 없음.",
        "entities": "앰버의 외모는 레퍼런스와 일치하며 얼굴에 땀과 눈물이 있으나 비교적 깨끗한 상태임.",
        "hard_violations": [],
        "physics": "땅에 발을 딛고 안정적으로 서 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 로케이션의 상점 건물을 배경에 정확히 배치하고, 사선으로 이어지는 도로 구도와 땀과 눈물로 엉망이 된 앰버의 지친 상태를 훌륭하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "요구된 화면 구도와 인물의 기본 외형은 일치하나, 로케이션 레퍼런스의 핵심 건축물이 누락되었고 지친 묘사가 다소 밋밋합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선과 주의가 프레임 왼쪽 뒤편을 향하고 있음.",
        "built_space": "아스팔트 도로 위에 서 있으며, 왼쪽에 레퍼런스 이미지와 동일한 낡은 상점 건물이 있고 오른쪽으로 도로가 사선으로 뻗어 있음.",
        "entities": "앰버의 이목구비, 금발이 레퍼런스와 일치하며, 땀과 눈물, 얼룩으로 얼굴이 엉망이 된 상태임.",
        "hard_violations": [],
        "physics": "땅에 발을 딛고 상체를 살짝 숙인 채 지친 자세를 스스로 지탱하고 있음."
       },
       {
        "label": "B",
        "direction": "시선은 프레임 왼쪽 밖을 멍하니 바라보고 있음.",
        "built_space": "아스팔트 도로 위에 서 있으며 도로가 오른쪽으로 뻗어 있으나, 배경은 텐트들만 보이고 지정된 상점 건물이 없음.",
        "entities": "앰버의 외모는 레퍼런스와 일치하며 얼굴에 땀과 눈물이 있으나 비교적 깨끗한 상태임.",
        "hard_violations": [],
        "physics": "땅에 발을 딛고 안정적으로 서 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "상체 미디엄 숏과 오른쪽의 빈 도로, 눈물은 충실하지만, 멈춰 선 탈진의 몸짓과 뒤돌아보는 동작이 약하고 지정 장소의 특징적인 상점이 확인되지 않는다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "오른쪽으로 비워 둔 도로와 상체 중심 구도 안에서 뒤를 살피며 지쳐 멈춘 순간을 구현하고, 참조 장소의 낡은 상점과 정착촌도 유지한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "몸통은 대체로 카메라 쪽이고 눈은 화면 왼쪽 밖을 향한다. 렌즈를 직접 응시하지는 않지만, 고개 회전이 작아 뒤쪽을 돌아보는 동작보다는 옆을 멍하게 보는 모습이다. 시선의 구체적인 대상은 화면에 없다.",
        "built_space": "아이는 화면 왼쪽 아스팔트 위에 있고, 오른쪽에는 차나 사람이 없는 도로가 사선으로 멀어진다. 양옆 연석과 보도, 왼쪽의 여러 천막, 원경 전신주가 보인다. 참조의 특징적인 단층 콘크리트 상점과 유리 출입구는 보이지 않아 정확한 장소 일치 근거가 약하다. 고정 시설의 명백한 중복이나 불가능한 반사는 없다.",
        "entities": "보이는 인물은 어린 여자아이 한 명이다. 약 10세의 체격, 금발, 둥근 얼굴, 큰 눈과 남색 반소매 셔츠는 앰버 참조에 가깝다. 혼혈 배경 자체는 외모만으로 확정할 수 없다. 이마에는 땀 광택이 있고 양 볼에는 눈물 자국이 보인다. 주변 도로는 비어 있고, 읽을 수 있는 글자나 추가 소품은 없다.",
        "hard_violations": [],
        "physics": "허리 부근에서 잘려 발의 접지는 보이지 않지만, 몸통은 자연스럽게 수직으로 이어져 서 있는 자세로 읽힌다. 공중에 뜬 몸이나 물건은 없다. 팔은 아래로 내려가 있고 어깨의 처짐은 약해, 탈진으로 동작이 중단된 순간보다는 비교적 반듯하게 서 있는 모습이다."
       },
       {
        "label": "B",
        "direction": "몸통의 방향과 달리 고개와 두 눈이 화면 왼쪽 바깥으로 뚜렷하게 돌아가 있다. 뒤쪽 상황을 다시 확인하는 시선으로 읽히며, 특정 대상이나 다른 사람은 보이지 않는다. 양손은 가슴 아래에서 느슨하게 굽혀져 있고 무언가를 겨누거나 가리키지 않는다.",
        "built_space": "아이는 도로의 왼쪽 전경에 있고, 넓게 빈 오른쪽 아스팔트가 사선으로 원경까지 이어진다. 왼쪽 상점에는 화면 가장자리에서 잘린 것까지 창 구획 세 개와 중앙 유리 양문 출입구 한 곳, 출입구 옆 작은 단말기가 보인다. 낡은 콘크리트 외벽, 연석과 배수구, 뒤편 철망과 천막, 전신주와 산이 참조 장소의 주요 구성을 유지한다. 불가능한 반사나 명백히 중복된 시설은 없다.",
        "entities": "앰버로 읽히는 어린 여자아이 한 명만 등장한다. 금발과 어린 체격, 둥근 얼굴 및 눈의 형태는 참조와 대체로 맞는다. 머리카락은 젖어 뭉쳤고 얼굴에는 땀, 눈물과 먼지가 섞여 있다. 반소매 셔츠는 참조의 남색보다 검고 더러워 보이지만 같은 계열의 의상이다. 빈 회색 도로와 정착촌이 보이며 읽을 수 있는 글자는 없다. 낮빛은 명확하나 참조보다 직사광과 그림자가 조금 강하다.",
        "hard_violations": [],
        "physics": "하체와 발은 프레임 밖이지만, 상체가 약간 앞으로 기울고 팔꿈치가 굽혀진 자세는 움직이다 지쳐 멈춘 상태로 성립한다. 손은 아무 물건도 들지 않으며 팔과 손목의 연결도 자연스럽다. 공중에 떠 있는 신체나 지지 없는 물건은 없고, 잘린 하체 때문에 부유한다고 볼 근거도 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "상체 미디엄 숏과 오른쪽의 빈 도로, 눈물은 충실하지만, 멈춰 선 탈진의 몸짓과 뒤돌아보는 동작이 약하고 지정 장소의 특징적인 상점이 확인되지 않는다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "오른쪽으로 비워 둔 도로와 상체 중심 구도 안에서 뒤를 살피며 지쳐 멈춘 순간을 구현하고, 참조 장소의 낡은 상점과 정착촌도 유지한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "몸통은 대체로 카메라 쪽이고 눈은 화면 왼쪽 밖을 향한다. 렌즈를 직접 응시하지는 않지만, 고개 회전이 작아 뒤쪽을 돌아보는 동작보다는 옆을 멍하게 보는 모습이다. 시선의 구체적인 대상은 화면에 없다.",
        "built_space": "아이는 화면 왼쪽 아스팔트 위에 있고, 오른쪽에는 차나 사람이 없는 도로가 사선으로 멀어진다. 양옆 연석과 보도, 왼쪽의 여러 천막, 원경 전신주가 보인다. 참조의 특징적인 단층 콘크리트 상점과 유리 출입구는 보이지 않아 정확한 장소 일치 근거가 약하다. 고정 시설의 명백한 중복이나 불가능한 반사는 없다.",
        "entities": "보이는 인물은 어린 여자아이 한 명이다. 약 10세의 체격, 금발, 둥근 얼굴, 큰 눈과 남색 반소매 셔츠는 앰버 참조에 가깝다. 혼혈 배경 자체는 외모만으로 확정할 수 없다. 이마에는 땀 광택이 있고 양 볼에는 눈물 자국이 보인다. 주변 도로는 비어 있고, 읽을 수 있는 글자나 추가 소품은 없다.",
        "hard_violations": [],
        "physics": "허리 부근에서 잘려 발의 접지는 보이지 않지만, 몸통은 자연스럽게 수직으로 이어져 서 있는 자세로 읽힌다. 공중에 뜬 몸이나 물건은 없다. 팔은 아래로 내려가 있고 어깨의 처짐은 약해, 탈진으로 동작이 중단된 순간보다는 비교적 반듯하게 서 있는 모습이다."
       },
       {
        "label": "A",
        "direction": "몸통의 방향과 달리 고개와 두 눈이 화면 왼쪽 바깥으로 뚜렷하게 돌아가 있다. 뒤쪽 상황을 다시 확인하는 시선으로 읽히며, 특정 대상이나 다른 사람은 보이지 않는다. 양손은 가슴 아래에서 느슨하게 굽혀져 있고 무언가를 겨누거나 가리키지 않는다.",
        "built_space": "아이는 도로의 왼쪽 전경에 있고, 넓게 빈 오른쪽 아스팔트가 사선으로 원경까지 이어진다. 왼쪽 상점에는 화면 가장자리에서 잘린 것까지 창 구획 세 개와 중앙 유리 양문 출입구 한 곳, 출입구 옆 작은 단말기가 보인다. 낡은 콘크리트 외벽, 연석과 배수구, 뒤편 철망과 천막, 전신주와 산이 참조 장소의 주요 구성을 유지한다. 불가능한 반사나 명백히 중복된 시설은 없다.",
        "entities": "앰버로 읽히는 어린 여자아이 한 명만 등장한다. 금발과 어린 체격, 둥근 얼굴 및 눈의 형태는 참조와 대체로 맞는다. 머리카락은 젖어 뭉쳤고 얼굴에는 땀, 눈물과 먼지가 섞여 있다. 반소매 셔츠는 참조의 남색보다 검고 더러워 보이지만 같은 계열의 의상이다. 빈 회색 도로와 정착촌이 보이며 읽을 수 있는 글자는 없다. 낮빛은 명확하나 참조보다 직사광과 그림자가 조금 강하다.",
        "hard_violations": [],
        "physics": "하체와 발은 프레임 밖이지만, 상체가 약간 앞으로 기울고 팔꿈치가 굽혀진 자세는 움직이다 지쳐 멈춘 상태로 성립한다. 손은 아무 물건도 들지 않으며 팔과 손목의 연결도 자연스럽다. 공중에 떠 있는 신체나 지지 없는 물건은 없고, 잘린 하체 때문에 부유한다고 볼 근거도 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.381
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.381
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1381
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지정된 로케이션의 상점 건물을 배경에 정확히 배치하고, 사선으로 이어지는 도로 구도와 땀과 눈물로 엉망이 된 앰버의 지친 상태를 훌륭하게 구현했습니다."
   },
   {
    "label": "B",
    "score": 1381,
    "verdict_ko": "요구된 화면 구도와 인물의 기본 외형은 일치하나, 로케이션 레퍼런스의 핵심 건축물이 누락되었고 지친 묘사가 다소 밋밋합니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L194B02.png",
    "asset_id": "2c19fbb8-19d6-4259-a297-df4959b88a58",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-8b07-7703-9c20-a957c4617739",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S40sh1__bgfirst_bg.png",
   "bg_asset_id": "3f8ff6ab-9364-4e5b-a480-4d9ff441d727",
   "bg_record_key": "S40sh1::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S40sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:18:31.783406+00:00",
  "fingerprint": "c581762f7b4b08a2c7818aca361a16d8b75e8d01f514db3997d6f38b127e64b7",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S40sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S40sh1_sel.png",
  "source_sha256": "50051dac51512f09dad90e6bc4241744237a16fb1e971edcb1364277e1f0110d",
  "file": "S40sh1_cine.png",
  "staged_sha256": "00a8e894dd8e3e347494bb79dc7b901632324a78bff2c373a5bac23b25876877",
  "latency_ms": 9677
 },
 "S40sh4::signage": {
  "fp": "6edb47a98acf4916",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S40sh4": {
  "input_fingerprint": "bc70b1887122c083",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 앰버를 향해 미간을 잔뜩 찌푸린 채 크게 입을 벌려 소리치는 현우의 절박한 얼굴.\n\nLOCATION (lock): On the exposed roadside at the refugee settlement's outskirts, where the exhausted child has stopped. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Street (No police are present in the street) — Only a restricted portion remains visible behind 현우; used as Soft peripheral context keeps the shot connected to the escape route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established daytime ambient light and restrained contrast so urgency comes from expression rather than a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the empty gray asphalt and the same daylight conditions at this roadside stopping point. Exclude floodwater, floating panels, and search boats from the earlier flooded district.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The escape continues through the daylight outskirts toward the city. 현우: He retains the facial bruises and untreated dog-bite wound on his leg. The intact contact card remains hidden inside his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 앰버를 향해 미간을 잔뜩 찌푸린 채 크게 입을 벌려 소리치는 현우의 절박한 얼굴.\n\nLOCATION (lock): On the exposed roadside at the refugee settlement's outskirts, where the exhausted child has stopped. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Street (No police are present in the street) — Only a restricted portion remains visible behind 현우; used as Soft peripheral context keeps the shot connected to the escape route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established daytime ambient light and restrained contrast so urgency comes from expression rather than a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the empty gray asphalt and the same daylight conditions at this roadside stopping point. Exclude floodwater, floating panels, and search boats from the earlier flooded district.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The escape continues through the daylight outskirts toward the city. 현우: He retains the facial bruises and untreated dog-bite wound on his leg. The intact contact card remains hidden inside his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 앰버를 향해 미간을 잔뜩 찌푸린 채 크게 입을 벌려 소리치는 현우의 절박한 얼굴.\n\nLOCATION (lock): On the exposed roadside at the refugee settlement's outskirts, where the exhausted child has stopped. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Street (No police are present in the street) — Only a restricted portion remains visible behind 현우; used as Soft peripheral context keeps the shot connected to the escape route.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established daytime ambient light and restrained contrast so urgency comes from expression rather than a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the empty gray asphalt and the same daylight conditions at this roadside stopping point. Exclude floodwater, floating panels, and search boats from the earlier flooded district.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The escape continues through the daylight outskirts toward the city. 현우: He retains the facial bruises and untreated dog-bite wound on his leg. The intact contact card remains hidden inside his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "인물의 시선과 외침이 프레임 밖의 대상(앰버)을 향하고 있습니다.",
    "built_space": "배경에 지정된 텅 빈 회색 아스팔트 도로가 얕은 심도로 보입니다.",
    "entities": "레퍼런스와 일치하는 현우의 외모, 얼굴의 타박상, 파란색 티셔츠가 보입니다.",
    "hard_violations": [],
    "physics": "클로즈업 샷으로 인물의 상체와 머리가 화면에 안정적으로 위치해 있습니다."
   },
   {
    "label": "B",
    "direction": "인물의 시선과 외침이 프레임 밖의 대상(앰버)을 향하고 있습니다.",
    "built_space": "배경에 지정된 텅 빈 회색 아스팔트 도로가 얕은 심도로 보입니다.",
    "entities": "레퍼런스와 일치하는 현우의 외모와 헝클어진 머리, 얼굴의 타박상, 흙먼지가 묻은 티셔츠가 보입니다.",
    "hard_violations": [],
    "physics": "클로즈업 샷으로 인물의 상체와 머리가 화면에 안정적으로 위치해 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "'미간을 잔뜩 찌푸린 채 크게 입을 벌려 소리치는 절박한 얼굴'이라는 핵심 표정 지시를 매우 사실적이고 강렬하게 잘 구현했습니다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "지시문대로 찌푸린 얼굴과 소리치는 입모양을 구현했으나, B 후보에 비해 미간의 주름과 절박한 감정 표현이 다소 약하게 나타납니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "인물의 시선과 외침이 프레임 밖의 대상(앰버)을 향하고 있습니다.",
        "built_space": "배경에 지정된 텅 빈 회색 아스팔트 도로가 얕은 심도로 보입니다.",
        "entities": "레퍼런스와 일치하는 현우의 외모, 얼굴의 타박상, 파란색 티셔츠가 보입니다.",
        "hard_violations": [],
        "physics": "클로즈업 샷으로 인물의 상체와 머리가 화면에 안정적으로 위치해 있습니다."
       },
       {
        "label": "B",
        "direction": "인물의 시선과 외침이 프레임 밖의 대상(앰버)을 향하고 있습니다.",
        "built_space": "배경에 지정된 텅 빈 회색 아스팔트 도로가 얕은 심도로 보입니다.",
        "entities": "레퍼런스와 일치하는 현우의 외모와 헝클어진 머리, 얼굴의 타박상, 흙먼지가 묻은 티셔츠가 보입니다.",
        "hard_violations": [],
        "physics": "클로즈업 샷으로 인물의 상체와 머리가 화면에 안정적으로 위치해 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "'미간을 잔뜩 찌푸린 채 크게 입을 벌려 소리치는 절박한 얼굴'이라는 핵심 표정 지시를 매우 사실적이고 강렬하게 잘 구현했습니다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "지시문대로 찌푸린 얼굴과 소리치는 입모양을 구현했으나, B 후보에 비해 미간의 주름과 절박한 감정 표현이 다소 약하게 나타납니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "인물의 시선과 외침이 프레임 밖의 대상(앰버)을 향하고 있습니다.",
        "built_space": "배경에 지정된 텅 빈 회색 아스팔트 도로가 얕은 심도로 보입니다.",
        "entities": "레퍼런스와 일치하는 현우의 외모, 얼굴의 타박상, 파란색 티셔츠가 보입니다.",
        "hard_violations": [],
        "physics": "클로즈업 샷으로 인물의 상체와 머리가 화면에 안정적으로 위치해 있습니다."
       },
       {
        "label": "B",
        "direction": "인물의 시선과 외침이 프레임 밖의 대상(앰버)을 향하고 있습니다.",
        "built_space": "배경에 지정된 텅 빈 회색 아스팔트 도로가 얕은 심도로 보입니다.",
        "entities": "레퍼런스와 일치하는 현우의 외모와 헝클어진 머리, 얼굴의 타박상, 흙먼지가 묻은 티셔츠가 보입니다.",
        "hard_violations": [],
        "physics": "클로즈업 샷으로 인물의 상체와 머리가 화면에 안정적으로 위치해 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "미간을 찌푸리고 크게 외치는 표정은 충실하지만, B보다 얼굴 비중이 작고 시선이 렌즈에 가까우며 티셔츠도 참조의 남색보다 검게 보인다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "얼굴 중심의 밀착 클로즈업, 화면 밖 상대를 향한 시선, 깊게 찌푸린 미간과 절박한 외침이 핵심 지시에 가장 충실하며 남색 의상도 참조와 부합한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴은 거의 정면이고 눈은 렌즈 가까운 화면 오른쪽을 향한다. 입을 크게 벌려 그 방향으로 외친다. 상대는 화면 밖이라 앰버인지 직접 확인할 수 없으며, 상대를 보는 방향성이 B보다 덜 분명하다. 무기나 방향성 있는 소품은 없다.",
        "built_space": "얼굴과 어깨 뒤로 빈 회색 아스팔트 도로 하나, 왼쪽 연석과 보도 한 줄, 상단의 흐릿한 낮은 건물 일부가 보인다. 도로 가장자리는 인물 뒤쪽에 있어 도로 위 인물 배치와 모순되지 않는다. 경찰이나 다른 사람, 침수 흔적은 없으며 중복 시설이나 반사 문제도 없다. B보다 도로와 상체가 조금 더 많이 포함된다.",
        "entities": "앳된 동아시아계 남성 한 명으로, 참조 현우와 유사한 얼굴과 헝클어진 검은 머리, 양 볼의 멍과 찰과상이 보인다. 한국계 미국인이라는 국적·배경은 영상만으로 확인할 수 없다. 둥근 목 티셔츠는 먼지 묻은 검정에 가까워 참조의 남색과 차이가 있다. 앰버나 이전 사진의 아이는 등장하지 않는다. 다리 상처와 신발 속 카드는 올바르게 화면 밖이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리는 목에, 목은 앞으로 기울인 상체에 자연스럽게 연결되어 있으며 외칠 때의 턱과 입 벌림도 가능한 범위다. 하체와 발은 프레임 밖이므로 지면 접촉은 확인할 수 없지만, 공중에 떠 있는 몸으로 묘사되지는 않는다. 떠 있거나 손 없이 들린 물체는 없다."
       },
       {
        "label": "B",
        "direction": "얼굴과 두 눈이 화면 왼쪽의 화면 밖 상대를 향하고, 그 방향으로 입을 크게 벌려 외친다. 렌즈를 직접 바라보지 않아 앰버를 향한 외침으로 자연스럽게 읽힌다. 다만 앰버는 보이지 않으므로 실제 시선 도착점은 확인할 수 없다. 무기나 지시 소품은 없다.",
        "built_space": "크게 잡힌 얼굴 뒤로 빈 회색 도로 하나, 왼쪽 연석과 보도 한 줄, 흐릿한 전신주 하나 및 먼 외곽 지형 일부가 보인다. 시설들은 배경 규모로 유지되고 도로 위 인물과 충돌하지 않는다. 거리의 경찰·다른 사람·홍수·배는 없다. 제한된 주변 배경과 낮빛이 이전 장소의 맥락을 유지하며, 불가능한 반사나 시설 중복은 보이지 않는다.",
        "entities": "참조와 유사한 얼굴형과 검은 헝클어진 머리를 가진 앳된 동아시아계 남성 한 명이다. 나이대와 남성 외형은 현우 설정에 부합하며, 국적은 외형만으로 검증할 수 없다. 양 볼의 멍과 상처, 먼지 묻은 남색 둥근 목 티셔츠가 보인다. 자연스러운 홍채와 동공이 유지된다. 다리와 신발은 클로즈업 밖이므로 상처와 숨긴 카드는 확인 대상이 아니며, 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "목과 어깨가 머리를 자연스럽게 지탱하고, 몸을 약간 앞으로 기울인 채 소리치는 얼굴 근육과 턱의 움직임이 해부학적으로 가능하다. 발과 지면 접촉은 촬영 범위 밖이나 부유를 나타내는 단서는 없다. 별도로 지지되지 않은 물체나 불가능한 신체 접촉도 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "미간을 찌푸리고 크게 외치는 표정은 충실하지만, B보다 얼굴 비중이 작고 시선이 렌즈에 가까우며 티셔츠도 참조의 남색보다 검게 보인다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "얼굴 중심의 밀착 클로즈업, 화면 밖 상대를 향한 시선, 깊게 찌푸린 미간과 절박한 외침이 핵심 지시에 가장 충실하며 남색 의상도 참조와 부합한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴은 거의 정면이고 눈은 렌즈 가까운 화면 오른쪽을 향한다. 입을 크게 벌려 그 방향으로 외친다. 상대는 화면 밖이라 앰버인지 직접 확인할 수 없으며, 상대를 보는 방향성이 B보다 덜 분명하다. 무기나 방향성 있는 소품은 없다.",
        "built_space": "얼굴과 어깨 뒤로 빈 회색 아스팔트 도로 하나, 왼쪽 연석과 보도 한 줄, 상단의 흐릿한 낮은 건물 일부가 보인다. 도로 가장자리는 인물 뒤쪽에 있어 도로 위 인물 배치와 모순되지 않는다. 경찰이나 다른 사람, 침수 흔적은 없으며 중복 시설이나 반사 문제도 없다. B보다 도로와 상체가 조금 더 많이 포함된다.",
        "entities": "앳된 동아시아계 남성 한 명으로, 참조 현우와 유사한 얼굴과 헝클어진 검은 머리, 양 볼의 멍과 찰과상이 보인다. 한국계 미국인이라는 국적·배경은 영상만으로 확인할 수 없다. 둥근 목 티셔츠는 먼지 묻은 검정에 가까워 참조의 남색과 차이가 있다. 앰버나 이전 사진의 아이는 등장하지 않는다. 다리 상처와 신발 속 카드는 올바르게 화면 밖이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리는 목에, 목은 앞으로 기울인 상체에 자연스럽게 연결되어 있으며 외칠 때의 턱과 입 벌림도 가능한 범위다. 하체와 발은 프레임 밖이므로 지면 접촉은 확인할 수 없지만, 공중에 떠 있는 몸으로 묘사되지는 않는다. 떠 있거나 손 없이 들린 물체는 없다."
       },
       {
        "label": "A",
        "direction": "얼굴과 두 눈이 화면 왼쪽의 화면 밖 상대를 향하고, 그 방향으로 입을 크게 벌려 외친다. 렌즈를 직접 바라보지 않아 앰버를 향한 외침으로 자연스럽게 읽힌다. 다만 앰버는 보이지 않으므로 실제 시선 도착점은 확인할 수 없다. 무기나 지시 소품은 없다.",
        "built_space": "크게 잡힌 얼굴 뒤로 빈 회색 도로 하나, 왼쪽 연석과 보도 한 줄, 흐릿한 전신주 하나 및 먼 외곽 지형 일부가 보인다. 시설들은 배경 규모로 유지되고 도로 위 인물과 충돌하지 않는다. 거리의 경찰·다른 사람·홍수·배는 없다. 제한된 주변 배경과 낮빛이 이전 장소의 맥락을 유지하며, 불가능한 반사나 시설 중복은 보이지 않는다.",
        "entities": "참조와 유사한 얼굴형과 검은 헝클어진 머리를 가진 앳된 동아시아계 남성 한 명이다. 나이대와 남성 외형은 현우 설정에 부합하며, 국적은 외형만으로 검증할 수 없다. 양 볼의 멍과 상처, 먼지 묻은 남색 둥근 목 티셔츠가 보인다. 자연스러운 홍채와 동공이 유지된다. 다리와 신발은 클로즈업 밖이므로 상처와 숨긴 카드는 확인 대상이 아니며, 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "목과 어깨가 머리를 자연스럽게 지탱하고, 몸을 약간 앞으로 기울인 채 소리치는 얼굴 근육과 턱의 움직임이 해부학적으로 가능하다. 발과 지면 접촉은 촬영 범위 밖이나 부유를 나타내는 단서는 없다. 별도로 지지되지 않은 물체나 불가능한 신체 접촉도 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.857,
    "B": 1.889
   },
   "adjusted": {
    "A": 1.857,
    "B": 1.889
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1889,
   "A": 1857
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1889,
    "verdict_ko": "'미간을 잔뜩 찌푸린 채 크게 입을 벌려 소리치는 절박한 얼굴'이라는 핵심 표정 지시를 매우 사실적이고 강렬하게 잘 구현했습니다."
   },
   {
    "label": "A",
    "score": 1857,
    "verdict_ko": "지시문대로 찌푸린 얼굴과 소리치는 입모양을 구현했으나, B 후보에 비해 미간의 주름과 절박한 감정 표현이 다소 약하게 나타납니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S40sh1_sel.png",
    "asset_id": "f2df1a6c-99b6-4789-8af0-4698ad69381b",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-8e63-766d-bd56-258361ca5f16",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S40sh1"
  }
 },
 "S40sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:19:25.476333+00:00",
  "fingerprint": "c550fa501b9917b5dc89d2a903bc07ef335406b61afe14bae5927f657e1d4ff0",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S40sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S40sh4_sel.png",
  "source_sha256": "60da2c55e5f02136c57303af007b9a18fe1a71770637d3dd10aeb29b611949e1",
  "file": "S40sh4_cine.png",
  "staged_sha256": "80f43711ea3eb6d1dc0687cb5c221701216b06a3432a8c640a092cd55e30bc0e",
  "latency_ms": 9393
 },
 "S40sh6::signage": {
  "fp": "6a851e20df9e5221",
  "inscriptions": [],
  "cues": [
   {
    "text_native": "",
    "source": "scene_text_implied",
    "source_quote": "네온 불빛이 켜진 무인 점포"
   }
  ],
  "dropped": []
 },
 "S40sh6": {
  "input_fingerprint": "35ff9988b706cdcf",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 네온 불빛이 켜진 무인 점포를 배경으로, 한쪽 발을 바닥에서 떼고 앞으로 달려나가는 자세의 현우와 앰버, 그리고 그 뒤를 묵묵히 따라 육중한 다리를 길게 내딛은 찰리의 뒷모습 전경.\n\nLOCATION (lock): On the street approaching an unattended shop, with its illuminated exterior signage ahead of the fleeing group. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Unattended shop across the street in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Unattended shop across the street (Visible across the street beyond the fleeing group) — Its street-facing exterior is seen obliquely in the upper-right background; used as Anchors the wider geography while remaining subordinate to the running figures; Street (No police are present) — The route extends away from the camera, with the shop across it; used as Provides separation between the camera, the three runners, and the opposite frontage.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light remains dominant, with the shop's stated neon contributing only a localized accent rather than recoloring the street.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): An unmanned store stands across the street along the daylight escape route. 앰버: She is moving again but remains exhausted and tear-streaked. 현우: He continues escaping with facial bruises and the untreated leg wound. The contact card remains concealed inside his shoe. 찰리: He follows along the street, still in his old coat-and-hat disguise over the worn metal body.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 네온 불빛이 켜진 무인 점포를 배경으로, 한쪽 발을 바닥에서 떼고 앞으로 달려나가는 자세의 현우와 앰버, 그리고 그 뒤를 묵묵히 따라 육중한 다리를 길게 내딛은 찰리의 뒷모습 전경.\n\nLOCATION (lock): On the street approaching an unattended shop, with its illuminated exterior signage ahead of the fleeing group. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Unattended shop across the street in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Unattended shop across the street (Visible across the street beyond the fleeing group) — Its street-facing exterior is seen obliquely in the upper-right background; used as Anchors the wider geography while remaining subordinate to the running figures; Street (No police are present) — The route extends away from the camera, with the shop across it; used as Provides separation between the camera, the three runners, and the opposite frontage.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light remains dominant, with the shop's stated neon contributing only a localized accent rather than recoloring the street.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): An unmanned store stands across the street along the daylight escape route. 앰버: She is moving again but remains exhausted and tear-streaked. 현우: He continues escaping with facial bruises and the untreated leg wound. The contact card remains concealed inside his shoe. 찰리: He follows along the street, still in his old coat-and-hat disguise over the worn metal body.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 네온 불빛이 켜진 무인 점포를 배경으로, 한쪽 발을 바닥에서 떼고 앞으로 달려나가는 자세의 현우와 앰버, 그리고 그 뒤를 묵묵히 따라 육중한 다리를 길게 내딛은 찰리의 뒷모습 전경.\n\nLOCATION (lock): On the street approaching an unattended shop, with its illuminated exterior signage ahead of the fleeing group. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Unattended shop across the street in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Unattended shop across the street (Visible across the street beyond the fleeing group) — Its street-facing exterior is seen obliquely in the upper-right background; used as Anchors the wider geography while remaining subordinate to the running figures; Street (No police are present) — The route extends away from the camera, with the shop across it; used as Provides separation between the camera, the three runners, and the opposite frontage.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light remains dominant, with the shop's stated neon contributing only a localized accent rather than recoloring the street.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): An unmanned store stands across the street along the daylight escape route. 앰버: She is moving again but remains exhausted and tear-streaked. 현우: He continues escaping with facial bruises and the untreated leg wound. The contact card remains concealed inside his shoe. 찰리: He follows along the street, still in his old coat-and-hat disguise over the worn metal body.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "현우와 앰버는 카메라를 향해 앞으로 달려오고 있으며, 찰리는 카메라를 등지고 있어 결과적으로 이들이 서로 마주 보고 있는 모순된 방향성을 가집니다.",
    "built_space": "도로는 카메라 뷰에서 멀어지는 방향으로 뻗어 있으며, 지시된 무인 상점은 프레임의 우측 상단 배경에 올바르게 배치되어 네온 불빛을 비추고 있습니다.",
    "entities": "찰리는 낡은 금속 몸체 위에 코트와 모자를 쓴 뒷모습으로 나타납니다. 현우는 멍든 얼굴을 하고 있으나, 앰버는 지치거나 눈물 자국이 있는 모습으로 묘사되지 않았습니다. 상점 간판에는 명확히 읽을 수 있는 한글 텍스트들이 존재합니다.",
    "hard_violations": [
     "[gemini-pro] 읽을 수 있는 텍스트 (상점 간판의 한글)",
     "[gemini-pro] 무대 연출 위반 (찰리가 일행을 뒤따라가야 하나, 현우와 앰버가 카메라를 향해 뛰어오며 찰리와 마주 보고 있음)",
     "[gpt-high] 찰리와 현우·앰버가 서로 마주 접근하도록 배치되어, 같은 방향으로 도망가는 두 사람을 찰리가 뒤따르는 필수 동선을 위반한다.",
     "[gpt-high] 큰 네온 간판과 거리 간판에 판독 가능한 한글이 노출되어 문자 금지를 위반한다."
    ],
    "physics": "현우와 앰버는 각각 한 발을 바닥에 딛고 다른 발을 공중에 띄워 달리는 자세를 물리적으로 무리 없이 취하고 있으며, 찰리 역시 바닥을 딛고 걷는 자세를 유지하고 있습니다."
   },
   {
    "label": "B",
    "direction": "현우와 앰버는 카메라 방향으로 띄어오고, 찰리는 카메라를 등지고 걷고 있어 지시된 추격/도망 구도와 달리 서로를 향해 걷고 있습니다.",
    "built_space": "도로가 배경으로 뻗어 있으나 상점이 우측 상단이 아닌 중앙 배경 쪽에 더 가깝게 배치되어 있습니다.",
    "entities": "앰버는 눈물 자국이 있는 지친 표정이고 현우는 멍이 있습니다. 찰리는 코트를 입고 있으나 거대한 체형이 어색하게 묘사되었습니다. 배경 상점 간판에 'OPEN' 등 읽을 수 있는 텍스트가 표시되어 있습니다.",
    "hard_violations": [
     "[gemini-pro] 읽을 수 있는 텍스트 (상점 간판의 'OPEN' 및 기타 글자)",
     "[gemini-pro] 무대 연출 위반 (일행과 찰리가 서로 마주 보는 방향으로 움직임)",
     "[gemini-pro] 물리적으로 불가능한 해부학 (찰리의 다리와 상체를 덮은 코트가 올바르게 연결되지 않은 채 허공에 떠 있는 것처럼 묘사됨)",
     "[gpt-high] 현우와 앰버가 찰리와 반대 방향으로 달려, 찰리가 두 사람 뒤를 따르는 필수 동선 및 배치가 성립하지 않는다.",
     "[gpt-high] 점포의 'OPEN'과 주변 간판에 읽을 수 있는 문자가 노출되어 문자 금지를 위반한다."
    ],
    "physics": "아이들은 바닥을 딛고 뛰는 자세를 취하고 있으나, 찰리의 경우 육중한 다리가 상체의 코트 아래에서 물리적인 지지나 구조적 연결 없이 분리된 것처럼 보입니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "프롬프트에 금지된 '읽을 수 있는 텍스트(간판 한글)'가 명확히 존재하며, 현우와 앰버가 카메라를 향해 달려오면서 찰리와 마주 보는 구도로 묘사되어 '일행을 뒤따라간다'는 핵심 연출 지시를 위반했습니다."
       },
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "금지된 텍스트('OPEN')가 포함되었을 뿐만 아니라, 일행과 찰리가 마주 보는 방향으로 렌더링되었고 찰리의 상체 코트와 다리가 물리적으로 연결되지 않은 해부학적 오류가 있어 사용할 수 없습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우와 앰버는 카메라를 향해 앞으로 달려오고 있으며, 찰리는 카메라를 등지고 있어 결과적으로 이들이 서로 마주 보고 있는 모순된 방향성을 가집니다.",
        "built_space": "도로는 카메라 뷰에서 멀어지는 방향으로 뻗어 있으며, 지시된 무인 상점은 프레임의 우측 상단 배경에 올바르게 배치되어 네온 불빛을 비추고 있습니다.",
        "entities": "찰리는 낡은 금속 몸체 위에 코트와 모자를 쓴 뒷모습으로 나타납니다. 현우는 멍든 얼굴을 하고 있으나, 앰버는 지치거나 눈물 자국이 있는 모습으로 묘사되지 않았습니다. 상점 간판에는 명확히 읽을 수 있는 한글 텍스트들이 존재합니다.",
        "hard_violations": [
         "읽을 수 있는 텍스트 (상점 간판의 한글)",
         "무대 연출 위반 (찰리가 일행을 뒤따라가야 하나, 현우와 앰버가 카메라를 향해 뛰어오며 찰리와 마주 보고 있음)"
        ],
        "physics": "현우와 앰버는 각각 한 발을 바닥에 딛고 다른 발을 공중에 띄워 달리는 자세를 물리적으로 무리 없이 취하고 있으며, 찰리 역시 바닥을 딛고 걷는 자세를 유지하고 있습니다."
       },
       {
        "label": "B",
        "direction": "현우와 앰버는 카메라 방향으로 띄어오고, 찰리는 카메라를 등지고 걷고 있어 지시된 추격/도망 구도와 달리 서로를 향해 걷고 있습니다.",
        "built_space": "도로가 배경으로 뻗어 있으나 상점이 우측 상단이 아닌 중앙 배경 쪽에 더 가깝게 배치되어 있습니다.",
        "entities": "앰버는 눈물 자국이 있는 지친 표정이고 현우는 멍이 있습니다. 찰리는 코트를 입고 있으나 거대한 체형이 어색하게 묘사되었습니다. 배경 상점 간판에 'OPEN' 등 읽을 수 있는 텍스트가 표시되어 있습니다.",
        "hard_violations": [
         "읽을 수 있는 텍스트 (상점 간판의 'OPEN' 및 기타 글자)",
         "무대 연출 위반 (일행과 찰리가 서로 마주 보는 방향으로 움직임)",
         "물리적으로 불가능한 해부학 (찰리의 다리와 상체를 덮은 코트가 올바르게 연결되지 않은 채 허공에 떠 있는 것처럼 묘사됨)"
        ],
        "physics": "아이들은 바닥을 딛고 뛰는 자세를 취하고 있으나, 찰리의 경우 육중한 다리가 상체의 코트 아래에서 물리적인 지지나 구조적 연결 없이 분리된 것처럼 보입니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "프롬프트에 금지된 '읽을 수 있는 텍스트(간판 한글)'가 명확히 존재하며, 현우와 앰버가 카메라를 향해 달려오면서 찰리와 마주 보는 구도로 묘사되어 '일행을 뒤따라간다'는 핵심 연출 지시를 위반했습니다."
       },
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "금지된 텍스트('OPEN')가 포함되었을 뿐만 아니라, 일행과 찰리가 마주 보는 방향으로 렌더링되었고 찰리의 상체 코트와 다리가 물리적으로 연결되지 않은 해부학적 오류가 있어 사용할 수 없습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "현우와 앰버는 카메라를 향해 앞으로 달려오고 있으며, 찰리는 카메라를 등지고 있어 결과적으로 이들이 서로 마주 보고 있는 모순된 방향성을 가집니다.",
        "built_space": "도로는 카메라 뷰에서 멀어지는 방향으로 뻗어 있으며, 지시된 무인 상점은 프레임의 우측 상단 배경에 올바르게 배치되어 네온 불빛을 비추고 있습니다.",
        "entities": "찰리는 낡은 금속 몸체 위에 코트와 모자를 쓴 뒷모습으로 나타납니다. 현우는 멍든 얼굴을 하고 있으나, 앰버는 지치거나 눈물 자국이 있는 모습으로 묘사되지 않았습니다. 상점 간판에는 명확히 읽을 수 있는 한글 텍스트들이 존재합니다.",
        "hard_violations": [
         "읽을 수 있는 텍스트 (상점 간판의 한글)",
         "무대 연출 위반 (찰리가 일행을 뒤따라가야 하나, 현우와 앰버가 카메라를 향해 뛰어오며 찰리와 마주 보고 있음)"
        ],
        "physics": "현우와 앰버는 각각 한 발을 바닥에 딛고 다른 발을 공중에 띄워 달리는 자세를 물리적으로 무리 없이 취하고 있으며, 찰리 역시 바닥을 딛고 걷는 자세를 유지하고 있습니다."
       },
       {
        "label": "B",
        "direction": "현우와 앰버는 카메라 방향으로 띄어오고, 찰리는 카메라를 등지고 걷고 있어 지시된 추격/도망 구도와 달리 서로를 향해 걷고 있습니다.",
        "built_space": "도로가 배경으로 뻗어 있으나 상점이 우측 상단이 아닌 중앙 배경 쪽에 더 가깝게 배치되어 있습니다.",
        "entities": "앰버는 눈물 자국이 있는 지친 표정이고 현우는 멍이 있습니다. 찰리는 코트를 입고 있으나 거대한 체형이 어색하게 묘사되었습니다. 배경 상점 간판에 'OPEN' 등 읽을 수 있는 텍스트가 표시되어 있습니다.",
        "hard_violations": [
         "읽을 수 있는 텍스트 (상점 간판의 'OPEN' 및 기타 글자)",
         "무대 연출 위반 (일행과 찰리가 서로 마주 보는 방향으로 움직임)",
         "물리적으로 불가능한 해부학 (찰리의 다리와 상체를 덮은 코트가 올바르게 연결되지 않은 채 허공에 떠 있는 것처럼 묘사됨)"
        ],
        "physics": "아이들은 바닥을 딛고 뛰는 자세를 취하고 있으나, 찰리의 경우 육중한 다리가 상체의 코트 아래에서 물리적인 지지나 구조적 연결 없이 분리된 것처럼 보입니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "둘 다 진행 방향과 문자 금지를 위반하지만, A가 우측 상단의 종속적인 점포 크기, 절제된 주간 네온, 찰리의 코트 위장을 더 잘 구현한다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "아스팔트 도로는 참조에 더 가깝지만, 일행이 서로 마주 달리고 점포가 배경을 과점유하며 선명한 간판과 도로의 과도한 네온 색 번짐이 요구와 어긋난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우와 앰버는 얼굴과 가슴을 카메라 쪽으로 향한 채 점포에서 멀어져 달린다. 찰리는 등을 보이고 두 사람 쪽으로 전진한다. 따라서 세 명이 같은 방향으로 탈출하고 찰리가 뒤따르는 관계가 아니라 서로 접근하는 관계다. 현우와 앰버의 시선도 전방의 찰리 또는 카메라 쪽이며, 점포를 향하지 않는다. 무기나 조준 물체는 없다.",
        "built_space": "넓은 화면의 우측 상단에 점포 한 곳이 있고, 중앙 유리 출입구와 양옆 진열창, 상부 가로 간판 하나가 보인다. 점포는 인물보다 뒤에 놓여 크기가 비교적 절제되지만, 외관은 요구한 사선보다 정면에 가깝다. 찰리는 왼쪽 전경, 두 사람은 중앙 중경에 있다. 도로 대부분이 사각 포장재여서 이전 장면의 아스팔트 노면과 다르다. 낮빛이 지배적이며 불가능한 반사는 보이지 않는다.",
        "entities": "현우, 앰버, 찰리에 해당하는 세 인물만 보이고 경찰은 없다. 현우는 젊은 동아시아계 남성으로 검은 헝클어진 머리, 더러운 검은 티셔츠와 얼굴의 상처가 참조에 대체로 맞는다. 앰버는 금발 어린 여자아이이며 남색 티셔츠와 지친 표정이 보이지만 혼혈 정체성과 눈물 자국은 확정하기 어렵다. 찰리는 낡은 모자와 코트 아래로 베이지색 기계 다리가 드러나 위장 상태에 부합한다. 흰 얼굴은 뒷모습이라 보이지 않는다. 현우의 다리 상처와 신발 속 카드는 가려져 확인할 수 없다. 점포의 'OPEN'과 주변 한글 간판은 읽을 수 있다.",
        "hard_violations": [
         "현우와 앰버가 찰리와 반대 방향으로 달려, 찰리가 두 사람 뒤를 따르는 필수 동선 및 배치가 성립하지 않는다.",
         "점포의 'OPEN'과 주변 간판에 읽을 수 있는 문자가 노출되어 문자 금지를 위반한다."
        ],
        "physics": "현우와 앰버는 각각 한 발을 노면에 대고 다른 발을 뒤로 들어 올렸으며, 맞잡은 손과 팔의 굽힘도 달리기 동작으로 가능하다. 찰리는 화면 오른쪽 발로 지면을 지지하고 반대 발을 들어 전진한다. 모자는 머리, 코트는 어깨와 몸통에 지지된다. 근거 없이 공중에 뜬 몸이나 물체는 없다."
       },
       {
        "label": "B",
        "direction": "현우와 앰버는 화면 오른쪽 전경의 찰리를 향해 달리며 얼굴을 카메라에 드러낸다. 찰리는 등을 보인 채 왼쪽 중경의 두 사람 쪽으로 발을 내딛는다. 세 인물의 이동 방향이 서로 마주 보므로 점포를 앞에 두고 함께 달아나는 후면 장면이 아니다. 두 사람의 얼굴과 시선은 대체로 찰리 쪽이다. 무기나 조준 물체는 없다.",
        "built_space": "우측 배경 점포 한 곳에 중앙 양개 유리 출입구 하나, 좌우 유리 진열창, 상단의 큰 발광 간판 두 구획이 보인다. 외관은 비스듬히 보이지만 화면 상부와 오른쪽을 크게 차지해 멀리 있는 종속적 배경이라는 요구에서 벗어난다. 찰리는 오른쪽 전경, 현우와 앰버는 중앙 왼쪽 중경에 있다. 갈라진 아스팔트와 연석은 이전 장면의 재질에 A보다 가깝다. 점포 앞 도로에는 넓은 분홍·청록색 빛이 번져 국소적인 네온 강조를 넘어선다. 인물의 불가능한 거울 반사는 없다.",
        "entities": "세 인물 외 추가 인물이나 경찰은 보이지 않는다. 현우의 검은 머리, 젊은 동아시아계 외모, 얼굴 상처와 더러운 검은 티셔츠는 참조와 대체로 맞는다. 앰버는 금발 어린 여자아이와 남색 티셔츠로 표현되지만 눈물 자국은 분명하지 않다. 찰리의 베이지색 장갑판과 긴 기계 팔은 참조 특징에 맞으나, 어깨와 팔 장갑이 코트 밖으로 크게 노출되어 금속 몸체 위에 코트를 입은 위장 상태의 충실도가 낮다. 모자는 착용했고 얼굴은 가려져 있다. 현우의 다리 상처와 숨긴 카드는 확인할 수 없다. 상단 및 거리의 여러 간판에는 읽을 수 있는 한글이 있다.",
        "hard_violations": [
         "찰리와 현우·앰버가 서로 마주 접근하도록 배치되어, 같은 방향으로 도망가는 두 사람을 찰리가 뒤따르는 필수 동선을 위반한다.",
         "큰 네온 간판과 거리 간판에 판독 가능한 한글이 노출되어 문자 금지를 위반한다."
        ],
        "physics": "현우와 앰버는 각각 한 발로 아스팔트를 딛고 반대 발을 들어 달리는 자세이며 손을 맞잡고 있다. 찰리는 화면 왼쪽 발을 지면에 두고 오른쪽 발을 들어 보행한다. 기계 관절과 몸의 무게 이동은 가능한 자세로 보인다. 코트와 모자도 몸에 지지되어 있으며, 지지나 추진 근거 없이 떠 있는 대상은 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "둘 다 진행 방향과 문자 금지를 위반하지만, A가 우측 상단의 종속적인 점포 크기, 절제된 주간 네온, 찰리의 코트 위장을 더 잘 구현한다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "아스팔트 도로는 참조에 더 가깝지만, 일행이 서로 마주 달리고 점포가 배경을 과점유하며 선명한 간판과 도로의 과도한 네온 색 번짐이 요구와 어긋난다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우와 앰버는 얼굴과 가슴을 카메라 쪽으로 향한 채 점포에서 멀어져 달린다. 찰리는 등을 보이고 두 사람 쪽으로 전진한다. 따라서 세 명이 같은 방향으로 탈출하고 찰리가 뒤따르는 관계가 아니라 서로 접근하는 관계다. 현우와 앰버의 시선도 전방의 찰리 또는 카메라 쪽이며, 점포를 향하지 않는다. 무기나 조준 물체는 없다.",
        "built_space": "넓은 화면의 우측 상단에 점포 한 곳이 있고, 중앙 유리 출입구와 양옆 진열창, 상부 가로 간판 하나가 보인다. 점포는 인물보다 뒤에 놓여 크기가 비교적 절제되지만, 외관은 요구한 사선보다 정면에 가깝다. 찰리는 왼쪽 전경, 두 사람은 중앙 중경에 있다. 도로 대부분이 사각 포장재여서 이전 장면의 아스팔트 노면과 다르다. 낮빛이 지배적이며 불가능한 반사는 보이지 않는다.",
        "entities": "현우, 앰버, 찰리에 해당하는 세 인물만 보이고 경찰은 없다. 현우는 젊은 동아시아계 남성으로 검은 헝클어진 머리, 더러운 검은 티셔츠와 얼굴의 상처가 참조에 대체로 맞는다. 앰버는 금발 어린 여자아이이며 남색 티셔츠와 지친 표정이 보이지만 혼혈 정체성과 눈물 자국은 확정하기 어렵다. 찰리는 낡은 모자와 코트 아래로 베이지색 기계 다리가 드러나 위장 상태에 부합한다. 흰 얼굴은 뒷모습이라 보이지 않는다. 현우의 다리 상처와 신발 속 카드는 가려져 확인할 수 없다. 점포의 'OPEN'과 주변 한글 간판은 읽을 수 있다.",
        "hard_violations": [
         "현우와 앰버가 찰리와 반대 방향으로 달려, 찰리가 두 사람 뒤를 따르는 필수 동선 및 배치가 성립하지 않는다.",
         "점포의 'OPEN'과 주변 간판에 읽을 수 있는 문자가 노출되어 문자 금지를 위반한다."
        ],
        "physics": "현우와 앰버는 각각 한 발을 노면에 대고 다른 발을 뒤로 들어 올렸으며, 맞잡은 손과 팔의 굽힘도 달리기 동작으로 가능하다. 찰리는 화면 오른쪽 발로 지면을 지지하고 반대 발을 들어 전진한다. 모자는 머리, 코트는 어깨와 몸통에 지지된다. 근거 없이 공중에 뜬 몸이나 물체는 없다."
       },
       {
        "label": "A",
        "direction": "현우와 앰버는 화면 오른쪽 전경의 찰리를 향해 달리며 얼굴을 카메라에 드러낸다. 찰리는 등을 보인 채 왼쪽 중경의 두 사람 쪽으로 발을 내딛는다. 세 인물의 이동 방향이 서로 마주 보므로 점포를 앞에 두고 함께 달아나는 후면 장면이 아니다. 두 사람의 얼굴과 시선은 대체로 찰리 쪽이다. 무기나 조준 물체는 없다.",
        "built_space": "우측 배경 점포 한 곳에 중앙 양개 유리 출입구 하나, 좌우 유리 진열창, 상단의 큰 발광 간판 두 구획이 보인다. 외관은 비스듬히 보이지만 화면 상부와 오른쪽을 크게 차지해 멀리 있는 종속적 배경이라는 요구에서 벗어난다. 찰리는 오른쪽 전경, 현우와 앰버는 중앙 왼쪽 중경에 있다. 갈라진 아스팔트와 연석은 이전 장면의 재질에 A보다 가깝다. 점포 앞 도로에는 넓은 분홍·청록색 빛이 번져 국소적인 네온 강조를 넘어선다. 인물의 불가능한 거울 반사는 없다.",
        "entities": "세 인물 외 추가 인물이나 경찰은 보이지 않는다. 현우의 검은 머리, 젊은 동아시아계 외모, 얼굴 상처와 더러운 검은 티셔츠는 참조와 대체로 맞는다. 앰버는 금발 어린 여자아이와 남색 티셔츠로 표현되지만 눈물 자국은 분명하지 않다. 찰리의 베이지색 장갑판과 긴 기계 팔은 참조 특징에 맞으나, 어깨와 팔 장갑이 코트 밖으로 크게 노출되어 금속 몸체 위에 코트를 입은 위장 상태의 충실도가 낮다. 모자는 착용했고 얼굴은 가려져 있다. 현우의 다리 상처와 숨긴 카드는 확인할 수 없다. 상단 및 거리의 여러 간판에는 읽을 수 있는 한글이 있다.",
        "hard_violations": [
         "찰리와 현우·앰버가 서로 마주 접근하도록 배치되어, 같은 방향으로 도망가는 두 사람을 찰리가 뒤따르는 필수 동선을 위반한다.",
         "큰 네온 간판과 거리 간판에 판독 가능한 한글이 노출되어 문자 금지를 위반한다."
        ],
        "physics": "현우와 앰버는 각각 한 발로 아스팔트를 딛고 반대 발을 들어 달리는 자세이며 손을 맞잡고 있다. 찰리는 화면 왼쪽 발을 지면에 두고 오른쪽 발을 들어 보행한다. 기계 관절과 몸의 무게 이동은 가능한 자세로 보인다. 코트와 모자도 몸에 지지되어 있으며, 지지나 추진 근거 없이 떠 있는 대상은 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.667,
    "B": 1.5
   },
   "adjusted": {
    "A": 1.417,
    "B": 1.25
   },
   "violations": {
    "A": [
     "[gemini-pro] 읽을 수 있는 텍스트 (상점 간판의 한글)",
     "[gemini-pro] 무대 연출 위반 (찰리가 일행을 뒤따라가야 하나, 현우와 앰버가 카메라를 향해 뛰어오며 찰리와 마주 보고 있음)",
     "[gpt-high] 찰리와 현우·앰버가 서로 마주 접근하도록 배치되어, 같은 방향으로 도망가는 두 사람을 찰리가 뒤따르는 필수 동선을 위반한다.",
     "[gpt-high] 큰 네온 간판과 거리 간판에 판독 가능한 한글이 노출되어 문자 금지를 위반한다."
    ],
    "B": [
     "[gemini-pro] 읽을 수 있는 텍스트 (상점 간판의 'OPEN' 및 기타 글자)",
     "[gemini-pro] 무대 연출 위반 (일행과 찰리가 서로 마주 보는 방향으로 움직임)",
     "[gemini-pro] 물리적으로 불가능한 해부학 (찰리의 다리와 상체를 덮은 코트가 올바르게 연결되지 않은 채 허공에 떠 있는 것처럼 묘사됨)",
     "[gpt-high] 현우와 앰버가 찰리와 반대 방향으로 달려, 찰리가 두 사람 뒤를 따르는 필수 동선 및 배치가 성립하지 않는다.",
     "[gpt-high] 점포의 'OPEN'과 주변 간판에 읽을 수 있는 문자가 노출되어 문자 금지를 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1417,
   "B": 1250
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1417,
    "verdict_ko": "프롬프트에 금지된 '읽을 수 있는 텍스트(간판 한글)'가 명확히 존재하며, 현우와 앰버가 카메라를 향해 달려오면서 찰리와 마주 보는 구도로 묘사되어 '일행을 뒤따라간다'는 핵심 연출 지시를 위반했습니다.  ★위반: [gemini-pro] 읽을 수 있는 텍스트 (상점 간판의 한글) / [gemini-pro] 무대 연출 위반 (찰리가 일행을 뒤따라가야 하나, 현우와 앰버가 카메라를 향해 뛰어오며 찰리와 마주 보고 있음) / [gpt-high] 찰리와 현우·앰버가 서로 마주 접근하도록 배치되어, 같은 방향으로 도망가는 두 사람을 찰리가 뒤따르는 필수 동선을 위반한다. / [gpt-high] 큰 네온 간판과 거리 간판에 판독 가능한 한글이 노출되어 문자 금지를 위반한다."
   },
   {
    "label": "B",
    "score": 1250,
    "verdict_ko": "금지된 텍스트('OPEN')가 포함되었을 뿐만 아니라, 일행과 찰리가 마주 보는 방향으로 렌더링되었고 찰리의 상체 코트와 다리가 물리적으로 연결되지 않은 해부학적 오류가 있어 사용할 수 없습니다.  ★위반: [gemini-pro] 읽을 수 있는 텍스트 (상점 간판의 'OPEN' 및 기타 글자) / [gemini-pro] 무대 연출 위반 (일행과 찰리가 서로 마주 보는 방향으로 움직임) / [gemini-pro] 물리적으로 불가능한 해부학 (찰리의 다리와 상체를 덮은 코트가 올바르게 연결되지 않은 채 허공에 떠 있는 것처럼 묘사됨) / [gpt-high] 현우와 앰버가 찰리와 반대 방향으로 달려, 찰리가 두 사람 뒤를 따르는 필수 동선 및 배치가 성립하지 않는다. / [gpt-high] 점포의 'OPEN'과 주변 간판에 읽을 수 있는 문자가 노출되어 문자 금지를 위반한다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S40sh4_sel.png",
    "asset_id": "fd8c895c-2b0b-4634-9ef2-25109e681b85",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-901d-7909-b608-c0dca9dcf63f",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S40sh4"
  },
  "lane_policy": "ab_select_bypass:prev"
 },
 "S40sh6::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:20:45.264427+00:00",
  "fingerprint": "8cdbd85997245e71c67832fef69f81d7bf310cd6fd3b9c812f44758089b6bdcf",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S40sh6_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S40sh6_sel.png",
  "source_sha256": "fa3a45a65a694f923f07451e077e4bdb7d00cb48f82d7b54dfa6221120fd5d7c",
  "file": "S40sh6_cine.png",
  "staged_sha256": "ae2aeba14e07ce94442dd766f062f5304d268207722fc64655023c3164c1d5be",
  "latency_ms": 10889
 },
 "S41sh14::signage": {
  "fp": "14c8720cad146e81",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S41sh14": {
  "input_fingerprint": "8926a458a0267fcd",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 거대한 금속 주먹이 현금인출기 앞면을 완전히 뚫고 들어간 파괴적인 찰나.\n\nLOCATION (lock): At the cash machine inside the unattended shop's sales floor, in daytime shop light near stocked aisles. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 현금인출기 (Front punctured by 찰리's fist; the fist has not yet withdrawn) — The front and a narrow adjacent side are visible obliquely, exposing the penetration rather than presenting the machine square-on; used as Impact surface and immediate spatial context for the extended arm; 점포 진열대 (Visible only as a narrow background fragment) — Seen obliquely beyond the cash machine; used as Preserves store context and prevents the impact from becoming an isolated effects image.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained ambient illumination appropriate to the store, with controlled contrast that keeps the fist and broken opening legible without added impact effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The ATM now has a fist-sized breach in its front. The scavenged duffel bag and backpack contain a disposable phone, a map, food, drinks and the remaining COPD medicine. 찰리: His metal fist is embedded in the ATM front. He still has the old coat and hat; the store blanket has not yet been handed over.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 거대한 금속 주먹이 현금인출기 앞면을 완전히 뚫고 들어간 파괴적인 찰나.\n\nLOCATION (lock): At the cash machine inside the unattended shop's sales floor, in daytime shop light near stocked aisles. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 현금인출기 (Front punctured by 찰리's fist; the fist has not yet withdrawn) — The front and a narrow adjacent side are visible obliquely, exposing the penetration rather than presenting the machine square-on; used as Impact surface and immediate spatial context for the extended arm; 점포 진열대 (Visible only as a narrow background fragment) — Seen obliquely beyond the cash machine; used as Preserves store context and prevents the impact from becoming an isolated effects image.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained ambient illumination appropriate to the store, with controlled contrast that keeps the fist and broken opening legible without added impact effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The ATM now has a fist-sized breach in its front. The scavenged duffel bag and backpack contain a disposable phone, a map, food, drinks and the remaining COPD medicine. 찰리: His metal fist is embedded in the ATM front. He still has the old coat and hat; the store blanket has not yet been handed over.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 거대한 금속 주먹이 현금인출기 앞면을 완전히 뚫고 들어간 파괴적인 찰나.\n\nLOCATION (lock): At the cash machine inside the unattended shop's sales floor, in daytime shop light near stocked aisles. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 현금인출기 (Front punctured by 찰리's fist; the fist has not yet withdrawn) — The front and a narrow adjacent side are visible obliquely, exposing the penetration rather than presenting the machine square-on; used as Impact surface and immediate spatial context for the extended arm; 점포 진열대 (Visible only as a narrow background fragment) — Seen obliquely beyond the cash machine; used as Preserves store context and prevents the impact from becoming an isolated effects image.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained ambient illumination appropriate to the store, with controlled contrast that keeps the fist and broken opening legible without added impact effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The ATM now has a fist-sized breach in its front. The scavenged duffel bag and backpack contain a disposable phone, a map, food, drinks and the remaining COPD medicine. 찰리: His metal fist is embedded in the ATM front. He still has the old coat and hat; the store blanket has not yet been handed over.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S41sh14__bgfirst_bg.png",
     "asset_id": "6398fb6a-6f79-4ad1-90ee-719e7320869a",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S41sh14.png",
     "asset_id": "2b2818ac-406a-4c9d-805c-3eb8df283def",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L195B01.png",
     "asset_id": "830dae63-3f71-4c94-9045-5c77a027e48e",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "왼쪽에서 뻗어나온 금속 주먹이 오른쪽의 현금인출기 앞면에 꽂혀 있음.",
    "built_space": "매장 내부와 인출기가 보이나, 배경의 유리문과 통로 배치가 기준 사진의 공간 구조와 일치하지 않음.",
    "entities": "찰리의 장갑판 팔과 코트, 인출기가 존재하나 카툰 렌더링 스타일로 묘사됨.",
    "hard_violations": [
     "[gemini-pro] 실사(Photorealistic) 지시를 위반하고 셀 셰이딩 및 일러스트 스타일의 그래픽 오버레이 형태로 렌더링됨"
    ],
    "physics": "주먹이 기기에 박혀 몸체의 지지를 받으나, 금속이 찢어진 단면이 2D 그림처럼 비현실적으로 묘사됨."
   },
   {
    "label": "B",
    "direction": "금속 주먹이 현금인출기 앞면을 타격하여 내부로 깊게 뚫고 들어감.",
    "built_space": "기준 사진과 동일한 벽면 포스터와 인출기, 뒤쪽으로 좁게 보이는 진열대가 정확한 각도로 배치됨.",
    "entities": "찰리의 모래색 기계 팔과 낡은 코트, 부서진 현금인출기 모두 프롬프트와 참조 이미지를 정확히 반영함.",
    "hard_violations": [
     "[gpt-high] 배경 벽 안내문의 한글 문구가 읽을 수 있게 남아 있어, 어디에도 판독 가능한 글자를 두지 말라는 금지 조건을 위반한다."
    ],
    "physics": "강한 타격으로 인해 철판이 구부러지고 파편이 튀는 물리적 현상이 사실적으로 지탱 및 묘사됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지정된 클로즈업 구도, 찰리의 낡은 코트와 기계 팔, 현금인출기 파괴 묘사 및 매장의 배경을 사실적으로 완벽하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "지정된 실사 렌더링이 아닌 일러스트 스타일로 표현되었으며, 매장 내부의 공간 배치가 기준 사진과 다릅니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "금속 주먹이 현금인출기 앞면을 타격하여 내부로 깊게 뚫고 들어감.",
        "built_space": "기준 사진과 동일한 벽면 포스터와 인출기, 뒤쪽으로 좁게 보이는 진열대가 정확한 각도로 배치됨.",
        "entities": "찰리의 모래색 기계 팔과 낡은 코트, 부서진 현금인출기 모두 프롬프트와 참조 이미지를 정확히 반영함.",
        "hard_violations": [],
        "physics": "강한 타격으로 인해 철판이 구부러지고 파편이 튀는 물리적 현상이 사실적으로 지탱 및 묘사됨."
       },
       {
        "label": "A",
        "direction": "왼쪽에서 뻗어나온 금속 주먹이 오른쪽의 현금인출기 앞면에 꽂혀 있음.",
        "built_space": "매장 내부와 인출기가 보이나, 배경의 유리문과 통로 배치가 기준 사진의 공간 구조와 일치하지 않음.",
        "entities": "찰리의 장갑판 팔과 코트, 인출기가 존재하나 카툰 렌더링 스타일로 묘사됨.",
        "hard_violations": [
         "실사(Photorealistic) 지시를 위반하고 셀 셰이딩 및 일러스트 스타일의 그래픽 오버레이 형태로 렌더링됨"
        ],
        "physics": "주먹이 기기에 박혀 몸체의 지지를 받으나, 금속이 찢어진 단면이 2D 그림처럼 비현실적으로 묘사됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지정된 클로즈업 구도, 찰리의 낡은 코트와 기계 팔, 현금인출기 파괴 묘사 및 매장의 배경을 사실적으로 완벽하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "지정된 실사 렌더링이 아닌 일러스트 스타일로 표현되었으며, 매장 내부의 공간 배치가 기준 사진과 다릅니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "금속 주먹이 현금인출기 앞면을 타격하여 내부로 깊게 뚫고 들어감.",
        "built_space": "기준 사진과 동일한 벽면 포스터와 인출기, 뒤쪽으로 좁게 보이는 진열대가 정확한 각도로 배치됨.",
        "entities": "찰리의 모래색 기계 팔과 낡은 코트, 부서진 현금인출기 모두 프롬프트와 참조 이미지를 정확히 반영함.",
        "hard_violations": [],
        "physics": "강한 타격으로 인해 철판이 구부러지고 파편이 튀는 물리적 현상이 사실적으로 지탱 및 묘사됨."
       },
       {
        "label": "A",
        "direction": "왼쪽에서 뻗어나온 금속 주먹이 오른쪽의 현금인출기 앞면에 꽂혀 있음.",
        "built_space": "매장 내부와 인출기가 보이나, 배경의 유리문과 통로 배치가 기준 사진의 공간 구조와 일치하지 않음.",
        "entities": "찰리의 장갑판 팔과 코트, 인출기가 존재하나 카툰 렌더링 스타일로 묘사됨.",
        "hard_violations": [
         "실사(Photorealistic) 지시를 위반하고 셀 셰이딩 및 일러스트 스타일의 그래픽 오버레이 형태로 렌더링됨"
        ],
        "physics": "주먹이 기기에 박혀 몸체의 지지를 받으나, 금속이 찢어진 단면이 2D 그림처럼 비현실적으로 묘사됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "금속 팔의 근접 구도는 맞지만 배경 안내문의 글자가 읽히며, 주먹 크기의 관통구보다 전면 패널 전체가 벌어진 파손으로 표현됐다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "주먹이 전면 관통구에 깊이 박힌 순간과 좁은 측면을 명확히 보여주고 장소도 잘 유지하지만, 배경과 몸통의 비중은 조금 크다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽에서 뻗은 금속 팔과 주먹이 오른쪽 현금인출기의 화면 옆 전면을 향한다. 주먹의 앞부분이 파손부와 겹치지만 손등과 접힌 손가락 상당 부분이 밖에 보여, 주먹 전체가 완전히 들어간 깊이는 다소 불명확하다. 얼굴과 시선은 화면 밖이다.",
        "built_space": "현금인출기 한 대의 화면, 하단 조작부와 배출구, 넓게 보이는 오른쪽 측면이 있다. 전면 패널이 본체에서 크게 벌어져 내부 구조가 노출된다. 뒤에는 출입문 일부와 상품 진열대의 좁은 조각이 보여 점포 맥락은 유지된다. 다만 요구한 좁은 인접 측면보다 검은 측면의 면적이 크고, 참고 장소의 국소적인 구멍보다 파손 범위가 훨씬 넓다.",
        "entities": "찰리의 일부 몸통, 어깨, 육중한 금속 팔과 주먹만 보이며 다른 인물은 없다. 각진 샌드 베이지 장갑판, 검은 관절과 낡은 외투는 참고 정체성에 부합한다. 얼굴, 모자와 가방은 이 구도에서 확인할 수 없다. 현금인출기와 상품 진열대는 식별되지만, 배경 벽 안내문에는 읽을 수 있는 한글이 남아 있다.",
        "hard_violations": [
         "배경 벽 안내문의 한글 문구가 읽을 수 있게 남아 있어, 어디에도 판독 가능한 글자를 두지 말라는 금지 조건을 위반한다."
        ],
        "physics": "주먹은 손목·전완·팔꿈치·상완으로 이어져 화면 왼쪽 몸통에 연결되므로 떠 있는 물체가 아니다. 파손된 금속판도 본체에 일부 붙어 있다. 작은 공중 파편은 타격으로 튄 것으로 해석 가능하다. 다만 주먹 주변뿐 아니라 긴 전면 패널까지 들려 있어 단일 주먹 크기 관통보다 광범위한 외장 분리에 가깝다."
       },
       {
        "label": "B",
        "direction": "왼쪽 어깨에서 오른쪽 아래로 뻗은 팔의 끝이 현금인출기 화면 오른쪽 전면의 구멍에 정확히 들어간다. 주먹 앞쪽은 기계 안으로 가려지고 손목에 가까운 부분만 외부에 남아 있어, 아직 빼지 않은 관통 순간이 명확하다. 얼굴과 눈은 보이지 않는다.",
        "built_space": "현금인출기 한 대의 전면과 오른쪽의 좁은 측면이 비스듬히 보인다. 화면 옆에 주먹을 둘러싼 파손구 하나가 있고 하단 패널은 대체로 유지된다. 뒤쪽 유리 출입문 한 개, 상부 창, 상품 선반 일부와 오른쪽 끝 담요 진열대가 참고 장소의 배치를 잘 보존한다. 다만 중앙 배경에 바닥과 창이 제법 넓게 드러나 충격부만을 압축한 근접 구도보다는 약간 여유롭다.",
        "entities": "찰리의 외투를 걸친 몸통 일부와 샌드 베이지 금속 팔, 거대한 주먹이 보인다. 육중한 기계 관절과 장갑판은 정체성에 부합하나 전완 장갑의 세부 형상은 참고와 다르다. 머리 대부분과 얼굴은 잘려 있어 마스크 얼굴은 평가 대상이 아니다. 현금인출기, 상품 및 담요 진열대가 있으며 추가 인물이나 명확히 읽히는 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "팔은 어깨에서 손목까지 연속적으로 연결되고 주먹은 찢어진 전면판 안에 박혀 있다. 따라서 팔과 주먹의 지지 및 접촉이 분명하다. 구멍 가장자리 금속 조각은 패널에 붙어 꺾여 있고, 떨어져 나온 작은 조각은 타격 직후의 비산으로 설명된다. 관통 동작을 깨는 부유나 불가능한 관절 연결은 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "금속 팔의 근접 구도는 맞지만 배경 안내문의 글자가 읽히며, 주먹 크기의 관통구보다 전면 패널 전체가 벌어진 파손으로 표현됐다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "주먹이 전면 관통구에 깊이 박힌 순간과 좁은 측면을 명확히 보여주고 장소도 잘 유지하지만, 배경과 몸통의 비중은 조금 크다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽에서 뻗은 금속 팔과 주먹이 오른쪽 현금인출기의 화면 옆 전면을 향한다. 주먹의 앞부분이 파손부와 겹치지만 손등과 접힌 손가락 상당 부분이 밖에 보여, 주먹 전체가 완전히 들어간 깊이는 다소 불명확하다. 얼굴과 시선은 화면 밖이다.",
        "built_space": "현금인출기 한 대의 화면, 하단 조작부와 배출구, 넓게 보이는 오른쪽 측면이 있다. 전면 패널이 본체에서 크게 벌어져 내부 구조가 노출된다. 뒤에는 출입문 일부와 상품 진열대의 좁은 조각이 보여 점포 맥락은 유지된다. 다만 요구한 좁은 인접 측면보다 검은 측면의 면적이 크고, 참고 장소의 국소적인 구멍보다 파손 범위가 훨씬 넓다.",
        "entities": "찰리의 일부 몸통, 어깨, 육중한 금속 팔과 주먹만 보이며 다른 인물은 없다. 각진 샌드 베이지 장갑판, 검은 관절과 낡은 외투는 참고 정체성에 부합한다. 얼굴, 모자와 가방은 이 구도에서 확인할 수 없다. 현금인출기와 상품 진열대는 식별되지만, 배경 벽 안내문에는 읽을 수 있는 한글이 남아 있다.",
        "hard_violations": [
         "배경 벽 안내문의 한글 문구가 읽을 수 있게 남아 있어, 어디에도 판독 가능한 글자를 두지 말라는 금지 조건을 위반한다."
        ],
        "physics": "주먹은 손목·전완·팔꿈치·상완으로 이어져 화면 왼쪽 몸통에 연결되므로 떠 있는 물체가 아니다. 파손된 금속판도 본체에 일부 붙어 있다. 작은 공중 파편은 타격으로 튄 것으로 해석 가능하다. 다만 주먹 주변뿐 아니라 긴 전면 패널까지 들려 있어 단일 주먹 크기 관통보다 광범위한 외장 분리에 가깝다."
       },
       {
        "label": "A",
        "direction": "왼쪽 어깨에서 오른쪽 아래로 뻗은 팔의 끝이 현금인출기 화면 오른쪽 전면의 구멍에 정확히 들어간다. 주먹 앞쪽은 기계 안으로 가려지고 손목에 가까운 부분만 외부에 남아 있어, 아직 빼지 않은 관통 순간이 명확하다. 얼굴과 눈은 보이지 않는다.",
        "built_space": "현금인출기 한 대의 전면과 오른쪽의 좁은 측면이 비스듬히 보인다. 화면 옆에 주먹을 둘러싼 파손구 하나가 있고 하단 패널은 대체로 유지된다. 뒤쪽 유리 출입문 한 개, 상부 창, 상품 선반 일부와 오른쪽 끝 담요 진열대가 참고 장소의 배치를 잘 보존한다. 다만 중앙 배경에 바닥과 창이 제법 넓게 드러나 충격부만을 압축한 근접 구도보다는 약간 여유롭다.",
        "entities": "찰리의 외투를 걸친 몸통 일부와 샌드 베이지 금속 팔, 거대한 주먹이 보인다. 육중한 기계 관절과 장갑판은 정체성에 부합하나 전완 장갑의 세부 형상은 참고와 다르다. 머리 대부분과 얼굴은 잘려 있어 마스크 얼굴은 평가 대상이 아니다. 현금인출기, 상품 및 담요 진열대가 있으며 추가 인물이나 명확히 읽히는 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "팔은 어깨에서 손목까지 연속적으로 연결되고 주먹은 찢어진 전면판 안에 박혀 있다. 따라서 팔과 주먹의 지지 및 접촉이 분명하다. 구멍 가장자리 금속 조각은 패널에 붙어 꺾여 있고, 떨어져 나온 작은 조각은 타격 직후의 비산으로 설명된다. 관통 동작을 깨는 부유나 불가능한 관절 연결은 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.571,
    "B": 1.5
   },
   "adjusted": {
    "A": 1.321,
    "B": 1.25
   },
   "violations": {
    "A": [
     "[gemini-pro] 실사(Photorealistic) 지시를 위반하고 셀 셰이딩 및 일러스트 스타일의 그래픽 오버레이 형태로 렌더링됨"
    ],
    "B": [
     "[gpt-high] 배경 벽 안내문의 한글 문구가 읽을 수 있게 남아 있어, 어디에도 판독 가능한 글자를 두지 말라는 금지 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1250,
   "A": 1321
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1250,
    "verdict_ko": "지정된 클로즈업 구도, 찰리의 낡은 코트와 기계 팔, 현금인출기 파괴 묘사 및 매장의 배경을 사실적으로 완벽하게 구현했습니다.  ★위반: [gpt-high] 배경 벽 안내문의 한글 문구가 읽을 수 있게 남아 있어, 어디에도 판독 가능한 글자를 두지 말라는 금지 조건을 위반한다."
   },
   {
    "label": "A",
    "score": 1321,
    "verdict_ko": "지정된 실사 렌더링이 아닌 일러스트 스타일로 표현되었으며, 매장 내부의 공간 배치가 기준 사진과 다릅니다.  ★위반: [gemini-pro] 실사(Photorealistic) 지시를 위반하고 셀 셰이딩 및 일러스트 스타일의 그래픽 오버레이 형태로 렌더링됨"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L195B01.png",
    "asset_id": "830dae63-3f71-4c94-9045-5c77a027e48e",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-91e9-7f6a-94a3-96b63a277952",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S41sh14__bgfirst_bg.png",
   "bg_asset_id": "6398fb6a-6f79-4ad1-90ee-719e7320869a",
   "bg_record_key": "S41sh14::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S41sh14::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:27:16.723638+00:00",
  "fingerprint": "bbdbe6e89cc2cdf88b421fdd433a0457e9e9293917c5eea4b33b5b5656ee113d",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S41sh14_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S41sh14_sel.png",
  "source_sha256": "df7bd611355071383fcb9694a15c562094ebde1b18b2f6b9884ef9468127c58e",
  "file": "S41sh14_cine.png",
  "staged_sha256": "b97a7a4a715d41545fe1531e8b6b3e641de6b7ca9bdf02f1e0370e189879c59f",
  "latency_ms": 10615
 },
 "S41sh22::signage": {
  "fp": "39e75a7ed2eaf5aa",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S41sh22": {
  "input_fingerprint": "be360ae99590e557",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 진열대 뒤에 웅크려 몸을 숨긴 채 바깥쪽을 긴장된 눈빛으로 엿보는 현우, 앰버, 찰리의 굳은 전신.\n\nLOCATION (lock): Behind a merchandise shelving unit inside the unattended shop, looking toward the entrance in daytime interior light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Display end leading toward the off-screen entrance in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: 몸을 숨긴 진열대 (Separates the crouching group from the entrance approach) — Its concealed side faces the camera; its end at screen right marks the route toward the entrance; used as Creates a visible boundary between safety and exposure without obscuring the group's full bodies.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the store's established ambient illumination and restrained contrast, allowing concealment to come from the display rather than an invented lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The ATM remains breached, with cash spilled from the opening and some already collected in a bag. The store entrance has been opened, and the scavenged supplies remain packed. 현우: He crouches behind a display shelf with the bag containing supplies and collected cash. His facial bruises and leg wound remain, and the contact card is still hidden in his shoe. 앰버: She crouches behind a display shelf, still wearing the scavenged bag and retaining the replacement shoes. 찰리: He is concealed behind a display shelf and now has the store blanket for covering his body, in addition to the earlier coat-and-hat disguise.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 진열대 뒤에 웅크려 몸을 숨긴 채 바깥쪽을 긴장된 눈빛으로 엿보는 현우, 앰버, 찰리의 굳은 전신.\n\nLOCATION (lock): Behind a merchandise shelving unit inside the unattended shop, looking toward the entrance in daytime interior light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Display end leading toward the off-screen entrance in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: 몸을 숨긴 진열대 (Separates the crouching group from the entrance approach) — Its concealed side faces the camera; its end at screen right marks the route toward the entrance; used as Creates a visible boundary between safety and exposure without obscuring the group's full bodies.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the store's established ambient illumination and restrained contrast, allowing concealment to come from the display rather than an invented lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The ATM remains breached, with cash spilled from the opening and some already collected in a bag. The store entrance has been opened, and the scavenged supplies remain packed. 현우: He crouches behind a display shelf with the bag containing supplies and collected cash. His facial bruises and leg wound remain, and the contact card is still hidden in his shoe. 앰버: She crouches behind a display shelf, still wearing the scavenged bag and retaining the replacement shoes. 찰리: He is concealed behind a display shelf and now has the store blanket for covering his body, in addition to the earlier coat-and-hat disguise.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 진열대 뒤에 웅크려 몸을 숨긴 채 바깥쪽을 긴장된 눈빛으로 엿보는 현우, 앰버, 찰리의 굳은 전신.\n\nLOCATION (lock): Behind a merchandise shelving unit inside the unattended shop, looking toward the entrance in daytime interior light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Display end leading toward the off-screen entrance in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: 몸을 숨긴 진열대 (Separates the crouching group from the entrance approach) — Its concealed side faces the camera; its end at screen right marks the route toward the entrance; used as Creates a visible boundary between safety and exposure without obscuring the group's full bodies.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the store's established ambient illumination and restrained contrast, allowing concealment to come from the display rather than an invented lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The ATM remains breached, with cash spilled from the opening and some already collected in a bag. The store entrance has been opened, and the scavenged supplies remain packed. 현우: He crouches behind a display shelf with the bag containing supplies and collected cash. His facial bruises and leg wound remain, and the contact card is still hidden in his shoe. 앰버: She crouches behind a display shelf, still wearing the scavenged bag and retaining the replacement shoes. 찰리: He is concealed behind a display shelf and now has the store blanket for covering his body, in addition to the earlier coat-and-hat disguise.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "세 인물이 모두 화면 우측의 출입구 방향을 향해 시선을 고정하고 있음.",
    "built_space": "지시대로 진열대가 좌측부터 중앙까지 이어지며 우측에 출입구 동선이 확보됨. 우측 벽에 이전 컷과 동일한 파손된 ATM이 존재함.",
    "entities": "현우(참조 일치, 현금 봉투 쥠), 앰버(참조 일치, 가방 착용), 찰리(로봇 참조 일치, 모자와 담요 착용) 모두 정확히 묘사됨.",
    "hard_violations": [
     "[gpt-high] 봉지 안 지폐에 액면 숫자와 문자가 판독 가능하게 보여, 읽을 수 있는 문자를 전면 금지한 조건을 위반한다."
    ],
    "physics": "세 인물 모두 바닥에 발과 다리를 딛고 안정적으로 웅크린 자세를 유지하고 있으며, 손에 든 물건들도 자연스럽게 지탱됨."
   },
   {
    "label": "B",
    "direction": "세 인물이 화면 좌측의 개방된 통로 방향을 향해 시선을 고정하고 있음.",
    "built_space": "진열대가 우측을 막고 좌측에 동선이 배치되어 프롬프트 지시와 반대 구조를 가짐. 매장 내 ATM은 보이지 않음.",
    "entities": "인물들의 외형은 참조와 일치하나, 현우가 미화 지폐를 들고 있고 찰리는 지시된 모자를 착용하지 않음.",
    "hard_violations": [
     "[gemini-pro] 화면 우측이어야 할 출입구 동선이 프롬프트 지시와 반대인 좌측으로 렌더링됨",
     "[gemini-pro] 이전 컷에서 반드시 유지되어야 할 파손된 ATM 환경이 완전히 누락됨"
    ],
    "physics": "인물들이 바닥에 무릎을 대고 웅크린 자세들이 물리적으로 자연스럽게 지탱됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "우측 출입구 동선과 이전 컷의 파손된 ATM 등 매장 구조를 정확히 구현했으며, 인물들의 소품과 자세 지침도 충실히 반영했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "진열대와 동선의 방향이 프롬프트와 반대이며, 이전 컷의 필수 요소인 파손된 ATM이 누락되어 지침을 크게 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "세 인물이 모두 화면 우측의 출입구 방향을 향해 시선을 고정하고 있음.",
        "built_space": "지시대로 진열대가 좌측부터 중앙까지 이어지며 우측에 출입구 동선이 확보됨. 우측 벽에 이전 컷과 동일한 파손된 ATM이 존재함.",
        "entities": "현우(참조 일치, 현금 봉투 쥠), 앰버(참조 일치, 가방 착용), 찰리(로봇 참조 일치, 모자와 담요 착용) 모두 정확히 묘사됨.",
        "hard_violations": [],
        "physics": "세 인물 모두 바닥에 발과 다리를 딛고 안정적으로 웅크린 자세를 유지하고 있으며, 손에 든 물건들도 자연스럽게 지탱됨."
       },
       {
        "label": "B",
        "direction": "세 인물이 화면 좌측의 개방된 통로 방향을 향해 시선을 고정하고 있음.",
        "built_space": "진열대가 우측을 막고 좌측에 동선이 배치되어 프롬프트 지시와 반대 구조를 가짐. 매장 내 ATM은 보이지 않음.",
        "entities": "인물들의 외형은 참조와 일치하나, 현우가 미화 지폐를 들고 있고 찰리는 지시된 모자를 착용하지 않음.",
        "hard_violations": [
         "화면 우측이어야 할 출입구 동선이 프롬프트 지시와 반대인 좌측으로 렌더링됨",
         "이전 컷에서 반드시 유지되어야 할 파손된 ATM 환경이 완전히 누락됨"
        ],
        "physics": "인물들이 바닥에 무릎을 대고 웅크린 자세들이 물리적으로 자연스럽게 지탱됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "우측 출입구 동선과 이전 컷의 파손된 ATM 등 매장 구조를 정확히 구현했으며, 인물들의 소품과 자세 지침도 충실히 반영했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "진열대와 동선의 방향이 프롬프트와 반대이며, 이전 컷의 필수 요소인 파손된 ATM이 누락되어 지침을 크게 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "세 인물이 모두 화면 우측의 출입구 방향을 향해 시선을 고정하고 있음.",
        "built_space": "지시대로 진열대가 좌측부터 중앙까지 이어지며 우측에 출입구 동선이 확보됨. 우측 벽에 이전 컷과 동일한 파손된 ATM이 존재함.",
        "entities": "현우(참조 일치, 현금 봉투 쥠), 앰버(참조 일치, 가방 착용), 찰리(로봇 참조 일치, 모자와 담요 착용) 모두 정확히 묘사됨.",
        "hard_violations": [],
        "physics": "세 인물 모두 바닥에 발과 다리를 딛고 안정적으로 웅크린 자세를 유지하고 있으며, 손에 든 물건들도 자연스럽게 지탱됨."
       },
       {
        "label": "B",
        "direction": "세 인물이 화면 좌측의 개방된 통로 방향을 향해 시선을 고정하고 있음.",
        "built_space": "진열대가 우측을 막고 좌측에 동선이 배치되어 프롬프트 지시와 반대 구조를 가짐. 매장 내 ATM은 보이지 않음.",
        "entities": "인물들의 외형은 참조와 일치하나, 현우가 미화 지폐를 들고 있고 찰리는 지시된 모자를 착용하지 않음.",
        "hard_violations": [
         "화면 우측이어야 할 출입구 동선이 프롬프트 지시와 반대인 좌측으로 렌더링됨",
         "이전 컷에서 반드시 유지되어야 할 파손된 ATM 환경이 완전히 누락됨"
        ],
        "physics": "인물들이 바닥에 무릎을 대고 웅크린 자세들이 물리적으로 자연스럽게 지탱됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "불투명 진열대 뒤에 숨는 관계는 성립하지만, 전신이 잘리고 진열대 끝과 현우의 시선이 지정된 오른쪽 출입구 방향과 반대다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "전신 와이드 구도와 찰리의 모자·담요는 더 충실하지만, 빈 진열대가 몸을 숨기지 못하고 지폐의 판독 가능한 숫자·문자가 무문자 조건을 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 진열대 왼쪽 가장자리 너머의 왼쪽 통로를 긴장하며 본다. 앰버는 화면 오른쪽 위를 보고, 찰리의 얼굴은 왼쪽 아래로 기울어 있다. 세 인물이 화면 중간 오른쪽의 프레임 밖 출입구를 함께 엿보는 방향 관계는 아니다. 무기나 이동 중인 물체는 없다.",
        "built_space": "왼쪽 상품 진열대, 중앙 뒤쪽 벽 진열대, 인물 옆의 큰 진열대가 보인다. 큰 진열대의 불투명 끝판은 세 인물 뒤에 있으며 통로와 몸 사이의 차폐물로 기능한다. 다만 노출 통로를 표시하는 끝은 인물들의 왼쪽에 있어 요구된 오른쪽 끝 배치와 반대다. 뒤쪽의 높은 창들과 유리문 일부, 회색 광택 타일, 천장 조명은 참고 장소와 대체로 이어진다. 현금인출기는 구도에 없으므로 파손 상태를 판단할 수 없다. 인물들이 오른쪽과 아래 가장자리에 몰려 전신이 온전히 들어오지 않는다.",
        "entities": "현우는 헝클어진 검은 머리와 앳된 동아시아계 남성 얼굴, 남색 상의, 볼의 멍을 갖추어 참고와 대체로 맞는다. 다리 상처는 확인되지 않고 신발 속 카드는 보이지 않는다. 양손에는 달러 계열로 보이는 지폐와 포장된 물품이 있으며, 물품과 현금을 함께 담은 가방이라는 상태는 명확하지 않다. 앰버는 금발의 어린 여자아이로 둥근 얼굴과 남색 상의, 배낭 끈이 보인다. 신발은 하단에서 일부 잘린다. 찰리는 모래색 장갑판과 흰 기계 얼굴, 큰 팔을 유지하고 천으로 몸을 덮었으나 모자는 보이지 않는다. 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "현우와 앰버는 무릎을 깊게 굽히고 바닥에 무릎이나 발을 대어 몸을 지탱한다. 현우의 지폐와 물품은 손으로 잡혀 있고 앰버의 배낭은 어깨끈으로 지지된다. 찰리의 천은 머리와 어깨에 걸려 아래로 드리워진다. 찰리 하체 일부가 가려지고 잘려 접지 전체는 확인할 수 없지만, 공중에 떠 있는 것으로 보이지는 않는다."
       },
       {
        "label": "B",
        "direction": "현우는 화면 오른쪽을, 앰버는 오른쪽 앞쪽을 바라본다. 찰리는 얼굴을 정면에서 약간 오른쪽으로 향한다. 두 사람의 경계 방향은 오른쪽 동선에 가깝지만, 실제로 보이는 열린 출입문은 현우 뒤쪽에 있어서 그 문 밖을 직접 엿보는 시선으로는 명확하지 않다. 무기나 이동 중인 물체는 없다.",
        "built_space": "카메라 앞에는 상품과 뒤판이 없는 금속 진열대 한 개가 있고, 뒤쪽과 왼쪽에는 상품 진열대가 있다. 오른쪽에는 파손된 회색 현금인출기 한 대와 그 뒤의 파란 기기 한 대가 있어 참고의 기기 배열과 대체로 맞는다. 유리 출입문 한 곳은 화면 중간 오른쪽에 직접 보이며, 프레임 밖이어야 한다는 조건과 다르다. 세 인물의 전신은 들어오지만 빈 진열대의 넓은 틈으로 몸이 노출되며, 특히 현우는 진열대 오른쪽 끝 밖에 나와 있다. 따라서 출입구 접근 방향과 인물 사이의 실질적인 은폐 경계가 성립하지 않는다.",
        "entities": "현우는 검은 헝클어진 머리, 어린 동아시아계 남성의 얼굴, 남색 상의와 볼의 멍을 갖추었다. 손에 지폐와 물품·현금이 든 투명 봉지를 들고 있지만 다리 상처는 드러나지 않는다. 앰버는 금발 여자아이와 남색 상의, 운동화로 참고에 대체로 맞으며, 가방은 몸 옆 바닥에 놓여 있어 계속 착용 중인 상태가 명확하지 않다. 찰리는 흰 기계 얼굴과 모래색 장갑, 긴 팔을 유지하고 챙 있는 모자와 회색 담요를 착용한다. 파손된 현금인출기 구멍과 바닥의 현금, 현금을 담은 가방이 보인다. 지폐에는 유로 계열의 도안과 판독 가능한 액면 숫자·문자가 나타난다. 추가 인물은 없다.",
        "hard_violations": [
         "봉지 안 지폐에 액면 숫자와 문자가 판독 가능하게 보여, 읽을 수 있는 문자를 전면 금지한 조건을 위반한다."
        ],
        "physics": "현우는 한쪽 무릎과 반대쪽 발로, 앰버는 굽힌 다리와 바닥에 닿은 신발로 몸을 지탱한다. 찰리는 굽힌 다리의 발과 바닥에 댄 큰 손으로 무게를 받는다. 담요는 찰리의 머리와 어깨에 걸려 있고 끝자락은 바닥에 닿는다. 투명 봉지는 현우가 손잡이를 잡고 있으며 바닥에도 닿고, 다른 가방들과 흩어진 지폐는 바닥에 놓여 있다. 현금인출기에서 나온 지폐는 파손 구멍 가장자리에 걸려 있어 근거 없이 떠 있는 물체는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "불투명 진열대 뒤에 숨는 관계는 성립하지만, 전신이 잘리고 진열대 끝과 현우의 시선이 지정된 오른쪽 출입구 방향과 반대다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "전신 와이드 구도와 찰리의 모자·담요는 더 충실하지만, 빈 진열대가 몸을 숨기지 못하고 지폐의 판독 가능한 숫자·문자가 무문자 조건을 위반한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 진열대 왼쪽 가장자리 너머의 왼쪽 통로를 긴장하며 본다. 앰버는 화면 오른쪽 위를 보고, 찰리의 얼굴은 왼쪽 아래로 기울어 있다. 세 인물이 화면 중간 오른쪽의 프레임 밖 출입구를 함께 엿보는 방향 관계는 아니다. 무기나 이동 중인 물체는 없다.",
        "built_space": "왼쪽 상품 진열대, 중앙 뒤쪽 벽 진열대, 인물 옆의 큰 진열대가 보인다. 큰 진열대의 불투명 끝판은 세 인물 뒤에 있으며 통로와 몸 사이의 차폐물로 기능한다. 다만 노출 통로를 표시하는 끝은 인물들의 왼쪽에 있어 요구된 오른쪽 끝 배치와 반대다. 뒤쪽의 높은 창들과 유리문 일부, 회색 광택 타일, 천장 조명은 참고 장소와 대체로 이어진다. 현금인출기는 구도에 없으므로 파손 상태를 판단할 수 없다. 인물들이 오른쪽과 아래 가장자리에 몰려 전신이 온전히 들어오지 않는다.",
        "entities": "현우는 헝클어진 검은 머리와 앳된 동아시아계 남성 얼굴, 남색 상의, 볼의 멍을 갖추어 참고와 대체로 맞는다. 다리 상처는 확인되지 않고 신발 속 카드는 보이지 않는다. 양손에는 달러 계열로 보이는 지폐와 포장된 물품이 있으며, 물품과 현금을 함께 담은 가방이라는 상태는 명확하지 않다. 앰버는 금발의 어린 여자아이로 둥근 얼굴과 남색 상의, 배낭 끈이 보인다. 신발은 하단에서 일부 잘린다. 찰리는 모래색 장갑판과 흰 기계 얼굴, 큰 팔을 유지하고 천으로 몸을 덮었으나 모자는 보이지 않는다. 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "현우와 앰버는 무릎을 깊게 굽히고 바닥에 무릎이나 발을 대어 몸을 지탱한다. 현우의 지폐와 물품은 손으로 잡혀 있고 앰버의 배낭은 어깨끈으로 지지된다. 찰리의 천은 머리와 어깨에 걸려 아래로 드리워진다. 찰리 하체 일부가 가려지고 잘려 접지 전체는 확인할 수 없지만, 공중에 떠 있는 것으로 보이지는 않는다."
       },
       {
        "label": "A",
        "direction": "현우는 화면 오른쪽을, 앰버는 오른쪽 앞쪽을 바라본다. 찰리는 얼굴을 정면에서 약간 오른쪽으로 향한다. 두 사람의 경계 방향은 오른쪽 동선에 가깝지만, 실제로 보이는 열린 출입문은 현우 뒤쪽에 있어서 그 문 밖을 직접 엿보는 시선으로는 명확하지 않다. 무기나 이동 중인 물체는 없다.",
        "built_space": "카메라 앞에는 상품과 뒤판이 없는 금속 진열대 한 개가 있고, 뒤쪽과 왼쪽에는 상품 진열대가 있다. 오른쪽에는 파손된 회색 현금인출기 한 대와 그 뒤의 파란 기기 한 대가 있어 참고의 기기 배열과 대체로 맞는다. 유리 출입문 한 곳은 화면 중간 오른쪽에 직접 보이며, 프레임 밖이어야 한다는 조건과 다르다. 세 인물의 전신은 들어오지만 빈 진열대의 넓은 틈으로 몸이 노출되며, 특히 현우는 진열대 오른쪽 끝 밖에 나와 있다. 따라서 출입구 접근 방향과 인물 사이의 실질적인 은폐 경계가 성립하지 않는다.",
        "entities": "현우는 검은 헝클어진 머리, 어린 동아시아계 남성의 얼굴, 남색 상의와 볼의 멍을 갖추었다. 손에 지폐와 물품·현금이 든 투명 봉지를 들고 있지만 다리 상처는 드러나지 않는다. 앰버는 금발 여자아이와 남색 상의, 운동화로 참고에 대체로 맞으며, 가방은 몸 옆 바닥에 놓여 있어 계속 착용 중인 상태가 명확하지 않다. 찰리는 흰 기계 얼굴과 모래색 장갑, 긴 팔을 유지하고 챙 있는 모자와 회색 담요를 착용한다. 파손된 현금인출기 구멍과 바닥의 현금, 현금을 담은 가방이 보인다. 지폐에는 유로 계열의 도안과 판독 가능한 액면 숫자·문자가 나타난다. 추가 인물은 없다.",
        "hard_violations": [
         "봉지 안 지폐에 액면 숫자와 문자가 판독 가능하게 보여, 읽을 수 있는 문자를 전면 금지한 조건을 위반한다."
        ],
        "physics": "현우는 한쪽 무릎과 반대쪽 발로, 앰버는 굽힌 다리와 바닥에 닿은 신발로 몸을 지탱한다. 찰리는 굽힌 다리의 발과 바닥에 댄 큰 손으로 무게를 받는다. 담요는 찰리의 머리와 어깨에 걸려 있고 끝자락은 바닥에 닿는다. 투명 봉지는 현우가 손잡이를 잡고 있으며 바닥에도 닿고, 다른 가방들과 흩어진 지폐는 바닥에 놓여 있다. 현금인출기에서 나온 지폐는 파손 구멍 가장자리에 걸려 있어 근거 없이 떠 있는 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.5,
    "B": 1.429
   },
   "adjusted": {
    "A": 1.25,
    "B": 1.179
   },
   "violations": {
    "B": [
     "[gemini-pro] 화면 우측이어야 할 출입구 동선이 프롬프트 지시와 반대인 좌측으로 렌더링됨",
     "[gemini-pro] 이전 컷에서 반드시 유지되어야 할 파손된 ATM 환경이 완전히 누락됨"
    ],
    "A": [
     "[gpt-high] 봉지 안 지폐에 액면 숫자와 문자가 판독 가능하게 보여, 읽을 수 있는 문자를 전면 금지한 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1250,
   "B": 1179
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1250,
    "verdict_ko": "우측 출입구 동선과 이전 컷의 파손된 ATM 등 매장 구조를 정확히 구현했으며, 인물들의 소품과 자세 지침도 충실히 반영했습니다.  ★위반: [gpt-high] 봉지 안 지폐에 액면 숫자와 문자가 판독 가능하게 보여, 읽을 수 있는 문자를 전면 금지한 조건을 위반한다."
   },
   {
    "label": "B",
    "score": 1179,
    "verdict_ko": "진열대와 동선의 방향이 프롬프트와 반대이며, 이전 컷의 필수 요소인 파손된 ATM이 누락되어 지침을 크게 위반했습니다.  ★위반: [gemini-pro] 화면 우측이어야 할 출입구 동선이 프롬프트 지시와 반대인 좌측으로 렌더링됨 / [gemini-pro] 이전 컷에서 반드시 유지되어야 할 파손된 ATM 환경이 완전히 누락됨"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S41sh14_sel.png",
    "asset_id": "8a6c9d28-eb19-4b6e-9403-c3983dbfc2a6",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-9549-7c90-b90d-f199f18d4951",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S41sh14"
  }
 },
 "S41sh22::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:29:11.236902+00:00",
  "fingerprint": "2ebfca887d0752e1b2cd0834156619f32ca9b511a749f52b8273bab5f78c5593",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S41sh22_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S41sh22_sel.png",
  "source_sha256": "4a2dda3ed858c5b7045014a5ba0624802bd021d6d258cb377b77c1405a134819",
  "file": "S41sh22_cine.png",
  "staged_sha256": "0bd5f094cff10bad980f5f9a34a2bce9e6439a6ac21fc83dd7ab4d36d24c6072",
  "latency_ms": 10597
 },
 "S41sh26::signage": {
  "fp": "27d5fb0abb4823ed",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S41sh26": {
  "input_fingerprint": "954bd8fca05a454d",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 서로를 꽉 껴안은 채 활짝 웃는 앰버와 라울의 환한 상체.\n\nLOCATION (lock): In the merchandise aisle inside the unattended shop, near the shelving used as cover, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 점포 진열대 (Remains behind the pair after 앰버 has left concealment) — An oblique end section remains peripheral behind 앰버; used as A subdued continuity marker connecting the embrace to the preceding hiding place.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the established store illumination unchanged, rendering both smiling faces with gentle tonal separation rather than introducing a warmer light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 앰버 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The breached ATM and spilled cash remain unchanged. Food, drinks, the disposable phone, map and COPD medicine remain packed in the scavenged bags. 앰버: She is out of hiding and visibly delighted, still wearing the scavenged bag and retaining the replacement shoes. 라울: He stands inside the store, delighted by the reunion, with his ponytail unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 서로를 꽉 껴안은 채 활짝 웃는 앰버와 라울의 환한 상체.\n\nLOCATION (lock): In the merchandise aisle inside the unattended shop, near the shelving used as cover, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 점포 진열대 (Remains behind the pair after 앰버 has left concealment) — An oblique end section remains peripheral behind 앰버; used as A subdued continuity marker connecting the embrace to the preceding hiding place.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the established store illumination unchanged, rendering both smiling faces with gentle tonal separation rather than introducing a warmer light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 앰버 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The breached ATM and spilled cash remain unchanged. Food, drinks, the disposable phone, map and COPD medicine remain packed in the scavenged bags. 앰버: She is out of hiding and visibly delighted, still wearing the scavenged bag and retaining the replacement shoes. 라울: He stands inside the store, delighted by the reunion, with his ponytail unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 서로를 꽉 껴안은 채 활짝 웃는 앰버와 라울의 환한 상체.\n\nLOCATION (lock): In the merchandise aisle inside the unattended shop, near the shelving used as cover, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 점포 진열대 (Remains behind the pair after 앰버 has left concealment) — An oblique end section remains peripheral behind 앰버; used as A subdued continuity marker connecting the embrace to the preceding hiding place.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the established store illumination unchanged, rendering both smiling faces with gentle tonal separation rather than introducing a warmer light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 앰버 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The breached ATM and spilled cash remain unchanged. Food, drinks, the disposable phone, map and COPD medicine remain packed in the scavenged bags. 앰버: She is out of hiding and visibly delighted, still wearing the scavenged bag and retaining the replacement shoes. 라울: He stands inside the store, delighted by the reunion, with his ponytail unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "B",
    "direction": "앰버는 카메라 쪽을 보며 웃고, 라울은 앰버 얼굴 가까운 왼쪽 아래로 시선을 둔다. 서로 몸을 향하고 양팔로 상대의 어깨와 등을 감싼다. 조준하거나 사용하는 소품은 없다.",
    "built_space": "왼쪽 벽면 상품 진열대 한 줄, 중앙 독립형 금속 진열대 한 줄, 오른쪽 현금인출기 두 대가 보인다. 오른쪽 기계 한 대에는 파손 구멍과 지폐가 있고 그 아래 가방이 놓여 있다. 뒤쪽 높은 창과 출입구, 회색 타일은 이전 장소와 연결된다. 두 아이는 진열대 옆 통로에 서 있지만, 중앙 진열대의 넓은 선반이 왼쪽 전경을 크게 차지해 작고 주변적인 연속성 표지라는 지시와 다르다.",
    "entities": "등장인물은 아이 두 명뿐이다. 앰버는 금발, 둥근 얼굴, 남색 반팔과 낡은 카키색 어깨가방을 유지한다. 라울은 갈색 피부의 어린 남자아이로, 뒤로 묶은 검은 머리와 남색 반팔이 참고 이미지에 부합한다. 혼혈 배경 자체는 외모만으로 확정할 수 없지만 두 인물의 외형은 참고와 대체로 맞는다. 두 사람 모두 이를 드러내며 크게 웃는다. 신발과 가방 속 물품은 프레임 밖이거나 가려져 있다. 읽을 수 있는 글자는 보이지 않는다.",
    "hard_violations": [],
    "physics": "두 몸통은 세워져 화면 아래로 이어지며 공중에 뜬 징후가 없다. 발은 잘렸지만 서서 포옹하는 자세와 모순되지 않는다. 손은 상대 어깨와 등에 닿고 팔의 연결도 자연스럽다. 앰버의 가방은 어깨끈에 매달려 있고 상품은 선반에, 현금인출기 옆 가방은 바닥에 놓여 있다."
   },
   {
    "label": "A",
    "direction": "앰버는 카메라 쪽을 보며 활짝 웃고, 라울은 앰버 쪽 아래로 눈길을 둔다. 두 아이는 볼과 상체를 붙이고 서로의 등과 어깨를 양팔로 감싸므로 재회의 포옹 대상이 분명하다. 방향을 확인해야 할 휴대 도구는 없다.",
    "built_space": "왼쪽 벽면 진열대 한 줄과 두 아이 뒤의 독립형 금속 진열대 한 줄, 오른쪽 가장자리에 현금인출기 두 대가 보인다. 가장 오른쪽 기계의 파손 부위와 지폐, 바닥의 가방과 흩어진 돈 일부도 보인다. 높은 창, 유리 출입구, 회색 타일과 중성적인 낮 조명이 이전 장소를 유지한다. 두 아이는 통로를 차지하고 진열대는 뒤에 남아 전경으로 돌출되지 않는다. 다만 보이는 진열대 끝은 앰버 뒤보다는 라울의 오른쪽 뒤에 가깝다.",
    "entities": "금발의 어린 앰버와 뒤로 묶은 검은 머리의 어린 라울, 두 사람만 등장한다. 앰버의 둥근 얼굴과 남색 반팔, 낡은 카키색 어깨가방이 유지되고 라울의 피부색, 얼굴 윤곽, 남색 반팔과 목 뒤 꽁지머리도 참고와 부합한다. 두 사람 모두 환하게 웃는다. 인종적 배경을 사진만으로 확정할 수는 없으나 참고 인물의 외형을 잘 따른다. 신발과 포장된 물품은 보이지 않아 판단 대상이 아니다. 판독 가능한 글자는 없다.",
    "hard_violations": [],
    "physics": "두 아이의 상체는 선 자세로 화면 아래까지 이어지며, 몸을 서로 기대고 팔로 밀착시키는 동작이 가능하다. 보이는 손들은 상대의 등과 옆구리에 닿아 있고 비정상적인 추가 팔다리는 없다. 발은 프레임 밖이며 부유를 시사하는 자세는 아니다. 앰버의 가방은 어깨끈으로 지지되고 진열 상품과 바닥 소품도 각각 선반과 바닥에 지지된다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": null,
     "normalized": null,
     "ok": false
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "포옹과 환한 웃음, 인물 외형은 충실하지만 진열대가 전경까지 크게 돌출되어 상체 중심 미디엄 숏과 주변부 배경 지시를 약화한다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "두 아이의 웃는 상체와 밀착된 포옹을 중심에 둔 미디엄 숏이 더 정확하며, 다만 진열대 끝부분이 앰버보다는 라울 뒤에 남는다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버는 카메라 쪽을 보며 웃고, 라울은 앰버 얼굴 가까운 왼쪽 아래로 시선을 둔다. 서로 몸을 향하고 양팔로 상대의 어깨와 등을 감싼다. 조준하거나 사용하는 소품은 없다.",
        "built_space": "왼쪽 벽면 상품 진열대 한 줄, 중앙 독립형 금속 진열대 한 줄, 오른쪽 현금인출기 두 대가 보인다. 오른쪽 기계 한 대에는 파손 구멍과 지폐가 있고 그 아래 가방이 놓여 있다. 뒤쪽 높은 창과 출입구, 회색 타일은 이전 장소와 연결된다. 두 아이는 진열대 옆 통로에 서 있지만, 중앙 진열대의 넓은 선반이 왼쪽 전경을 크게 차지해 작고 주변적인 연속성 표지라는 지시와 다르다.",
        "entities": "등장인물은 아이 두 명뿐이다. 앰버는 금발, 둥근 얼굴, 남색 반팔과 낡은 카키색 어깨가방을 유지한다. 라울은 갈색 피부의 어린 남자아이로, 뒤로 묶은 검은 머리와 남색 반팔이 참고 이미지에 부합한다. 혼혈 배경 자체는 외모만으로 확정할 수 없지만 두 인물의 외형은 참고와 대체로 맞는다. 두 사람 모두 이를 드러내며 크게 웃는다. 신발과 가방 속 물품은 프레임 밖이거나 가려져 있다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "두 몸통은 세워져 화면 아래로 이어지며 공중에 뜬 징후가 없다. 발은 잘렸지만 서서 포옹하는 자세와 모순되지 않는다. 손은 상대 어깨와 등에 닿고 팔의 연결도 자연스럽다. 앰버의 가방은 어깨끈에 매달려 있고 상품은 선반에, 현금인출기 옆 가방은 바닥에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "앰버는 카메라 쪽을 보며 활짝 웃고, 라울은 앰버 쪽 아래로 눈길을 둔다. 두 아이는 볼과 상체를 붙이고 서로의 등과 어깨를 양팔로 감싸므로 재회의 포옹 대상이 분명하다. 방향을 확인해야 할 휴대 도구는 없다.",
        "built_space": "왼쪽 벽면 진열대 한 줄과 두 아이 뒤의 독립형 금속 진열대 한 줄, 오른쪽 가장자리에 현금인출기 두 대가 보인다. 가장 오른쪽 기계의 파손 부위와 지폐, 바닥의 가방과 흩어진 돈 일부도 보인다. 높은 창, 유리 출입구, 회색 타일과 중성적인 낮 조명이 이전 장소를 유지한다. 두 아이는 통로를 차지하고 진열대는 뒤에 남아 전경으로 돌출되지 않는다. 다만 보이는 진열대 끝은 앰버 뒤보다는 라울의 오른쪽 뒤에 가깝다.",
        "entities": "금발의 어린 앰버와 뒤로 묶은 검은 머리의 어린 라울, 두 사람만 등장한다. 앰버의 둥근 얼굴과 남색 반팔, 낡은 카키색 어깨가방이 유지되고 라울의 피부색, 얼굴 윤곽, 남색 반팔과 목 뒤 꽁지머리도 참고와 부합한다. 두 사람 모두 환하게 웃는다. 인종적 배경을 사진만으로 확정할 수는 없으나 참고 인물의 외형을 잘 따른다. 신발과 포장된 물품은 보이지 않아 판단 대상이 아니다. 판독 가능한 글자는 없다.",
        "hard_violations": [],
        "physics": "두 아이의 상체는 선 자세로 화면 아래까지 이어지며, 몸을 서로 기대고 팔로 밀착시키는 동작이 가능하다. 보이는 손들은 상대의 등과 옆구리에 닿아 있고 비정상적인 추가 팔다리는 없다. 발은 프레임 밖이며 부유를 시사하는 자세는 아니다. 앰버의 가방은 어깨끈으로 지지되고 진열 상품과 바닥 소품도 각각 선반과 바닥에 지지된다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "포옹과 환한 웃음, 인물 외형은 충실하지만 진열대가 전경까지 크게 돌출되어 상체 중심 미디엄 숏과 주변부 배경 지시를 약화한다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "두 아이의 웃는 상체와 밀착된 포옹을 중심에 둔 미디엄 숏이 더 정확하며, 다만 진열대 끝부분이 앰버보다는 라울 뒤에 남는다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "앰버는 카메라 쪽을 보며 웃고, 라울은 앰버 얼굴 가까운 왼쪽 아래로 시선을 둔다. 서로 몸을 향하고 양팔로 상대의 어깨와 등을 감싼다. 조준하거나 사용하는 소품은 없다.",
        "built_space": "왼쪽 벽면 상품 진열대 한 줄, 중앙 독립형 금속 진열대 한 줄, 오른쪽 현금인출기 두 대가 보인다. 오른쪽 기계 한 대에는 파손 구멍과 지폐가 있고 그 아래 가방이 놓여 있다. 뒤쪽 높은 창과 출입구, 회색 타일은 이전 장소와 연결된다. 두 아이는 진열대 옆 통로에 서 있지만, 중앙 진열대의 넓은 선반이 왼쪽 전경을 크게 차지해 작고 주변적인 연속성 표지라는 지시와 다르다.",
        "entities": "등장인물은 아이 두 명뿐이다. 앰버는 금발, 둥근 얼굴, 남색 반팔과 낡은 카키색 어깨가방을 유지한다. 라울은 갈색 피부의 어린 남자아이로, 뒤로 묶은 검은 머리와 남색 반팔이 참고 이미지에 부합한다. 혼혈 배경 자체는 외모만으로 확정할 수 없지만 두 인물의 외형은 참고와 대체로 맞는다. 두 사람 모두 이를 드러내며 크게 웃는다. 신발과 가방 속 물품은 프레임 밖이거나 가려져 있다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "두 몸통은 세워져 화면 아래로 이어지며 공중에 뜬 징후가 없다. 발은 잘렸지만 서서 포옹하는 자세와 모순되지 않는다. 손은 상대 어깨와 등에 닿고 팔의 연결도 자연스럽다. 앰버의 가방은 어깨끈에 매달려 있고 상품은 선반에, 현금인출기 옆 가방은 바닥에 놓여 있다."
       },
       {
        "label": "A",
        "direction": "앰버는 카메라 쪽을 보며 활짝 웃고, 라울은 앰버 쪽 아래로 눈길을 둔다. 두 아이는 볼과 상체를 붙이고 서로의 등과 어깨를 양팔로 감싸므로 재회의 포옹 대상이 분명하다. 방향을 확인해야 할 휴대 도구는 없다.",
        "built_space": "왼쪽 벽면 진열대 한 줄과 두 아이 뒤의 독립형 금속 진열대 한 줄, 오른쪽 가장자리에 현금인출기 두 대가 보인다. 가장 오른쪽 기계의 파손 부위와 지폐, 바닥의 가방과 흩어진 돈 일부도 보인다. 높은 창, 유리 출입구, 회색 타일과 중성적인 낮 조명이 이전 장소를 유지한다. 두 아이는 통로를 차지하고 진열대는 뒤에 남아 전경으로 돌출되지 않는다. 다만 보이는 진열대 끝은 앰버 뒤보다는 라울의 오른쪽 뒤에 가깝다.",
        "entities": "금발의 어린 앰버와 뒤로 묶은 검은 머리의 어린 라울, 두 사람만 등장한다. 앰버의 둥근 얼굴과 남색 반팔, 낡은 카키색 어깨가방이 유지되고 라울의 피부색, 얼굴 윤곽, 남색 반팔과 목 뒤 꽁지머리도 참고와 부합한다. 두 사람 모두 환하게 웃는다. 인종적 배경을 사진만으로 확정할 수는 없으나 참고 인물의 외형을 잘 따른다. 신발과 포장된 물품은 보이지 않아 판단 대상이 아니다. 판독 가능한 글자는 없다.",
        "hard_violations": [],
        "physics": "두 아이의 상체는 선 자세로 화면 아래까지 이어지며, 몸을 서로 기대고 팔로 밀착시키는 동작이 가능하다. 보이는 손들은 상대의 등과 옆구리에 닿아 있고 비정상적인 추가 팔다리는 없다. 발은 프레임 밖이며 부유를 시사하는 자세는 아니다. 앰버의 가방은 어깨끈으로 지지되고 진열 상품과 바닥 소품도 각각 선반과 바닥에 지지된다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gemini-pro"
   ],
   "route": "single_reverse"
  },
  "totals": {
   "B": 7,
   "A": 9
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 7,
    "verdict_ko": "포옹과 환한 웃음, 인물 외형은 충실하지만 진열대가 전경까지 크게 돌출되어 상체 중심 미디엄 숏과 주변부 배경 지시를 약화한다."
   },
   {
    "label": "A",
    "score": 9,
    "verdict_ko": "두 아이의 웃는 상체와 밀착된 포옹을 중심에 둔 미디엄 숏이 더 정확하며, 다만 진열대 끝부분이 앰버보다는 라울 뒤에 남는다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 앰버 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S41sh22_sel.png",
    "asset_id": "80c1fc28-e5fa-4594-b0dd-ea822b418ec9",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163202>",
    "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-9717-7625-9fbe-44cd1d6d1316",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S41sh22"
  }
 },
 "S41sh26::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:30:09.520363+00:00",
  "fingerprint": "6ff39f554d603cb8b7cd082f949864b736ed1dcb00ddc2cf62cd0a8536df25ab",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S41sh26_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S41sh26_sel.png",
  "source_sha256": "5865a8636efd6685654e185be4b84af94eacee815265c959b231b624085b804e",
  "file": "S41sh26_cine.png",
  "staged_sha256": "e6f17b17514593ee05f0149cbb67c235778f4eb15837efcca1f48a2e21c05681",
  "latency_ms": 9653
 },
 "S42sh2::signage": {
  "fp": "bdbf3f72217f9441",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S42sh2": {
  "input_fingerprint": "202b35ddd6f2ce5b",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 보디백 안, 핏기 없이 죽은 구도환의 창백한 얼굴을 내려다보는 시점 쇼트.\n\nLOCATION (lock): At ground level near the refugee settlement entrance at night, looking into an opened body bag. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 보디백 (Unzipped around 구도환's exposed face) — The opened interior and near rim are seen from above; used as Frames the face with evidence of death while remaining subordinate to the human subject.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued nighttime ambient illumination sufficient to reveal 구도환's bloodless pallor, with no dreamlike distortion or invented visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Gu Dohwan's dead body is enclosed in a body bag with the zipper opened to expose his face for inspection from above. His torso and limbs remain concealed, and the source does not specify his head's turn or the surface supporting the bag.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The body bag's zipper is open at the face end. It is night at the refugee-settlement entrance. 구도환: He is dead and is enclosed in a body bag with the zipper opened to expose his face for inspection from above.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 보디백 안, 핏기 없이 죽은 구도환의 창백한 얼굴을 내려다보는 시점 쇼트.\n\nLOCATION (lock): At ground level near the refugee settlement entrance at night, looking into an opened body bag. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 보디백 (Unzipped around 구도환's exposed face) — The opened interior and near rim are seen from above; used as Frames the face with evidence of death while remaining subordinate to the human subject.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued nighttime ambient illumination sufficient to reveal 구도환's bloodless pallor, with no dreamlike distortion or invented visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Gu Dohwan's dead body is enclosed in a body bag with the zipper opened to expose his face for inspection from above. His torso and limbs remain concealed, and the source does not specify his head's turn or the surface supporting the bag.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The body bag's zipper is open at the face end. It is night at the refugee-settlement entrance. 구도환: He is dead and is enclosed in a body bag with the zipper opened to expose his face for inspection from above.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 보디백 안, 핏기 없이 죽은 구도환의 창백한 얼굴을 내려다보는 시점 쇼트.\n\nLOCATION (lock): At ground level near the refugee settlement entrance at night, looking into an opened body bag. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 보디백 (Unzipped around 구도환's exposed face) — The opened interior and near rim are seen from above; used as Frames the face with evidence of death while remaining subordinate to the human subject.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued nighttime ambient illumination sufficient to reveal 구도환's bloodless pallor, with no dreamlike distortion or invented visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Gu Dohwan's dead body is enclosed in a body bag with the zipper opened to expose his face for inspection from above. His torso and limbs remain concealed, and the source does not specify his head's turn or the surface supporting the bag.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The body bag's zipper is open at the face end. It is night at the refugee-settlement entrance. 구도환: He is dead and is enclosed in a body bag with the zipper opened to expose his face for inspection from above.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 구도환 (한국인 남성, 50대, 중년의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S42sh2__bgfirst_bg.png",
     "asset_id": "0428d63b-de1f-42c5-ad4e-e5abb4f95632",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S42sh2.png",
     "asset_id": "ed925a67-f2a3-41f5-9413-65abf6516887",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1311816>",
     "asset_id": "623f0421-67dc-4592-a009-148d8e957276",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_camp_gate_b39179.png",
     "asset_id": "83651a7a-615d-46c7-ba87-2459a6bc8bf2",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1311816>",
     "asset_id": "623f0421-67dc-4592-a009-148d8e957276",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "낮은 위치에서 얼굴을 올려다보며 배경의 수용소 게이트를 향함.",
    "built_space": "수용소 입구 구조물이 배경에 있으나, 원근감 오류로 보디백이 거대한 동굴처럼 보임.",
    "entities": "구도환의 얼굴과 창백한 피부 톤이 묘사됨. 열린 보디백이 있음.",
    "hard_violations": [],
    "physics": "보디백의 천이 중력을 무시하고 뻣뻣한 아치 형태를 유지함."
   },
   {
    "label": "B",
    "direction": "약간 높은 위치에서 얼굴을 내려다보며 배경의 게이트를 향함.",
    "built_space": "수용소 입구 장벽과 천막이 제자리에 있으며 보디백이 젖은 바닥에 놓임.",
    "entities": "구도환의 창백하고 죽은 얼굴이 잘 묘사됨. 지퍼가 열린 보디백이 얼굴을 감쌈.",
    "hard_violations": [],
    "physics": "신체와 보디백이 중력에 맞게 바닥에 자연스럽게 밀착되어 처져 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지시문이 요구한 내려다보는 앵글과 핏기 없는 죽은 얼굴을 사실적인 보디백 질감과 함께 잘 구현했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "카메라가 피사체를 오히려 올려다보는 구도를 취했으며, 보디백의 형태와 재질이 비현실적으로 뻣뻣합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "낮은 위치에서 얼굴을 올려다보며 배경의 수용소 게이트를 향함.",
        "built_space": "수용소 입구 구조물이 배경에 있으나, 원근감 오류로 보디백이 거대한 동굴처럼 보임.",
        "entities": "구도환의 얼굴과 창백한 피부 톤이 묘사됨. 열린 보디백이 있음.",
        "hard_violations": [],
        "physics": "보디백의 천이 중력을 무시하고 뻣뻣한 아치 형태를 유지함."
       },
       {
        "label": "B",
        "direction": "약간 높은 위치에서 얼굴을 내려다보며 배경의 게이트를 향함.",
        "built_space": "수용소 입구 장벽과 천막이 제자리에 있으며 보디백이 젖은 바닥에 놓임.",
        "entities": "구도환의 창백하고 죽은 얼굴이 잘 묘사됨. 지퍼가 열린 보디백이 얼굴을 감쌈.",
        "hard_violations": [],
        "physics": "신체와 보디백이 중력에 맞게 바닥에 자연스럽게 밀착되어 처져 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지시문이 요구한 내려다보는 앵글과 핏기 없는 죽은 얼굴을 사실적인 보디백 질감과 함께 잘 구현했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "카메라가 피사체를 오히려 올려다보는 구도를 취했으며, 보디백의 형태와 재질이 비현실적으로 뻣뻣합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "낮은 위치에서 얼굴을 올려다보며 배경의 수용소 게이트를 향함.",
        "built_space": "수용소 입구 구조물이 배경에 있으나, 원근감 오류로 보디백이 거대한 동굴처럼 보임.",
        "entities": "구도환의 얼굴과 창백한 피부 톤이 묘사됨. 열린 보디백이 있음.",
        "hard_violations": [],
        "physics": "보디백의 천이 중력을 무시하고 뻣뻣한 아치 형태를 유지함."
       },
       {
        "label": "B",
        "direction": "약간 높은 위치에서 얼굴을 내려다보며 배경의 게이트를 향함.",
        "built_space": "수용소 입구 장벽과 천막이 제자리에 있으며 보디백이 젖은 바닥에 놓임.",
        "entities": "구도환의 창백하고 죽은 얼굴이 잘 묘사됨. 지퍼가 열린 보디백이 얼굴을 감쌈.",
        "hard_violations": [],
        "physics": "신체와 보디백이 중력에 맞게 바닥에 자연스럽게 밀착되어 처져 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "창백한 사망자와 얼굴 쪽이 열린 보디백은 맞지만, 얼굴보다 보디백과 주변 도로가 크게 보여 요구한 얼굴 클로즈업과 위에서 내려다보는 시점이 약하다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "얼굴을 크게 잡아 클로즈업과 핏기 없는 사망 상태를 더 충실히 구현하지만, 카메라가 턱 쪽에서 낮게 바라봐 명시된 내려다보는 시점에는 미달한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "남성은 눈을 감고 얼굴을 위로 향하고 있어 시선의 대상은 없다. 카메라는 몸통 쪽에서 얼굴을 비스듬히 바라보며, 보디백 내부는 보이지만 얼굴 바로 위에서 내려다보는 시점보다 낮다. 무기나 방향을 가진 동작은 없다.",
        "built_space": "젖은 도로 위 보디백 한 개 뒤로 중앙 철문 한 조, 양옆 출입구 기둥 두 개, 왼쪽 콘크리트 벽, 오른쪽 철망과 천막이 보인다. 출입구 상부는 프레임 밖으로 잘린다. 참고 장소의 재료와 배치는 대체로 맞고 젖은 지면의 빛 반사도 가능하다. 다만 도로와 보디백 몸통 부분이 넓게 들어와 얼굴의 화면 점유율이 작다.",
        "entities": "보이는 인물은 검은 머리의 중년 동아시아계 남성 한 명이며, 참고 인물의 머리 모양과 얼굴 윤곽에 대체로 부합한다. 한국인이라는 설정과 시각적으로 충돌하지 않는다. 눈을 감은 회백색 얼굴과 탈색된 입술은 사망 상태를 표현한다. 검은 보디백 한 개의 지퍼가 얼굴 주변에서 열려 있고 몸통과 팔다리는 가려져 있다. 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "보디백은 도로에 놓여 있고 뒤통수는 내부 바닥과 주름진 안감에 받쳐져 있다. 얼굴이나 목을 스스로 들어 올리는 명백한 긴장은 없으며, 보이는 신체에 지지 없는 부유는 없다. 열린 테두리는 접힌 두꺼운 소재로 형태를 유지한다."
       },
       {
        "label": "B",
        "direction": "남성은 눈을 감은 채 얼굴을 위로 향하며 어떤 대상을 응시하지 않는다. 카메라는 턱 쪽의 낮은 사선에서 얼굴을 향해 콧구멍과 턱 아래를 드러낸다. 얼굴 위에서 보디백 안을 내려다보는 검사 시점은 충분히 구현되지 않았다.",
        "built_space": "중앙 출입구 구조 한 개에 기둥 두 개와 가로보 하나가 있고, 철문은 전경 보디백에 일부 가려져 있다. 왼쪽 콘크리트 벽, 오른쪽 철망과 천막, 양쪽 확성기 기둥 두 개가 참고 장소와 같은 관계로 보인다. 젖은 도로의 반사도 자연스럽다. 얼굴과 보디백 안감이 전경을 크게 차지해 A보다 클로즈업에 가깝지만, 하늘과 출입구가 보이는 낮은 카메라 축이 두드러진다.",
        "entities": "검은 머리의 중년 동아시아계 남성 한 명만 보이며 참고 인물의 헤어라인, 코와 입, 턱 윤곽이 대체로 유지된다. 한국인 남성이라는 설정과 충돌하는 외형은 없다. 닫힌 눈과 창백한 피부, 혈색이 적은 입술이 사망 상태에 부합한다. 보디백 한 개의 열린 지퍼와 안감이 얼굴을 둘러싸고 몸통과 팔다리는 숨겨져 있다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리는 주름진 보디백 안감에 누워 있고 목 아래는 덮개에 가려진다. 보디백은 지면에 닿아 있으며 열린 부분은 양옆으로 접혀 지지된다. 사망자가 근력으로 신체를 들고 있는 모습이나 지지 없이 떠 있는 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "창백한 사망자와 얼굴 쪽이 열린 보디백은 맞지만, 얼굴보다 보디백과 주변 도로가 크게 보여 요구한 얼굴 클로즈업과 위에서 내려다보는 시점이 약하다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "얼굴을 크게 잡아 클로즈업과 핏기 없는 사망 상태를 더 충실히 구현하지만, 카메라가 턱 쪽에서 낮게 바라봐 명시된 내려다보는 시점에는 미달한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "남성은 눈을 감고 얼굴을 위로 향하고 있어 시선의 대상은 없다. 카메라는 몸통 쪽에서 얼굴을 비스듬히 바라보며, 보디백 내부는 보이지만 얼굴 바로 위에서 내려다보는 시점보다 낮다. 무기나 방향을 가진 동작은 없다.",
        "built_space": "젖은 도로 위 보디백 한 개 뒤로 중앙 철문 한 조, 양옆 출입구 기둥 두 개, 왼쪽 콘크리트 벽, 오른쪽 철망과 천막이 보인다. 출입구 상부는 프레임 밖으로 잘린다. 참고 장소의 재료와 배치는 대체로 맞고 젖은 지면의 빛 반사도 가능하다. 다만 도로와 보디백 몸통 부분이 넓게 들어와 얼굴의 화면 점유율이 작다.",
        "entities": "보이는 인물은 검은 머리의 중년 동아시아계 남성 한 명이며, 참고 인물의 머리 모양과 얼굴 윤곽에 대체로 부합한다. 한국인이라는 설정과 시각적으로 충돌하지 않는다. 눈을 감은 회백색 얼굴과 탈색된 입술은 사망 상태를 표현한다. 검은 보디백 한 개의 지퍼가 얼굴 주변에서 열려 있고 몸통과 팔다리는 가려져 있다. 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "보디백은 도로에 놓여 있고 뒤통수는 내부 바닥과 주름진 안감에 받쳐져 있다. 얼굴이나 목을 스스로 들어 올리는 명백한 긴장은 없으며, 보이는 신체에 지지 없는 부유는 없다. 열린 테두리는 접힌 두꺼운 소재로 형태를 유지한다."
       },
       {
        "label": "A",
        "direction": "남성은 눈을 감은 채 얼굴을 위로 향하며 어떤 대상을 응시하지 않는다. 카메라는 턱 쪽의 낮은 사선에서 얼굴을 향해 콧구멍과 턱 아래를 드러낸다. 얼굴 위에서 보디백 안을 내려다보는 검사 시점은 충분히 구현되지 않았다.",
        "built_space": "중앙 출입구 구조 한 개에 기둥 두 개와 가로보 하나가 있고, 철문은 전경 보디백에 일부 가려져 있다. 왼쪽 콘크리트 벽, 오른쪽 철망과 천막, 양쪽 확성기 기둥 두 개가 참고 장소와 같은 관계로 보인다. 젖은 도로의 반사도 자연스럽다. 얼굴과 보디백 안감이 전경을 크게 차지해 A보다 클로즈업에 가깝지만, 하늘과 출입구가 보이는 낮은 카메라 축이 두드러진다.",
        "entities": "검은 머리의 중년 동아시아계 남성 한 명만 보이며 참고 인물의 헤어라인, 코와 입, 턱 윤곽이 대체로 유지된다. 한국인 남성이라는 설정과 충돌하는 외형은 없다. 닫힌 눈과 창백한 피부, 혈색이 적은 입술이 사망 상태에 부합한다. 보디백 한 개의 열린 지퍼와 안감이 얼굴을 둘러싸고 몸통과 팔다리는 숨겨져 있다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리는 주름진 보디백 안감에 누워 있고 목 아래는 덮개에 가려진다. 보디백은 지면에 닿아 있으며 열린 부분은 양옆으로 접혀 지지된다. 사망자가 근력으로 신체를 들고 있는 모습이나 지지 없이 떠 있는 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.571,
    "B": 1.714
   },
   "adjusted": {
    "A": 1.571,
    "B": 1.714
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1714,
   "A": 1571
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1714,
    "verdict_ko": "지시문이 요구한 내려다보는 앵글과 핏기 없는 죽은 얼굴을 사실적인 보디백 질감과 함께 잘 구현했습니다."
   },
   {
    "label": "A",
    "score": 1571,
    "verdict_ko": "카메라가 피사체를 오히려 올려다보는 구도를 취했으며, 보디백의 형태와 재질이 비현실적으로 뻣뻣합니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_camp_gate_b39179.png",
    "asset_id": "83651a7a-615d-46c7-ba87-2459a6bc8bf2",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 구도환: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1311816>",
    "asset_id": "623f0421-67dc-4592-a009-148d8e957276",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-98e2-7856-b2f6-c53d54bf9db2",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S42sh2__bgfirst_bg.png",
   "bg_asset_id": "0428d63b-de1f-42c5-ad4e-e5abb4f95632",
   "bg_record_key": "S42sh2::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "camp_gate",
   "groupbg_asset_id": "83651a7a-615d-46c7-ba87-2459a6bc8bf2"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S42sh2::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:32:12.679545+00:00",
  "fingerprint": "9361f8d33ab2b3d637edd9c7b7f8ea6dd7b1c7c0f6cf606440e5df1b19ca99d9",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S42sh2_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S42sh2_sel.png",
  "source_sha256": "d50230d119f29bec9ecc6a0714ae2678bf1e4bfd58d66ce8f3f2fd8cb5c8756a",
  "file": "S42sh2_cine.png",
  "staged_sha256": "659d5dad424988908bf51b72e2b545e1d61ed13b0ba03f88f85a4a9e2ab055ef",
  "latency_ms": 8841
 },
 "S42sh11::signage": {
  "fp": "5bb1fa41ba640e00",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S42sh11": {
  "input_fingerprint": "8b6a7cd5f488693a",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 열린 차 문을 한 손으로 꽉 틀어쥔 채 단호하게 막아선 박철진의 팽팽한 자세.\n\nLOCATION (lock): Outside the open door of a sedan stopped on the muddy approach to the refugee settlement entrance. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 고급 세단의 열린 차 문 (Open and held by 박철진, preventing departure) — The edge and an oblique portion of the inner face are visible at the left side; used as Connects the gripping hand to the blocked doorway while occupying less than two-fifths of the frame; 세단 출입구 (Open, with 박철진 blocking its use) — Viewed diagonally from outside; the occupant remains beyond the crop; used as Provides the spatial evidence for his obstructing posture and inward-directed gaze.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained nighttime ambient illumination, preserving readable tension in the gripping hand and face without specifying an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The sedan is stopped with its door open. The body bag remains open at the face end nearby. 박철진: He stands at the sedan with one hand holding the open door.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 박철진 right now, so 박철진's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 박철진: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 열린 차 문을 한 손으로 꽉 틀어쥔 채 단호하게 막아선 박철진의 팽팽한 자세.\n\nLOCATION (lock): Outside the open door of a sedan stopped on the muddy approach to the refugee settlement entrance. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 고급 세단의 열린 차 문 (Open and held by 박철진, preventing departure) — The edge and an oblique portion of the inner face are visible at the left side; used as Connects the gripping hand to the blocked doorway while occupying less than two-fifths of the frame; 세단 출입구 (Open, with 박철진 blocking its use) — Viewed diagonally from outside; the occupant remains beyond the crop; used as Provides the spatial evidence for his obstructing posture and inward-directed gaze.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained nighttime ambient illumination, preserving readable tension in the gripping hand and face without specifying an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The sedan is stopped with its door open. The body bag remains open at the face end nearby. 박철진: He stands at the sedan with one hand holding the open door.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 박철진 right now, so 박철진's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 박철진: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 열린 차 문을 한 손으로 꽉 틀어쥔 채 단호하게 막아선 박철진의 팽팽한 자세.\n\nLOCATION (lock): Outside the open door of a sedan stopped on the muddy approach to the refugee settlement entrance. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 고급 세단의 열린 차 문 (Open and held by 박철진, preventing departure) — The edge and an oblique portion of the inner face are visible at the left side; used as Connects the gripping hand to the blocked doorway while occupying less than two-fifths of the frame; 세단 출입구 (Open, with 박철진 blocking its use) — Viewed diagonally from outside; the occupant remains beyond the crop; used as Provides the spatial evidence for his obstructing posture and inward-directed gaze.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained nighttime ambient illumination, preserving readable tension in the gripping hand and face without specifying an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The sedan is stopped with its door open. The body bag remains open at the face end nearby. 박철진: He stands at the sedan with one hand holding the open door.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 박철진 right now, so 박철진's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 박철진: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "박철진의 시선은 굳게 닫힌 입술과 함께 열린 차 문 안쪽을 매섭게 향하고 있습니다.",
    "built_space": "이전 샷과 동일한 진흙탕 길 위에 고급 세단이 왼쪽에 정차해 있으며, 문이 열려 화면의 좌측 일부를 차지합니다. 배경에는 난민촌 입구의 철책과 가로등이 올바르게 배치되어 있습니다.",
    "entities": "레퍼런스와 일치하는 외모의 박철진이 유일한 인물로 등장합니다. 짙은 색 세단과 우측 하단 바닥에 놓인 검은색 시신 가방이 존재합니다.",
    "hard_violations": [],
    "physics": "두 발은 보이지 않으나 안정적으로 땅을 딛고 선 자세이며, 오른손은 열린 차 문의 창문 틀 가장자리를 자연스럽고 단단하게 움켜쥐고 있습니다."
   },
   {
    "label": "B",
    "direction": "박철진의 시선이 차 안쪽을 향하며 경계하는 듯한 모습을 보입니다.",
    "built_space": "진흙탕 바닥과 난민촌 입구 철문이 배경에 있으며, 좌측에 세단 문이 열려 있습니다.",
    "entities": "박철진의 얼굴과 복장은 레퍼런스와 일치하나, 배경에 프롬프트에 명시되지 않은 여러 명의 사람(추가 인물)이 서성이고 있습니다. 명시된 시신 가방은 보이지 않습니다.",
    "hard_violations": [
     "[gemini-pro] 프롬프트에 없는 인물들(invented people)이 배경에 다수 추가됨",
     "[gpt-high] 박철진 외에는 누구도 등장하지 말라는 지시와 달리 배경에 사람 다섯 명이 추가되었다.",
     "[gpt-high] 장소 참조에 없던 콘크리트 차단물들이 진입로에 추가되어 고정된 공간 구성이 바뀌었다."
    ],
    "physics": "오른손으로 차 문틀의 상단을 잡고 체중을 지탱하며 딛고 서 있는 자세입니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "지시된 유일한 인물(박철진)만을 정확히 프레임에 담았으며, 차 문을 잡고 있는 손의 위치와 배경의 시신 가방 등 디테일을 훌륭하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "프롬프트에서 엄격하게 금지한 추가 인물들이 배경에 다수 등장하여 치명적인 오류를 범했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진의 시선은 굳게 닫힌 입술과 함께 열린 차 문 안쪽을 매섭게 향하고 있습니다.",
        "built_space": "이전 샷과 동일한 진흙탕 길 위에 고급 세단이 왼쪽에 정차해 있으며, 문이 열려 화면의 좌측 일부를 차지합니다. 배경에는 난민촌 입구의 철책과 가로등이 올바르게 배치되어 있습니다.",
        "entities": "레퍼런스와 일치하는 외모의 박철진이 유일한 인물로 등장합니다. 짙은 색 세단과 우측 하단 바닥에 놓인 검은색 시신 가방이 존재합니다.",
        "hard_violations": [],
        "physics": "두 발은 보이지 않으나 안정적으로 땅을 딛고 선 자세이며, 오른손은 열린 차 문의 창문 틀 가장자리를 자연스럽고 단단하게 움켜쥐고 있습니다."
       },
       {
        "label": "B",
        "direction": "박철진의 시선이 차 안쪽을 향하며 경계하는 듯한 모습을 보입니다.",
        "built_space": "진흙탕 바닥과 난민촌 입구 철문이 배경에 있으며, 좌측에 세단 문이 열려 있습니다.",
        "entities": "박철진의 얼굴과 복장은 레퍼런스와 일치하나, 배경에 프롬프트에 명시되지 않은 여러 명의 사람(추가 인물)이 서성이고 있습니다. 명시된 시신 가방은 보이지 않습니다.",
        "hard_violations": [
         "프롬프트에 없는 인물들(invented people)이 배경에 다수 추가됨"
        ],
        "physics": "오른손으로 차 문틀의 상단을 잡고 체중을 지탱하며 딛고 서 있는 자세입니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "지시된 유일한 인물(박철진)만을 정확히 프레임에 담았으며, 차 문을 잡고 있는 손의 위치와 배경의 시신 가방 등 디테일을 훌륭하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "프롬프트에서 엄격하게 금지한 추가 인물들이 배경에 다수 등장하여 치명적인 오류를 범했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "박철진의 시선은 굳게 닫힌 입술과 함께 열린 차 문 안쪽을 매섭게 향하고 있습니다.",
        "built_space": "이전 샷과 동일한 진흙탕 길 위에 고급 세단이 왼쪽에 정차해 있으며, 문이 열려 화면의 좌측 일부를 차지합니다. 배경에는 난민촌 입구의 철책과 가로등이 올바르게 배치되어 있습니다.",
        "entities": "레퍼런스와 일치하는 외모의 박철진이 유일한 인물로 등장합니다. 짙은 색 세단과 우측 하단 바닥에 놓인 검은색 시신 가방이 존재합니다.",
        "hard_violations": [],
        "physics": "두 발은 보이지 않으나 안정적으로 땅을 딛고 선 자세이며, 오른손은 열린 차 문의 창문 틀 가장자리를 자연스럽고 단단하게 움켜쥐고 있습니다."
       },
       {
        "label": "B",
        "direction": "박철진의 시선이 차 안쪽을 향하며 경계하는 듯한 모습을 보입니다.",
        "built_space": "진흙탕 바닥과 난민촌 입구 철문이 배경에 있으며, 좌측에 세단 문이 열려 있습니다.",
        "entities": "박철진의 얼굴과 복장은 레퍼런스와 일치하나, 배경에 프롬프트에 명시되지 않은 여러 명의 사람(추가 인물)이 서성이고 있습니다. 명시된 시신 가방은 보이지 않습니다.",
        "hard_violations": [
         "프롬프트에 없는 인물들(invented people)이 배경에 다수 추가됨"
        ],
        "physics": "오른손으로 차 문틀의 상단을 잡고 체중을 지탱하며 딛고 서 있는 자세입니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "차 안을 향한 시선과 힘준 자세는 맞지만, 금지된 배경 인물 다섯 명이 추가되었고 문이 지나치게 크게 배치되며 참조의 정장도 바뀌었다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "박철진 혼자 문을 붙들고 차 안을 응시하는 동작은 더 충실하지만, 문 안쪽 대신 외판을 크게 보여주고 참조의 정장 대신 야전 재킷을 입혔다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진의 얼굴과 눈은 화면 왼쪽의 열린 차량 출입구를 향한다. 카메라를 보지 않으며, 화면 밖 탑승자를 주시하는 방향으로 읽힌다. 한 손은 왼쪽 아래에서 올라오는 문틀을 움켜쥐고 다른 손은 몸 옆에서 주먹을 쥔다.",
        "built_space": "왼쪽에 세단 한 대와 좌석 일부, 내장 패널이 보이는 열린 문 하나가 있다. 여기에 전경을 대각선으로 가르는 별도의 문틀 같은 구조가 겹쳐 잡은 부분과 출입구의 연결이 불명확하다. 문 관련 구조가 화면 왼쪽 절반 이상을 차지해 지정된 5분의 2 미만 조건을 벗어난다. 배경의 콘크리트 벽, 중앙 철문, 오른쪽 천막과 야간 조명은 장소 참조와 대체로 이어지지만 도로의 콘크리트 차단물은 새로 추가되었다.",
        "entities": "주인공은 짧은 검은 머리와 작은 귀걸이를 한 중년 한국인 남성으로 보이며 얼굴은 박철진 참조와 유사하다. 그러나 남색 정장·흰 셔츠·줄무늬 넥타이 대신 어두운 야전 재킷과 니트를 입었다. 오른쪽 배경에는 허용되지 않은 사람 다섯 명이 보인다. 세단은 보이며 시신 가방은 프레임에서 확인되지 않는다. 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [
         "박철진 외에는 누구도 등장하지 말라는 지시와 달리 배경에 사람 다섯 명이 추가되었다.",
         "장소 참조에 없던 콘크리트 차단물들이 진입로에 추가되어 고정된 공간 구성이 바뀌었다."
        ],
        "physics": "박철진은 상체를 앞으로 기울이고 다리를 벌린 채 서 있으며 하체는 화면 아래로 이어진다. 발은 잘렸지만 공중에 뜬 자세는 아니다. 손가락이 문틀을 감싸 쥐는 접촉은 보인다. 다만 잡은 대각선 문틀과 뒤쪽 열린 문의 구조적 연결은 화면에서 명료하지 않다."
       },
       {
        "label": "B",
        "direction": "박철진은 화면 왼쪽 아래의 차량 실내 쪽으로 시선을 내린다. 한 손을 뻗어 열린 문의 세로 끝단을 잡고 있어, 화면 밖 탑승자에게 향한 주의와 문을 놓아주지 않는 행동이 읽힌다.",
        "built_space": "세단 한 대의 앞문 하나가 열려 있고 뒤문 하나는 닫혀 있다. 박철진은 열린 문 끝 바깥에 서서 팔로 출입 방향을 가로막는다. 출입구는 왼쪽에 일부 보이고 탑승자는 나오지 않는다. 다만 카메라에는 요구된 문의 안쪽 면이 아니라 외부 손잡이가 달린 외판이 보이며, 문이 화면의 약 절반을 차지한다. 중앙 철문, 왼쪽 콘크리트 벽, 오른쪽 천막, 젖은 진입로와 절제된 야간 조명은 참조 장소에 가깝다.",
        "entities": "보이는 사람은 박철진 한 명뿐이다. 중년 한국인 남성으로 보이는 얼굴, 짧은 검은 머리, 체격과 귀걸이는 참조에 가깝다. 흰 셔츠와 넥타이 일부는 보이지만 남색 정장 상의가 어두운 야전 재킷으로 바뀌었다. 세단과 오른쪽 아래 바닥의 열린 검은 가방 일부가 보이며, 가방의 얼굴 쪽 개방 상태는 잘린 화면만으로 확인할 수 없다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "손가락과 엄지가 문의 세로 모서리를 감싸고 손목과 팔이 자연스럽게 이어져 실제로 문을 붙드는 접촉이 성립한다. 문은 차체에 연결되어 열린 상태로 읽힌다. 몸은 지면에 선 자세로 화면 아래까지 이어지며, 잘린 발을 제외한 부분에서 부유나 불가능한 관절은 보이지 않는다. 검은 가방은 도로 위에 놓여 있다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "차 안을 향한 시선과 힘준 자세는 맞지만, 금지된 배경 인물 다섯 명이 추가되었고 문이 지나치게 크게 배치되며 참조의 정장도 바뀌었다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "박철진 혼자 문을 붙들고 차 안을 응시하는 동작은 더 충실하지만, 문 안쪽 대신 외판을 크게 보여주고 참조의 정장 대신 야전 재킷을 입혔다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "박철진의 얼굴과 눈은 화면 왼쪽의 열린 차량 출입구를 향한다. 카메라를 보지 않으며, 화면 밖 탑승자를 주시하는 방향으로 읽힌다. 한 손은 왼쪽 아래에서 올라오는 문틀을 움켜쥐고 다른 손은 몸 옆에서 주먹을 쥔다.",
        "built_space": "왼쪽에 세단 한 대와 좌석 일부, 내장 패널이 보이는 열린 문 하나가 있다. 여기에 전경을 대각선으로 가르는 별도의 문틀 같은 구조가 겹쳐 잡은 부분과 출입구의 연결이 불명확하다. 문 관련 구조가 화면 왼쪽 절반 이상을 차지해 지정된 5분의 2 미만 조건을 벗어난다. 배경의 콘크리트 벽, 중앙 철문, 오른쪽 천막과 야간 조명은 장소 참조와 대체로 이어지지만 도로의 콘크리트 차단물은 새로 추가되었다.",
        "entities": "주인공은 짧은 검은 머리와 작은 귀걸이를 한 중년 한국인 남성으로 보이며 얼굴은 박철진 참조와 유사하다. 그러나 남색 정장·흰 셔츠·줄무늬 넥타이 대신 어두운 야전 재킷과 니트를 입었다. 오른쪽 배경에는 허용되지 않은 사람 다섯 명이 보인다. 세단은 보이며 시신 가방은 프레임에서 확인되지 않는다. 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [
         "박철진 외에는 누구도 등장하지 말라는 지시와 달리 배경에 사람 다섯 명이 추가되었다.",
         "장소 참조에 없던 콘크리트 차단물들이 진입로에 추가되어 고정된 공간 구성이 바뀌었다."
        ],
        "physics": "박철진은 상체를 앞으로 기울이고 다리를 벌린 채 서 있으며 하체는 화면 아래로 이어진다. 발은 잘렸지만 공중에 뜬 자세는 아니다. 손가락이 문틀을 감싸 쥐는 접촉은 보인다. 다만 잡은 대각선 문틀과 뒤쪽 열린 문의 구조적 연결은 화면에서 명료하지 않다."
       },
       {
        "label": "A",
        "direction": "박철진은 화면 왼쪽 아래의 차량 실내 쪽으로 시선을 내린다. 한 손을 뻗어 열린 문의 세로 끝단을 잡고 있어, 화면 밖 탑승자에게 향한 주의와 문을 놓아주지 않는 행동이 읽힌다.",
        "built_space": "세단 한 대의 앞문 하나가 열려 있고 뒤문 하나는 닫혀 있다. 박철진은 열린 문 끝 바깥에 서서 팔로 출입 방향을 가로막는다. 출입구는 왼쪽에 일부 보이고 탑승자는 나오지 않는다. 다만 카메라에는 요구된 문의 안쪽 면이 아니라 외부 손잡이가 달린 외판이 보이며, 문이 화면의 약 절반을 차지한다. 중앙 철문, 왼쪽 콘크리트 벽, 오른쪽 천막, 젖은 진입로와 절제된 야간 조명은 참조 장소에 가깝다.",
        "entities": "보이는 사람은 박철진 한 명뿐이다. 중년 한국인 남성으로 보이는 얼굴, 짧은 검은 머리, 체격과 귀걸이는 참조에 가깝다. 흰 셔츠와 넥타이 일부는 보이지만 남색 정장 상의가 어두운 야전 재킷으로 바뀌었다. 세단과 오른쪽 아래 바닥의 열린 검은 가방 일부가 보이며, 가방의 얼굴 쪽 개방 상태는 잘린 화면만으로 확인할 수 없다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "손가락과 엄지가 문의 세로 모서리를 감싸고 손목과 팔이 자연스럽게 이어져 실제로 문을 붙드는 접촉이 성립한다. 문은 차체에 연결되어 열린 상태로 읽힌다. 몸은 지면에 선 자세로 화면 아래까지 이어지며, 잘린 발을 제외한 부분에서 부유나 불가능한 관절은 보이지 않는다. 검은 가방은 도로 위에 놓여 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.667
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.417
   },
   "violations": {
    "B": [
     "[gemini-pro] 프롬프트에 없는 인물들(invented people)이 배경에 다수 추가됨",
     "[gpt-high] 박철진 외에는 누구도 등장하지 말라는 지시와 달리 배경에 사람 다섯 명이 추가되었다.",
     "[gpt-high] 장소 참조에 없던 콘크리트 차단물들이 진입로에 추가되어 고정된 공간 구성이 바뀌었다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 417
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지시된 유일한 인물(박철진)만을 정확히 프레임에 담았으며, 차 문을 잡고 있는 손의 위치와 배경의 시신 가방 등 디테일을 훌륭하게 구현했습니다."
   },
   {
    "label": "B",
    "score": 417,
    "verdict_ko": "프롬프트에서 엄격하게 금지한 추가 인물들이 배경에 다수 등장하여 치명적인 오류를 범했습니다.  ★위반: [gemini-pro] 프롬프트에 없는 인물들(invented people)이 배경에 다수 추가됨 / [gpt-high] 박철진 외에는 누구도 등장하지 말라는 지시와 달리 배경에 사람 다섯 명이 추가되었다. / [gpt-high] 장소 참조에 없던 콘크리트 차단물들이 진입로에 추가되어 고정된 공간 구성이 바뀌었다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S42sh2_sel.png",
    "asset_id": "76146db3-5884-4ab8-9a93-e223ac96f3f9",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1401722>",
    "asset_id": "fee7383c-fb61-4b3a-ba7c-79f2555de00b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-9dc7-72e2-8af3-55f5f63a7e28",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S42sh2"
  },
  "lane_policy": "ab_select_bypass:prev"
 },
 "S42sh11::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:33:10.794747+00:00",
  "fingerprint": "66794824f1fe3b50ef14f67cf686f4bb63be10e6a8a34b9a84aa728e7e581f64",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S42sh11_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S42sh11_sel.png",
  "source_sha256": "771c1a101c469ab74d43633ca4419ca6226dfb5a7ae1bc7220f42a38ba82f12e",
  "file": "S42sh11_cine.png",
  "staged_sha256": "eb8f2554a11c70d3b59c05510e8ae27eae21fcec4a9610c48c4a1186f13888eb",
  "latency_ms": 9584
 },
 "S42sh19::signage": {
  "fp": "8f3b3442a49ea966",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S42sh19": {
  "input_fingerprint": "f77a13f1940e3f97",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 핸드폰을 주시한 채 충격을 받은 듯 눈이 동그랗게 커진 박철진의 얼어붙은 얼굴 클로즈업.\n\nLOCATION (lock): On the muddy roadside at the refugee settlement entrance at night, after the sedan has departed. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the established nighttime ambient treatment, with controlled facial contrast that makes the sudden loss of composure legible.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the muddy entrance ground, nearby structures, and nighttime lighting. Exclude the luxury sedan and its open door, since the vehicle has departed.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The sedan has left with its door closed. The opened body bag remains at the nighttime settlement entrance. 박철진: He remains at the entrance, his smile gone and his face now tense.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 박철진 right now, so 박철진's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 박철진: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 핸드폰을 주시한 채 충격을 받은 듯 눈이 동그랗게 커진 박철진의 얼어붙은 얼굴 클로즈업.\n\nLOCATION (lock): On the muddy roadside at the refugee settlement entrance at night, after the sedan has departed. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the established nighttime ambient treatment, with controlled facial contrast that makes the sudden loss of composure legible.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the muddy entrance ground, nearby structures, and nighttime lighting. Exclude the luxury sedan and its open door, since the vehicle has departed.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The sedan has left with its door closed. The opened body bag remains at the nighttime settlement entrance. 박철진: He remains at the entrance, his smile gone and his face now tense.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 박철진 right now, so 박철진's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 박철진: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 핸드폰을 주시한 채 충격을 받은 듯 눈이 동그랗게 커진 박철진의 얼어붙은 얼굴 클로즈업.\n\nLOCATION (lock): On the muddy roadside at the refugee settlement entrance at night, after the sedan has departed. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the established nighttime ambient treatment, with controlled facial contrast that makes the sudden loss of composure legible.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the muddy entrance ground, nearby structures, and nighttime lighting. Exclude the luxury sedan and its open door, since the vehicle has departed.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The sedan has left with its door closed. The opened body bag remains at the nighttime settlement entrance. 박철진: He remains at the entrance, his smile gone and his face now tense.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 박철진 right now, so 박철진's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 박철진: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "시선은 양손으로 들고 있는 핸드폰 화면을 정확히 향하고 있음.",
    "built_space": "진흙길과 난민촌 입구 구조물 등 이전 샷의 배경을 동일하게 묘사하였으며 세단은 제외됨.",
    "entities": "박철진의 인물 특징이 일치하며, 이전 샷의 어두운 셔츠와 재킷 복장을 정확히 유지함. 놀라 커진 눈의 표정이 잘 드러남.",
    "hard_violations": [],
    "physics": "양손이 핸드폰 하단과 측면을 자연스럽게 쥐고 지지함."
   },
   {
    "label": "B",
    "direction": "시선은 한 손으로 든 핸드폰 화면을 향하고 있음.",
    "built_space": "진흙길과 조명탑 등 배경 요소가 잘 나타나며, 지시된 바디백이 뒤쪽에 보임. 세단은 제외됨.",
    "entities": "박철진의 얼굴은 일치하나, 이전 샷 복장이 아닌 캐릭터 레퍼런스의 흰 셔츠와 넥타이를 착용하여 복장 유지 지시를 위반함.",
    "hard_violations": [],
    "physics": "한 손으로 핸드폰을 안정적으로 쥐고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "클로즈업 프레이밍을 완벽히 구현했으며, 이전 샷과 동일한 복장을 유지하라는 지시를 정확히 준수하여 가장 우수함."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "표정 연출과 배경은 적절하나, 이전 샷 복장 유지 지시를 위반하고 캐릭터 레퍼런스의 흰 셔츠와 넥타이를 차용하여 감점됨."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 양손으로 들고 있는 핸드폰 화면을 정확히 향하고 있음.",
        "built_space": "진흙길과 난민촌 입구 구조물 등 이전 샷의 배경을 동일하게 묘사하였으며 세단은 제외됨.",
        "entities": "박철진의 인물 특징이 일치하며, 이전 샷의 어두운 셔츠와 재킷 복장을 정확히 유지함. 놀라 커진 눈의 표정이 잘 드러남.",
        "hard_violations": [],
        "physics": "양손이 핸드폰 하단과 측면을 자연스럽게 쥐고 지지함."
       },
       {
        "label": "B",
        "direction": "시선은 한 손으로 든 핸드폰 화면을 향하고 있음.",
        "built_space": "진흙길과 조명탑 등 배경 요소가 잘 나타나며, 지시된 바디백이 뒤쪽에 보임. 세단은 제외됨.",
        "entities": "박철진의 얼굴은 일치하나, 이전 샷 복장이 아닌 캐릭터 레퍼런스의 흰 셔츠와 넥타이를 착용하여 복장 유지 지시를 위반함.",
        "hard_violations": [],
        "physics": "한 손으로 핸드폰을 안정적으로 쥐고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "클로즈업 프레이밍을 완벽히 구현했으며, 이전 샷과 동일한 복장을 유지하라는 지시를 정확히 준수하여 가장 우수함."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "표정 연출과 배경은 적절하나, 이전 샷 복장 유지 지시를 위반하고 캐릭터 레퍼런스의 흰 셔츠와 넥타이를 차용하여 감점됨."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 양손으로 들고 있는 핸드폰 화면을 정확히 향하고 있음.",
        "built_space": "진흙길과 난민촌 입구 구조물 등 이전 샷의 배경을 동일하게 묘사하였으며 세단은 제외됨.",
        "entities": "박철진의 인물 특징이 일치하며, 이전 샷의 어두운 셔츠와 재킷 복장을 정확히 유지함. 놀라 커진 눈의 표정이 잘 드러남.",
        "hard_violations": [],
        "physics": "양손이 핸드폰 하단과 측면을 자연스럽게 쥐고 지지함."
       },
       {
        "label": "B",
        "direction": "시선은 한 손으로 든 핸드폰 화면을 향하고 있음.",
        "built_space": "진흙길과 조명탑 등 배경 요소가 잘 나타나며, 지시된 바디백이 뒤쪽에 보임. 세단은 제외됨.",
        "entities": "박철진의 얼굴은 일치하나, 이전 샷 복장이 아닌 캐릭터 레퍼런스의 흰 셔츠와 넥타이를 착용하여 복장 유지 지시를 위반함.",
        "hard_violations": [],
        "physics": "한 손으로 핸드폰을 안정적으로 쥐고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "휴대폰을 응시하는 충격과 의상 연속성은 잘 살렸지만, 시신 가방과 도로에 화면을 많이 할애하여 요청한 얼어붙은 얼굴 클로즈업보다 느슨하다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "얼굴을 더 크게 잡아 휴대폰을 보며 눈이 커진 순간을 충실하게 구현했으며, 다만 재킷 안쪽 옷차림은 이전 장면과 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진의 두 눈은 얼굴 아래쪽에 든 휴대폰 방향을 향한다. 휴대폰 화면은 본인 쪽으로, 후면 카메라는 관객 쪽으로 향해 사용 방향이 맞는다.",
        "built_space": "왼쪽의 긴 벽체, 뒤쪽 출입구 한 곳, 오른쪽 천막 구조물과 여러 야간 조명, 젖은 진흙 도로가 보인다. 열린 검은 시신 가방 한 개가 인물 뒤쪽 바닥에 놓여 있으며 차량과 열린 차문은 없다. 장소의 주요 재질과 배치는 참조와 대체로 이어지지만 배경과 상반신의 비중이 커 얼굴 중심 클로즈업이 느슨해졌다.",
        "entities": "중년 한국인 남성 한 명만 보이며 짧은 검은 머리, 얼굴 윤곽, 작은 귀걸이가 박철진 참조와 대체로 맞는다. 회색 계열 재킷과 안쪽 셔츠·넥타이는 이전 장면의 차림에 가깝다. 휴대폰 한 대와 열린 시신 가방이 보이고 다른 사람이나 읽을 수 있는 글자는 없다. 웃음 없이 눈을 크게 뜨고 입술을 벌린 표정으로 충격을 표현한다.",
        "hard_violations": [],
        "physics": "휴대폰은 손가락과 엄지 사이에 실제로 잡혀 있고 손은 화면 아래로 이어지는 팔에 연결된다. 시신 가방은 진흙 바닥에 펼쳐져 지지된다. 하체는 프레임 밖이라 발의 접지는 확인할 수 없지만, 보이는 상체나 소품에 부유 또는 불가능한 지지 관계는 없다."
       },
       {
        "label": "B",
        "direction": "커진 두 눈의 시선은 얼굴 앞 아래쪽의 휴대폰으로 모인다. 관객에게는 휴대폰 뒷면이 보이고 화면은 박철진의 눈을 향하므로 응시 대상과 기기 방향이 일치한다.",
        "built_space": "왼쪽 벽체와 뒤쪽 출입구 일부, 오른쪽 천막 구조물, 젖은 진흙길이 흐릿하게 보인다. 조명은 중앙 왼쪽의 단일 광점 두 곳과 오른쪽 위의 쌍등 한 묶음이 뚜렷하다. 차량과 차문은 없다. 얼굴이 화면 높이 대부분을 차지하는 밀착된 구도이며, 시신 가방은 이 프레임에서 식별되지 않지만 이를 보여주려고 구도를 넓히지 않은 점은 요청에 맞는다.",
        "entities": "중년 한국인 남성 한 명의 얼굴, 짧은 검은 머리, 귀걸이와 회색 재킷이 참조 인물에 대체로 부합한다. 다만 재킷 안에는 어두운 셔츠가 보여 이전 장면의 밝은 셔츠·넥타이 차림과 차이가 있다. 휴대폰 한 대를 본인의 손으로 잡고 있으며 다른 사람이나 읽을 수 있는 문구는 없다. 자연스러운 홍채와 동공을 유지하면서 눈꺼풀을 크게 열고 얼굴을 굳혀 충격을 드러낸다.",
        "hard_violations": [],
        "physics": "휴대폰의 아래와 옆면을 손가락과 엄지가 감싸 지지한다. 손의 크기와 피부는 보이는 중년 남성의 얼굴에 어울리며 손목 방향도 상체와 자연스럽게 이어진다. 하체와 지면 접점은 클로즈업 밖에 있지만, 보이는 부분에서 근거 없는 부유나 불가능한 자세는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "휴대폰을 응시하는 충격과 의상 연속성은 잘 살렸지만, 시신 가방과 도로에 화면을 많이 할애하여 요청한 얼어붙은 얼굴 클로즈업보다 느슨하다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "얼굴을 더 크게 잡아 휴대폰을 보며 눈이 커진 순간을 충실하게 구현했으며, 다만 재킷 안쪽 옷차림은 이전 장면과 다르다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "박철진의 두 눈은 얼굴 아래쪽에 든 휴대폰 방향을 향한다. 휴대폰 화면은 본인 쪽으로, 후면 카메라는 관객 쪽으로 향해 사용 방향이 맞는다.",
        "built_space": "왼쪽의 긴 벽체, 뒤쪽 출입구 한 곳, 오른쪽 천막 구조물과 여러 야간 조명, 젖은 진흙 도로가 보인다. 열린 검은 시신 가방 한 개가 인물 뒤쪽 바닥에 놓여 있으며 차량과 열린 차문은 없다. 장소의 주요 재질과 배치는 참조와 대체로 이어지지만 배경과 상반신의 비중이 커 얼굴 중심 클로즈업이 느슨해졌다.",
        "entities": "중년 한국인 남성 한 명만 보이며 짧은 검은 머리, 얼굴 윤곽, 작은 귀걸이가 박철진 참조와 대체로 맞는다. 회색 계열 재킷과 안쪽 셔츠·넥타이는 이전 장면의 차림에 가깝다. 휴대폰 한 대와 열린 시신 가방이 보이고 다른 사람이나 읽을 수 있는 글자는 없다. 웃음 없이 눈을 크게 뜨고 입술을 벌린 표정으로 충격을 표현한다.",
        "hard_violations": [],
        "physics": "휴대폰은 손가락과 엄지 사이에 실제로 잡혀 있고 손은 화면 아래로 이어지는 팔에 연결된다. 시신 가방은 진흙 바닥에 펼쳐져 지지된다. 하체는 프레임 밖이라 발의 접지는 확인할 수 없지만, 보이는 상체나 소품에 부유 또는 불가능한 지지 관계는 없다."
       },
       {
        "label": "A",
        "direction": "커진 두 눈의 시선은 얼굴 앞 아래쪽의 휴대폰으로 모인다. 관객에게는 휴대폰 뒷면이 보이고 화면은 박철진의 눈을 향하므로 응시 대상과 기기 방향이 일치한다.",
        "built_space": "왼쪽 벽체와 뒤쪽 출입구 일부, 오른쪽 천막 구조물, 젖은 진흙길이 흐릿하게 보인다. 조명은 중앙 왼쪽의 단일 광점 두 곳과 오른쪽 위의 쌍등 한 묶음이 뚜렷하다. 차량과 차문은 없다. 얼굴이 화면 높이 대부분을 차지하는 밀착된 구도이며, 시신 가방은 이 프레임에서 식별되지 않지만 이를 보여주려고 구도를 넓히지 않은 점은 요청에 맞는다.",
        "entities": "중년 한국인 남성 한 명의 얼굴, 짧은 검은 머리, 귀걸이와 회색 재킷이 참조 인물에 대체로 부합한다. 다만 재킷 안에는 어두운 셔츠가 보여 이전 장면의 밝은 셔츠·넥타이 차림과 차이가 있다. 휴대폰 한 대를 본인의 손으로 잡고 있으며 다른 사람이나 읽을 수 있는 문구는 없다. 자연스러운 홍채와 동공을 유지하면서 눈꺼풀을 크게 열고 얼굴을 굳혀 충격을 드러낸다.",
        "hard_violations": [],
        "physics": "휴대폰의 아래와 옆면을 손가락과 엄지가 감싸 지지한다. 손의 크기와 피부는 보이는 중년 남성의 얼굴에 어울리며 손목 방향도 상체와 자연스럽게 이어진다. 하체와 지면 접점은 클로즈업 밖에 있지만, 보이는 부분에서 근거 없는 부유나 불가능한 자세는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.589
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.589
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1589
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "클로즈업 프레이밍을 완벽히 구현했으며, 이전 샷과 동일한 복장을 유지하라는 지시를 정확히 준수하여 가장 우수함."
   },
   {
    "label": "B",
    "score": 1589,
    "verdict_ko": "표정 연출과 배경은 적절하나, 이전 샷 복장 유지 지시를 위반하고 캐릭터 레퍼런스의 흰 셔츠와 넥타이를 차용하여 감점됨."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S42sh11_sel.png",
    "asset_id": "492e7b83-1888-4fa1-aae1-13fc9bd863fa",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1401722>",
    "asset_id": "fee7383c-fb61-4b3a-ba7c-79f2555de00b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-9f96-7dd2-9cd9-7bc5fc51be5d",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S42sh11"
  }
 },
 "S42sh19::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:34:23.538102+00:00",
  "fingerprint": "a59e403ad18c81a540ed5f2d5e0f5627a116e054b0884105181f69d954efa73e",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S42sh19_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S42sh19_sel.png",
  "source_sha256": "c3a068681f8c34e01336f9d33e11a21e3297051016bfc609423a857e69ff190f",
  "file": "S42sh19_cine.png",
  "staged_sha256": "5eacbb9386fee051f6d3257d823a23b1415f5d83a6e60093434a60ea3105d113",
  "latency_ms": 9548
 },
 "S43sh2::signage": {
  "fp": "1d602c0a44fb091c",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::5c5e686b8cecba11": {
  "subjects": [],
  "subject_text": "국방장관 집무실\n집무용 책상과 의자, 전화가 놓인 사무 공간. 책상 앞쪽에는 대면할 수 있는 좌석 공간이 마련돼 있다.",
  "identity": "canonical",
  "scope_id": "L196",
  "scope_role": "location_interior",
  "scope_sha": "eb6522f45277b7b9"
 },
 "S43sh2::bgfirst_bg": {
  "input_fingerprint": "158b48a517dc33f0",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 집무실 소파에 깊숙이 기댄 채 서늘한 눈빛으로 허공을 응시하는 국방장관의 여유로운 상체.\n\nLOCATION (lock): At a sofa inside the defense minister's executive office, under ordinary office lighting during the phone call.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 집무실 소파 (Occupied by the deeply reclining 국방장관) — An oblique section of its seat and back is visible beneath and behind him; used as Supports the backward body weight and visually substantiates his ease.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral ambient illumination appropriate to the office, with measured contrast and no unsupported source or color accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 집무실 소파에 깊숙이 기댄 채 서늘한 눈빛으로 허공을 응시하는 국방장관의 여유로운 상체.\n\nLOCATION (lock): At a sofa inside the defense minister's executive office, under ordinary office lighting during the phone call.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 집무실 소파 (Occupied by the deeply reclining 국방장관) — An oblique section of its seat and back is visible beneath and behind him; used as Supports the backward body weight and visually substantiates his ease.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral ambient illumination appropriate to the office, with measured contrast and no unsupported source or color accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S43sh2__bgfirst_bg.png",
  "asset_id": "a14ea04b-9a30-4065-8253-7e970e47fb07",
  "input_asset_ids": [
   "b8079413-7782-40e9-9e87-17589712c08c",
   "345f0672-c175-4e82-98e7-7fba2f349457"
  ]
 },
 "S43sh2": {
  "input_fingerprint": "f35ecf3938a446d9",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 집무실 소파에 깊숙이 기댄 채 서늘한 눈빛으로 허공을 응시하는 국방장관의 여유로운 상체.\n\nLOCATION (lock): At a sofa inside the defense minister's executive office, under ordinary office lighting during the phone call. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 집무실 소파 (Occupied by the deeply reclining 국방장관) — An oblique section of its seat and back is visible beneath and behind him; used as Supports the backward body weight and visually substantiates his ease.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral ambient illumination appropriate to the office, with measured contrast and no unsupported source or color accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The knife previously thrown into the militia-office wall remains embedded there. 국방장관: He is in his own office, engaged in the same telephone call.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 국방장관 (한국인, 성인 남성, 단정한 짧은 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 집무실 소파에 깊숙이 기댄 채 서늘한 눈빛으로 허공을 응시하는 국방장관의 여유로운 상체.\n\nLOCATION (lock): At a sofa inside the defense minister's executive office, under ordinary office lighting during the phone call. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 집무실 소파 (Occupied by the deeply reclining 국방장관) — An oblique section of its seat and back is visible beneath and behind him; used as Supports the backward body weight and visually substantiates his ease.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral ambient illumination appropriate to the office, with measured contrast and no unsupported source or color accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The knife previously thrown into the militia-office wall remains embedded there. 국방장관: He is in his own office, engaged in the same telephone call.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 국방장관 (한국인, 성인 남성, 단정한 짧은 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 집무실 소파에 깊숙이 기댄 채 서늘한 눈빛으로 허공을 응시하는 국방장관의 여유로운 상체.\n\nLOCATION (lock): At a sofa inside the defense minister's executive office, under ordinary office lighting during the phone call. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: 집무실 소파 (Occupied by the deeply reclining 국방장관) — An oblique section of its seat and back is visible beneath and behind him; used as Supports the backward body weight and visually substantiates his ease.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral ambient illumination appropriate to the office, with measured contrast and no unsupported source or color accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The knife previously thrown into the militia-office wall remains embedded there. 국방장관: He is in his own office, engaged in the same telephone call.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 국방장관 (한국인, 성인 남성, 단정한 짧은 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S43sh2__bgfirst_bg.png",
     "asset_id": "a14ea04b-9a30-4065-8253-7e970e47fb07",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S43sh2.png",
     "asset_id": "b8079413-7782-40e9-9e87-17589712c08c",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 국방장관: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1183281>",
     "asset_id": "65f118fd-34f5-434a-8c3f-1378f0e8a758",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L196B01.png",
     "asset_id": "345f0672-c175-4e82-98e7-7fba2f349457",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 국방장관: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1183281>",
     "asset_id": "65f118fd-34f5-434a-8c3f-1378f0e8a758",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "시선은 프레임 왼쪽 바깥 허공을 향하며, 왼손으로 스마트폰을 귀에 대고 있음.",
    "built_space": "장관 집무실. 왼쪽 벽의 산맥 그림, 중앙 기둥의 한반도 지도, 우측의 책상과 깃발 등 기준 사진의 가구 배치를 완벽히 재현함.",
    "entities": "국방장관(한국인 남성, 단정한 머리) 본인이 회색 셔츠를 입고 앉아 있음. 텍스트가 모두 알아볼 수 없게 처리됨.",
    "hard_violations": [],
    "physics": "소파에 체중을 싣고 안정적으로 앉아 있으며 꼰 다리와 팔걸이에 얹은 팔, 전화기를 든 손 모두 자연스럽게 지지됨."
   },
   {
    "label": "B",
    "direction": "고개를 약간 들고 왼쪽 상단 허공을 응시함.",
    "built_space": "장관 집무실 내부이나, 소파 바로 뒤에 책상과 깃발이 배치되어 기준 사진의 올바른 공간 구조(측면 벽에 소파 위치)와 크게 어긋남.",
    "entities": "국방장관(한국인 남성)이 남색 셔츠를 입고 앉아 있음. 통화 중인 상태가 아니며 우측 책상 위 명패에 한글이 명확히 보임.",
    "hard_violations": [
     "[gemini-pro] 화면 우측 명패에 '국방장관' 등 읽을 수 있는 텍스트가 노출되어 텍스트 금지 규칙 위반",
     "[gpt-high] 배경 명패에 '국방장관'이라는 읽을 수 있는 문자가 남아 있어, 이미지 어디에도 읽히는 글자가 없어야 한다는 조건을 위반한다."
    ],
    "physics": "소파 등받이에 상체를 깊게 기대어 안정적으로 체중을 지탱하고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "미디엄 샷 지정에도 불구하고 전신에 가까운 넓은 프레임을 연출한 점은 감점이나, 통화 중인 상태와 기준 공간 구조를 정확히 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "텍스트 노출 금지 규칙을 어기고 명패에 글씨가 명확히 보이며, 필수적인 통화 중 상태 누락 및 공간 구조 왜곡이 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 프레임 왼쪽 바깥 허공을 향하며, 왼손으로 스마트폰을 귀에 대고 있음.",
        "built_space": "장관 집무실. 왼쪽 벽의 산맥 그림, 중앙 기둥의 한반도 지도, 우측의 책상과 깃발 등 기준 사진의 가구 배치를 완벽히 재현함.",
        "entities": "국방장관(한국인 남성, 단정한 머리) 본인이 회색 셔츠를 입고 앉아 있음. 텍스트가 모두 알아볼 수 없게 처리됨.",
        "hard_violations": [],
        "physics": "소파에 체중을 싣고 안정적으로 앉아 있으며 꼰 다리와 팔걸이에 얹은 팔, 전화기를 든 손 모두 자연스럽게 지지됨."
       },
       {
        "label": "B",
        "direction": "고개를 약간 들고 왼쪽 상단 허공을 응시함.",
        "built_space": "장관 집무실 내부이나, 소파 바로 뒤에 책상과 깃발이 배치되어 기준 사진의 올바른 공간 구조(측면 벽에 소파 위치)와 크게 어긋남.",
        "entities": "국방장관(한국인 남성)이 남색 셔츠를 입고 앉아 있음. 통화 중인 상태가 아니며 우측 책상 위 명패에 한글이 명확히 보임.",
        "hard_violations": [
         "화면 우측 명패에 '국방장관' 등 읽을 수 있는 텍스트가 노출되어 텍스트 금지 규칙 위반"
        ],
        "physics": "소파 등받이에 상체를 깊게 기대어 안정적으로 체중을 지탱하고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "미디엄 샷 지정에도 불구하고 전신에 가까운 넓은 프레임을 연출한 점은 감점이나, 통화 중인 상태와 기준 공간 구조를 정확히 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "텍스트 노출 금지 규칙을 어기고 명패에 글씨가 명확히 보이며, 필수적인 통화 중 상태 누락 및 공간 구조 왜곡이 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 프레임 왼쪽 바깥 허공을 향하며, 왼손으로 스마트폰을 귀에 대고 있음.",
        "built_space": "장관 집무실. 왼쪽 벽의 산맥 그림, 중앙 기둥의 한반도 지도, 우측의 책상과 깃발 등 기준 사진의 가구 배치를 완벽히 재현함.",
        "entities": "국방장관(한국인 남성, 단정한 머리) 본인이 회색 셔츠를 입고 앉아 있음. 텍스트가 모두 알아볼 수 없게 처리됨.",
        "hard_violations": [],
        "physics": "소파에 체중을 싣고 안정적으로 앉아 있으며 꼰 다리와 팔걸이에 얹은 팔, 전화기를 든 손 모두 자연스럽게 지지됨."
       },
       {
        "label": "B",
        "direction": "고개를 약간 들고 왼쪽 상단 허공을 응시함.",
        "built_space": "장관 집무실 내부이나, 소파 바로 뒤에 책상과 깃발이 배치되어 기준 사진의 올바른 공간 구조(측면 벽에 소파 위치)와 크게 어긋남.",
        "entities": "국방장관(한국인 남성)이 남색 셔츠를 입고 앉아 있음. 통화 중인 상태가 아니며 우측 책상 위 명패에 한글이 명확히 보임.",
        "hard_violations": [
         "화면 우측 명패에 '국방장관' 등 읽을 수 있는 텍스트가 노출되어 텍스트 금지 규칙 위반"
        ],
        "physics": "소파 등받이에 상체를 깊게 기대어 안정적으로 체중을 지탱하고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "여유롭게 기댄 상체의 미디엄 숏과 남색 셔츠는 잘 맞지만, 배경 명패에 읽히는 한글이 남아 무문자 조건을 위반한다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "참조 집무실의 소파 배치와 야간 통화는 충실하지만, 다리와 집무실까지 넓힌 구도는 상체 중심 미디엄 숏에서 벗어나며 셔츠 색도 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "장관은 턱을 조금 들고 화면 왼쪽 위의 허공을 바라본다. 카메라나 특정 물체를 응시하지 않아 서늘하게 허공을 응시하라는 지시와 맞는다. 겨누는 물체나 이동하는 몸은 없다.",
        "built_space": "장관 뒤와 아래에 검은 가죽 등받이와 좌석이 있고 양옆에 목재 팔걸이가 보인다. 왼쪽 뒤로 별도의 좌석 등받이가 이어져 장관의 자리는 긴 소파보다는 독립 안락의자처럼 읽히는 면이 있다. 배경에는 창 구획, 지도 액자 일부 1개, 태극기 1개, 짙은색 깃발 1개, 책장 일부와 목재 업무 가구가 보인다. 참조의 재료와 집무실 요소는 유지하지만, 벽을 따라 놓인 참조 소파와 책상의 위치 관계는 덜 명확하다. 불가능한 반사는 보이지 않는다.",
        "entities": "인물은 단정하게 넘긴 짧은 검은 머리의 성인 동아시아계 남성 1명이며, 참조 국방장관의 얼굴과 체격에 가깝다. 남색 셔츠도 참조와 일치한다. 눈은 정상적인 홍채와 동공을 갖는다. 전화기는 보이지 않아 통화 중이라는 상태는 시각적으로 확인되지 않는다. 다른 장소에 남아 있어야 하는 칼은 등장하지 않는다. 배경 명패에는 '국방장관'이라는 한글이 읽힌다.",
        "hard_violations": [
         "배경 명패에 '국방장관'이라는 읽을 수 있는 문자가 남아 있어, 이미지 어디에도 읽히는 글자가 없어야 한다는 조건을 위반한다."
        ],
        "physics": "상체는 뒤로 기울어 가죽 등받이에 지지되고, 골반은 화면 아래 좌석에 놓인다. 화면 왼쪽 팔과 손은 목재 팔걸이에 자연스럽게 얹혀 있다. 보이는 자세에 부유하거나 지지 없이 매달린 신체는 없으며, 깊이 기대는 동작 자체는 물리적으로 가능하다."
       },
       {
        "label": "B",
        "direction": "장관은 화면 오른쪽의 비어 있는 공간을 수평으로 바라보며 카메라를 보지 않는다. 손에 든 전화기는 귀에 붙어 있고 화면을 관객에게 제시하지 않는다. 허공을 바라보며 통화하는 관계가 명확하다.",
        "built_space": "장관은 왼쪽 벽을 따라 놓인 긴 소파에 앉아 있다. 중앙에 안락의자 1개, 오른쪽 앞에 커피 테이블 1개, 오른쪽 뒤에 업무 책상 1개와 사무용 의자 1개가 보인다. 뒤쪽의 큰 창 영역 2곳 사이에는 지도 액자 1개가 있고, 왼쪽에는 산 풍경 액자 1개와 스탠드 1개, 오른쪽에는 책장과 깃발 2개가 있다. 참조 집무실의 주요 배치와 재료를 잘 유지한다. 다만 넓은 공간과 다리까지 포함해 요구된 상체 중심 미디엄 숏보다 넓다.",
        "entities": "인물은 짧고 단정한 검은 머리의 성인 동아시아계 남성 1명으로 참조 인물과 대체로 유사하다. 셔츠는 참조의 남색이 아니라 회색이다. 검은 전화기를 귀에 대고 있어 동일 통화의 지속 상태가 드러난다. 검은 가죽 소파, 태극기와 짙은색 깃발이 보이며 추가 인물은 없다. 명패와 책의 표기는 판독되지 않는다. 창밖은 어둡고 도시 조명이 켜져 있어 밤에 부합한다.",
        "hard_violations": [],
        "physics": "골반과 허벅지가 소파 좌석에 실리고 등은 등받이에 기대어 있다. 한쪽 팔은 소파 등받이 위에 걸쳐 지지되며, 다른 손은 전화기를 직접 잡는다. 교차한 다리는 아래쪽 다리와 좌석의 지지를 받는 자연스러운 앉은 자세다. 깊숙이 젖힌 정도는 약하지만 물리적으로 불가능하거나 무지지 상태인 부분은 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "여유롭게 기댄 상체의 미디엄 숏과 남색 셔츠는 잘 맞지만, 배경 명패에 읽히는 한글이 남아 무문자 조건을 위반한다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "참조 집무실의 소파 배치와 야간 통화는 충실하지만, 다리와 집무실까지 넓힌 구도는 상체 중심 미디엄 숏에서 벗어나며 셔츠 색도 다르다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "장관은 턱을 조금 들고 화면 왼쪽 위의 허공을 바라본다. 카메라나 특정 물체를 응시하지 않아 서늘하게 허공을 응시하라는 지시와 맞는다. 겨누는 물체나 이동하는 몸은 없다.",
        "built_space": "장관 뒤와 아래에 검은 가죽 등받이와 좌석이 있고 양옆에 목재 팔걸이가 보인다. 왼쪽 뒤로 별도의 좌석 등받이가 이어져 장관의 자리는 긴 소파보다는 독립 안락의자처럼 읽히는 면이 있다. 배경에는 창 구획, 지도 액자 일부 1개, 태극기 1개, 짙은색 깃발 1개, 책장 일부와 목재 업무 가구가 보인다. 참조의 재료와 집무실 요소는 유지하지만, 벽을 따라 놓인 참조 소파와 책상의 위치 관계는 덜 명확하다. 불가능한 반사는 보이지 않는다.",
        "entities": "인물은 단정하게 넘긴 짧은 검은 머리의 성인 동아시아계 남성 1명이며, 참조 국방장관의 얼굴과 체격에 가깝다. 남색 셔츠도 참조와 일치한다. 눈은 정상적인 홍채와 동공을 갖는다. 전화기는 보이지 않아 통화 중이라는 상태는 시각적으로 확인되지 않는다. 다른 장소에 남아 있어야 하는 칼은 등장하지 않는다. 배경 명패에는 '국방장관'이라는 한글이 읽힌다.",
        "hard_violations": [
         "배경 명패에 '국방장관'이라는 읽을 수 있는 문자가 남아 있어, 이미지 어디에도 읽히는 글자가 없어야 한다는 조건을 위반한다."
        ],
        "physics": "상체는 뒤로 기울어 가죽 등받이에 지지되고, 골반은 화면 아래 좌석에 놓인다. 화면 왼쪽 팔과 손은 목재 팔걸이에 자연스럽게 얹혀 있다. 보이는 자세에 부유하거나 지지 없이 매달린 신체는 없으며, 깊이 기대는 동작 자체는 물리적으로 가능하다."
       },
       {
        "label": "A",
        "direction": "장관은 화면 오른쪽의 비어 있는 공간을 수평으로 바라보며 카메라를 보지 않는다. 손에 든 전화기는 귀에 붙어 있고 화면을 관객에게 제시하지 않는다. 허공을 바라보며 통화하는 관계가 명확하다.",
        "built_space": "장관은 왼쪽 벽을 따라 놓인 긴 소파에 앉아 있다. 중앙에 안락의자 1개, 오른쪽 앞에 커피 테이블 1개, 오른쪽 뒤에 업무 책상 1개와 사무용 의자 1개가 보인다. 뒤쪽의 큰 창 영역 2곳 사이에는 지도 액자 1개가 있고, 왼쪽에는 산 풍경 액자 1개와 스탠드 1개, 오른쪽에는 책장과 깃발 2개가 있다. 참조 집무실의 주요 배치와 재료를 잘 유지한다. 다만 넓은 공간과 다리까지 포함해 요구된 상체 중심 미디엄 숏보다 넓다.",
        "entities": "인물은 짧고 단정한 검은 머리의 성인 동아시아계 남성 1명으로 참조 인물과 대체로 유사하다. 셔츠는 참조의 남색이 아니라 회색이다. 검은 전화기를 귀에 대고 있어 동일 통화의 지속 상태가 드러난다. 검은 가죽 소파, 태극기와 짙은색 깃발이 보이며 추가 인물은 없다. 명패와 책의 표기는 판독되지 않는다. 창밖은 어둡고 도시 조명이 켜져 있어 밤에 부합한다.",
        "hard_violations": [],
        "physics": "골반과 허벅지가 소파 좌석에 실리고 등은 등받이에 기대어 있다. 한쪽 팔은 소파 등받이 위에 걸쳐 지지되며, 다른 손은 전화기를 직접 잡는다. 교차한 다리는 아래쪽 다리와 좌석의 지지를 받는 자연스러운 앉은 자세다. 깊숙이 젖힌 정도는 약하지만 물리적으로 불가능하거나 무지지 상태인 부분은 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.929
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.679
   },
   "violations": {
    "B": [
     "[gemini-pro] 화면 우측 명패에 '국방장관' 등 읽을 수 있는 텍스트가 노출되어 텍스트 금지 규칙 위반",
     "[gpt-high] 배경 명패에 '국방장관'이라는 읽을 수 있는 문자가 남아 있어, 이미지 어디에도 읽히는 글자가 없어야 한다는 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 679
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "미디엄 샷 지정에도 불구하고 전신에 가까운 넓은 프레임을 연출한 점은 감점이나, 통화 중인 상태와 기준 공간 구조를 정확히 구현했습니다."
   },
   {
    "label": "B",
    "score": 679,
    "verdict_ko": "텍스트 노출 금지 규칙을 어기고 명패에 글씨가 명확히 보이며, 필수적인 통화 중 상태 누락 및 공간 구조 왜곡이 발생했습니다.  ★위반: [gemini-pro] 화면 우측 명패에 '국방장관' 등 읽을 수 있는 텍스트가 노출되어 텍스트 금지 규칙 위반 / [gpt-high] 배경 명패에 '국방장관'이라는 읽을 수 있는 문자가 남아 있어, 이미지 어디에도 읽히는 글자가 없어야 한다는 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L196B01.png",
    "asset_id": "345f0672-c175-4e82-98e7-7fba2f349457",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 국방장관: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1183281>",
    "asset_id": "65f118fd-34f5-434a-8c3f-1378f0e8a758",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-a150-7374-b393-75008fa74e1f",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S43sh2__bgfirst_bg.png",
   "bg_asset_id": "a14ea04b-9a30-4065-8253-7e970e47fb07",
   "bg_record_key": "S43sh2::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S43sh2::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:30:43.813176+00:00",
  "fingerprint": "f4e729ab4a357a13f5ae7a772673a63667b3aeb079254225466f606b97b53fd7",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S43sh2_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S43sh2_sel.png",
  "source_sha256": "8dea1910e48163843d0bc49c38b6111b0aece0e7e7173e9067e7ddf140a88fea",
  "file": "S43sh2_cine.png",
  "staged_sha256": "35f1c1f72713a057f787de99c66e14911dc64e67b58d398bde2bf53229e09bf7",
  "latency_ms": 14870
 },
 "S43sh4::signage": {
  "fp": "f657f571eefa4ecc",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S43sh4": {
  "input_fingerprint": "b6add6db4afd43ce",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 핸드폰을 귀에 댄 채 허리를 90도로 굽히고 비굴하게 고개를 숙인 박철진의 전신.\n\nLOCATION (lock): Inside the militia commander's own office, in the open standing area under ordinary office lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 핸드폰 (Held to 박철진's ear during the call) — Its edge and back are visible beside his lowered head; no screen content is presented; used as Links the otherwise unseen authority to his full-body submission.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral office ambient illumination with restrained contrast, expressing submission through framing rather than a fabricated lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The knife remains embedded in the militia-office wall. 박철진: He remains on the telephone, bowing submissively with a trembling hand.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 박철진 right now, so 박철진's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 박철진: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 핸드폰을 귀에 댄 채 허리를 90도로 굽히고 비굴하게 고개를 숙인 박철진의 전신.\n\nLOCATION (lock): Inside the militia commander's own office, in the open standing area under ordinary office lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 핸드폰 (Held to 박철진's ear during the call) — Its edge and back are visible beside his lowered head; no screen content is presented; used as Links the otherwise unseen authority to his full-body submission.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral office ambient illumination with restrained contrast, expressing submission through framing rather than a fabricated lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The knife remains embedded in the militia-office wall. 박철진: He remains on the telephone, bowing submissively with a trembling hand.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 박철진 right now, so 박철진's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 박철진: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 핸드폰을 귀에 댄 채 허리를 90도로 굽히고 비굴하게 고개를 숙인 박철진의 전신.\n\nLOCATION (lock): Inside the militia commander's own office, in the open standing area under ordinary office lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 핸드폰 (Held to 박철진's ear during the call) — Its edge and back are visible beside his lowered head; no screen content is presented; used as Links the otherwise unseen authority to his full-body submission.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral office ambient illumination with restrained contrast, expressing submission through framing rather than a fabricated lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The knife remains embedded in the militia-office wall. 박철진: He remains on the telephone, bowing submissively with a trembling hand.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 박철진 right now, so 박철진's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 박철진: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "인물은 고개를 숙인 채 바닥을 향해 시선을 두고 있으며, 왼쪽 귀에 휴대폰을 대고 있음.",
    "built_space": "이전 샷과 동일한 구조의 사무실로, 왼쪽 벽에 꽂힌 칼, 뒷배경의 철제 책상, 의자, 캐비닛, 문 등이 올바른 비례와 위치에 배치됨.",
    "entities": "박철진의 얼굴 윤곽, 헤어스타일, 정장 복장이 레퍼런스와 일치하며, 벽에 꽂힌 칼과 귀에 댄 휴대폰이 정확히 묘사됨.",
    "hard_violations": [
     "[gpt-high] 왼쪽 위 액자에 읽을 수 있는 한글 문구가 노출되어, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 명시적 금지 조건을 위반한다."
    ],
    "physics": "인물은 하반신을 통해 바닥에 안정적으로 지탱하며 90도 인사를 하고 있고, 손은 휴대폰을 귀에 댄 상태로 무리 없이 유지됨."
   },
   {
    "label": "B",
    "direction": "인물은 고개를 숙이고 시선을 아래로 향한 상태에서 휴대폰을 귀에 밀착하고 있음.",
    "built_space": "이전 샷의 요소를 반영하여 전경 책상 위의 회의용 전화기와 후경의 가구들, 벽에 꽂힌 칼을 배치함.",
    "entities": "박철진의 인상착의와 복장, 벽에 꽂힌 칼 등은 프롬프트와 레퍼런스를 따르고 있음.",
    "hard_violations": [
     "[gemini-pro] 왼쪽 어깨에서 나온 팔이 휴대폰을 들고 있음에도, 같은 쪽 몸통 아래로 주먹을 쥔 또 다른 팔이 늘어져 있는 신체 중복(팔 3개) 오류"
    ],
    "physics": "왼쪽 팔이 두 개로 분할되어 허공과 몸통 쪽에 동시에 존재하는 물리적, 해부학적 불가능 상태."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "요구된 '전신(wide shot)'에 가깝게 앵글을 잡아냈으며, 신체 오류 없이 90도로 굽힌 비굴한 자세와 배경의 연속성을 충실히 구현했습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "같은 쪽(왼쪽)에 휴대폰을 든 팔과 아래로 늘어뜨린 팔이 동시에 존재하는 치명적인 신체 중복 오류가 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "인물은 고개를 숙인 채 바닥을 향해 시선을 두고 있으며, 왼쪽 귀에 휴대폰을 대고 있음.",
        "built_space": "이전 샷과 동일한 구조의 사무실로, 왼쪽 벽에 꽂힌 칼, 뒷배경의 철제 책상, 의자, 캐비닛, 문 등이 올바른 비례와 위치에 배치됨.",
        "entities": "박철진의 얼굴 윤곽, 헤어스타일, 정장 복장이 레퍼런스와 일치하며, 벽에 꽂힌 칼과 귀에 댄 휴대폰이 정확히 묘사됨.",
        "hard_violations": [],
        "physics": "인물은 하반신을 통해 바닥에 안정적으로 지탱하며 90도 인사를 하고 있고, 손은 휴대폰을 귀에 댄 상태로 무리 없이 유지됨."
       },
       {
        "label": "B",
        "direction": "인물은 고개를 숙이고 시선을 아래로 향한 상태에서 휴대폰을 귀에 밀착하고 있음.",
        "built_space": "이전 샷의 요소를 반영하여 전경 책상 위의 회의용 전화기와 후경의 가구들, 벽에 꽂힌 칼을 배치함.",
        "entities": "박철진의 인상착의와 복장, 벽에 꽂힌 칼 등은 프롬프트와 레퍼런스를 따르고 있음.",
        "hard_violations": [
         "왼쪽 어깨에서 나온 팔이 휴대폰을 들고 있음에도, 같은 쪽 몸통 아래로 주먹을 쥔 또 다른 팔이 늘어져 있는 신체 중복(팔 3개) 오류"
        ],
        "physics": "왼쪽 팔이 두 개로 분할되어 허공과 몸통 쪽에 동시에 존재하는 물리적, 해부학적 불가능 상태."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "요구된 '전신(wide shot)'에 가깝게 앵글을 잡아냈으며, 신체 오류 없이 90도로 굽힌 비굴한 자세와 배경의 연속성을 충실히 구현했습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "같은 쪽(왼쪽)에 휴대폰을 든 팔과 아래로 늘어뜨린 팔이 동시에 존재하는 치명적인 신체 중복 오류가 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "인물은 고개를 숙인 채 바닥을 향해 시선을 두고 있으며, 왼쪽 귀에 휴대폰을 대고 있음.",
        "built_space": "이전 샷과 동일한 구조의 사무실로, 왼쪽 벽에 꽂힌 칼, 뒷배경의 철제 책상, 의자, 캐비닛, 문 등이 올바른 비례와 위치에 배치됨.",
        "entities": "박철진의 얼굴 윤곽, 헤어스타일, 정장 복장이 레퍼런스와 일치하며, 벽에 꽂힌 칼과 귀에 댄 휴대폰이 정확히 묘사됨.",
        "hard_violations": [],
        "physics": "인물은 하반신을 통해 바닥에 안정적으로 지탱하며 90도 인사를 하고 있고, 손은 휴대폰을 귀에 댄 상태로 무리 없이 유지됨."
       },
       {
        "label": "B",
        "direction": "인물은 고개를 숙이고 시선을 아래로 향한 상태에서 휴대폰을 귀에 밀착하고 있음.",
        "built_space": "이전 샷의 요소를 반영하여 전경 책상 위의 회의용 전화기와 후경의 가구들, 벽에 꽂힌 칼을 배치함.",
        "entities": "박철진의 인상착의와 복장, 벽에 꽂힌 칼 등은 프롬프트와 레퍼런스를 따르고 있음.",
        "hard_violations": [
         "왼쪽 어깨에서 나온 팔이 휴대폰을 들고 있음에도, 같은 쪽 몸통 아래로 주먹을 쥔 또 다른 팔이 늘어져 있는 신체 중복(팔 3개) 오류"
        ],
        "physics": "왼쪽 팔이 두 개로 분할되어 허공과 몸통 쪽에 동시에 존재하는 물리적, 해부학적 불가능 상태."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "통화 방향과 장소·복장은 부합하고 명백한 금지 위반은 없지만, 하체가 잘린 구도와 부족한 굽힘 각도로 ‘90도로 숙인 전신 와이드 숏’을 충족하지 못한다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "A보다 넓게 담았지만 발이 잘리고 허리도 90도에 못 미치며, 왼쪽 위 액자의 읽을 수 있는 한글이 명시적인 문자 금지 조건을 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진은 화면 오른쪽 아래의 바닥을 향해 얼굴과 눈을 내리고 있다. 손으로 잡은 휴대전화의 화면 쪽은 귀를 향하며, 카메라에는 뒷면과 가장자리가 보인다. 왼쪽 벽의 칼은 날이 벽 안으로 들어가고 손잡이가 실내 쪽으로 돌출되어 있다.",
        "built_space": "벗겨진 회백색 투톤 벽, 뒤쪽 금속 책상 하나와 검은 사무용 의자 하나, 오른쪽 뒤 문 하나, 왼쪽 뒤 수납장 하나가 보인다. 앞쪽에는 회의용 전화기가 놓인 책상 하나가 있고, 인물 뒤 왼쪽에는 접이식 의자 일부, 오른쪽 아래에는 다른 의자 일부가 보인다. 벽의 칼은 하나이며 참고 이미지와 같은 쪽에 있다. 박철진은 책상 앞 빈 공간에 서 있다. 기존 장소의 주요 재질과 배치를 대체로 유지하며 불가능한 반사는 없다.",
        "entities": "등장인물은 중년 한국인 남성으로 보이는 박철진 한 명이다. 짧게 넘긴 검은 머리, 남색 정장, 흰 셔츠와 사선무늬 넥타이가 참고 이미지에 부합한다. 숙인 얼굴의 보이는 부분도 참고 인물과 대체로 일치한다. 귀에 댄 휴대전화 하나와 벽에 박힌 칼 하나가 있다. 읽을 수 있는 화면 내용이나 뚜렷한 문자는 보이지 않는다. 무릎 아래와 발은 프레임 밖이므로 요청한 전신이 아니다.",
        "hard_violations": [],
        "physics": "휴대전화는 박철진의 손가락과 손바닥에 잡혀 귀에 밀착되어 있다. 칼은 벽에 박힌 날이 지지한다. 몸은 아래쪽 프레임 밖으로 이어지는 다리로 지탱하는 자연스러운 선 자세로 보이며, 공중에 떠 있다는 증거는 없다. 다만 발의 접지는 화면에서 확인할 수 없다. 허리를 굽히고 고개를 내린 동작은 가능하지만 몸통이 수평까지 내려가지 않아 명시한 90도 인사와 다르다. 손의 떨림은 정지 화면에서 뚜렷하게 드러나지 않는다."
       },
       {
        "label": "B",
        "direction": "박철진의 고개와 시선은 오른쪽 아래 바닥을 향한다. 휴대전화는 귀에 붙어 있고 후면 카메라가 있는 뒷면이 관객을 향하여 통화 용도에 맞는다. 왼쪽 벽의 칼은 날이 벽을 향하고 손잡이가 실내로 나와 있다.",
        "built_space": "뒤쪽 금속 책상 하나와 그 뒤 회전의자 하나, 오른쪽 뒤 문 하나, 왼쪽 뒤 높은 수납장 하나, 오른쪽 낮은 수납장 하나가 보인다. 오른쪽 아래에는 앞쪽 책상 모서리만 들어온다. 왼쪽 벽에는 칼 하나와 액자 두 개, 종이 게시물 하나가 보이고 오른쪽 벽에도 액자 일부가 있다. 인물은 책상 앞의 열린 바닥 공간에 서 있다. 낡은 투톤 벽과 금속 가구는 참고 장소와 대체로 맞지만, 왼쪽 위 액자의 문구가 읽히도록 노출되어 있다.",
        "entities": "박철진으로 보이는 중년 한국인 남성 한 명만 등장한다. 검은 머리와 남색 정장, 흰 셔츠는 참고와 일치하며 얼굴은 숙인 측면으로 일부만 확인된다. 넥타이는 자세와 손에 가려 세부 비교가 어렵다. 휴대전화 하나와 벽에 박힌 칼 하나가 존재한다. A보다 다리를 많이 보여 주지만 발과 신발은 하단 밖으로 잘려 전신 조건을 충족하지 않는다. 왼쪽 위 액자에는 식별 가능한 한글 문구가 있다.",
        "hard_violations": [
         "왼쪽 위 액자에 읽을 수 있는 한글 문구가 노출되어, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 명시적 금지 조건을 위반한다."
        ],
        "physics": "휴대전화는 귀 옆에서 손으로 확실히 잡고 있으며, 벽의 칼은 박힌 날로 지지된다. 인물의 다리는 바닥 방향으로 이어져 정상적으로 서서 굽힌 자세로 보인다. 발은 잘려 접지점을 직접 확인할 수 없지만 무지지 부유나 불가능한 관절은 보이지 않는다. 몸통은 여전히 대각선으로 기울어 있어 허리를 90도로 접은 동작에 못 미친다. 손 떨림도 명확하게 확인되지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "통화 방향과 장소·복장은 부합하고 명백한 금지 위반은 없지만, 하체가 잘린 구도와 부족한 굽힘 각도로 ‘90도로 숙인 전신 와이드 숏’을 충족하지 못한다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "A보다 넓게 담았지만 발이 잘리고 허리도 90도에 못 미치며, 왼쪽 위 액자의 읽을 수 있는 한글이 명시적인 문자 금지 조건을 위반한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "박철진은 화면 오른쪽 아래의 바닥을 향해 얼굴과 눈을 내리고 있다. 손으로 잡은 휴대전화의 화면 쪽은 귀를 향하며, 카메라에는 뒷면과 가장자리가 보인다. 왼쪽 벽의 칼은 날이 벽 안으로 들어가고 손잡이가 실내 쪽으로 돌출되어 있다.",
        "built_space": "벗겨진 회백색 투톤 벽, 뒤쪽 금속 책상 하나와 검은 사무용 의자 하나, 오른쪽 뒤 문 하나, 왼쪽 뒤 수납장 하나가 보인다. 앞쪽에는 회의용 전화기가 놓인 책상 하나가 있고, 인물 뒤 왼쪽에는 접이식 의자 일부, 오른쪽 아래에는 다른 의자 일부가 보인다. 벽의 칼은 하나이며 참고 이미지와 같은 쪽에 있다. 박철진은 책상 앞 빈 공간에 서 있다. 기존 장소의 주요 재질과 배치를 대체로 유지하며 불가능한 반사는 없다.",
        "entities": "등장인물은 중년 한국인 남성으로 보이는 박철진 한 명이다. 짧게 넘긴 검은 머리, 남색 정장, 흰 셔츠와 사선무늬 넥타이가 참고 이미지에 부합한다. 숙인 얼굴의 보이는 부분도 참고 인물과 대체로 일치한다. 귀에 댄 휴대전화 하나와 벽에 박힌 칼 하나가 있다. 읽을 수 있는 화면 내용이나 뚜렷한 문자는 보이지 않는다. 무릎 아래와 발은 프레임 밖이므로 요청한 전신이 아니다.",
        "hard_violations": [],
        "physics": "휴대전화는 박철진의 손가락과 손바닥에 잡혀 귀에 밀착되어 있다. 칼은 벽에 박힌 날이 지지한다. 몸은 아래쪽 프레임 밖으로 이어지는 다리로 지탱하는 자연스러운 선 자세로 보이며, 공중에 떠 있다는 증거는 없다. 다만 발의 접지는 화면에서 확인할 수 없다. 허리를 굽히고 고개를 내린 동작은 가능하지만 몸통이 수평까지 내려가지 않아 명시한 90도 인사와 다르다. 손의 떨림은 정지 화면에서 뚜렷하게 드러나지 않는다."
       },
       {
        "label": "A",
        "direction": "박철진의 고개와 시선은 오른쪽 아래 바닥을 향한다. 휴대전화는 귀에 붙어 있고 후면 카메라가 있는 뒷면이 관객을 향하여 통화 용도에 맞는다. 왼쪽 벽의 칼은 날이 벽을 향하고 손잡이가 실내로 나와 있다.",
        "built_space": "뒤쪽 금속 책상 하나와 그 뒤 회전의자 하나, 오른쪽 뒤 문 하나, 왼쪽 뒤 높은 수납장 하나, 오른쪽 낮은 수납장 하나가 보인다. 오른쪽 아래에는 앞쪽 책상 모서리만 들어온다. 왼쪽 벽에는 칼 하나와 액자 두 개, 종이 게시물 하나가 보이고 오른쪽 벽에도 액자 일부가 있다. 인물은 책상 앞의 열린 바닥 공간에 서 있다. 낡은 투톤 벽과 금속 가구는 참고 장소와 대체로 맞지만, 왼쪽 위 액자의 문구가 읽히도록 노출되어 있다.",
        "entities": "박철진으로 보이는 중년 한국인 남성 한 명만 등장한다. 검은 머리와 남색 정장, 흰 셔츠는 참고와 일치하며 얼굴은 숙인 측면으로 일부만 확인된다. 넥타이는 자세와 손에 가려 세부 비교가 어렵다. 휴대전화 하나와 벽에 박힌 칼 하나가 존재한다. A보다 다리를 많이 보여 주지만 발과 신발은 하단 밖으로 잘려 전신 조건을 충족하지 않는다. 왼쪽 위 액자에는 식별 가능한 한글 문구가 있다.",
        "hard_violations": [
         "왼쪽 위 액자에 읽을 수 있는 한글 문구가 노출되어, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 명시적 금지 조건을 위반한다."
        ],
        "physics": "휴대전화는 귀 옆에서 손으로 확실히 잡고 있으며, 벽의 칼은 박힌 날로 지지된다. 인물의 다리는 바닥 방향으로 이어져 정상적으로 서서 굽힌 자세로 보인다. 발은 잘려 접지점을 직접 확인할 수 없지만 무지지 부유나 불가능한 관절은 보이지 않는다. 몸통은 여전히 대각선으로 기울어 있어 허리를 90도로 접은 동작에 못 미친다. 손 떨림도 명확하게 확인되지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.4,
    "B": 1.286
   },
   "adjusted": {
    "A": 1.15,
    "B": 1.036
   },
   "violations": {
    "B": [
     "[gemini-pro] 왼쪽 어깨에서 나온 팔이 휴대폰을 들고 있음에도, 같은 쪽 몸통 아래로 주먹을 쥔 또 다른 팔이 늘어져 있는 신체 중복(팔 3개) 오류"
    ],
    "A": [
     "[gpt-high] 왼쪽 위 액자에 읽을 수 있는 한글 문구가 노출되어, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 명시적 금지 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1150,
   "B": 1036
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1150,
    "verdict_ko": "요구된 '전신(wide shot)'에 가깝게 앵글을 잡아냈으며, 신체 오류 없이 90도로 굽힌 비굴한 자세와 배경의 연속성을 충실히 구현했습니다.  ★위반: [gpt-high] 왼쪽 위 액자에 읽을 수 있는 한글 문구가 노출되어, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 명시적 금지 조건을 위반한다."
   },
   {
    "label": "B",
    "score": 1036,
    "verdict_ko": "같은 쪽(왼쪽)에 휴대폰을 든 팔과 아래로 늘어뜨린 팔이 동시에 존재하는 치명적인 신체 중복 오류가 발생했습니다.  ★위반: [gemini-pro] 왼쪽 어깨에서 나온 팔이 휴대폰을 들고 있음에도, 같은 쪽 몸통 아래로 주먹을 쥔 또 다른 팔이 늘어져 있는 신체 중복(팔 3개) 오류"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S18sh6_sel.png",
    "asset_id": "fa1b24a9-e6f4-40dc-be6c-1c60928703e0",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1401722>",
    "asset_id": "fee7383c-fb61-4b3a-ba7c-79f2555de00b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-a4a7-71de-ab8f-a2c435009b9e",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S18sh6"
  }
 },
 "S43sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:31:48.747623+00:00",
  "fingerprint": "078abc2f538218a6b0d9cf05f9a13f777676c2ed7df71ea353965b8cef4a6a14",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S43sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S43sh4_sel.png",
  "source_sha256": "e5ea70c8bf7f8ab1378ebb0f63c6bff8d0c721e46a1234f7d18dbf0834f4c262",
  "file": "S43sh4_cine.png",
  "staged_sha256": "7a4417bc170e7d2cf88229b3d81276668e13f460a2d658924e17b93997c60e98",
  "latency_ms": 14232
 },
 "S43sh9::signage": {
  "fp": "b893f85bc7bfe54d",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S43sh9": {
  "input_fingerprint": "d00749d1ee061d0d",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 야욕에 찬 핏발 선 눈으로 사악한 미소를 띤 박철진의 얼굴 클로즈업.\n\nLOCATION (lock): Inside the militia commander's office, near his subordinate after the phone call, under ordinary office lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the office's neutral ambient treatment, keeping the eyes and forming smile legible through controlled tonal contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the militia commander's office surfaces and nighttime interior lighting visible in the reference. Exclude the minister's sofa and furnishings from the other end of the telephone conversation.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The knife remains embedded in the militia-office wall. 박철진: He has ended the telephone call and remains in his office with an ambitious smile.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 야욕에 찬 핏발 선 눈으로 사악한 미소를 띤 박철진의 얼굴 클로즈업.\n\nLOCATION (lock): Inside the militia commander's office, near his subordinate after the phone call, under ordinary office lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the office's neutral ambient treatment, keeping the eyes and forming smile legible through controlled tonal contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the militia commander's office surfaces and nighttime interior lighting visible in the reference. Exclude the minister's sofa and furnishings from the other end of the telephone conversation.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The knife remains embedded in the militia-office wall. 박철진: He has ended the telephone call and remains in his office with an ambitious smile.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 야욕에 찬 핏발 선 눈으로 사악한 미소를 띤 박철진의 얼굴 클로즈업.\n\nLOCATION (lock): Inside the militia commander's office, near his subordinate after the phone call, under ordinary office lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the office's neutral ambient treatment, keeping the eyes and forming smile legible through controlled tonal contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the militia commander's office surfaces and nighttime interior lighting visible in the reference. Exclude the minister's sofa and furnishings from the other end of the telephone conversation.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The knife remains embedded in the militia-office wall. 박철진: He has ended the telephone call and remains in his office with an ambitious smile.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "인물의 시선이 정면의 카메라를 향하고 있음.",
    "built_space": "투톤 벽면의 사무실 내부. 왼쪽 뒤에 철제 캐비닛, 오른쪽에 액자가 보이나 공간의 연속성이 부족함.",
    "entities": "박철진. 핏발 선 눈과 크게 웃는 미소. 남색 정장과 흰 셔츠를 입었으나 넥타이가 없음.",
    "hard_violations": [
     "[gemini-pro] 이전 샷 및 캐릭터 레퍼런스로 고정된 복장 중 넥타이 누락"
    ],
    "physics": "자연스러운 상체의 자세를 유지하며 지지 문제가 없음."
   },
   {
    "label": "B",
    "direction": "인물의 시선이 정면의 카메라를 똑바로 응시함.",
    "built_space": "사무실 내부. 왼쪽 투톤 벽에 꽂힌 칼, 오른쪽 배경의 문과 책상 등 이전 샷의 공간 구조와 가구 배치가 일치함.",
    "entities": "박철진. 레퍼런스와 일치하는 이목구비, 핏발 선 눈, 사악한 미소. 남색 정장, 흰 셔츠, 넥타이, 귀걸이 등 복장이 완벽히 일치함.",
    "hard_violations": [],
    "physics": "자연스럽게 서 있는 상체 자세로 특별한 지지 문제가 없음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "클로즈업 프레임 안에서 복장 고정 지침을 정확히 지켰으며, 이전 샷의 배경(벽에 꽂힌 칼, 구조)과 요구된 표정 연기를 완벽하게 재현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "표정은 프롬프트를 따랐으나, 고정 지침인 넥타이가 누락되었고 유지되어야 할 배경 소품(칼)이 프레임에 나타나지 않았습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "인물의 시선이 정면의 카메라를 똑바로 응시함.",
        "built_space": "사무실 내부. 왼쪽 투톤 벽에 꽂힌 칼, 오른쪽 배경의 문과 책상 등 이전 샷의 공간 구조와 가구 배치가 일치함.",
        "entities": "박철진. 레퍼런스와 일치하는 이목구비, 핏발 선 눈, 사악한 미소. 남색 정장, 흰 셔츠, 넥타이, 귀걸이 등 복장이 완벽히 일치함.",
        "hard_violations": [],
        "physics": "자연스럽게 서 있는 상체 자세로 특별한 지지 문제가 없음."
       },
       {
        "label": "A",
        "direction": "인물의 시선이 정면의 카메라를 향하고 있음.",
        "built_space": "투톤 벽면의 사무실 내부. 왼쪽 뒤에 철제 캐비닛, 오른쪽에 액자가 보이나 공간의 연속성이 부족함.",
        "entities": "박철진. 핏발 선 눈과 크게 웃는 미소. 남색 정장과 흰 셔츠를 입었으나 넥타이가 없음.",
        "hard_violations": [
         "이전 샷 및 캐릭터 레퍼런스로 고정된 복장 중 넥타이 누락"
        ],
        "physics": "자연스러운 상체의 자세를 유지하며 지지 문제가 없음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "클로즈업 프레임 안에서 복장 고정 지침을 정확히 지켰으며, 이전 샷의 배경(벽에 꽂힌 칼, 구조)과 요구된 표정 연기를 완벽하게 재현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "표정은 프롬프트를 따랐으나, 고정 지침인 넥타이가 누락되었고 유지되어야 할 배경 소품(칼)이 프레임에 나타나지 않았습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "인물의 시선이 정면의 카메라를 똑바로 응시함.",
        "built_space": "사무실 내부. 왼쪽 투톤 벽에 꽂힌 칼, 오른쪽 배경의 문과 책상 등 이전 샷의 공간 구조와 가구 배치가 일치함.",
        "entities": "박철진. 레퍼런스와 일치하는 이목구비, 핏발 선 눈, 사악한 미소. 남색 정장, 흰 셔츠, 넥타이, 귀걸이 등 복장이 완벽히 일치함.",
        "hard_violations": [],
        "physics": "자연스럽게 서 있는 상체 자세로 특별한 지지 문제가 없음."
       },
       {
        "label": "A",
        "direction": "인물의 시선이 정면의 카메라를 향하고 있음.",
        "built_space": "투톤 벽면의 사무실 내부. 왼쪽 뒤에 철제 캐비닛, 오른쪽에 액자가 보이나 공간의 연속성이 부족함.",
        "entities": "박철진. 핏발 선 눈과 크게 웃는 미소. 남색 정장과 흰 셔츠를 입었으나 넥타이가 없음.",
        "hard_violations": [
         "이전 샷 및 캐릭터 레퍼런스로 고정된 복장 중 넥타이 누락"
        ],
        "physics": "자연스러운 상체의 자세를 유지하며 지지 문제가 없음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "인물과 사무실의 연속성은 좋지만, B보다 얼굴 비중이 작고 충혈과 사악한 미소가 약해 핵심 표정 지시를 덜 충실하게 구현한다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "얼굴을 밀착해 담은 클로즈업에서 자연스럽게 충혈된 눈과 야욕 어린 미소가 선명하며, 인물 외형과 중립적인 사무실 조명도 유지한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진은 얼굴을 정면으로 두고 렌즈 쪽을 바라본다. 시선이 향하는 별도 인물은 보이지 않으며, 원문도 특정 시선 대상을 지정하지 않는다. 왼쪽 벽의 칼은 날이 벽에 들어가고 손잡이가 오른쪽으로 돌출된 모습이다.",
        "built_space": "왼쪽 벽에는 잘린 액자 하나와 세로 게시판 하나, 흐릿한 종이와 꽂힌 칼 하나가 보인다. 뒤에는 회색 문 하나, 책상 일부와 검은 의자 등받이 하나, 오른쪽 수납장 일부가 있으며 천장 조명 하나가 보인다. 박철진은 책상 앞쪽에 있고, 낡은 미색·회색 벽과 비품의 배치는 이전 장면과 대체로 연결된다. 반사상이나 명백히 중복된 설비는 없다.",
        "entities": "보이는 사람은 박철진 한 명이다. 한국인 중년 남성 설정에 부합하는 외형이며, 짧게 넘긴 검은 머리, 얼굴 윤곽, 작은 귀걸이, 남색 정장과 흰 셔츠, 사선무늬 넥타이가 참조와 잘 맞는다. 눈은 정상적인 홍채와 동공을 유지하나 충혈이 약하고, 입을 다문 미소는 사악함보다는 절제된 만족감에 가깝다. 전화기는 보이지 않으며 통화 중인 동작도 없다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리는 목과 어깨에 자연스럽게 연결되어 지지된다. 하체와 손은 클로즈업 밖이므로 지면 접촉이나 소지 상태는 확인할 수 없다. 칼은 벽에 박힌 날로 지지되며, 배경 비품에도 공중 부유나 불가능한 접촉은 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "얼굴은 거의 정면이고 두 눈은 렌즈 가까운 방향을 응시한다. 화면 안에 별도 시선 대상은 없으며, 미소와 집중된 눈빛 때문에 멍한 정면 응시로 읽히지는 않는다. 겨누거나 움직이는 물체는 보이지 않는다.",
        "built_space": "배경에는 미색 상부와 회색 하부로 나뉜 낡은 벽, 왼쪽 철제 수납장 하나의 일부, 오른쪽 세로 게시판 하나가 보인다. 참조의 재료와 비품 종류는 이어지지만, 좁은 배경만으로 정확한 상대 배치까지 확정하기는 어렵다. 문·책상·의자·벽의 칼은 프레임 밖이며, 이를 담기 위해 얼굴 클로즈업을 넓히지 않았다. 중복 설비나 부자연스러운 반사는 없다.",
        "entities": "박철진 한 명만 등장한다. 참조와 닮은 중년 한국인 남성의 얼굴, 짧은 검은 머리, 남색 정장과 흰 셔츠가 보인다. 넥타이와 귀 장식은 이 구도에서 확인하기 어렵다. 정상적인 홍채와 동공 주변의 붉어진 흰자와 눈가가 선명하고, 치아를 조금 드러낸 미소가 야욕과 사악함을 더 직접적으로 전달한다. 피부와 의복은 실물 질감이며, 게시판 글씨는 읽을 수 없고 통화 중인 소품이나 추가 인물도 없다.",
        "hard_violations": [],
        "physics": "얼굴 아래 목과 어깨가 이어져 머리를 정상적으로 지지한다. 눈꺼풀과 입 주변 근육으로 만들어진 미소는 실제 배우가 가능한 표정이다. 손과 하체는 화면 밖이고, 보이는 수납장과 벽 게시판에도 지지 없이 떠 있는 부분은 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "인물과 사무실의 연속성은 좋지만, B보다 얼굴 비중이 작고 충혈과 사악한 미소가 약해 핵심 표정 지시를 덜 충실하게 구현한다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "얼굴을 밀착해 담은 클로즈업에서 자연스럽게 충혈된 눈과 야욕 어린 미소가 선명하며, 인물 외형과 중립적인 사무실 조명도 유지한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "박철진은 얼굴을 정면으로 두고 렌즈 쪽을 바라본다. 시선이 향하는 별도 인물은 보이지 않으며, 원문도 특정 시선 대상을 지정하지 않는다. 왼쪽 벽의 칼은 날이 벽에 들어가고 손잡이가 오른쪽으로 돌출된 모습이다.",
        "built_space": "왼쪽 벽에는 잘린 액자 하나와 세로 게시판 하나, 흐릿한 종이와 꽂힌 칼 하나가 보인다. 뒤에는 회색 문 하나, 책상 일부와 검은 의자 등받이 하나, 오른쪽 수납장 일부가 있으며 천장 조명 하나가 보인다. 박철진은 책상 앞쪽에 있고, 낡은 미색·회색 벽과 비품의 배치는 이전 장면과 대체로 연결된다. 반사상이나 명백히 중복된 설비는 없다.",
        "entities": "보이는 사람은 박철진 한 명이다. 한국인 중년 남성 설정에 부합하는 외형이며, 짧게 넘긴 검은 머리, 얼굴 윤곽, 작은 귀걸이, 남색 정장과 흰 셔츠, 사선무늬 넥타이가 참조와 잘 맞는다. 눈은 정상적인 홍채와 동공을 유지하나 충혈이 약하고, 입을 다문 미소는 사악함보다는 절제된 만족감에 가깝다. 전화기는 보이지 않으며 통화 중인 동작도 없다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리는 목과 어깨에 자연스럽게 연결되어 지지된다. 하체와 손은 클로즈업 밖이므로 지면 접촉이나 소지 상태는 확인할 수 없다. 칼은 벽에 박힌 날로 지지되며, 배경 비품에도 공중 부유나 불가능한 접촉은 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "얼굴은 거의 정면이고 두 눈은 렌즈 가까운 방향을 응시한다. 화면 안에 별도 시선 대상은 없으며, 미소와 집중된 눈빛 때문에 멍한 정면 응시로 읽히지는 않는다. 겨누거나 움직이는 물체는 보이지 않는다.",
        "built_space": "배경에는 미색 상부와 회색 하부로 나뉜 낡은 벽, 왼쪽 철제 수납장 하나의 일부, 오른쪽 세로 게시판 하나가 보인다. 참조의 재료와 비품 종류는 이어지지만, 좁은 배경만으로 정확한 상대 배치까지 확정하기는 어렵다. 문·책상·의자·벽의 칼은 프레임 밖이며, 이를 담기 위해 얼굴 클로즈업을 넓히지 않았다. 중복 설비나 부자연스러운 반사는 없다.",
        "entities": "박철진 한 명만 등장한다. 참조와 닮은 중년 한국인 남성의 얼굴, 짧은 검은 머리, 남색 정장과 흰 셔츠가 보인다. 넥타이와 귀 장식은 이 구도에서 확인하기 어렵다. 정상적인 홍채와 동공 주변의 붉어진 흰자와 눈가가 선명하고, 치아를 조금 드러낸 미소가 야욕과 사악함을 더 직접적으로 전달한다. 피부와 의복은 실물 질감이며, 게시판 글씨는 읽을 수 없고 통화 중인 소품이나 추가 인물도 없다.",
        "hard_violations": [],
        "physics": "얼굴 아래 목과 어깨가 이어져 머리를 정상적으로 지지한다. 눈꺼풀과 입 주변 근육으로 만들어진 미소는 실제 배우가 가능한 표정이다. 손과 하체는 화면 밖이고, 보이는 수납장과 벽 게시판에도 지지 없이 떠 있는 부분은 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.429,
    "B": 1.778
   },
   "adjusted": {
    "A": 1.179,
    "B": 1.778
   },
   "violations": {
    "A": [
     "[gemini-pro] 이전 샷 및 캐릭터 레퍼런스로 고정된 복장 중 넥타이 누락"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1778,
   "A": 1179
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1778,
    "verdict_ko": "클로즈업 프레임 안에서 복장 고정 지침을 정확히 지켰으며, 이전 샷의 배경(벽에 꽂힌 칼, 구조)과 요구된 표정 연기를 완벽하게 재현했습니다."
   },
   {
    "label": "A",
    "score": 1179,
    "verdict_ko": "표정은 프롬프트를 따랐으나, 고정 지침인 넥타이가 누락되었고 유지되어야 할 배경 소품(칼)이 프레임에 나타나지 않았습니다.  ★위반: [gemini-pro] 이전 샷 및 캐릭터 레퍼런스로 고정된 복장 중 넥타이 누락"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S43sh4_sel.png",
    "asset_id": "ba00d484-5923-40d8-a2ef-153c8bc84add",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1401722>",
    "asset_id": "fee7383c-fb61-4b3a-ba7c-79f2555de00b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-a65f-79df-a524-a7600cec7eb9",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S43sh4"
  }
 },
 "S43sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:32:51.896904+00:00",
  "fingerprint": "93382f00917140cb359e5aeb90226f89252c428f8a6392f4f002469624e1150f",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S43sh9_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S43sh9_sel.png",
  "source_sha256": "a29abc3129d6071d641b4a6a0e83afb96d48bf376392d2b003264915b37623d2",
  "file": "S43sh9_cine.png",
  "staged_sha256": "373b5d24ba7e5b4e39cf6322ed8a227877f10e665efc7ebd8a7bc924d4a87a5e",
  "latency_ms": 21790
 },
 "S44sh4::signage": {
  "fp": "94498938fa0032fc",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S44sh4": {
  "input_fingerprint": "658978ed3c0a8c9c",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 창고 어둠 속에서 먼지를 뒤집어쓴 채 서 있는 아담한 소형 캠핑카의 실루엣 전경.\n\nLOCATION (lock): In the dark vehicle-storage bay of a warehouse beside an abandoned factory, behind its raised shutter. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 소형 캠핑카 (Stationary inside the dark warehouse and covered in dust; its rear door has not yet been opened) — Seen diagonally with a side and the rear end readable, preserving the route toward the rear door; used as Primary reveal subject, contained within surrounding negative space rather than enlarged to fill the image; 창고 셔터 입구 (Raised to admit the group and camera) — Only a peripheral portion of the opening remains visible from just inside the entrance; used as Marks the completed threshold crossing and anchors the reveal to the approach path.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the warehouse darkness, using only enough ambient tonal separation to read the camper's silhouette and dust-covered surfaces without inventing a beam or fixture.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The warehouse shutter has been raised, revealing a small camper in the darkness. The scavenged bags still contain supplies, medicine and the collected cash.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 창고 어둠 속에서 먼지를 뒤집어쓴 채 서 있는 아담한 소형 캠핑카의 실루엣 전경.\n\nLOCATION (lock): In the dark vehicle-storage bay of a warehouse beside an abandoned factory, behind its raised shutter. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 소형 캠핑카 (Stationary inside the dark warehouse and covered in dust; its rear door has not yet been opened) — Seen diagonally with a side and the rear end readable, preserving the route toward the rear door; used as Primary reveal subject, contained within surrounding negative space rather than enlarged to fill the image; 창고 셔터 입구 (Raised to admit the group and camera) — Only a peripheral portion of the opening remains visible from just inside the entrance; used as Marks the completed threshold crossing and anchors the reveal to the approach path.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the warehouse darkness, using only enough ambient tonal separation to read the camper's silhouette and dust-covered surfaces without inventing a beam or fixture.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The warehouse shutter has been raised, revealing a small camper in the darkness. The scavenged bags still contain supplies, medicine and the collected cash.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 창고 어둠 속에서 먼지를 뒤집어쓴 채 서 있는 아담한 소형 캠핑카의 실루엣 전경.\n\nLOCATION (lock): In the dark vehicle-storage bay of a warehouse beside an abandoned factory, behind its raised shutter. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 소형 캠핑카 (Stationary inside the dark warehouse and covered in dust; its rear door has not yet been opened) — Seen diagonally with a side and the rear end readable, preserving the route toward the rear door; used as Primary reveal subject, contained within surrounding negative space rather than enlarged to fill the image; 창고 셔터 입구 (Raised to admit the group and camera) — Only a peripheral portion of the opening remains visible from just inside the entrance; used as Marks the completed threshold crossing and anchors the reveal to the approach path.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the warehouse darkness, using only enough ambient tonal separation to read the camper's silhouette and dust-covered surfaces without inventing a beam or fixture.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The warehouse shutter has been raised, revealing a small camper in the darkness. The scavenged bags still contain supplies, medicine and the collected cash.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라가 창고 내부에서 캠핑카의 측면과 후면을 대각선으로 향해 바라봄.",
    "built_space": "어두운 창고 내부, 왼쪽 가장자리에 열린 셔터의 일부가 보이며 바닥에 가방들이 놓여 있음.",
    "entities": "먼지를 뒤집어쓴 소형 캠핑카, 바닥의 가방 묶음. 지시대로 인물 없음.",
    "hard_violations": [
     "[gpt-high] 장소 설명과 샷에 없는 다수의 상자, 수납장, 작업대 및 적치물을 만들어 넣어, 명시된 요소 외에는 발명하지 말라는 장소 제한을 위반한다."
    ],
    "physics": "캠핑카와 가방들이 콘크리트 바닥에 자연스럽게 지지되어 있음."
   },
   {
    "label": "B",
    "direction": "카메라가 창고 외부(또는 경계)에서 내부의 캠핑카를 정면으로 바라봄.",
    "built_space": "화면 테두리를 감싸고 있는 열린 셔터, 창고 내부 공간.",
    "entities": "먼지를 뒤집어쓴 소형 캠핑카. 프롬프트에 명시된 가방이 보이지 않음. 인물 없음.",
    "hard_violations": [
     "[gemini-pro] 카메라 위치 위반: '입구 바로 안쪽에서 주변부만 보여야 하는' 셔터를 외부에서 정면 전체 프레임으로 렌더링함.",
     "[gpt-high] 입구 바로 안쪽으로 지정된 카메라를 창고 외부에 배치하여, 문턱을 이미 통과했다는 고정된 시점을 위반한다."
    ],
    "physics": "캠핑카가 바닥에 안정적으로 정차해 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "카메라 위치(내부에서 주변부 셔터만 보임), 캠핑카의 대각선 구도 및 먼지 덮인 질감, 바닥의 가방 등 모든 프롬프트 요구사항을 훌륭하게 충족함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "지정된 카메라 위치(내부)를 어기고 외부에서 정면으로 촬영했으며, 필수 요소인 바닥의 가방이 누락됨."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라가 창고 내부에서 캠핑카의 측면과 후면을 대각선으로 향해 바라봄.",
        "built_space": "어두운 창고 내부, 왼쪽 가장자리에 열린 셔터의 일부가 보이며 바닥에 가방들이 놓여 있음.",
        "entities": "먼지를 뒤집어쓴 소형 캠핑카, 바닥의 가방 묶음. 지시대로 인물 없음.",
        "hard_violations": [],
        "physics": "캠핑카와 가방들이 콘크리트 바닥에 자연스럽게 지지되어 있음."
       },
       {
        "label": "B",
        "direction": "카메라가 창고 외부(또는 경계)에서 내부의 캠핑카를 정면으로 바라봄.",
        "built_space": "화면 테두리를 감싸고 있는 열린 셔터, 창고 내부 공간.",
        "entities": "먼지를 뒤집어쓴 소형 캠핑카. 프롬프트에 명시된 가방이 보이지 않음. 인물 없음.",
        "hard_violations": [
         "카메라 위치 위반: '입구 바로 안쪽에서 주변부만 보여야 하는' 셔터를 외부에서 정면 전체 프레임으로 렌더링함."
        ],
        "physics": "캠핑카가 바닥에 안정적으로 정차해 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "카메라 위치(내부에서 주변부 셔터만 보임), 캠핑카의 대각선 구도 및 먼지 덮인 질감, 바닥의 가방 등 모든 프롬프트 요구사항을 훌륭하게 충족함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "지정된 카메라 위치(내부)를 어기고 외부에서 정면으로 촬영했으며, 필수 요소인 바닥의 가방이 누락됨."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라가 창고 내부에서 캠핑카의 측면과 후면을 대각선으로 향해 바라봄.",
        "built_space": "어두운 창고 내부, 왼쪽 가장자리에 열린 셔터의 일부가 보이며 바닥에 가방들이 놓여 있음.",
        "entities": "먼지를 뒤집어쓴 소형 캠핑카, 바닥의 가방 묶음. 지시대로 인물 없음.",
        "hard_violations": [],
        "physics": "캠핑카와 가방들이 콘크리트 바닥에 자연스럽게 지지되어 있음."
       },
       {
        "label": "B",
        "direction": "카메라가 창고 외부(또는 경계)에서 내부의 캠핑카를 정면으로 바라봄.",
        "built_space": "화면 테두리를 감싸고 있는 열린 셔터, 창고 내부 공간.",
        "entities": "먼지를 뒤집어쓴 소형 캠핑카. 프롬프트에 명시된 가방이 보이지 않음. 인물 없음.",
        "hard_violations": [
         "카메라 위치 위반: '입구 바로 안쪽에서 주변부만 보여야 하는' 셔터를 외부에서 정면 전체 프레임으로 렌더링함."
        ],
        "physics": "캠핑카가 바닥에 안정적으로 정차해 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "어둠 속 작은 캠핑카와 닫힌 후면 문은 잘 보이지만, 카메라가 창고 밖에서 입구 전체를 바라봐 입구를 이미 통과한 내부 시점을 위반한다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "입구 안쪽 시점과 가장자리에 남은 셔터는 더 정확하지만, 지시되지 않은 상자·수납장·작업대 등으로 공간을 채워 엄격한 장소 제한을 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "캠핑카 앞부분은 화면 왼쪽 안쪽을 향하고, 후면은 오른쪽의 카메라 쪽으로 향한다. 측면과 닫힌 후면 문이 함께 보이며 문 앞으로 접근할 바닥도 비어 있다. 사람이나 시선, 조준 대상은 없다.",
        "built_space": "셔터 입구 하나의 양쪽 기둥과 상단 셔터가 모두 보인다. 전경 바닥에서 문턱 너머 차량 보관실을 바라보는 외부 시점으로, 입구 안쪽에서 개구부 일부만 보이라는 조건과 다르다. 캠핑카 한 대 주변에는 충분한 어두운 여백이 있다.",
        "entities": "먼지와 때가 덮인 소형 캠핑카 한 대, 닫힌 후면 출입문 하나, 올라간 셔터 하나가 보인다. 사람은 없고 읽을 수 있는 글자도 없다. 별도 광원이나 빛줄기 없이 어두운 차체가 구분된다. 물품 가방과 그 내용물은 보이지 않으며, 이 구도에서 반드시 보여야 하는 대상은 아니다.",
        "hard_violations": [
         "입구 바로 안쪽으로 지정된 카메라를 창고 외부에 배치하여, 문턱을 이미 통과했다는 고정된 시점을 위반한다."
        ],
        "physics": "보이는 앞뒤 타이어가 콘크리트 바닥에 닿아 차체를 지지한다. 차량은 정지해 있으며 떠 있는 물체나 불가능한 자세는 없다. 셔터는 입구의 측면 레일과 상부 구조에 지지되어 있다."
       },
       {
        "label": "B",
        "direction": "캠핑카 앞부분은 왼쪽 안쪽으로, 후면은 오른쪽의 카메라 방향으로 향한다. 측면과 후면이 동시에 읽히며 후면으로 향하는 바닥 통로가 남아 있다. 다만 후면의 넓은 닫힌 패널은 출입문인지 창 덮개인지 명확하지 않다. 사람이나 조준 대상은 없다.",
        "built_space": "카메라는 창고 내부에 있고, 올라간 셔터 입구 하나가 왼쪽 가장자리에 일부 보인다. 외부의 어두운 하늘도 그 개구부로 보인다. 내부에는 철골 지붕과 벽 외에 여러 벽면 창, 수납장, 상자 더미와 작업대가 추가되어 있다. 캠핑카 주변 여백은 있지만 A보다 차체가 조금 크고 배경이 복잡하다.",
        "entities": "먼지로 덮인 소형 캠핑카 한 대와 올라간 셔터 하나가 보이고, 사람과 읽을 수 있는 글자는 없다. 왼쪽 전경에는 여러 가방이 놓여 있으나 약품·현금 등 내용물은 확인할 수 없다. 가방 자체는 지문에 언급되지만, 다수의 상자와 수납장 및 작업대는 지시되지 않은 추가 물체다. 밤의 어둠과 약한 주변광은 유지된다.",
        "hard_violations": [
         "장소 설명과 샷에 없는 다수의 상자, 수납장, 작업대 및 적치물을 만들어 넣어, 명시된 요소 외에는 발명하지 말라는 장소 제한을 위반한다."
        ],
        "physics": "캠핑카의 보이는 앞뒤 바퀴가 바닥에 닿아 차량을 지지한다. 가방은 바닥에 놓여 있고, 상자와 적치물도 바닥이나 가구 위에 지지되어 있다. 셔터는 측면 레일과 상부 구조에 연결되어 있으며, 공중에 무지지 상태로 떠 있는 대상은 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "어둠 속 작은 캠핑카와 닫힌 후면 문은 잘 보이지만, 카메라가 창고 밖에서 입구 전체를 바라봐 입구를 이미 통과한 내부 시점을 위반한다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "입구 안쪽 시점과 가장자리에 남은 셔터는 더 정확하지만, 지시되지 않은 상자·수납장·작업대 등으로 공간을 채워 엄격한 장소 제한을 위반한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "캠핑카 앞부분은 화면 왼쪽 안쪽을 향하고, 후면은 오른쪽의 카메라 쪽으로 향한다. 측면과 닫힌 후면 문이 함께 보이며 문 앞으로 접근할 바닥도 비어 있다. 사람이나 시선, 조준 대상은 없다.",
        "built_space": "셔터 입구 하나의 양쪽 기둥과 상단 셔터가 모두 보인다. 전경 바닥에서 문턱 너머 차량 보관실을 바라보는 외부 시점으로, 입구 안쪽에서 개구부 일부만 보이라는 조건과 다르다. 캠핑카 한 대 주변에는 충분한 어두운 여백이 있다.",
        "entities": "먼지와 때가 덮인 소형 캠핑카 한 대, 닫힌 후면 출입문 하나, 올라간 셔터 하나가 보인다. 사람은 없고 읽을 수 있는 글자도 없다. 별도 광원이나 빛줄기 없이 어두운 차체가 구분된다. 물품 가방과 그 내용물은 보이지 않으며, 이 구도에서 반드시 보여야 하는 대상은 아니다.",
        "hard_violations": [
         "입구 바로 안쪽으로 지정된 카메라를 창고 외부에 배치하여, 문턱을 이미 통과했다는 고정된 시점을 위반한다."
        ],
        "physics": "보이는 앞뒤 타이어가 콘크리트 바닥에 닿아 차체를 지지한다. 차량은 정지해 있으며 떠 있는 물체나 불가능한 자세는 없다. 셔터는 입구의 측면 레일과 상부 구조에 지지되어 있다."
       },
       {
        "label": "A",
        "direction": "캠핑카 앞부분은 왼쪽 안쪽으로, 후면은 오른쪽의 카메라 방향으로 향한다. 측면과 후면이 동시에 읽히며 후면으로 향하는 바닥 통로가 남아 있다. 다만 후면의 넓은 닫힌 패널은 출입문인지 창 덮개인지 명확하지 않다. 사람이나 조준 대상은 없다.",
        "built_space": "카메라는 창고 내부에 있고, 올라간 셔터 입구 하나가 왼쪽 가장자리에 일부 보인다. 외부의 어두운 하늘도 그 개구부로 보인다. 내부에는 철골 지붕과 벽 외에 여러 벽면 창, 수납장, 상자 더미와 작업대가 추가되어 있다. 캠핑카 주변 여백은 있지만 A보다 차체가 조금 크고 배경이 복잡하다.",
        "entities": "먼지로 덮인 소형 캠핑카 한 대와 올라간 셔터 하나가 보이고, 사람과 읽을 수 있는 글자는 없다. 왼쪽 전경에는 여러 가방이 놓여 있으나 약품·현금 등 내용물은 확인할 수 없다. 가방 자체는 지문에 언급되지만, 다수의 상자와 수납장 및 작업대는 지시되지 않은 추가 물체다. 밤의 어둠과 약한 주변광은 유지된다.",
        "hard_violations": [
         "장소 설명과 샷에 없는 다수의 상자, 수납장, 작업대 및 적치물을 만들어 넣어, 명시된 요소 외에는 발명하지 말라는 장소 제한을 위반한다."
        ],
        "physics": "캠핑카의 보이는 앞뒤 바퀴가 바닥에 닿아 차량을 지지한다. 가방은 바닥에 놓여 있고, 상자와 적치물도 바닥이나 가구 위에 지지되어 있다. 셔터는 측면 레일과 상부 구조에 연결되어 있으며, 공중에 무지지 상태로 떠 있는 대상은 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.125
   },
   "adjusted": {
    "A": 1.75,
    "B": 0.875
   },
   "violations": {
    "B": [
     "[gemini-pro] 카메라 위치 위반: '입구 바로 안쪽에서 주변부만 보여야 하는' 셔터를 외부에서 정면 전체 프레임으로 렌더링함.",
     "[gpt-high] 입구 바로 안쪽으로 지정된 카메라를 창고 외부에 배치하여, 문턱을 이미 통과했다는 고정된 시점을 위반한다."
    ],
    "A": [
     "[gpt-high] 장소 설명과 샷에 없는 다수의 상자, 수납장, 작업대 및 적치물을 만들어 넣어, 명시된 요소 외에는 발명하지 말라는 장소 제한을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 875
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "카메라 위치(내부에서 주변부 셔터만 보임), 캠핑카의 대각선 구도 및 먼지 덮인 질감, 바닥의 가방 등 모든 프롬프트 요구사항을 훌륭하게 충족함.  ★위반: [gpt-high] 장소 설명과 샷에 없는 다수의 상자, 수납장, 작업대 및 적치물을 만들어 넣어, 명시된 요소 외에는 발명하지 말라는 장소 제한을 위반한다."
   },
   {
    "label": "B",
    "score": 875,
    "verdict_ko": "지정된 카메라 위치(내부)를 어기고 외부에서 정면으로 촬영했으며, 필수 요소인 바닥의 가방이 누락됨.  ★위반: [gemini-pro] 카메라 위치 위반: '입구 바로 안쪽에서 주변부만 보여야 하는' 셔터를 외부에서 정면 전체 프레임으로 렌더링함. / [gpt-high] 입구 바로 안쪽으로 지정된 카메라를 창고 외부에 배치하여, 문턱을 이미 통과했다는 고정된 시점을 위반한다."
   }
  ],
  "refs": [],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-a809-7331-a3b9-92e53a9def6b",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S44sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:33:51.057437+00:00",
  "fingerprint": "77ec558b343017c29b5c0545c688bbc51ed5c28fc6cd86c0f04cab329ee9e339",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S44sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S44sh4_sel.png",
  "source_sha256": "6564b3d5174018a84441ae5bd4988e0c21093aa89a54685f6163cbfd1eb0e165",
  "file": "S44sh4_cine.png",
  "staged_sha256": "b1faf371ad06f14e387b2ca93c408b78e8cd95dbaf9903c7ebadfdc22c49616b",
  "latency_ms": 15186
 },
 "S44sh7::signage": {
  "fp": "caad744326f137ae",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S44sh7": {
  "input_fingerprint": "3b74cdc96c8f7470",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 열린 캠핑카 안으로 시선을 고정한 채 안도하는 표정으로 미소 짓고 있는 현우의 옆얼굴.\n\nLOCATION (lock): Inside the dark warehouse beside the abandoned factory, immediately outside the camper's open rear doorway. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Camper rear doorway (Open after 현우 has opened the rear door) — Seen obliquely beside his profile, with the opening leading toward the interior; used as Provides looking room and a narrow spatial boundary beside his face; Camper bed (Present inside the camper) — Only a small portion is visible through the rear opening; used as Gives concrete context to his relief without competing with his expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the warehouse's established darkness while keeping the small change in his expression legible through controlled tonal separation.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the dusty camper exterior and the surrounding dark storage warehouse. Exclude any assumption that the camper door remains closed; the rear door is now open.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The warehouse shutter remains raised, and the camper's rear door is now open. Its interior includes a bed and enough room for four. 현우: He stands at the open rear of the camper with his facial bruises and leg wound unchanged. The contact card remains concealed inside his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 열린 캠핑카 안으로 시선을 고정한 채 안도하는 표정으로 미소 짓고 있는 현우의 옆얼굴.\n\nLOCATION (lock): Inside the dark warehouse beside the abandoned factory, immediately outside the camper's open rear doorway. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Camper rear doorway (Open after 현우 has opened the rear door) — Seen obliquely beside his profile, with the opening leading toward the interior; used as Provides looking room and a narrow spatial boundary beside his face; Camper bed (Present inside the camper) — Only a small portion is visible through the rear opening; used as Gives concrete context to his relief without competing with his expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the warehouse's established darkness while keeping the small change in his expression legible through controlled tonal separation.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the dusty camper exterior and the surrounding dark storage warehouse. Exclude any assumption that the camper door remains closed; the rear door is now open.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The warehouse shutter remains raised, and the camper's rear door is now open. Its interior includes a bed and enough room for four. 현우: He stands at the open rear of the camper with his facial bruises and leg wound unchanged. The contact card remains concealed inside his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 열린 캠핑카 안으로 시선을 고정한 채 안도하는 표정으로 미소 짓고 있는 현우의 옆얼굴.\n\nLOCATION (lock): Inside the dark warehouse beside the abandoned factory, immediately outside the camper's open rear doorway. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Camper rear doorway (Open after 현우 has opened the rear door) — Seen obliquely beside his profile, with the opening leading toward the interior; used as Provides looking room and a narrow spatial boundary beside his face; Camper bed (Present inside the camper) — Only a small portion is visible through the rear opening; used as Gives concrete context to his relief without competing with his expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the warehouse's established darkness while keeping the small change in his expression legible through controlled tonal separation.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the dusty camper exterior and the surrounding dark storage warehouse. Exclude any assumption that the camper door remains closed; the rear door is now open.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The warehouse shutter remains raised, and the camper's rear door is now open. Its interior includes a bed and enough room for four. 현우: He stands at the open rear of the camper with his facial bruises and leg wound unchanged. The contact card remains concealed inside his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선이 화면 왼쪽의 열린 캠핑카 내부를 향하고 있습니다.",
    "built_space": "어두운 창고 안, 캠핑카의 측면 전체가 보이며 측면 문이 열려 있고 그 안으로 침대가 보입니다.",
    "entities": "현우의 얼굴과 헤어스타일이 레퍼런스와 일치하며, 뺨에 상처가 있고 안도하는 미소를 짓고 있습니다.",
    "hard_violations": [
     "[gpt-high] 후면의 열린 출입구 바로 바깥이라는 지정 위치를 측면 출입구와 그 앞쪽의 떨어진 전경으로 바꾸어, 현우와 차량의 필수 공간 배치를 어겼다."
    ],
    "physics": "현우가 지면에 안정적으로 서 있으며, 다른 물리적 오류는 없습니다."
   },
   {
    "label": "B",
    "direction": "현우의 시선이 화면 왼쪽 프레임 가장자리에 있는 열린 캠핑카 내부를 향하고 있습니다.",
    "built_space": "어두운 창고 안, 열린 문틀이 현우의 얼굴 바로 옆에서 좁은 경계를 형성하고 있으며 안쪽에 침대 일부가 보입니다.",
    "entities": "현우의 얼굴과 헤어스타일이 레퍼런스와 일치하며, 안도하며 미소 짓는 표정이 잘 드러납니다.",
    "hard_violations": [],
    "physics": "현우가 지면에 안정적으로 서 있으며, 다른 물리적 오류는 없습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "얼굴 바로 옆에 열린 캠핑카 문을 배치하여 시선 공간과 공간적 경계를 만들라는 클로즈업 프레이밍 지시를 완벽하게 준수했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "프레이밍이 지시된 클로즈업보다 훨씬 넓으며, 후방 문이 아닌 캠핑카의 측면 문을 보여주어 공간적 경계 조건을 충족하지 못했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선이 화면 왼쪽의 열린 캠핑카 내부를 향하고 있습니다.",
        "built_space": "어두운 창고 안, 캠핑카의 측면 전체가 보이며 측면 문이 열려 있고 그 안으로 침대가 보입니다.",
        "entities": "현우의 얼굴과 헤어스타일이 레퍼런스와 일치하며, 뺨에 상처가 있고 안도하는 미소를 짓고 있습니다.",
        "hard_violations": [],
        "physics": "현우가 지면에 안정적으로 서 있으며, 다른 물리적 오류는 없습니다."
       },
       {
        "label": "B",
        "direction": "현우의 시선이 화면 왼쪽 프레임 가장자리에 있는 열린 캠핑카 내부를 향하고 있습니다.",
        "built_space": "어두운 창고 안, 열린 문틀이 현우의 얼굴 바로 옆에서 좁은 경계를 형성하고 있으며 안쪽에 침대 일부가 보입니다.",
        "entities": "현우의 얼굴과 헤어스타일이 레퍼런스와 일치하며, 안도하며 미소 짓는 표정이 잘 드러납니다.",
        "hard_violations": [],
        "physics": "현우가 지면에 안정적으로 서 있으며, 다른 물리적 오류는 없습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "얼굴 바로 옆에 열린 캠핑카 문을 배치하여 시선 공간과 공간적 경계를 만들라는 클로즈업 프레이밍 지시를 완벽하게 준수했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "프레이밍이 지시된 클로즈업보다 훨씬 넓으며, 후방 문이 아닌 캠핑카의 측면 문을 보여주어 공간적 경계 조건을 충족하지 못했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선이 화면 왼쪽의 열린 캠핑카 내부를 향하고 있습니다.",
        "built_space": "어두운 창고 안, 캠핑카의 측면 전체가 보이며 측면 문이 열려 있고 그 안으로 침대가 보입니다.",
        "entities": "현우의 얼굴과 헤어스타일이 레퍼런스와 일치하며, 뺨에 상처가 있고 안도하는 미소를 짓고 있습니다.",
        "hard_violations": [],
        "physics": "현우가 지면에 안정적으로 서 있으며, 다른 물리적 오류는 없습니다."
       },
       {
        "label": "B",
        "direction": "현우의 시선이 화면 왼쪽 프레임 가장자리에 있는 열린 캠핑카 내부를 향하고 있습니다.",
        "built_space": "어두운 창고 안, 열린 문틀이 현우의 얼굴 바로 옆에서 좁은 경계를 형성하고 있으며 안쪽에 침대 일부가 보입니다.",
        "entities": "현우의 얼굴과 헤어스타일이 레퍼런스와 일치하며, 안도하며 미소 짓는 표정이 잘 드러납니다.",
        "hard_violations": [],
        "physics": "현우가 지면에 안정적으로 서 있으며, 다른 물리적 오류는 없습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "열린 문 바로 옆에서 내부를 바라보며 안도하는 옆얼굴 클로즈업을 충실히 구현했지만, 얼굴의 부상 흔적은 불분명하다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "표정과 얼굴 상처는 잘 보이지만, 후면 대신 측면 출입구를 보여주고 현우의 시선도 뒤쪽 개구부가 아닌 화면 왼쪽으로 지나간다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 화면 오른쪽에서 왼쪽의 열린 출입구를 향해 얼굴과 눈을 돌리고 있다. 시선이 문틀 사이의 어두운 내부로 이어져 캠핑카 안을 응시한다는 지시와 부합한다. 입꼬리가 올라가고 치아가 조금 보여 안도하며 미소 짓는 순간으로 읽힌다.",
        "built_space": "왼쪽에 외벽 일부, 중앙에 열린 출입구 하나, 그 오른쪽에 경첩으로 연결된 문짝 하나가 보인다. 내부 아래쪽에는 침대 한 개의 끝부분만 드러난다. 현우는 문틀 바로 바깥에 있고, 개구부가 얼굴 앞의 시선 여백을 만든다. 후면 전체와 창고 셔터는 클로즈업 밖이라 확인할 수 없다. 참조의 낡은 캠핑카와 어두운 창고 분위기는 유지되지만 외벽 표면의 세부 질감은 다소 다르다.",
        "entities": "인물은 한 명이며, 앳된 동아시아계 남성의 외모와 헝클어진 검은 머리, 남색 둥근 목 티셔츠가 현우의 참조와 대체로 일치한다. 국적은 외모만으로 확인할 수 없다. 얼굴의 멍이나 상처는 어둠 속에서 뚜렷하게 식별되지 않는다. 캠핑카와 침대가 보이며 추가 인물이나 읽을 수 있는 글자는 없다. 다리 상처와 신발 속 카드는 프레임 밖이다.",
        "hard_violations": [],
        "physics": "머리와 목, 어깨가 자연스럽게 연결되어 서 있는 상체로 읽힌다. 발은 프레임 밖이므로 지면 접촉은 확인되지 않지만 공중에 뜬 징후는 없다. 문짝은 보이는 경첩으로 차체에 연결되고, 침구는 침대 받침 위에 놓여 있다. 지지 없이 떠 있는 물체는 없다."
       },
       {
        "label": "B",
        "direction": "현우의 얼굴과 눈은 화면 왼쪽을 향하지만, 열린 출입구는 그의 뒤쪽 중경에 있다. 보이는 시선은 개구부 안으로 꺾이지 않고 차량 앞을 지나 화면 왼쪽 밖으로 이어지는 것으로 읽힌다. 미소는 분명하지만 캠핑카 내부를 응시하는 관계는 성립하지 않는다.",
        "built_space": "차량 측면의 창 하나, 환기구 두 개, 바퀴 하나, 낮은 수납 패널들, 열린 출입구 하나와 문짝 하나가 보인다. 출입구는 측면 바퀴와 창이 있는 같은 차체 면에 설치되어 후면 출입구로 읽히지 않는다. 내부에는 침대 일부와 상부 수납장이 보인다. 현우는 출입구 바로 옆보다 카메라 가까운 전경에 떨어져 있고, 차량 측면과 창고 바닥까지 상당 부분 드러나 문이 얼굴 옆의 좁은 경계가 되어야 한다는 구성에서 벗어난다.",
        "entities": "인물은 한 명이며 젊은 동아시아계 남성의 외모, 검은 흐트러진 머리와 남색 티셔츠가 참조에 대체로 맞는다. 뺨과 관자놀이에 상처가 뚜렷하다. 캠핑카와 침대는 식별되며 추가 인물이나 읽을 수 있는 문자는 없다. 다리와 신발은 보이지 않아 다리 상처와 숨긴 카드는 평가할 수 없다.",
        "hard_violations": [
         "후면의 열린 출입구 바로 바깥이라는 지정 위치를 측면 출입구와 그 앞쪽의 떨어진 전경으로 바꾸어, 현우와 차량의 필수 공간 배치를 어겼다."
        ],
        "physics": "현우의 목과 어깨는 자연스럽고 하체는 프레임 밖이다. 차량은 보이는 타이어를 통해 바닥에 지지되며, 열린 문은 경첩으로 차체에 연결된다. 침구는 침대 위에 놓여 있다. 부유하거나 해부학적으로 불가능한 요소는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "열린 문 바로 옆에서 내부를 바라보며 안도하는 옆얼굴 클로즈업을 충실히 구현했지만, 얼굴의 부상 흔적은 불분명하다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "표정과 얼굴 상처는 잘 보이지만, 후면 대신 측면 출입구를 보여주고 현우의 시선도 뒤쪽 개구부가 아닌 화면 왼쪽으로 지나간다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 화면 오른쪽에서 왼쪽의 열린 출입구를 향해 얼굴과 눈을 돌리고 있다. 시선이 문틀 사이의 어두운 내부로 이어져 캠핑카 안을 응시한다는 지시와 부합한다. 입꼬리가 올라가고 치아가 조금 보여 안도하며 미소 짓는 순간으로 읽힌다.",
        "built_space": "왼쪽에 외벽 일부, 중앙에 열린 출입구 하나, 그 오른쪽에 경첩으로 연결된 문짝 하나가 보인다. 내부 아래쪽에는 침대 한 개의 끝부분만 드러난다. 현우는 문틀 바로 바깥에 있고, 개구부가 얼굴 앞의 시선 여백을 만든다. 후면 전체와 창고 셔터는 클로즈업 밖이라 확인할 수 없다. 참조의 낡은 캠핑카와 어두운 창고 분위기는 유지되지만 외벽 표면의 세부 질감은 다소 다르다.",
        "entities": "인물은 한 명이며, 앳된 동아시아계 남성의 외모와 헝클어진 검은 머리, 남색 둥근 목 티셔츠가 현우의 참조와 대체로 일치한다. 국적은 외모만으로 확인할 수 없다. 얼굴의 멍이나 상처는 어둠 속에서 뚜렷하게 식별되지 않는다. 캠핑카와 침대가 보이며 추가 인물이나 읽을 수 있는 글자는 없다. 다리 상처와 신발 속 카드는 프레임 밖이다.",
        "hard_violations": [],
        "physics": "머리와 목, 어깨가 자연스럽게 연결되어 서 있는 상체로 읽힌다. 발은 프레임 밖이므로 지면 접촉은 확인되지 않지만 공중에 뜬 징후는 없다. 문짝은 보이는 경첩으로 차체에 연결되고, 침구는 침대 받침 위에 놓여 있다. 지지 없이 떠 있는 물체는 없다."
       },
       {
        "label": "A",
        "direction": "현우의 얼굴과 눈은 화면 왼쪽을 향하지만, 열린 출입구는 그의 뒤쪽 중경에 있다. 보이는 시선은 개구부 안으로 꺾이지 않고 차량 앞을 지나 화면 왼쪽 밖으로 이어지는 것으로 읽힌다. 미소는 분명하지만 캠핑카 내부를 응시하는 관계는 성립하지 않는다.",
        "built_space": "차량 측면의 창 하나, 환기구 두 개, 바퀴 하나, 낮은 수납 패널들, 열린 출입구 하나와 문짝 하나가 보인다. 출입구는 측면 바퀴와 창이 있는 같은 차체 면에 설치되어 후면 출입구로 읽히지 않는다. 내부에는 침대 일부와 상부 수납장이 보인다. 현우는 출입구 바로 옆보다 카메라 가까운 전경에 떨어져 있고, 차량 측면과 창고 바닥까지 상당 부분 드러나 문이 얼굴 옆의 좁은 경계가 되어야 한다는 구성에서 벗어난다.",
        "entities": "인물은 한 명이며 젊은 동아시아계 남성의 외모, 검은 흐트러진 머리와 남색 티셔츠가 참조에 대체로 맞는다. 뺨과 관자놀이에 상처가 뚜렷하다. 캠핑카와 침대는 식별되며 추가 인물이나 읽을 수 있는 문자는 없다. 다리와 신발은 보이지 않아 다리 상처와 숨긴 카드는 평가할 수 없다.",
        "hard_violations": [
         "후면의 열린 출입구 바로 바깥이라는 지정 위치를 측면 출입구와 그 앞쪽의 떨어진 전경으로 바꾸어, 현우와 차량의 필수 공간 배치를 어겼다."
        ],
        "physics": "현우의 목과 어깨는 자연스럽고 하체는 프레임 밖이다. 차량은 보이는 타이어를 통해 바닥에 지지되며, 열린 문은 경첩으로 차체에 연결된다. 침구는 침대 위에 놓여 있다. 부유하거나 해부학적으로 불가능한 요소는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.946,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.696,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gpt-high] 후면의 열린 출입구 바로 바깥이라는 지정 위치를 측면 출입구와 그 앞쪽의 떨어진 전경으로 바꾸어, 현우와 차량의 필수 공간 배치를 어겼다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 696
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "얼굴 바로 옆에 열린 캠핑카 문을 배치하여 시선 공간과 공간적 경계를 만들라는 클로즈업 프레이밍 지시를 완벽하게 준수했습니다."
   },
   {
    "label": "A",
    "score": 696,
    "verdict_ko": "프레이밍이 지시된 클로즈업보다 훨씬 넓으며, 후방 문이 아닌 캠핑카의 측면 문을 보여주어 공간적 경계 조건을 충족하지 못했습니다.  ★위반: [gpt-high] 후면의 열린 출입구 바로 바깥이라는 지정 위치를 측면 출입구와 그 앞쪽의 떨어진 전경으로 바꾸어, 현우와 차량의 필수 공간 배치를 어겼다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S44sh4_sel.png",
    "asset_id": "c2695594-fdd0-410f-9f81-0a1a52a57d8f",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-a9b2-7f98-8d7f-fad3a9a72b33",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S44sh4"
  }
 },
 "S44sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:34:48.581879+00:00",
  "fingerprint": "5b7a61fb4add5702cef3988faaa1ff7cc9af206dfaac90237ace604d89a8f2f0",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S44sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S44sh7_sel.png",
  "source_sha256": "02615ff972d0961d1782b7cca54e6f21dd00896663d72a61ba8167cdeee587a0",
  "file": "S44sh7_cine.png",
  "staged_sha256": "b00621f0260cbbbb9520b82cc816256c1ec2c786ad49b4d88cdeac26ff3ac884",
  "latency_ms": 17353
 },
 "S45sh1::signage": {
  "fp": "fdd9b2611d824981",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S45sh1": {
  "input_fingerprint": "4732513fd5f88956",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 거센 비가 내리는 텅 빈 도로 위를 헤드라이트가 꺼진 채 달리는 낡은 캠핑카의 외관.\n\nLOCATION (lock): On a nearly deserted road at the city's outskirts at night, where the camper travels through heavy rain without headlights. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Old camper (Traveling through heavy rain with its headlights off) — Front and passenger-side quarter visible from above; used as Small moving subject surrounded by exposed road space; Outlying road (Empty within the frame during the downpour) — Extends diagonally past the camper in its direction of travel; used as Establishes isolation and the forward travel axis.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the rainy night dark and the headlights unlit, using restrained tonal separation to distinguish the moving camper from the road.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper travels through heavy nighttime rain with its headlights off and one small rear-compartment lamp lit. The scavenged luggage and supplies are stowed inside.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 거센 비가 내리는 텅 빈 도로 위를 헤드라이트가 꺼진 채 달리는 낡은 캠핑카의 외관.\n\nLOCATION (lock): On a nearly deserted road at the city's outskirts at night, where the camper travels through heavy rain without headlights. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Old camper (Traveling through heavy rain with its headlights off) — Front and passenger-side quarter visible from above; used as Small moving subject surrounded by exposed road space; Outlying road (Empty within the frame during the downpour) — Extends diagonally past the camper in its direction of travel; used as Establishes isolation and the forward travel axis.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the rainy night dark and the headlights unlit, using restrained tonal separation to distinguish the moving camper from the road.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper travels through heavy nighttime rain with its headlights off and one small rear-compartment lamp lit. The scavenged luggage and supplies are stowed inside.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 거센 비가 내리는 텅 빈 도로 위를 헤드라이트가 꺼진 채 달리는 낡은 캠핑카의 외관.\n\nLOCATION (lock): On a nearly deserted road at the city's outskirts at night, where the camper travels through heavy rain without headlights. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Old camper (Traveling through heavy rain with its headlights off) — Front and passenger-side quarter visible from above; used as Small moving subject surrounded by exposed road space; Outlying road (Empty within the frame during the downpour) — Extends diagonally past the camper in its direction of travel; used as Establishes isolation and the forward travel axis.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the rainy night dark and the headlights unlit, using restrained tonal separation to distinguish the moving camper from the road.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper travels through heavy nighttime rain with its headlights off and one small rear-compartment lamp lit. The scavenged luggage and supplies are stowed inside.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "캠핑카가 프레임을 가로지르는 대각선 도로를 따라 앞을 향해 이동 중임.",
    "built_space": "가드레일과 우측 하단 구조물이 있는 도로가 높은 곳에서 내려다보이는 부감 샷으로 구성됨.",
    "entities": "낡은 캠핑카의 정면과 우측(조수석) 측면이 보이며, 헤드라이트가 꺼진 상태로 내부 뒷면 램프만 켜져 있음.",
    "hard_violations": [],
    "physics": "차량 바퀴가 젖은 도로 노면에 자연스럽게 닿아 무게감을 지탱함."
   },
   {
    "label": "B",
    "direction": "캠핑카가 도로의 차선을 따라 전방으로 주행하고 있음.",
    "built_space": "지상 눈높이에서 굽은 도로와 가드레일, 멀리 배경의 도시 조명이 보임.",
    "entities": "캠핑카의 정면과 좌측(운전석) 측면이 크게 노출되었으며, 헤드라이트는 꺼져 있고 내부 조명이 켜짐.",
    "hard_violations": [
     "[gpt-high] 앞 번호판에 읽을 수 있는 숫자열이 노출되어 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
    ],
    "physics": "캠핑카가 빗물이 고인 도로 위에 안정적으로 서서 이동함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "요청된 '위에서 내려다본 조수석 측면 뷰'와 '넓은 여백 속 작은 피사체'라는 프레이밍 지침을 정확히 구현하여 높은 점수를 부여함."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "'위에서 내려다본 구도'와 '조수석 측면 노출'이라는 명시적인 카메라 및 앵글 지시를 위반하고 피사체를 너무 크게 배치함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "캠핑카가 프레임을 가로지르는 대각선 도로를 따라 앞을 향해 이동 중임.",
        "built_space": "가드레일과 우측 하단 구조물이 있는 도로가 높은 곳에서 내려다보이는 부감 샷으로 구성됨.",
        "entities": "낡은 캠핑카의 정면과 우측(조수석) 측면이 보이며, 헤드라이트가 꺼진 상태로 내부 뒷면 램프만 켜져 있음.",
        "hard_violations": [],
        "physics": "차량 바퀴가 젖은 도로 노면에 자연스럽게 닿아 무게감을 지탱함."
       },
       {
        "label": "B",
        "direction": "캠핑카가 도로의 차선을 따라 전방으로 주행하고 있음.",
        "built_space": "지상 눈높이에서 굽은 도로와 가드레일, 멀리 배경의 도시 조명이 보임.",
        "entities": "캠핑카의 정면과 좌측(운전석) 측면이 크게 노출되었으며, 헤드라이트는 꺼져 있고 내부 조명이 켜짐.",
        "hard_violations": [],
        "physics": "캠핑카가 빗물이 고인 도로 위에 안정적으로 서서 이동함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "요청된 '위에서 내려다본 조수석 측면 뷰'와 '넓은 여백 속 작은 피사체'라는 프레이밍 지침을 정확히 구현하여 높은 점수를 부여함."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "'위에서 내려다본 구도'와 '조수석 측면 노출'이라는 명시적인 카메라 및 앵글 지시를 위반하고 피사체를 너무 크게 배치함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "캠핑카가 프레임을 가로지르는 대각선 도로를 따라 앞을 향해 이동 중임.",
        "built_space": "가드레일과 우측 하단 구조물이 있는 도로가 높은 곳에서 내려다보이는 부감 샷으로 구성됨.",
        "entities": "낡은 캠핑카의 정면과 우측(조수석) 측면이 보이며, 헤드라이트가 꺼진 상태로 내부 뒷면 램프만 켜져 있음.",
        "hard_violations": [],
        "physics": "차량 바퀴가 젖은 도로 노면에 자연스럽게 닿아 무게감을 지탱함."
       },
       {
        "label": "B",
        "direction": "캠핑카가 도로의 차선을 따라 전방으로 주행하고 있음.",
        "built_space": "지상 눈높이에서 굽은 도로와 가드레일, 멀리 배경의 도시 조명이 보임.",
        "entities": "캠핑카의 정면과 좌측(운전석) 측면이 크게 노출되었으며, 헤드라이트는 꺼져 있고 내부 조명이 켜짐.",
        "hard_violations": [],
        "physics": "캠핑카가 빗물이 고인 도로 위에 안정적으로 서서 이동함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "폭우와 소등 상태는 맞지만, 차량이 크고 카메라가 낮아 지정된 하이앵글 와이드 구도에서 벗어나며 번호판 숫자가 읽힌다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "높은 시점에서 작은 캠핑카와 넓은 대각선 도로를 담아 핵심 구도를 가장 충실히 구현했지만, 보이는 측면은 요구된 조수석 쪽과 반대다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "캠핑카 앞부분은 화면 왼쪽 아래를 향하고 도로는 오른쪽 뒤로 이어져 진행축이 일치한다. 전조등의 전방 광선은 없다. 사람의 시선은 식별되지 않는다. 앞부분과 함께 보이는 차체 측면은 통상적인 한국 좌핸들 차량의 운전석 쪽에 해당해 조수석 측면 요구와 다르다.",
        "built_space": "노란 중앙선이 있는 2차로 도로, 양쪽 가장자리의 가드레일 두 줄, 여러 전신주와 전선, 수목이 보인다. 참조의 젖은 아스팔트와 금속 가드레일은 유지하지만 멀리 도시 건물과 다수의 불빛이 더 두드러진다. 차량은 화면 폭의 약 3분의 1을 차지하며 지붕이 조금만 보여, 도로에 둘러싸인 작은 피사체를 위에서 보는 구도와 거리가 있다.",
        "entities": "낡고 얼룩진 흰색 오버캡 캠핑카 한 대가 있으며 다른 차량이나 식별 가능한 사람은 없다. 헤드라이트는 꺼져 있다. 객실 창 안에 따뜻한 갓등 하나가 뚜렷하게 보이고 뒤쪽 창에도 약한 빛이 있다. 외부에 실린 짐은 없다. 밤과 강한 비는 구현되었으나 앞 번호판의 숫자열은 읽을 수 있어 무문자 조건에 어긋난다.",
        "hard_violations": [
         "앞 번호판에 읽을 수 있는 숫자열이 노출되어 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
        ],
        "physics": "보이는 앞뒤 타이어가 노면에 닿아 차체를 지지한다. 바퀴 주변의 물보라와 젖은 노면은 빗속 주행으로 가능한 모습이다. 실내 갓등은 창 아래에 놓인 것으로 보이며 공중에 떠 있다는 증거는 없다. 노면의 반사도 젖은 표면에서 가능한 범위다."
       },
       {
        "label": "B",
        "direction": "캠핑카는 화면 왼쪽 아래를 향하고, 도로 역시 그 방향으로 대각선으로 뻗어 진행축이 맞는다. 전조등은 빛을 쏘지 않는다. 사람의 시선은 식별되지 않는다. 다만 앞부분과 함께 노출된 측면은 통상적인 한국 좌핸들 차량의 운전석 쪽이어서 조수석 측면 요구와 반대다.",
        "built_space": "노란 중앙선으로 나뉜 2차로 도로와 양쪽 가드레일 두 줄, 도로변 전신주 여러 개와 전선, 어두운 수목이 보인다. 참조의 주요 도로 재료와 시설은 대체로 유지된다. 오른쪽 가장자리는 참조보다 높고 교량 난간처럼 보이는 차이가 있다. 높은 외부 시점에서 지붕과 앞부분을 내려다보며, 차량 주변으로 넓은 빈 노면이 확보되어 지정된 와이드 구도에 잘 맞는다.",
        "entities": "낡은 흰색 오버캡 캠핑카 한 대만 보이며 다른 차량이나 식별 가능한 사람은 없다. 앞 헤드라이트는 꺼져 있고 후방 객실의 작은 창에 광원 하나가 켜져 있다. 옆의 큰 창에는 약한 실내 빛만 보여 별도의 추가 램프로 단정할 수 없다. 짐과 보급품이 외부에 노출되지 않는다. 강한 비와 어두운 밤이 표현되었으며 확실히 읽히는 문자는 없다.",
        "hard_violations": [],
        "physics": "앞뒤 바퀴가 도로에 닿아 차량을 지지하며 차체가 떠 있지 않다. 타이어 주변의 물보라가 빗길 주행을 뒷받침한다. 객실 불빛은 창 안의 실내 광원으로 읽힌다. 빗줄기, 물웅덩이와 노면 반사는 물리적으로 가능한 모습이다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "폭우와 소등 상태는 맞지만, 차량이 크고 카메라가 낮아 지정된 하이앵글 와이드 구도에서 벗어나며 번호판 숫자가 읽힌다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "높은 시점에서 작은 캠핑카와 넓은 대각선 도로를 담아 핵심 구도를 가장 충실히 구현했지만, 보이는 측면은 요구된 조수석 쪽과 반대다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "캠핑카 앞부분은 화면 왼쪽 아래를 향하고 도로는 오른쪽 뒤로 이어져 진행축이 일치한다. 전조등의 전방 광선은 없다. 사람의 시선은 식별되지 않는다. 앞부분과 함께 보이는 차체 측면은 통상적인 한국 좌핸들 차량의 운전석 쪽에 해당해 조수석 측면 요구와 다르다.",
        "built_space": "노란 중앙선이 있는 2차로 도로, 양쪽 가장자리의 가드레일 두 줄, 여러 전신주와 전선, 수목이 보인다. 참조의 젖은 아스팔트와 금속 가드레일은 유지하지만 멀리 도시 건물과 다수의 불빛이 더 두드러진다. 차량은 화면 폭의 약 3분의 1을 차지하며 지붕이 조금만 보여, 도로에 둘러싸인 작은 피사체를 위에서 보는 구도와 거리가 있다.",
        "entities": "낡고 얼룩진 흰색 오버캡 캠핑카 한 대가 있으며 다른 차량이나 식별 가능한 사람은 없다. 헤드라이트는 꺼져 있다. 객실 창 안에 따뜻한 갓등 하나가 뚜렷하게 보이고 뒤쪽 창에도 약한 빛이 있다. 외부에 실린 짐은 없다. 밤과 강한 비는 구현되었으나 앞 번호판의 숫자열은 읽을 수 있어 무문자 조건에 어긋난다.",
        "hard_violations": [
         "앞 번호판에 읽을 수 있는 숫자열이 노출되어 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
        ],
        "physics": "보이는 앞뒤 타이어가 노면에 닿아 차체를 지지한다. 바퀴 주변의 물보라와 젖은 노면은 빗속 주행으로 가능한 모습이다. 실내 갓등은 창 아래에 놓인 것으로 보이며 공중에 떠 있다는 증거는 없다. 노면의 반사도 젖은 표면에서 가능한 범위다."
       },
       {
        "label": "A",
        "direction": "캠핑카는 화면 왼쪽 아래를 향하고, 도로 역시 그 방향으로 대각선으로 뻗어 진행축이 맞는다. 전조등은 빛을 쏘지 않는다. 사람의 시선은 식별되지 않는다. 다만 앞부분과 함께 노출된 측면은 통상적인 한국 좌핸들 차량의 운전석 쪽이어서 조수석 측면 요구와 반대다.",
        "built_space": "노란 중앙선으로 나뉜 2차로 도로와 양쪽 가드레일 두 줄, 도로변 전신주 여러 개와 전선, 어두운 수목이 보인다. 참조의 주요 도로 재료와 시설은 대체로 유지된다. 오른쪽 가장자리는 참조보다 높고 교량 난간처럼 보이는 차이가 있다. 높은 외부 시점에서 지붕과 앞부분을 내려다보며, 차량 주변으로 넓은 빈 노면이 확보되어 지정된 와이드 구도에 잘 맞는다.",
        "entities": "낡은 흰색 오버캡 캠핑카 한 대만 보이며 다른 차량이나 식별 가능한 사람은 없다. 앞 헤드라이트는 꺼져 있고 후방 객실의 작은 창에 광원 하나가 켜져 있다. 옆의 큰 창에는 약한 실내 빛만 보여 별도의 추가 램프로 단정할 수 없다. 짐과 보급품이 외부에 노출되지 않는다. 강한 비와 어두운 밤이 표현되었으며 확실히 읽히는 문자는 없다.",
        "hard_violations": [],
        "physics": "앞뒤 바퀴가 도로에 닿아 차량을 지지하며 차체가 떠 있지 않다. 타이어 주변의 물보라가 빗길 주행을 뒷받침한다. 객실 불빛은 창 안의 실내 광원으로 읽힌다. 빗줄기, 물웅덩이와 노면 반사는 물리적으로 가능한 모습이다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.946
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.696
   },
   "violations": {
    "B": [
     "[gpt-high] 앞 번호판에 읽을 수 있는 숫자열이 노출되어 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 696
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "요청된 '위에서 내려다본 조수석 측면 뷰'와 '넓은 여백 속 작은 피사체'라는 프레이밍 지침을 정확히 구현하여 높은 점수를 부여함."
   },
   {
    "label": "B",
    "score": 696,
    "verdict_ko": "'위에서 내려다본 구도'와 '조수석 측면 노출'이라는 명시적인 카메라 및 앵글 지시를 위반하고 피사체를 너무 크게 배치함.  ★위반: [gpt-high] 앞 번호판에 읽을 수 있는 숫자열이 노출되어 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L199B04.png",
    "asset_id": "4c3502ac-ef92-4502-9e15-01f19a5f689e",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-ab60-72ea-abd6-c9f8fb6cb12f",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S45sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:35:53.678191+00:00",
  "fingerprint": "249c067adb1bc5f50ed135e0b915f3e9e0f4bcad8f458d119a3b851066081f44",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S45sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S45sh1_sel.png",
  "source_sha256": "141ec6e614c0d5532afcf6a0d18e3d54a2f846f3037baa6c51191aaaa26af67e",
  "file": "S45sh1_cine.png",
  "staged_sha256": "b7efd4c215214e712189839560f568aa85e5f0ad057d510fe1883490157d73d1",
  "latency_ms": 17020
 },
 "S45sh4::confined_fp_apt": {
  "applies": true,
  "reason_ko": "캠핑카 운전석이라는 제한된 내부 공간에서 운전자인 현우가 백미러를 통해 뒷좌석을 주시하는 상황입니다. 운전석, 백미러, 뒷좌석의 정확한 위치 관계와 시선 방향을 올바르게 연출해야만 하는 샷이므로 평면도 보조가 유용합니다.",
  "input_fingerprint": "6de6aca07def2239"
 },
 "S45sh4::signage": {
  "fp": "ac469a439fca3474",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "confinedfp::c0923a6850ca": {
  "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/confinedfp_base_c0923a6850ca.png",
  "place_text": "Inside the camper's compact driving cab at night, at the steering wheel and rearview mirror, with a small rear-compartment lamp behind.",
  "input_fingerprint": "76c4295ae003be69"
 },
 "S45sh4::confined_fp": {
  "reads": {
   "controls": "The steering wheel is located at the front left seat.",
   "mirrors": "A rearview mirror is mounted at the top center of the windshield area, facing backward toward the rear cabin.",
   "camera": "The camera is located between the driver and passenger seats, pointing forward and upward directly at the rearview mirror.",
   "occupants": "현우 occupies the front left driver's seat. The passenger seat is empty."
  },
  "mismatches": [],
  "scene_description_en": "The camera is positioned between the front seats, aiming forward and slightly upward. The rearview mirror occupies the upper center of the frame, its reflective glass facing directly back at the lens. From this angle, the mirror's surface physically reflects the face of 현우, who sits in the front left driver's seat. The back of 현우's head and the steering wheel are situated in the lower left foreground. The unoccupied passenger seat is visible in the lower right foreground. The rear lamp is located behind the camera and does not appear in this forward-facing shot.",
  "fixed": false,
  "input_fingerprint": "a95da29a919b3e74"
 },
 "era_assess::1e11ec86026bde7c": {
  "subjects": [],
  "subject_text": "캠핑카 내부\n앞쪽 운전석과 조수석 뒤로 침대와 뒷좌석이 이어지는 소형 이동식 주거 공간. 작은 전등과 측면 창문, 뒤쪽 출입문이 있다.",
  "identity": "canonical",
  "scope_id": "L199",
  "scope_role": "location_interior",
  "scope_sha": "efdbc327adcdaedd"
 },
 "S45sh4": {
  "input_fingerprint": "c48370caa9aec348",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 백미러를 통해 뒷좌석을 차갑게 노려보는 현우의 매서운 눈매 클로즈업.\n\nLOCATION (lock): Inside the camper's compact driving cab at night, at the steering wheel and rearview mirror, with a small rear-compartment lamp behind. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Rearview mirror (Reflecting 현우's narrowed eyes) — Its reflective face is visible with a small rim retained; the reflected face is mirror-reversed and its eyeline runs toward the rear seats rather than the lens; used as Frames the indirect view and separates his hostile scrutiny from direct audience address; Front cabin (Occupied driving compartment) — Only soft peripheral fragments remain around the mirror; used as Maintains an interior spatial reference around the reflected close-up.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the dark cabin and restrained illumination from the small rear-seat light, allowing the reflected eyes to remain readable without an added source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Heavy rain continues around the moving camper; its headlights remain off and the small rear lamp remains lit. The luggage and supplies remain inside. 현우: He remains at the wheel with facial bruises and the untreated leg wound. Yoon's contact card remains hidden in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera is positioned between the front seats, aiming forward and slightly upward. The rearview mirror occupies the upper center of the frame, its reflective glass facing directly back at the lens. From this angle, the mirror's surface physically reflects the face of 현우, who sits in the front left driver's seat. The back of 현우's head and the steering wheel are situated in the lower left foreground. The unoccupied passenger seat is visible in the lower right foreground. The rear lamp is located behind the camera and does not appear in this forward-facing shot.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 백미러를 통해 뒷좌석을 차갑게 노려보는 현우의 매서운 눈매 클로즈업.\n\nLOCATION (lock): Inside the camper's compact driving cab at night, at the steering wheel and rearview mirror, with a small rear-compartment lamp behind. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the dark cabin and restrained illumination from the small rear-seat light, allowing the reflected eyes to remain readable without an added source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Heavy rain continues around the moving camper; its headlights remain off and the small rear lamp remains lit. The luggage and supplies remain inside. 현우: He remains at the wheel with facial bruises and the untreated leg wound. Yoon's contact card remains hidden in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera is positioned between the front seats, aiming forward and slightly upward. The rearview mirror occupies the upper center of the frame, its reflective glass facing directly back at the lens. From this angle, the mirror's surface physically reflects the face of 현우, who sits in the front left driver's seat. The back of 현우's head and the steering wheel are situated in the lower left foreground. The unoccupied passenger seat is visible in the lower right foreground. The rear lamp is located behind the camera and does not appear in this forward-facing shot.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 백미러를 통해 뒷좌석을 차갑게 노려보는 현우의 매서운 눈매 클로즈업.\n\nLOCATION (lock): Inside the camper's compact driving cab at night, at the steering wheel and rearview mirror, with a small rear-compartment lamp behind. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the dark cabin and restrained illumination from the small rear-seat light, allowing the reflected eyes to remain readable without an added source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Heavy rain continues around the moving camper; its headlights remain off and the small rear lamp remains lit. The luggage and supplies remain inside. 현우: He remains at the wheel with facial bruises and the untreated leg wound. Yoon's contact card remains hidden in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S45sh4_confinedfp.png",
     "asset_id": null,
     "role": null
    },
    {
     "label": "현우",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S45sh4_confinedfp.png",
     "asset_id": null,
     "role": null
    },
    {
     "label": "현우",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "백미러에 비친 현우의 시선이 뒷좌석(카메라 렌즈 방향)을 명확히 향하고 있음.",
    "built_space": "차량 운전석 구조가 확인되며, 스티어링 휠과 백미러, 조수석이 올바른 위치에 배치됨.",
    "entities": "현우의 외모는 레퍼런스와 일치하나, 요구된 얼굴의 멍 자국은 보이지 않음.",
    "hard_violations": [],
    "physics": "인물이 운전석에 자연스럽게 앉아 자세를 지탱하고 있음."
   },
   {
    "label": "B",
    "direction": "백미러 속 시선이 뒷좌석 쪽을 향하고 있음.",
    "built_space": "대시보드와 차량 전면부가 넓게 보이며 백미러와 좌석이 위치하지만, 요구된 클로즈업보다 배경이 너무 많이 노출됨.",
    "entities": "현우의 외모가 일치하며 얼굴에 멍 자국이 표현됨.",
    "hard_violations": [
     "[gemini-pro] 차창 밖의 비가 차량 내부인 운전자 뒷머리와 좌석 위에 텍스처 오버레이처럼 잘못 겹쳐져 렌더링됨."
    ],
    "physics": "인물은 운전석에 앉아 있으나, 실내 허공과 물체 표면에 빗방울이 떠 있어 물리적으로 불가능한 상태임."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "요구된 클로즈업 프레이밍에 더 가깝게 연출되었으나, 프롬프트에 명시된 얼굴의 멍과 내리는 비의 묘사가 누락되었습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "얼굴의 멍은 묘사되었으나 프레이밍이 너무 넓고, 차량 내부에 빗방울이 맺히는 치명적인 렌더링 오류가 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "백미러에 비친 현우의 시선이 뒷좌석(카메라 렌즈 방향)을 명확히 향하고 있음.",
        "built_space": "차량 운전석 구조가 확인되며, 스티어링 휠과 백미러, 조수석이 올바른 위치에 배치됨.",
        "entities": "현우의 외모는 레퍼런스와 일치하나, 요구된 얼굴의 멍 자국은 보이지 않음.",
        "hard_violations": [],
        "physics": "인물이 운전석에 자연스럽게 앉아 자세를 지탱하고 있음."
       },
       {
        "label": "B",
        "direction": "백미러 속 시선이 뒷좌석 쪽을 향하고 있음.",
        "built_space": "대시보드와 차량 전면부가 넓게 보이며 백미러와 좌석이 위치하지만, 요구된 클로즈업보다 배경이 너무 많이 노출됨.",
        "entities": "현우의 외모가 일치하며 얼굴에 멍 자국이 표현됨.",
        "hard_violations": [
         "차창 밖의 비가 차량 내부인 운전자 뒷머리와 좌석 위에 텍스처 오버레이처럼 잘못 겹쳐져 렌더링됨."
        ],
        "physics": "인물은 운전석에 앉아 있으나, 실내 허공과 물체 표면에 빗방울이 떠 있어 물리적으로 불가능한 상태임."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "요구된 클로즈업 프레이밍에 더 가깝게 연출되었으나, 프롬프트에 명시된 얼굴의 멍과 내리는 비의 묘사가 누락되었습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "얼굴의 멍은 묘사되었으나 프레이밍이 너무 넓고, 차량 내부에 빗방울이 맺히는 치명적인 렌더링 오류가 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "백미러에 비친 현우의 시선이 뒷좌석(카메라 렌즈 방향)을 명확히 향하고 있음.",
        "built_space": "차량 운전석 구조가 확인되며, 스티어링 휠과 백미러, 조수석이 올바른 위치에 배치됨.",
        "entities": "현우의 외모는 레퍼런스와 일치하나, 요구된 얼굴의 멍 자국은 보이지 않음.",
        "hard_violations": [],
        "physics": "인물이 운전석에 자연스럽게 앉아 자세를 지탱하고 있음."
       },
       {
        "label": "B",
        "direction": "백미러 속 시선이 뒷좌석 쪽을 향하고 있음.",
        "built_space": "대시보드와 차량 전면부가 넓게 보이며 백미러와 좌석이 위치하지만, 요구된 클로즈업보다 배경이 너무 많이 노출됨.",
        "entities": "현우의 외모가 일치하며 얼굴에 멍 자국이 표현됨.",
        "hard_violations": [
         "차창 밖의 비가 차량 내부인 운전자 뒷머리와 좌석 위에 텍스처 오버레이처럼 잘못 겹쳐져 렌더링됨."
        ],
        "physics": "인물은 운전석에 앉아 있으나, 실내 허공과 물체 표면에 빗방울이 떠 있어 물리적으로 불가능한 상태임."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "야간 폭우와 얼굴 타박상은 보이지만, 운전실 전체를 넓게 담아 핵심인 백미러 속 눈매 클로즈업을 놓쳤다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "거울 속 눈매를 더 크게 담고 작은 후방등과 어두운 실내를 살렸지만, 여전히 구도가 넓고 시선도 렌즈와 충분히 분리되지 않는다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 전방을 향해 앉아 있고, 거울 속 좁힌 눈은 거울을 통해 차량 뒤쪽 중앙, 카메라에 가까운 방향을 바라본다. 화면 밖 뒷좌석의 특정 지점을 겨냥한다기보다 관객을 정면으로 보는 인상이 강하다. 차량의 실제 진행 방향은 정지 화면만으로 확인되지 않는다.",
        "built_space": "왼쪽 운전석과 오른쪽 빈 조수석, 운전대 하나, 중앙 백미러 하나, 상단 선바이저 두 개, 중앙 송풍구 두 개가 보인다. 좌측 운전석 배치는 평면도와 부합하고, 중앙에서 거울에 운전자의 얼굴이 보이는 반사는 가능한 배치다. 다만 운전자의 뒤통수와 좌석, 대시보드, 앞유리가 넓게 노출되어 운전실이 부드러운 주변 파편으로만 남아야 한다는 지시와 다르다. 후방의 작은 등은 식별되지 않는다.",
        "entities": "인물은 현우 한 명이며 거울상은 추가 인물이 아니다. 검은 헝클어진 머리와 동아시아계 젊은 남성의 외형, 어두운 상의가 참조와 대체로 맞고 거울 속 볼에 붉은 타박상이 보인다. 정확한 나이나 한국계 미국인이라는 국적·배경은 외형만으로 확인할 수 없다. 백미러와 빗물이 흐르는 앞유리가 보이며 밤으로 읽힌다. 다리 상처, 신발 속 카드, 짐은 구도 밖이라 판단하지 않는다. 판독 가능한 문구는 없다.",
        "hard_violations": [],
        "physics": "현우의 몸은 운전석에 앉아 지지되고 머리는 목과 몸통에 연결되어 있다. 백미러는 위쪽 장착대로 고정되고 운전대와 좌석도 차체에 연결되어 있다. 손은 명확히 보이지 않아 운전대를 잡았는지는 확인할 수 없지만, 떠 있는 신체나 무지지 물체는 확인되지 않는다."
       },
       {
        "label": "B",
        "direction": "현우는 운전석에서 앞을 향하고 거울 속 눈은 약간 화면 왼쪽으로 치우친 채 차량 뒤쪽 중앙을 주시한다. 차갑게 좁힌 눈매는 보이지만 시선의 도착점이 카메라 근처여서, 렌즈가 아닌 뒷좌석을 노려본다는 구분은 명확하지 않다. 차량의 이동 자체는 확인되지 않는다.",
        "built_space": "왼쪽 운전석과 오른쪽 빈 조수석, 일부 노출된 운전대 하나, 중앙 백미러 하나, 상단 선바이저 두 개가 보인다. 거울 오른쪽에는 켜진 작은 후방등 하나가 반사되어 있다. 좌측 운전자와 중앙 거울의 배치는 평면도와 양립하며 반사도 불가능하다고 볼 근거는 없다. 거울은 A보다 크게 잡혔지만 뒤통수와 두 좌석, 앞유리가 여전히 화면 상당 부분을 차지해 눈매 중심 클로즈업에는 못 미친다.",
        "entities": "현우 한 명의 뒤통수와 거울 속 얼굴이 보인다. 앳된 동아시아계 남성의 얼굴, 헝클어진 검은 머리는 인물 참조와 대체로 일치한다. 의복은 어둡고 흐린 일부만 보여 정확한 색과 형태를 확정하기 어렵다. 눈은 정상적인 사람의 눈이며 표정으로 경계심을 표현한다. 뚜렷한 얼굴 타박상은 A보다 약하고, 폭우도 명확하게 드러나지 않는다. 작은 후방등과 어두운 야간 실내는 보인다. 상처 난 다리와 숨겨진 카드, 짐은 구도 밖이며 판독 가능한 글자는 없다.",
        "hard_violations": [],
        "physics": "현우는 운전석에 앉아 있고 목과 머리의 연결도 자연스럽다. 거울 위의 어두운 연결부가 차체 쪽으로 이어지며, 좌석과 운전대는 고정된 차량 부품으로 보인다. 거울 속 등은 후방 벽면에 붙은 조명으로 읽힌다. 손의 운전대 접촉은 화면에서 확인되지 않지만, 지지 없이 떠 있는 물체나 불가능한 신체 자세는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "야간 폭우와 얼굴 타박상은 보이지만, 운전실 전체를 넓게 담아 핵심인 백미러 속 눈매 클로즈업을 놓쳤다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "거울 속 눈매를 더 크게 담고 작은 후방등과 어두운 실내를 살렸지만, 여전히 구도가 넓고 시선도 렌즈와 충분히 분리되지 않는다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 전방을 향해 앉아 있고, 거울 속 좁힌 눈은 거울을 통해 차량 뒤쪽 중앙, 카메라에 가까운 방향을 바라본다. 화면 밖 뒷좌석의 특정 지점을 겨냥한다기보다 관객을 정면으로 보는 인상이 강하다. 차량의 실제 진행 방향은 정지 화면만으로 확인되지 않는다.",
        "built_space": "왼쪽 운전석과 오른쪽 빈 조수석, 운전대 하나, 중앙 백미러 하나, 상단 선바이저 두 개, 중앙 송풍구 두 개가 보인다. 좌측 운전석 배치는 평면도와 부합하고, 중앙에서 거울에 운전자의 얼굴이 보이는 반사는 가능한 배치다. 다만 운전자의 뒤통수와 좌석, 대시보드, 앞유리가 넓게 노출되어 운전실이 부드러운 주변 파편으로만 남아야 한다는 지시와 다르다. 후방의 작은 등은 식별되지 않는다.",
        "entities": "인물은 현우 한 명이며 거울상은 추가 인물이 아니다. 검은 헝클어진 머리와 동아시아계 젊은 남성의 외형, 어두운 상의가 참조와 대체로 맞고 거울 속 볼에 붉은 타박상이 보인다. 정확한 나이나 한국계 미국인이라는 국적·배경은 외형만으로 확인할 수 없다. 백미러와 빗물이 흐르는 앞유리가 보이며 밤으로 읽힌다. 다리 상처, 신발 속 카드, 짐은 구도 밖이라 판단하지 않는다. 판독 가능한 문구는 없다.",
        "hard_violations": [],
        "physics": "현우의 몸은 운전석에 앉아 지지되고 머리는 목과 몸통에 연결되어 있다. 백미러는 위쪽 장착대로 고정되고 운전대와 좌석도 차체에 연결되어 있다. 손은 명확히 보이지 않아 운전대를 잡았는지는 확인할 수 없지만, 떠 있는 신체나 무지지 물체는 확인되지 않는다."
       },
       {
        "label": "A",
        "direction": "현우는 운전석에서 앞을 향하고 거울 속 눈은 약간 화면 왼쪽으로 치우친 채 차량 뒤쪽 중앙을 주시한다. 차갑게 좁힌 눈매는 보이지만 시선의 도착점이 카메라 근처여서, 렌즈가 아닌 뒷좌석을 노려본다는 구분은 명확하지 않다. 차량의 이동 자체는 확인되지 않는다.",
        "built_space": "왼쪽 운전석과 오른쪽 빈 조수석, 일부 노출된 운전대 하나, 중앙 백미러 하나, 상단 선바이저 두 개가 보인다. 거울 오른쪽에는 켜진 작은 후방등 하나가 반사되어 있다. 좌측 운전자와 중앙 거울의 배치는 평면도와 양립하며 반사도 불가능하다고 볼 근거는 없다. 거울은 A보다 크게 잡혔지만 뒤통수와 두 좌석, 앞유리가 여전히 화면 상당 부분을 차지해 눈매 중심 클로즈업에는 못 미친다.",
        "entities": "현우 한 명의 뒤통수와 거울 속 얼굴이 보인다. 앳된 동아시아계 남성의 얼굴, 헝클어진 검은 머리는 인물 참조와 대체로 일치한다. 의복은 어둡고 흐린 일부만 보여 정확한 색과 형태를 확정하기 어렵다. 눈은 정상적인 사람의 눈이며 표정으로 경계심을 표현한다. 뚜렷한 얼굴 타박상은 A보다 약하고, 폭우도 명확하게 드러나지 않는다. 작은 후방등과 어두운 야간 실내는 보인다. 상처 난 다리와 숨겨진 카드, 짐은 구도 밖이며 판독 가능한 글자는 없다.",
        "hard_violations": [],
        "physics": "현우는 운전석에 앉아 있고 목과 머리의 연결도 자연스럽다. 거울 위의 어두운 연결부가 차체 쪽으로 이어지며, 좌석과 운전대는 고정된 차량 부품으로 보인다. 거울 속 등은 후방 벽면에 붙은 조명으로 읽힌다. 손의 운전대 접촉은 화면에서 확인되지 않지만, 지지 없이 떠 있는 물체나 불가능한 신체 자세는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.167
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.917
   },
   "violations": {
    "B": [
     "[gemini-pro] 차창 밖의 비가 차량 내부인 운전자 뒷머리와 좌석 위에 텍스처 오버레이처럼 잘못 겹쳐져 렌더링됨."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 917
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "요구된 클로즈업 프레이밍에 더 가깝게 연출되었으나, 프롬프트에 명시된 얼굴의 멍과 내리는 비의 묘사가 누락되었습니다."
   },
   {
    "label": "B",
    "score": 917,
    "verdict_ko": "얼굴의 멍은 묘사되었으나 프레이밍이 너무 넓고, 차량 내부에 빗방울이 맺히는 치명적인 렌더링 오류가 발생했습니다.  ★위반: [gemini-pro] 차창 밖의 비가 차량 내부인 운전자 뒷머리와 좌석 위에 텍스처 오버레이처럼 잘못 겹쳐져 렌더링됨."
   }
  ],
  "refs": [
   {
    "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S45sh4_confinedfp.png",
    "asset_id": null,
    "role": null
   },
   {
    "label": "현우",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-ad26-7623-9f86-b4aa4ac8eb7f",
  "confined_fp": {
   "base_key": "confinedfp::c0923a6850ca",
   "apt_reason": "캠핑카 운전석이라는 제한된 내부 공간에서 운전자인 현우가 백미러를 통해 뒷좌석을 주시하는 상황입니다. 운전석, 백미러, 뒷좌석의 정확한 위치 관계와 시선 방향을 올바르게 연출해야만 하는 샷이므로 평면도 보조가 유용합니다.",
   "fixed": false,
   "mismatches": []
  },
  "ref_mode": "confined_fp: 도면+장면설명+엔티티",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S45sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:38:16.303796+00:00",
  "fingerprint": "67ccc9251672d609d2e07e25dc1e4aff686e3349ab392e8e0b07fb0e3a552200",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S45sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S45sh4_sel.png",
  "source_sha256": "a0b1518b5b6a5cef0e5cfa74be48fa6fb15e8f4e0dee623dcf2a870aa5e1a8a7",
  "file": "S45sh4_cine.png",
  "staged_sha256": "c186db603dbde576ccca476b765a590a6e82926d6dbf5fe00af8e84987fbfee3",
  "latency_ms": 11821
 },
 "S45sh9::signage": {
  "fp": "40fbd5b51c65c2e1",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S45sh9::bgfirst_bg": {
  "input_fingerprint": "e4e360d1aa9b3fbf",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 창문 틈새로 들이치는 빗물을 막기 위해 헝겊을 강하게 누르고 있는 앰버의 찡그린 상체.\n\nLOCATION (lock): Inside the camper's rear seating area, beside a leaking window under the single small rear lamp.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Rear-seat window frame and gap (Rain entering through the gap) — Interior edge seen obliquely beside 앰버's hands; used as Keeps the source of her effort visible beside her grimacing profile; Cloth (Pressed firmly against the leaking window gap) — Compressed between her hands and the window frame; used as Small contact detail linking her upper-body tension to the leak.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the established small rear-seat light within the dark cabin, retaining readable facial strain and incoming rain without adding a new source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 창문 틈새로 들이치는 빗물을 막기 위해 헝겊을 강하게 누르고 있는 앰버의 찡그린 상체.\n\nLOCATION (lock): Inside the camper's rear seating area, beside a leaking window under the single small rear lamp.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Rear-seat window frame and gap (Rain entering through the gap) — Interior edge seen obliquely beside 앰버's hands; used as Keeps the source of her effort visible beside her grimacing profile; Cloth (Pressed firmly against the leaking window gap) — Compressed between her hands and the window frame; used as Small contact detail linking her upper-body tension to the leak.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the established small rear-seat light within the dark cabin, retaining readable facial strain and incoming rain without adding a new source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S45sh9__bgfirst_bg.png",
  "asset_id": "dc0bb94a-03a6-4948-805a-1b44a4d89d2d",
  "input_asset_ids": [
   "b921f333-63fe-415e-a03d-af2cd1337f45",
   "4be906f8-5377-4122-9528-ee6aaccbd804"
  ]
 },
 "S45sh9": {
  "input_fingerprint": "7a2e7fd87a606767",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 창문 틈새로 들이치는 빗물을 막기 위해 헝겊을 강하게 누르고 있는 앰버의 찡그린 상체.\n\nLOCATION (lock): Inside the camper's rear seating area, beside a leaking window under the single small rear lamp. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Rear-seat window frame and gap (Rain entering through the gap) — Interior edge seen obliquely beside 앰버's hands; used as Keeps the source of her effort visible beside her grimacing profile; Cloth (Pressed firmly against the leaking window gap) — Compressed between her hands and the window frame; used as Small contact detail linking her upper-body tension to the leak.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the established small rear-seat light within the dark cabin, retaining readable facial strain and incoming rain without adding a new source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Rainwater leaks through the camper's roof and window frames despite cloth pressed against the openings. The headlights remain off and the small rear lamp remains lit. 앰버: She presses clothing against a leaking window frame.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 앰버 right now, so 앰버's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 앰버: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 창문 틈새로 들이치는 빗물을 막기 위해 헝겊을 강하게 누르고 있는 앰버의 찡그린 상체.\n\nLOCATION (lock): Inside the camper's rear seating area, beside a leaking window under the single small rear lamp. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Rear-seat window frame and gap (Rain entering through the gap) — Interior edge seen obliquely beside 앰버's hands; used as Keeps the source of her effort visible beside her grimacing profile; Cloth (Pressed firmly against the leaking window gap) — Compressed between her hands and the window frame; used as Small contact detail linking her upper-body tension to the leak.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the established small rear-seat light within the dark cabin, retaining readable facial strain and incoming rain without adding a new source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Rainwater leaks through the camper's roof and window frames despite cloth pressed against the openings. The headlights remain off and the small rear lamp remains lit. 앰버: She presses clothing against a leaking window frame.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 앰버 right now, so 앰버's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 앰버: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 창문 틈새로 들이치는 빗물을 막기 위해 헝겊을 강하게 누르고 있는 앰버의 찡그린 상체.\n\nLOCATION (lock): Inside the camper's rear seating area, beside a leaking window under the single small rear lamp. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Rear-seat window frame and gap (Rain entering through the gap) — Interior edge seen obliquely beside 앰버's hands; used as Keeps the source of her effort visible beside her grimacing profile; Cloth (Pressed firmly against the leaking window gap) — Compressed between her hands and the window frame; used as Small contact detail linking her upper-body tension to the leak.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the established small rear-seat light within the dark cabin, retaining readable facial strain and incoming rain without adding a new source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Rainwater leaks through the camper's roof and window frames despite cloth pressed against the openings. The headlights remain off and the small rear lamp remains lit. 앰버: She presses clothing against a leaking window frame.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 앰버 right now, so 앰버's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 앰버: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S45sh9__bgfirst_bg.png",
     "asset_id": "dc0bb94a-03a6-4948-805a-1b44a4d89d2d",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S45sh9.png",
     "asset_id": "b921f333-63fe-415e-a03d-af2cd1337f45",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L199B03.png",
     "asset_id": "4be906f8-5377-4122-9528-ee6aaccbd804",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "앰버의 시선은 자신의 손이 헝겊을 누르고 있는 오른쪽 창문 틈새를 향하고 있습니다.",
    "built_space": "캠핑카 내부의 오른쪽 창문 앞입니다. 배경에 소파가 있고, 인물 바로 뒤쪽 벽면에 레퍼런스와 일치하는 작은 사각형 조명이 위치해 있어 제시된 공간 구조와 정확히 부합합니다.",
    "entities": "앰버는 금발 머리, 둥근 얼굴, 10세 소녀의 외양으로 캐릭터 레퍼런스와 일치하며 찡그린 표정을 짓고 있습니다.",
    "hard_violations": [],
    "physics": "상체를 앞으로 숙이고 무릎이나 하체로 체중을 지탱하고 있으며, 두 손은 창틀에 헝겊을 강하게 밀착시켜 자연스러운 힘의 작용을 보여줍니다."
   },
   {
    "label": "B",
    "direction": "앰버는 왼쪽 창문 쪽에 기대어 손이 있는 곳을 바라보고 있으며, 물이 가로로 강하게 뿜어져 들어옵니다.",
    "built_space": "캠핑카 내부의 왼쪽 창문 앞입니다. 소파와 작은 조명이 프레임 우측 배경에 멀리 떨어져 있어, '작은 조명 아래'라는 프롬프트의 위치 지정과 어긋납니다.",
    "entities": "앰버의 인상착의는 캐릭터 레퍼런스와 일치하며, 눈을 질끈 감고 찡그린 표정을 짓고 있습니다.",
    "hard_violations": [
     "[gemini-pro] 지정된 위치(조명 아래 창문)가 아닌 맞은편 창문에 인물을 배치하여 스테이징을 위반함."
    ],
    "physics": "손으로 헝겊을 누르고 있으나 창문에 밀착되는 힘의 방향이 다소 어색하며, 창문에서 뿜어져 나오는 물줄기가 비가 새는 것이라기엔 물리적으로 과장되어 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "지정된 위치(작은 조명 아래 창문)에서 비가 새는 것을 막는 앰버의 모습과 표정을 훌륭하게 구현했으나, 새 조명을 추가하지 말라는 지시와 달리 얼굴 쪽에 인위적인 빛이 들어갔습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "지정된 조명 아래 창문이 아닌 반대편 창문에 인물을 배치하여 위치 지시를 어겼으며, 들이치는 물줄기의 물리적 묘사가 부자연스럽습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버의 시선은 자신의 손이 헝겊을 누르고 있는 오른쪽 창문 틈새를 향하고 있습니다.",
        "built_space": "캠핑카 내부의 오른쪽 창문 앞입니다. 배경에 소파가 있고, 인물 바로 뒤쪽 벽면에 레퍼런스와 일치하는 작은 사각형 조명이 위치해 있어 제시된 공간 구조와 정확히 부합합니다.",
        "entities": "앰버는 금발 머리, 둥근 얼굴, 10세 소녀의 외양으로 캐릭터 레퍼런스와 일치하며 찡그린 표정을 짓고 있습니다.",
        "hard_violations": [],
        "physics": "상체를 앞으로 숙이고 무릎이나 하체로 체중을 지탱하고 있으며, 두 손은 창틀에 헝겊을 강하게 밀착시켜 자연스러운 힘의 작용을 보여줍니다."
       },
       {
        "label": "B",
        "direction": "앰버는 왼쪽 창문 쪽에 기대어 손이 있는 곳을 바라보고 있으며, 물이 가로로 강하게 뿜어져 들어옵니다.",
        "built_space": "캠핑카 내부의 왼쪽 창문 앞입니다. 소파와 작은 조명이 프레임 우측 배경에 멀리 떨어져 있어, '작은 조명 아래'라는 프롬프트의 위치 지정과 어긋납니다.",
        "entities": "앰버의 인상착의는 캐릭터 레퍼런스와 일치하며, 눈을 질끈 감고 찡그린 표정을 짓고 있습니다.",
        "hard_violations": [
         "지정된 위치(조명 아래 창문)가 아닌 맞은편 창문에 인물을 배치하여 스테이징을 위반함."
        ],
        "physics": "손으로 헝겊을 누르고 있으나 창문에 밀착되는 힘의 방향이 다소 어색하며, 창문에서 뿜어져 나오는 물줄기가 비가 새는 것이라기엔 물리적으로 과장되어 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "지정된 위치(작은 조명 아래 창문)에서 비가 새는 것을 막는 앰버의 모습과 표정을 훌륭하게 구현했으나, 새 조명을 추가하지 말라는 지시와 달리 얼굴 쪽에 인위적인 빛이 들어갔습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "지정된 조명 아래 창문이 아닌 반대편 창문에 인물을 배치하여 위치 지시를 어겼으며, 들이치는 물줄기의 물리적 묘사가 부자연스럽습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "앰버의 시선은 자신의 손이 헝겊을 누르고 있는 오른쪽 창문 틈새를 향하고 있습니다.",
        "built_space": "캠핑카 내부의 오른쪽 창문 앞입니다. 배경에 소파가 있고, 인물 바로 뒤쪽 벽면에 레퍼런스와 일치하는 작은 사각형 조명이 위치해 있어 제시된 공간 구조와 정확히 부합합니다.",
        "entities": "앰버는 금발 머리, 둥근 얼굴, 10세 소녀의 외양으로 캐릭터 레퍼런스와 일치하며 찡그린 표정을 짓고 있습니다.",
        "hard_violations": [],
        "physics": "상체를 앞으로 숙이고 무릎이나 하체로 체중을 지탱하고 있으며, 두 손은 창틀에 헝겊을 강하게 밀착시켜 자연스러운 힘의 작용을 보여줍니다."
       },
       {
        "label": "B",
        "direction": "앰버는 왼쪽 창문 쪽에 기대어 손이 있는 곳을 바라보고 있으며, 물이 가로로 강하게 뿜어져 들어옵니다.",
        "built_space": "캠핑카 내부의 왼쪽 창문 앞입니다. 소파와 작은 조명이 프레임 우측 배경에 멀리 떨어져 있어, '작은 조명 아래'라는 프롬프트의 위치 지정과 어긋납니다.",
        "entities": "앰버의 인상착의는 캐릭터 레퍼런스와 일치하며, 눈을 질끈 감고 찡그린 표정을 짓고 있습니다.",
        "hard_violations": [
         "지정된 위치(조명 아래 창문)가 아닌 맞은편 창문에 인물을 배치하여 스테이징을 위반함."
        ],
        "physics": "손으로 헝겊을 누르고 있으나 창문에 밀착되는 힘의 방향이 다소 어색하며, 창문에서 뿜어져 나오는 물줄기가 비가 새는 것이라기엔 물리적으로 과장되어 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "찡그린 상체 중심의 미디엄 숏에서 두 손으로 창틀 틈의 헝겊을 압박하는 동작과 유입되는 빗물이 명확하며, 남색 반소매 의상도 참조와 일치한다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "옆얼굴의 긴장과 누르는 동작은 좋지만 허벅지까지 담은 구도가 상체 중심 지시에서 벗어나고, 헝겊의 압박점이 틈보다 유리 쪽에 치우치며 의상도 참조와 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버는 화면 왼쪽 창문과 손의 접촉점을 향해 얼굴을 돌리고 눈을 찡그린다. 두 팔의 힘은 헝겊이 놓인 창문 하단 틈을 향하며, 그 부근에서 물방울이 실내 쪽으로 튄다. 시선과 압박 방향 모두 누수 지점을 대상으로 한다.",
        "built_space": "왼쪽에 창문 한 개와 비스듬한 하단 창틀, 그 뒤에 커튼 한 폭, 오른쪽 배경에 후방 벤치 한 개와 켜진 작은 벽등 한 개가 보인다. 앰버는 벤치 앞에서 창틀에 손이 닿는 거리에 있다. 낡은 벽판과 체크무늬 좌석, 벽등의 배치가 장소 참조에 가깝다. 유리의 희미한 손 반사는 유리 바로 옆에 놓인 손과 양립하는 위치이며, 불가능한 반사는 보이지 않는다.",
        "entities": "인물은 어린 여자아이 한 명뿐이다. 젖은 금발, 둥근 얼굴, 아동의 체격과 손 크기는 앰버의 참조에 가깝고 남색 반소매 티셔츠도 일치한다. 눈은 찡그려 큰 눈의 형태를 완전히 비교하기 어렵고, 한국계·백인 혼혈이라는 배경은 외관만으로 확정할 수 없다. 회색 헝겊, 젖은 창틀, 빗방울, 작은 후방등이 보이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "양손이 헝겊을 붙잡아 단단한 하단 창틀에 누르고 있고, 팔과 어깨의 전방 긴장이 그 힘을 뒷받침한다. 헝겊 윗부분은 손과 창틀 사이에 눌리고 나머지는 중력 방향으로 늘어진다. 틈 부근의 튀는 물과 아래로 떨어지는 물방울이 누수 동작에 부합한다. 하체 지지는 프레임 밖이지만 상체가 공중에 떠 있다는 징후는 없다."
       },
       {
        "label": "B",
        "direction": "앰버의 옆얼굴과 시선은 화면 오른쪽 손과 창문을 향한다. 두 손은 헝겊을 창문 쪽으로 밀지만 압박 중심이 하단 틈보다 조금 위의 유리면에 놓여 보인다. 헝겊 아래쪽은 창틀에 걸쳐 있으나 틈을 직접 막는 접촉은 A보다 불명확하다.",
        "built_space": "오른쪽 근경 창문 한 개, 왼쪽 배경 창문 한 개, 후방 벤치 한 개, 아래쪽 전경 좌석 일부와 켜진 작은 벽등 한 개가 보인다. 창가 커튼과 낡은 벽판도 있다. 앰버는 전경 좌석에 앉아 오른쪽 창문으로 몸을 기울인다. 캠퍼 후방 공간의 재질과 기본 설비는 유사하지만 창문과 수납부의 화면상 배열은 장소 참조와 차이가 있다. 창유리의 약한 등 반사에 명백한 광학적 모순은 없다.",
        "entities": "금발의 어린 여자아이 한 명으로, 나이와 체격 및 둥근 얼굴은 앰버 설정과 대체로 맞는다. 옆얼굴과 찡그린 눈 때문에 참조의 눈 형태를 자세히 확인하기는 어렵고 혈통도 시각적으로 확정할 수 없다. 참조의 남색 반소매 대신 회갈색 긴소매 상의를 입었다. 밝은 회색 헝겊, 젖은 창문과 벽면의 물줄기, 작은 후방등이 보인다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "엉덩이와 허벅지는 아래쪽 좌석 쿠션에 지지되고, 앞으로 기울인 몸의 힘이 두 팔과 손을 거쳐 창문으로 전달된다. 양손이 헝겊을 실제로 붙잡고 있으며 아래 자락은 창틀 아래로 처진다. 손이나 몸이 지지 없이 떠 있지는 않다. 다만 물은 주로 유리와 벽면을 따라 흐르는 모습으로, 압박 중인 틈에서 실내로 들이치는 빗물은 뚜렷하지 않다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "찡그린 상체 중심의 미디엄 숏에서 두 손으로 창틀 틈의 헝겊을 압박하는 동작과 유입되는 빗물이 명확하며, 남색 반소매 의상도 참조와 일치한다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "옆얼굴의 긴장과 누르는 동작은 좋지만 허벅지까지 담은 구도가 상체 중심 지시에서 벗어나고, 헝겊의 압박점이 틈보다 유리 쪽에 치우치며 의상도 참조와 다르다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "앰버는 화면 왼쪽 창문과 손의 접촉점을 향해 얼굴을 돌리고 눈을 찡그린다. 두 팔의 힘은 헝겊이 놓인 창문 하단 틈을 향하며, 그 부근에서 물방울이 실내 쪽으로 튄다. 시선과 압박 방향 모두 누수 지점을 대상으로 한다.",
        "built_space": "왼쪽에 창문 한 개와 비스듬한 하단 창틀, 그 뒤에 커튼 한 폭, 오른쪽 배경에 후방 벤치 한 개와 켜진 작은 벽등 한 개가 보인다. 앰버는 벤치 앞에서 창틀에 손이 닿는 거리에 있다. 낡은 벽판과 체크무늬 좌석, 벽등의 배치가 장소 참조에 가깝다. 유리의 희미한 손 반사는 유리 바로 옆에 놓인 손과 양립하는 위치이며, 불가능한 반사는 보이지 않는다.",
        "entities": "인물은 어린 여자아이 한 명뿐이다. 젖은 금발, 둥근 얼굴, 아동의 체격과 손 크기는 앰버의 참조에 가깝고 남색 반소매 티셔츠도 일치한다. 눈은 찡그려 큰 눈의 형태를 완전히 비교하기 어렵고, 한국계·백인 혼혈이라는 배경은 외관만으로 확정할 수 없다. 회색 헝겊, 젖은 창틀, 빗방울, 작은 후방등이 보이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "양손이 헝겊을 붙잡아 단단한 하단 창틀에 누르고 있고, 팔과 어깨의 전방 긴장이 그 힘을 뒷받침한다. 헝겊 윗부분은 손과 창틀 사이에 눌리고 나머지는 중력 방향으로 늘어진다. 틈 부근의 튀는 물과 아래로 떨어지는 물방울이 누수 동작에 부합한다. 하체 지지는 프레임 밖이지만 상체가 공중에 떠 있다는 징후는 없다."
       },
       {
        "label": "A",
        "direction": "앰버의 옆얼굴과 시선은 화면 오른쪽 손과 창문을 향한다. 두 손은 헝겊을 창문 쪽으로 밀지만 압박 중심이 하단 틈보다 조금 위의 유리면에 놓여 보인다. 헝겊 아래쪽은 창틀에 걸쳐 있으나 틈을 직접 막는 접촉은 A보다 불명확하다.",
        "built_space": "오른쪽 근경 창문 한 개, 왼쪽 배경 창문 한 개, 후방 벤치 한 개, 아래쪽 전경 좌석 일부와 켜진 작은 벽등 한 개가 보인다. 창가 커튼과 낡은 벽판도 있다. 앰버는 전경 좌석에 앉아 오른쪽 창문으로 몸을 기울인다. 캠퍼 후방 공간의 재질과 기본 설비는 유사하지만 창문과 수납부의 화면상 배열은 장소 참조와 차이가 있다. 창유리의 약한 등 반사에 명백한 광학적 모순은 없다.",
        "entities": "금발의 어린 여자아이 한 명으로, 나이와 체격 및 둥근 얼굴은 앰버 설정과 대체로 맞는다. 옆얼굴과 찡그린 눈 때문에 참조의 눈 형태를 자세히 확인하기는 어렵고 혈통도 시각적으로 확정할 수 없다. 참조의 남색 반소매 대신 회갈색 긴소매 상의를 입었다. 밝은 회색 헝겊, 젖은 창문과 벽면의 물줄기, 작은 후방등이 보인다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "엉덩이와 허벅지는 아래쪽 좌석 쿠션에 지지되고, 앞으로 기울인 몸의 힘이 두 팔과 손을 거쳐 창문으로 전달된다. 양손이 헝겊을 실제로 붙잡고 있으며 아래 자락은 창틀 아래로 처진다. 손이나 몸이 지지 없이 떠 있지는 않다. 다만 물은 주로 유리와 벽면을 따라 흐르는 모습으로, 압박 중인 틈에서 실내로 들이치는 빗물은 뚜렷하지 않다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.778,
    "B": 1.5
   },
   "adjusted": {
    "A": 1.778,
    "B": 1.25
   },
   "violations": {
    "B": [
     "[gemini-pro] 지정된 위치(조명 아래 창문)가 아닌 맞은편 창문에 인물을 배치하여 스테이징을 위반함."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1778,
   "B": 1250
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1778,
    "verdict_ko": "지정된 위치(작은 조명 아래 창문)에서 비가 새는 것을 막는 앰버의 모습과 표정을 훌륭하게 구현했으나, 새 조명을 추가하지 말라는 지시와 달리 얼굴 쪽에 인위적인 빛이 들어갔습니다."
   },
   {
    "label": "B",
    "score": 1250,
    "verdict_ko": "지정된 조명 아래 창문이 아닌 반대편 창문에 인물을 배치하여 위치 지시를 어겼으며, 들이치는 물줄기의 물리적 묘사가 부자연스럽습니다.  ★위반: [gemini-pro] 지정된 위치(조명 아래 창문)가 아닌 맞은편 창문에 인물을 배치하여 스테이징을 위반함."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L199B03.png",
    "asset_id": "4be906f8-5377-4122-9528-ee6aaccbd804",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-b066-7490-9cfd-9499604c443a",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S45sh9__bgfirst_bg.png",
   "bg_asset_id": "dc0bb94a-03a6-4948-805a-1b44a4d89d2d",
   "bg_record_key": "S45sh9::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S45sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:40:23.084924+00:00",
  "fingerprint": "ccb2642530e65180d9c0d2373c981fc08e98942c3bd79b402e5ca0dff2d16a0e",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S45sh9_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S45sh9_sel.png",
  "source_sha256": "fe5bfbd87eb28b07506bc965a2f2e466e9d1ffdcd20cb4a75b53aa00fe230187",
  "file": "S45sh9_cine.png",
  "staged_sha256": "f1bb0e1b61ae7cbe2b57e1a0ebcb3ce6da13ffa765c4b4d02a6fa943abb7a279",
  "latency_ms": 11562
 },
 "S46sh1::signage": {
  "fp": "014702876d03f367",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S46sh1": {
  "input_fingerprint": "56f98fbf45873d89",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 거대한 금이 가고 마른 넝쿨이 잔뜩 얽힌 낡은 미추홀도서관 외벽 앞에 멈춰 선 캠핑카 전경.\n\nLOCATION (lock): Outside the abandoned library, in front of its cracked, vine-covered facade, where the camper stops in nighttime rain. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Old camper (Stopped in front of the library) — Seen obliquely from above, with its side and roof visible; used as Foreground scale reference and connection to the previous road scene; Library facade (Large cracks across the walls, with dry vines clinging densely to them) — The building front recedes obliquely across the upper frame; used as Establishes the abandoned destination without overwhelming the vehicle-ground relationship; Library sign (Old) — The lettered face is visible to camera and reads '미추홀도서관, 인천'; used as Identifies the destination within the establishing composition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the rainy nighttime setting with subdued ambient exposure and controlled contrast, without specifying an unsupported exterior light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The leaking camper is stopped outside the library in nighttime rain. Dry vines cover the deeply cracked exterior beneath the aged sign reading “미추홀도서관, 인천.”\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 거대한 금이 가고 마른 넝쿨이 잔뜩 얽힌 낡은 미추홀도서관 외벽 앞에 멈춰 선 캠핑카 전경.\n\nLOCATION (lock): Outside the abandoned library, in front of its cracked, vine-covered facade, where the camper stops in nighttime rain. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Old camper (Stopped in front of the library) — Seen obliquely from above, with its side and roof visible; used as Foreground scale reference and connection to the previous road scene; Library facade (Large cracks across the walls, with dry vines clinging densely to them) — The building front recedes obliquely across the upper frame; used as Establishes the abandoned destination without overwhelming the vehicle-ground relationship; Library sign (Old) — The lettered face is visible to camera and reads '미추홀도서관, 인천'; used as Identifies the destination within the establishing composition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the rainy nighttime setting with subdued ambient exposure and controlled contrast, without specifying an unsupported exterior light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The leaking camper is stopped outside the library in nighttime rain. Dry vines cover the deeply cracked exterior beneath the aged sign reading “미추홀도서관, 인천.”\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 거대한 금이 가고 마른 넝쿨이 잔뜩 얽힌 낡은 미추홀도서관 외벽 앞에 멈춰 선 캠핑카 전경.\n\nLOCATION (lock): Outside the abandoned library, in front of its cracked, vine-covered facade, where the camper stops in nighttime rain. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Old camper (Stopped in front of the library) — Seen obliquely from above, with its side and roof visible; used as Foreground scale reference and connection to the previous road scene; Library facade (Large cracks across the walls, with dry vines clinging densely to them) — The building front recedes obliquely across the upper frame; used as Establishes the abandoned destination without overwhelming the vehicle-ground relationship; Library sign (Old) — The lettered face is visible to camera and reads '미추홀도서관, 인천'; used as Identifies the destination within the establishing composition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the rainy nighttime setting with subdued ambient exposure and controlled contrast, without specifying an unsupported exterior light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The leaking camper is stopped outside the library in nighttime rain. Dry vines cover the deeply cracked exterior beneath the aged sign reading “미추홀도서관, 인천.”\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라가 도서관 입구와 주차된 캠핑카를 비스듬히 위에서 내려다보고 있음.",
    "built_space": "레퍼런스와 일치하는 도서관 건물, 유리문 입구, 주차장 구획선이 올바르게 배치됨.",
    "entities": "비 오는 밤, 금이 가고 넝쿨이 얽힌 도서관 외벽, 지붕과 측면이 보이는 낡은 캠핑카(모터홈). 단, 간판 텍스트는 프롬프트와 달리 '시립도서관'으로 나타남.",
    "hard_violations": [
     "[gpt-high] 간판의 ‘시립도서관’이 판독 가능하여 마지막의 읽을 수 있는 문자 금지 지시를 위반합니다."
    ],
    "physics": "캠핑카가 젖은 아스팔트 바닥에 안정적으로 정차해 있으며 빗방울과 바닥 반사가 자연스러움."
   },
   {
    "label": "B",
    "direction": "카메라가 도서관 건물과 트레일러를 비스듬히 위에서 내려다보고 있음.",
    "built_space": "도서관 건물 구조는 유사하나, 좌측 원경의 도시 건물들이 레퍼런스보다 과장되게 가깝게 변경됨.",
    "entities": "비 오는 밤, 넝쿨이 얽힌 도서관 외벽. 운전석이 없는 트레일러 형태의 캠핑카. 간판은 역시 '시립도서관'으로 출력됨.",
    "hard_violations": [
     "[gpt-high] 간판의 ‘시립도서관’이 판독 가능하여 마지막의 읽을 수 있는 문자 금지 지시를 위반합니다."
    ],
    "physics": "트레일러가 바닥에 놓여 있으나, 정차한 상태를 설명할 만한 견인 차량이 없음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "요구된 간판 텍스트('미추홀도서관, 인천') 대신 레퍼런스의 글씨('시립도서관')가 그대로 출력되었으나, 지정된 부감 앵글과 낡은 캠핑카(모터홈)의 묘사가 프롬프트의 지시를 충실히 따랐습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "간판 텍스트 요구사항을 실패했으며, 캠핑카가 견인차 없는 트레일러 형태로 묘사되어 '도착해서 멈춰 선' 상황에 어울리지 않고 배경 원경이 레퍼런스에 비해 과장되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라가 도서관 입구와 주차된 캠핑카를 비스듬히 위에서 내려다보고 있음.",
        "built_space": "레퍼런스와 일치하는 도서관 건물, 유리문 입구, 주차장 구획선이 올바르게 배치됨.",
        "entities": "비 오는 밤, 금이 가고 넝쿨이 얽힌 도서관 외벽, 지붕과 측면이 보이는 낡은 캠핑카(모터홈). 단, 간판 텍스트는 프롬프트와 달리 '시립도서관'으로 나타남.",
        "hard_violations": [],
        "physics": "캠핑카가 젖은 아스팔트 바닥에 안정적으로 정차해 있으며 빗방울과 바닥 반사가 자연스러움."
       },
       {
        "label": "B",
        "direction": "카메라가 도서관 건물과 트레일러를 비스듬히 위에서 내려다보고 있음.",
        "built_space": "도서관 건물 구조는 유사하나, 좌측 원경의 도시 건물들이 레퍼런스보다 과장되게 가깝게 변경됨.",
        "entities": "비 오는 밤, 넝쿨이 얽힌 도서관 외벽. 운전석이 없는 트레일러 형태의 캠핑카. 간판은 역시 '시립도서관'으로 출력됨.",
        "hard_violations": [],
        "physics": "트레일러가 바닥에 놓여 있으나, 정차한 상태를 설명할 만한 견인 차량이 없음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "요구된 간판 텍스트('미추홀도서관, 인천') 대신 레퍼런스의 글씨('시립도서관')가 그대로 출력되었으나, 지정된 부감 앵글과 낡은 캠핑카(모터홈)의 묘사가 프롬프트의 지시를 충실히 따랐습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "간판 텍스트 요구사항을 실패했으며, 캠핑카가 견인차 없는 트레일러 형태로 묘사되어 '도착해서 멈춰 선' 상황에 어울리지 않고 배경 원경이 레퍼런스에 비해 과장되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라가 도서관 입구와 주차된 캠핑카를 비스듬히 위에서 내려다보고 있음.",
        "built_space": "레퍼런스와 일치하는 도서관 건물, 유리문 입구, 주차장 구획선이 올바르게 배치됨.",
        "entities": "비 오는 밤, 금이 가고 넝쿨이 얽힌 도서관 외벽, 지붕과 측면이 보이는 낡은 캠핑카(모터홈). 단, 간판 텍스트는 프롬프트와 달리 '시립도서관'으로 나타남.",
        "hard_violations": [],
        "physics": "캠핑카가 젖은 아스팔트 바닥에 안정적으로 정차해 있으며 빗방울과 바닥 반사가 자연스러움."
       },
       {
        "label": "B",
        "direction": "카메라가 도서관 건물과 트레일러를 비스듬히 위에서 내려다보고 있음.",
        "built_space": "도서관 건물 구조는 유사하나, 좌측 원경의 도시 건물들이 레퍼런스보다 과장되게 가깝게 변경됨.",
        "entities": "비 오는 밤, 넝쿨이 얽힌 도서관 외벽. 운전석이 없는 트레일러 형태의 캠핑카. 간판은 역시 '시립도서관'으로 출력됨.",
        "hard_violations": [],
        "physics": "트레일러가 바닥에 놓여 있으나, 정차한 상태를 설명할 만한 견인 차량이 없음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "야간 비와 외벽의 균열·넝쿨은 맞지만, 차량 하부가 크게 잘려 캠핑카와 지면의 관계가 약하고 판독 가능한 간판이 최종 문자 금지 지시를 위반합니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "캠핑카의 운전석·측면·지붕과 접지 바퀴를 함께 보여 지정된 부감 전경에 더 충실하지만, 읽히는 간판 때문에 최종 문자 금지 지시는 충족하지 못합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "인물이나 조준 물체는 없습니다. 차량의 넓은 측면과 둥근 끝부분이 카메라를 향하지만 운전석이 없어 진행 방향은 확정하기 어렵습니다. 도서관 정면은 화면 오른쪽 가까운 곳에서 왼쪽 뒤로 물러나며, 간판의 글자 면은 카메라에 노출되어 있습니다.",
        "built_space": "참조의 2층 콘크리트 외벽, 왼쪽 수직 창탑 하나, 가로 창열 두 층, 중앙 오른쪽 유리 출입구 한 구역과 그 위 녹슨 간판 하나, 앞쪽 화단이 보입니다. 출입 계단은 차량에 상당 부분 가려집니다. 차량은 화면 하단을 크게 차지하고 하부가 프레임 밖으로 잘려, 요구한 차량과 지면의 관계보다 건물과 차량 상부가 강조됩니다.",
        "entities": "사람은 없으므로 불필요한 인물 추가는 없습니다. 낡고 얼룩진 숙박용 차량 한 대는 보이지만 동력 캠핑카인지 견인식 카라반인지 확인하기 어렵습니다. 외벽의 큰 균열과 빽빽한 마른 넝쿨, 비 내리는 밤과 젖은 포장은 일치합니다. 간판은 요청된 ‘미추홀도서관, 인천’이 아니라 참조의 ‘시립도서관’으로 읽히며, 마지막의 모든 문자 비가독화 지시에도 어긋납니다. 실내 누수는 확인되지 않습니다.",
        "hard_violations": [
         "간판의 ‘시립도서관’이 판독 가능하여 마지막의 읽을 수 있는 문자 금지 지시를 위반합니다."
        ],
        "physics": "차량 바퀴와 지면 접점은 하단 크롭 밖이라 지지 상태를 직접 확인할 수 없지만, 공중에 떠 있다고 볼 시각적 근거도 없습니다. 지붕 설비는 지붕에 붙어 있고 간판은 출입구 구조에 고정되어 있습니다. 젖은 포장 반사는 가능한 모습이며 움직임이나 충돌은 보이지 않습니다."
       },
       {
        "label": "B",
        "direction": "인물이나 조준 물체는 없습니다. 캠핑카 운전석과 앞바퀴는 화면 왼쪽을 향하고 차량은 도서관 앞에 정지해 있습니다. 카메라는 비스듬한 위쪽에서 차량 측면과 지붕을 봅니다. 도서관 정면은 화면 상부에서 왼쪽 뒤로 물러나고 간판 전면은 카메라를 향합니다.",
        "built_space": "참조와 대응하는 왼쪽 창탑 하나, 두 층의 가로 창열, 중앙 오른쪽 유리 출입구 한 구역, 그 위 녹슨 간판 하나, 전면 화단과 출입 계단·난간이 보입니다. 차량은 오른쪽 아래 전경에 놓이고 왼쪽에는 넓은 젖은 주차장이 남아 건물·차량·지면의 크기 관계가 드러납니다. 차량 뒤쪽과 하부 일부는 잘렸지만 앞바퀴의 접지는 보입니다.",
        "entities": "인물은 없습니다. 운전실과 상부 침상 돌출부를 갖춘 낡은 캠핑카 한 대가 명확합니다. 외벽의 깊은 균열, 마른 넝쿨, 노후한 간판, 야간 빗줄기와 물 고인 포장이 보입니다. 간판은 ‘시립도서관’으로 읽혀 지정 명칭과 다르고 최종 문자 비가독화 지시도 어깁니다. 지붕의 물 고임은 보이지만 실내 누수까지 확인되지는 않습니다.",
        "hard_violations": [
         "간판의 ‘시립도서관’이 판독 가능하여 마지막의 읽을 수 있는 문자 금지 지시를 위반합니다."
        ],
        "physics": "보이는 앞바퀴가 포장면에 닿아 차량을 지지하며 정차 상태가 자연스럽습니다. 지붕 설비와 난간은 차체에 부착되어 있고 물은 지붕과 포장면에 고여 있습니다. 출입구 간판은 건물 구조가 지지합니다. 떠 있는 물체나 불가능한 반사는 보이지 않습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "야간 비와 외벽의 균열·넝쿨은 맞지만, 차량 하부가 크게 잘려 캠핑카와 지면의 관계가 약하고 판독 가능한 간판이 최종 문자 금지 지시를 위반합니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "캠핑카의 운전석·측면·지붕과 접지 바퀴를 함께 보여 지정된 부감 전경에 더 충실하지만, 읽히는 간판 때문에 최종 문자 금지 지시는 충족하지 못합니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "인물이나 조준 물체는 없습니다. 차량의 넓은 측면과 둥근 끝부분이 카메라를 향하지만 운전석이 없어 진행 방향은 확정하기 어렵습니다. 도서관 정면은 화면 오른쪽 가까운 곳에서 왼쪽 뒤로 물러나며, 간판의 글자 면은 카메라에 노출되어 있습니다.",
        "built_space": "참조의 2층 콘크리트 외벽, 왼쪽 수직 창탑 하나, 가로 창열 두 층, 중앙 오른쪽 유리 출입구 한 구역과 그 위 녹슨 간판 하나, 앞쪽 화단이 보입니다. 출입 계단은 차량에 상당 부분 가려집니다. 차량은 화면 하단을 크게 차지하고 하부가 프레임 밖으로 잘려, 요구한 차량과 지면의 관계보다 건물과 차량 상부가 강조됩니다.",
        "entities": "사람은 없으므로 불필요한 인물 추가는 없습니다. 낡고 얼룩진 숙박용 차량 한 대는 보이지만 동력 캠핑카인지 견인식 카라반인지 확인하기 어렵습니다. 외벽의 큰 균열과 빽빽한 마른 넝쿨, 비 내리는 밤과 젖은 포장은 일치합니다. 간판은 요청된 ‘미추홀도서관, 인천’이 아니라 참조의 ‘시립도서관’으로 읽히며, 마지막의 모든 문자 비가독화 지시에도 어긋납니다. 실내 누수는 확인되지 않습니다.",
        "hard_violations": [
         "간판의 ‘시립도서관’이 판독 가능하여 마지막의 읽을 수 있는 문자 금지 지시를 위반합니다."
        ],
        "physics": "차량 바퀴와 지면 접점은 하단 크롭 밖이라 지지 상태를 직접 확인할 수 없지만, 공중에 떠 있다고 볼 시각적 근거도 없습니다. 지붕 설비는 지붕에 붙어 있고 간판은 출입구 구조에 고정되어 있습니다. 젖은 포장 반사는 가능한 모습이며 움직임이나 충돌은 보이지 않습니다."
       },
       {
        "label": "A",
        "direction": "인물이나 조준 물체는 없습니다. 캠핑카 운전석과 앞바퀴는 화면 왼쪽을 향하고 차량은 도서관 앞에 정지해 있습니다. 카메라는 비스듬한 위쪽에서 차량 측면과 지붕을 봅니다. 도서관 정면은 화면 상부에서 왼쪽 뒤로 물러나고 간판 전면은 카메라를 향합니다.",
        "built_space": "참조와 대응하는 왼쪽 창탑 하나, 두 층의 가로 창열, 중앙 오른쪽 유리 출입구 한 구역, 그 위 녹슨 간판 하나, 전면 화단과 출입 계단·난간이 보입니다. 차량은 오른쪽 아래 전경에 놓이고 왼쪽에는 넓은 젖은 주차장이 남아 건물·차량·지면의 크기 관계가 드러납니다. 차량 뒤쪽과 하부 일부는 잘렸지만 앞바퀴의 접지는 보입니다.",
        "entities": "인물은 없습니다. 운전실과 상부 침상 돌출부를 갖춘 낡은 캠핑카 한 대가 명확합니다. 외벽의 깊은 균열, 마른 넝쿨, 노후한 간판, 야간 빗줄기와 물 고인 포장이 보입니다. 간판은 ‘시립도서관’으로 읽혀 지정 명칭과 다르고 최종 문자 비가독화 지시도 어깁니다. 지붕의 물 고임은 보이지만 실내 누수까지 확인되지는 않습니다.",
        "hard_violations": [
         "간판의 ‘시립도서관’이 판독 가능하여 마지막의 읽을 수 있는 문자 금지 지시를 위반합니다."
        ],
        "physics": "보이는 앞바퀴가 포장면에 닿아 차량을 지지하며 정차 상태가 자연스럽습니다. 지붕 설비와 난간은 차체에 부착되어 있고 물은 지붕과 포장면에 고여 있습니다. 출입구 간판은 건물 구조가 지지합니다. 떠 있는 물체나 불가능한 반사는 보이지 않습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.417
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.167
   },
   "violations": {
    "B": [
     "[gpt-high] 간판의 ‘시립도서관’이 판독 가능하여 마지막의 읽을 수 있는 문자 금지 지시를 위반합니다."
    ],
    "A": [
     "[gpt-high] 간판의 ‘시립도서관’이 판독 가능하여 마지막의 읽을 수 있는 문자 금지 지시를 위반합니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 1167
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "요구된 간판 텍스트('미추홀도서관, 인천') 대신 레퍼런스의 글씨('시립도서관')가 그대로 출력되었으나, 지정된 부감 앵글과 낡은 캠핑카(모터홈)의 묘사가 프롬프트의 지시를 충실히 따랐습니다.  ★위반: [gpt-high] 간판의 ‘시립도서관’이 판독 가능하여 마지막의 읽을 수 있는 문자 금지 지시를 위반합니다."
   },
   {
    "label": "B",
    "score": 1167,
    "verdict_ko": "간판 텍스트 요구사항을 실패했으며, 캠핑카가 견인차 없는 트레일러 형태로 묘사되어 '도착해서 멈춰 선' 상황에 어울리지 않고 배경 원경이 레퍼런스에 비해 과장되었습니다.  ★위반: [gpt-high] 간판의 ‘시립도서관’이 판독 가능하여 마지막의 읽을 수 있는 문자 금지 지시를 위반합니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L202B03.png",
    "asset_id": "1b1737dc-7622-466c-bcf0-357121348d7e",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-b3ba-755d-a4b6-4be46cc8b3b5",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S46sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:41:29.922748+00:00",
  "fingerprint": "3e22f9109e14f9305fcb372f0ac4f3480faf25d2dd9b9e9ecb5f513bd3810291",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S46sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S46sh1_sel.png",
  "source_sha256": "49341a5cef17c7fc9ca0be998e239607602606e8e8633428bf3649835f2a9d9f",
  "file": "S46sh1_cine.png",
  "staged_sha256": "2da0014c8e824495f79baea7bad1eecafb6be941e913bea23364a3b7723e5487",
  "latency_ms": 10363
 },
 "S46sh17::signage": {
  "fp": "affc0ecbd612238c",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::9e4648ab22ad1f33": {
  "subjects": [
   {
    "subject_native": "인천 미추홀도서관 열람실 및 천창",
    "search_terms_native": [
     "미추홀도서관 열람실",
     "미추홀도서관 내부",
     "미추홀도서관 종합자료실",
     "인천 미추홀도서관 천창"
    ],
    "language_lock_native": "모든 검색어는 반드시 한국어로만 작성해야 하며 다른 언어로 번역하거나 추가해서는 안 됩니다.",
    "reason_ko": "인천 미추홀도서관 특유의 현대식 천창 구조와 한국 공공도서관 열람실 형태 대신 서구식 고전 도서관 폐허로 왜곡될 가능성이 큼."
   }
  ],
  "subject_text": "미추홀도서관 열람실\n먼지와 거미줄이 가득한 버려진 열람실. 부서진 책장과 낡은 책, 고물 컴퓨터가 남아 있고 깨진 천창으로 잿빛 햇살이 들어온다.",
  "identity": "canonical",
  "scope_id": "L202",
  "scope_role": "location_interior",
  "scope_sha": "b3dc6c070f833894"
 },
 "era_fail::06265b5c57a63a76": {
  "stage": "research",
  "subject": "인천 미추홀도서관 열람실 및 천창",
  "terms": [
   "미추홀도서관 열람실",
   "미추홀도서관 내부",
   "미추홀도서관 종합자료실",
   "인천 미추홀도서관 천창"
  ],
  "status": "no_usable",
  "queries": [
   [
    "인천 미추홀도서관 내부 종합자료실 열람실 천창"
   ],
   [
    "\"미추홀도서관\" \"천창\"",
    "\"미추홀도서관\" \"종합자료실\" \"내부\""
   ],
   [
    "미추홀도서관 내부 사진",
    "미추홀도서관 천장",
    "미추홀도서관 시설현황"
   ]
  ],
  "candidate_urls": [
   "https://vmspace.com/ActiveFile/spacem.org/board_img/20619362075f698adf0ed6e.jpg",
   "https://images.squarespace-cdn.com/content/v1/5ac2ddf285ede15e39a57666/1523432273711-PPBJHS6C6W1DB5SDD130/Hannae-B-3.jpg",
   "https://cdn.welfarehello.com/naver-blog/production/tong_namgu/2023-06/223117359642/tong_namgu_223117359642_9.png?f=webp&q=80&w=800",
   "https://cdn.welfarehello.com/naver-blog/production/tong_namgu/2024-03/223396831400/tong_namgu_223396831400_6.jpg"
  ],
  "coarse": {
   "eligible": [],
   "chosen_index": 0,
   "reason": "종류·보임·기준을 다 만족하는 후보가 없다 — 이 라운드에선 안 고른다",
   "single_judge": true,
   "rejected_judges": {}
  },
  "verdicts": [
   {
    "index": 1,
    "object_type_match": "yes",
    "visible": true,
    "criteria_match": "unsure",
    "similarity": 95
   },
   {
    "index": 2,
    "object_type_match": "yes",
    "visible": true,
    "criteria_match": "unsure",
    "similarity": 85
   },
   {
    "index": 3,
    "object_type_match": "no",
    "visible": true,
    "criteria_match": "unsure",
    "similarity": 40
   },
   {
    "index": 4,
    "object_type_match": "no",
    "visible": true,
    "criteria_match": "unsure",
    "similarity": 30
   }
  ],
  "chosen_reason_ko": "",
  "attempts": 8
 },
 "S46sh17::bgfirst_bg": {
  "input_fingerprint": "45002603facd4e3e",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리의 손에 들린 책을 향해 거칠게 팔을 뻗어 타격하는 mid-impact 순간의 현우.\n\nLOCATION (lock): At the group's makeshift resting spot inside the dusty abandoned-library reading room, in the dim nighttime interior.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Presented robotics book (Open to the next-generation robot chapter and still at 찰리's hands during impact) — The open printed pages are angled toward 현우, with part of the chapter visible obliquely to camera; used as Small central contact point joining the extended forearm and 찰리's hands; Reading-room bookshelves (Disordered and dusty) — Partial shelf faces remain behind the two figures; used as Soft spatial context without cluttering the impact silhouette.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained ambient illumination appropriate to the nighttime reading room, with enough tonal separation to read the arm, hands, and book without adding a source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리의 손에 들린 책을 향해 거칠게 팔을 뻗어 타격하는 mid-impact 순간의 현우.\n\nLOCATION (lock): At the group's makeshift resting spot inside the dusty abandoned-library reading room, in the dim nighttime interior.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Presented robotics book (Open to the next-generation robot chapter and still at 찰리's hands during impact) — The open printed pages are angled toward 현우, with part of the chapter visible obliquely to camera; used as Small central contact point joining the extended forearm and 찰리's hands; Reading-room bookshelves (Disordered and dusty) — Partial shelf faces remain behind the two figures; used as Soft spatial context without cluttering the impact silhouette.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained ambient illumination appropriate to the nighttime reading room, with enough tonal separation to read the arm, hands, and book without adding a source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S46sh17__bgfirst_bg.png",
  "asset_id": "423c83f7-5368-4b62-8c21-328accdd47f4",
  "input_asset_ids": [
   "c2cd0399-5f87-4274-8987-1f4261def294",
   "247ce4f2-a8a9-4191-9f73-361363704c8d"
  ]
 },
 "S46sh17": {
  "input_fingerprint": "4cb2a561168111a0",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 찰리의 손에 들린 책을 향해 거칠게 팔을 뻗어 타격하는 mid-impact 순간의 현우.\n\nLOCATION (lock): At the group's makeshift resting spot inside the dusty abandoned-library reading room, in the dim nighttime interior. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Presented robotics book (Open to the next-generation robot chapter and still at 찰리's hands during impact) — The open printed pages are angled toward 현우, with part of the chapter visible obliquely to camera; used as Small central contact point joining the extended forearm and 찰리's hands; Reading-room bookshelves (Disordered and dusty) — Partial shelf faces remain behind the two figures; used as Soft spatial context without cluttering the impact silhouette.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained ambient illumination appropriate to the nighttime reading room, with enough tonal separation to read the arm, hands, and book without adding a source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The abandoned reading room remains dusty and cobwebbed, with disordered shelves and books. Opened canned food and the scavenged bags remain at the makeshift eating place, while the Ubik book is open to its next-generation robot chapter. 현우: He remains at the eating place with facial bruises and the untreated leg wound. The contact card is still inside his shoe, and the scavenged disposable phone remains among his belongings. 찰리: He holds the Ubik book open for display, retaining his worn body, earlier disguise and store blanket.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리, 현우 right now, so 찰리, 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리, 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 찰리의 손에 들린 책을 향해 거칠게 팔을 뻗어 타격하는 mid-impact 순간의 현우.\n\nLOCATION (lock): At the group's makeshift resting spot inside the dusty abandoned-library reading room, in the dim nighttime interior. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Presented robotics book (Open to the next-generation robot chapter and still at 찰리's hands during impact) — The open printed pages are angled toward 현우, with part of the chapter visible obliquely to camera; used as Small central contact point joining the extended forearm and 찰리's hands; Reading-room bookshelves (Disordered and dusty) — Partial shelf faces remain behind the two figures; used as Soft spatial context without cluttering the impact silhouette.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained ambient illumination appropriate to the nighttime reading room, with enough tonal separation to read the arm, hands, and book without adding a source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The abandoned reading room remains dusty and cobwebbed, with disordered shelves and books. Opened canned food and the scavenged bags remain at the makeshift eating place, while the Ubik book is open to its next-generation robot chapter. 현우: He remains at the eating place with facial bruises and the untreated leg wound. The contact card is still inside his shoe, and the scavenged disposable phone remains among his belongings. 찰리: He holds the Ubik book open for display, retaining his worn body, earlier disguise and store blanket.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리, 현우 right now, so 찰리, 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리, 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 찰리의 손에 들린 책을 향해 거칠게 팔을 뻗어 타격하는 mid-impact 순간의 현우.\n\nLOCATION (lock): At the group's makeshift resting spot inside the dusty abandoned-library reading room, in the dim nighttime interior. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Presented robotics book (Open to the next-generation robot chapter and still at 찰리's hands during impact) — The open printed pages are angled toward 현우, with part of the chapter visible obliquely to camera; used as Small central contact point joining the extended forearm and 찰리's hands; Reading-room bookshelves (Disordered and dusty) — Partial shelf faces remain behind the two figures; used as Soft spatial context without cluttering the impact silhouette.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained ambient illumination appropriate to the nighttime reading room, with enough tonal separation to read the arm, hands, and book without adding a source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The abandoned reading room remains dusty and cobwebbed, with disordered shelves and books. Opened canned food and the scavenged bags remain at the makeshift eating place, while the Ubik book is open to its next-generation robot chapter. 현우: He remains at the eating place with facial bruises and the untreated leg wound. The contact card is still inside his shoe, and the scavenged disposable phone remains among his belongings. 찰리: He holds the Ubik book open for display, retaining his worn body, earlier disguise and store blanket.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리, 현우 right now, so 찰리, 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리, 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S46sh17__bgfirst_bg.png",
     "asset_id": "423c83f7-5368-4b62-8c21-328accdd47f4",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S46sh17.png",
     "asset_id": "c2cd0399-5f87-4274-8987-1f4261def294",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 유빅사 책: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:1126012>",
     "asset_id": "6f976831-f189-4330-b56f-fdd0f9951db8",
     "role": "prop_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L202B04.png",
     "asset_id": "247ce4f2-a8a9-4191-9f73-361363704c8d",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 유빅사 책: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:1126012>",
     "asset_id": "6f976831-f189-4330-b56f-fdd0f9951db8",
     "role": "prop_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선과 뻗은 주먹이 찰리가 들고 있는 책의 펼쳐진 페이지를 정확히 향하고 있으며, 책이 주먹과 로봇 손 사이의 접촉점이 되어 타격이 이루어지는 순간을 보여줍니다.",
    "built_space": "어두운 야간의 도서관 내부로, 위치 레퍼런스와 동일한 낡은 책장들이 배경에 있습니다. 앞쪽 테이블에는 통조림, 컵, 스마트폰 등 지시된 소품들이 놓여 있습니다.",
    "entities": "현우는 레퍼런스의 외모와 일치하며 얼굴에 멍이 있습니다. 찰리는 로봇 본체의 디테일이 레퍼런스와 일치하지만, 프롬프트에서 요구된 담요(store blanket)를 두르지 않았습니다. 책은 펼쳐져 있으나 레퍼런스의 고릴라 로봇 그림과 다릅니다.",
    "hard_violations": [],
    "physics": "현우는 테이블에 앉은 채 상체를 내밀어 타격하고 있으며, 찰리는 양손으로 책을 안정적으로 받치고 있어 모든 자세와 사물의 지지가 물리적으로 자연스럽습니다."
   },
   {
    "label": "B",
    "direction": "현우의 주먹이 지시된 책이 아닌 찰리의 가슴 부위를 향해 뻗어 타격하고 있으며, 책은 타격 지점 아래에 들려 있습니다.",
    "built_space": "도서관 내부로, 배경에 책장이 위치하고 우측에 통조림이 놓인 테이블이 보입니다. 공간의 구성은 레퍼런스와 부합합니다.",
    "entities": "현우와 찰리의 얼굴 및 외형이 레퍼런스와 일치합니다. 찰리는 요구된 담요를 두르고 있으며, 책에는 레퍼런스와 유사한 고릴라 로봇 형태가 그려져 있습니다.",
    "hard_violations": [],
    "physics": "현우가 서서 주먹을 뻗는 자세를 취하고 있으며, 찰리는 로봇 팔로 책을 들고 있어 중력이나 지지에 어긋나는 부분은 없습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "핵심 동작인 책을 타격하는 순간을 프롬프트의 지시대로 정확하게 구현했으나, 찰리가 두르고 있어야 할 담요와 책 안의 세부 일러스트가 일치하지 않는 점이 감점 요소입니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "찰리의 담요와 책의 디테일은 레퍼런스에 가깝게 표현되었으나, 샷 텍스트의 가장 중요한 지시인 '책을 타격하는 동작' 대신 로봇의 몸을 치고 있어 연출 우선순위에서 크게 어긋납니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선과 뻗은 주먹이 찰리가 들고 있는 책의 펼쳐진 페이지를 정확히 향하고 있으며, 책이 주먹과 로봇 손 사이의 접촉점이 되어 타격이 이루어지는 순간을 보여줍니다.",
        "built_space": "어두운 야간의 도서관 내부로, 위치 레퍼런스와 동일한 낡은 책장들이 배경에 있습니다. 앞쪽 테이블에는 통조림, 컵, 스마트폰 등 지시된 소품들이 놓여 있습니다.",
        "entities": "현우는 레퍼런스의 외모와 일치하며 얼굴에 멍이 있습니다. 찰리는 로봇 본체의 디테일이 레퍼런스와 일치하지만, 프롬프트에서 요구된 담요(store blanket)를 두르지 않았습니다. 책은 펼쳐져 있으나 레퍼런스의 고릴라 로봇 그림과 다릅니다.",
        "hard_violations": [],
        "physics": "현우는 테이블에 앉은 채 상체를 내밀어 타격하고 있으며, 찰리는 양손으로 책을 안정적으로 받치고 있어 모든 자세와 사물의 지지가 물리적으로 자연스럽습니다."
       },
       {
        "label": "B",
        "direction": "현우의 주먹이 지시된 책이 아닌 찰리의 가슴 부위를 향해 뻗어 타격하고 있으며, 책은 타격 지점 아래에 들려 있습니다.",
        "built_space": "도서관 내부로, 배경에 책장이 위치하고 우측에 통조림이 놓인 테이블이 보입니다. 공간의 구성은 레퍼런스와 부합합니다.",
        "entities": "현우와 찰리의 얼굴 및 외형이 레퍼런스와 일치합니다. 찰리는 요구된 담요를 두르고 있으며, 책에는 레퍼런스와 유사한 고릴라 로봇 형태가 그려져 있습니다.",
        "hard_violations": [],
        "physics": "현우가 서서 주먹을 뻗는 자세를 취하고 있으며, 찰리는 로봇 팔로 책을 들고 있어 중력이나 지지에 어긋나는 부분은 없습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "핵심 동작인 책을 타격하는 순간을 프롬프트의 지시대로 정확하게 구현했으나, 찰리가 두르고 있어야 할 담요와 책 안의 세부 일러스트가 일치하지 않는 점이 감점 요소입니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "찰리의 담요와 책의 디테일은 레퍼런스에 가깝게 표현되었으나, 샷 텍스트의 가장 중요한 지시인 '책을 타격하는 동작' 대신 로봇의 몸을 치고 있어 연출 우선순위에서 크게 어긋납니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선과 뻗은 주먹이 찰리가 들고 있는 책의 펼쳐진 페이지를 정확히 향하고 있으며, 책이 주먹과 로봇 손 사이의 접촉점이 되어 타격이 이루어지는 순간을 보여줍니다.",
        "built_space": "어두운 야간의 도서관 내부로, 위치 레퍼런스와 동일한 낡은 책장들이 배경에 있습니다. 앞쪽 테이블에는 통조림, 컵, 스마트폰 등 지시된 소품들이 놓여 있습니다.",
        "entities": "현우는 레퍼런스의 외모와 일치하며 얼굴에 멍이 있습니다. 찰리는 로봇 본체의 디테일이 레퍼런스와 일치하지만, 프롬프트에서 요구된 담요(store blanket)를 두르지 않았습니다. 책은 펼쳐져 있으나 레퍼런스의 고릴라 로봇 그림과 다릅니다.",
        "hard_violations": [],
        "physics": "현우는 테이블에 앉은 채 상체를 내밀어 타격하고 있으며, 찰리는 양손으로 책을 안정적으로 받치고 있어 모든 자세와 사물의 지지가 물리적으로 자연스럽습니다."
       },
       {
        "label": "B",
        "direction": "현우의 주먹이 지시된 책이 아닌 찰리의 가슴 부위를 향해 뻗어 타격하고 있으며, 책은 타격 지점 아래에 들려 있습니다.",
        "built_space": "도서관 내부로, 배경에 책장이 위치하고 우측에 통조림이 놓인 테이블이 보입니다. 공간의 구성은 레퍼런스와 부합합니다.",
        "entities": "현우와 찰리의 얼굴 및 외형이 레퍼런스와 일치합니다. 찰리는 요구된 담요를 두르고 있으며, 책에는 레퍼런스와 유사한 고릴라 로봇 형태가 그려져 있습니다.",
        "hard_violations": [],
        "physics": "현우가 서서 주먹을 뻗는 자세를 취하고 있으며, 찰리는 로봇 팔로 책을 들고 있어 중력이나 지지에 어긋나는 부분은 없습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "담요와 인물 외형은 잘 유지했지만, 현우의 주먹이 책이 아니라 찰리의 가슴을 향해 핵심인 책 타격 순간을 구현하지 못했다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "현우가 찰리의 손에 남아 있는 펼친 책을 실제로 타격하는 접점을 구현했으며, 다만 담요가 없고 구도가 요구보다 다소 넓다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 찰리의 얼굴 쪽을 바라보며 오른팔을 왼쪽으로 뻗고 있다. 주먹의 진행 방향과 끝점은 찰리의 가슴 장갑이며, 펼친 책은 주먹보다 아래이자 화면 오른쪽에 있어 타격 경로에서 벗어나 있다. 찰리의 얼굴은 아래쪽 책과 팔 쪽으로 기울어 있다. 책의 인쇄면은 위로 열려 현우와 카메라 양쪽에서 비스듬히 보인다.",
        "built_space": "왼쪽 가장자리의 높은 서가 한 면, 중앙 기둥 옆 서가 한 면, 오른쪽 뒤로 이어지는 서가들이 보인다. 어두운 창, 벗겨진 기둥과 낡은 바닥은 참조 열람실과 부합한다. 오른쪽 식사 탁자 하나와 그 뒤 의자 등받이 하나가 명확히 보이며, 두 인물은 탁자 옆에 있다. 상체 중심의 미디엄 구도지만 책은 팔과 손을 잇는 중앙 접점이 되지 못한다. 불가능한 반사나 명백한 고정 시설 중복은 보이지 않는다.",
        "entities": "현우는 검은 헝클어진 머리와 앳된 동아시아계 남성 얼굴, 볼의 상처를 갖추며 참조와 대체로 닮았다. 남색 상의 위 회색 셔츠가 추가되어 있다. 찰리는 샌드 베이지 장갑, 흰 기계 얼굴, 긴 기계 팔을 갖추고 체크 담요를 두르고 있다. 책 한 권은 펼쳐져 로봇 그림과 인쇄 문단을 보여 주지만 참조의 고릴라형 로봇 도판 구성과는 다르다. 탁자에는 가방, 여러 통조림, 금속 컵과 천이 있다. 통조림이 개봉된 상태인지는 명확하지 않다. 다리 상처와 신발 속 카드는 프레임 밖이므로 판단하지 않는다.",
        "hard_violations": [],
        "physics": "찰리의 기계 손들이 책의 아래쪽과 옆쪽을 받치고 있어 책은 공중에 떠 있지 않다. 현우의 뻗은 팔도 어깨부터 주먹까지 자연스럽게 연결된다. 하체 지지점은 잘렸지만 공중에 뜬 자세는 아니다. 다만 주먹과 책 사이에 거리가 있고 책에 충격이 전달되는 접촉도 없어, 물리적으로 보이는 동작은 책 타격이 아니다."
       },
       {
        "label": "B",
        "direction": "현우의 시선은 펼친 책과 주먹의 접촉부를 향한다. 오른팔이 화면 오른쪽으로 뻗으며 주먹이 책의 왼쪽 펼침면에 닿아 있어 지정된 목표에 타격이 도달한다. 찰리도 고개를 내려 책 쪽을 보고 있다. 책은 현우 쪽으로 기울어진 채 인쇄면 일부를 카메라에 비스듬히 드러낸다.",
        "built_space": "왼쪽 벽 서가 한 면, 중앙의 높은 서가 열 두 개, 뒤쪽의 작은 서가 한 면이 구별된다. 오른쪽의 검은 창틀, 낡은 기둥, 천장 조명 기구와 서가 사이 통로가 참조 장소의 구조를 잘 유지한다. 전경 식사 탁자 하나와 왼쪽·아래쪽 의자 등받이가 보인다. 현우는 탁자 뒤 왼쪽에서 몸을 내밀고 찰리는 오른쪽에 있어 책을 사이에 둔 동작이 성립한다. 다만 탁자 면적과 찰리의 허벅지까지 포함해 요구한 미디엄보다 조금 넓게 느껴진다.",
        "entities": "현우는 참조와 유사한 검은 머리의 앳된 동아시아계 남성이며 남색 상의와 볼의 멍을 유지한다. 찰리의 흰 얼굴판, 베이지 장갑, 육중한 몸통과 긴 팔은 참조에 가깝지만 요구된 담요와 이전 위장은 보이지 않는다. 펼친 책에는 로봇 도해와 인쇄 문단이 있으나 참조의 고릴라형 로봇 대신 직립형 로봇 선화가 들어 있다. 식사 자리에는 가방, 통조림 여러 개, 금속 컵, 천과 휴대전화 한 개가 보인다. 통조림의 개봉 상태는 분명하지 않으며 책의 문장은 판독하기 어렵다.",
        "hard_violations": [],
        "physics": "찰리의 화면 오른쪽 기계 손이 책의 아래 모서리를 받치고 있어 책의 지지가 보인다. 현우의 주먹은 펼친 면에 닿고 접촉 주변의 흐림과 먼지가 타격 순간을 뒷받침한다. 어깨와 몸통을 앞으로 기울여 팔을 뻗는 자세도 가능한 동작이다. 발은 프레임 밖이지만 두 몸통은 아래로 이어지며, 지지 없이 떠 있는 인물이나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "담요와 인물 외형은 잘 유지했지만, 현우의 주먹이 책이 아니라 찰리의 가슴을 향해 핵심인 책 타격 순간을 구현하지 못했다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "현우가 찰리의 손에 남아 있는 펼친 책을 실제로 타격하는 접점을 구현했으며, 다만 담요가 없고 구도가 요구보다 다소 넓다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 찰리의 얼굴 쪽을 바라보며 오른팔을 왼쪽으로 뻗고 있다. 주먹의 진행 방향과 끝점은 찰리의 가슴 장갑이며, 펼친 책은 주먹보다 아래이자 화면 오른쪽에 있어 타격 경로에서 벗어나 있다. 찰리의 얼굴은 아래쪽 책과 팔 쪽으로 기울어 있다. 책의 인쇄면은 위로 열려 현우와 카메라 양쪽에서 비스듬히 보인다.",
        "built_space": "왼쪽 가장자리의 높은 서가 한 면, 중앙 기둥 옆 서가 한 면, 오른쪽 뒤로 이어지는 서가들이 보인다. 어두운 창, 벗겨진 기둥과 낡은 바닥은 참조 열람실과 부합한다. 오른쪽 식사 탁자 하나와 그 뒤 의자 등받이 하나가 명확히 보이며, 두 인물은 탁자 옆에 있다. 상체 중심의 미디엄 구도지만 책은 팔과 손을 잇는 중앙 접점이 되지 못한다. 불가능한 반사나 명백한 고정 시설 중복은 보이지 않는다.",
        "entities": "현우는 검은 헝클어진 머리와 앳된 동아시아계 남성 얼굴, 볼의 상처를 갖추며 참조와 대체로 닮았다. 남색 상의 위 회색 셔츠가 추가되어 있다. 찰리는 샌드 베이지 장갑, 흰 기계 얼굴, 긴 기계 팔을 갖추고 체크 담요를 두르고 있다. 책 한 권은 펼쳐져 로봇 그림과 인쇄 문단을 보여 주지만 참조의 고릴라형 로봇 도판 구성과는 다르다. 탁자에는 가방, 여러 통조림, 금속 컵과 천이 있다. 통조림이 개봉된 상태인지는 명확하지 않다. 다리 상처와 신발 속 카드는 프레임 밖이므로 판단하지 않는다.",
        "hard_violations": [],
        "physics": "찰리의 기계 손들이 책의 아래쪽과 옆쪽을 받치고 있어 책은 공중에 떠 있지 않다. 현우의 뻗은 팔도 어깨부터 주먹까지 자연스럽게 연결된다. 하체 지지점은 잘렸지만 공중에 뜬 자세는 아니다. 다만 주먹과 책 사이에 거리가 있고 책에 충격이 전달되는 접촉도 없어, 물리적으로 보이는 동작은 책 타격이 아니다."
       },
       {
        "label": "A",
        "direction": "현우의 시선은 펼친 책과 주먹의 접촉부를 향한다. 오른팔이 화면 오른쪽으로 뻗으며 주먹이 책의 왼쪽 펼침면에 닿아 있어 지정된 목표에 타격이 도달한다. 찰리도 고개를 내려 책 쪽을 보고 있다. 책은 현우 쪽으로 기울어진 채 인쇄면 일부를 카메라에 비스듬히 드러낸다.",
        "built_space": "왼쪽 벽 서가 한 면, 중앙의 높은 서가 열 두 개, 뒤쪽의 작은 서가 한 면이 구별된다. 오른쪽의 검은 창틀, 낡은 기둥, 천장 조명 기구와 서가 사이 통로가 참조 장소의 구조를 잘 유지한다. 전경 식사 탁자 하나와 왼쪽·아래쪽 의자 등받이가 보인다. 현우는 탁자 뒤 왼쪽에서 몸을 내밀고 찰리는 오른쪽에 있어 책을 사이에 둔 동작이 성립한다. 다만 탁자 면적과 찰리의 허벅지까지 포함해 요구한 미디엄보다 조금 넓게 느껴진다.",
        "entities": "현우는 참조와 유사한 검은 머리의 앳된 동아시아계 남성이며 남색 상의와 볼의 멍을 유지한다. 찰리의 흰 얼굴판, 베이지 장갑, 육중한 몸통과 긴 팔은 참조에 가깝지만 요구된 담요와 이전 위장은 보이지 않는다. 펼친 책에는 로봇 도해와 인쇄 문단이 있으나 참조의 고릴라형 로봇 대신 직립형 로봇 선화가 들어 있다. 식사 자리에는 가방, 통조림 여러 개, 금속 컵, 천과 휴대전화 한 개가 보인다. 통조림의 개봉 상태는 분명하지 않으며 책의 문장은 판독하기 어렵다.",
        "hard_violations": [],
        "physics": "찰리의 화면 오른쪽 기계 손이 책의 아래 모서리를 받치고 있어 책의 지지가 보인다. 현우의 주먹은 펼친 면에 닿고 접촉 주변의 흐림과 먼지가 타격 순간을 뒷받침한다. 어깨와 몸통을 앞으로 기울여 팔을 뻗는 자세도 가능한 동작이다. 발은 프레임 밖이지만 두 몸통은 아래로 이어지며, 지지 없이 떠 있는 인물이나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.875
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.875
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 875
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "핵심 동작인 책을 타격하는 순간을 프롬프트의 지시대로 정확하게 구현했으나, 찰리가 두르고 있어야 할 담요와 책 안의 세부 일러스트가 일치하지 않는 점이 감점 요소입니다."
   },
   {
    "label": "B",
    "score": 875,
    "verdict_ko": "찰리의 담요와 책의 디테일은 레퍼런스에 가깝게 표현되었으나, 샷 텍스트의 가장 중요한 지시인 '책을 타격하는 동작' 대신 로봇의 몸을 치고 있어 연출 우선순위에서 크게 어긋납니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L202B04.png",
    "asset_id": "247ce4f2-a8a9-4191-9f73-361363704c8d",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   },
   {
    "label": "PROP REFERENCE — 유빅사 책: the exact object appearing in this shot; match its look, material and wear exactly.",
    "path": "<bytes:1126012>",
    "asset_id": "6f976831-f189-4330-b56f-fdd0f9951db8",
    "role": "prop_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae927-b570-7728-b6ac-7fb37dee007b",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S46sh17__bgfirst_bg.png",
   "bg_asset_id": "423c83f7-5368-4b62-8c21-328accdd47f4",
   "bg_record_key": "S46sh17::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S46sh17::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:44:22.548573+00:00",
  "fingerprint": "118ee82c69792d6cc6ca00d1a1ccfd7798222b36f28aa19e9fc8ca0d44964c9b",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S46sh17_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S46sh17_sel.png",
  "source_sha256": "c8db1328204b387b17ceaf651c4ddce0bda302c9408008fd8e118d76ad88cc76",
  "file": "S46sh17_cine.png",
  "staged_sha256": "2a98bee69e9146bfd1233782adef582ccf94b8936ea923e7c5364c2e49680ea3",
  "latency_ms": 11964
 },
 "S46sh28::signage": {
  "fp": "9d890661c8b50cc9",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S46sh28": {
  "input_fingerprint": "18fcb28001004262",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 모니터 화면의 불빛이 비친 채 입 모양을 크고 둥글게 벌리고 집중하는 찰리의 낡은 금속 얼굴.\n\nLOCATION (lock): At an old computer among the abandoned library's dusty bookshelves, with monitor light illuminating the otherwise dark area. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old computer monitor (In use during 찰리's speech practice) — Only a narrow side edge is visible; the display face points toward 찰리 outside the crop; used as Locates his attention without showing invented screen content; Library books and shelving (Dust-covered) — Indistinct fragments remain behind his head; used as Retains the abandoned reading-room context at shallow focus.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the monitor's established light articulate 찰리's worn metal face against the subdued room, without assigning the display an unsupported color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The dusty library shelves and books remain disordered, and an old computer is now running. Boxes and books have been arranged as makeshift bedding in one corner. 찰리: He is awake at the old computer, practicing speech after gathering data. His worn body and retained disguise are unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 모니터 화면의 불빛이 비친 채 입 모양을 크고 둥글게 벌리고 집중하는 찰리의 낡은 금속 얼굴.\n\nLOCATION (lock): At an old computer among the abandoned library's dusty bookshelves, with monitor light illuminating the otherwise dark area. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old computer monitor (In use during 찰리's speech practice) — Only a narrow side edge is visible; the display face points toward 찰리 outside the crop; used as Locates his attention without showing invented screen content; Library books and shelving (Dust-covered) — Indistinct fragments remain behind his head; used as Retains the abandoned reading-room context at shallow focus.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the monitor's established light articulate 찰리's worn metal face against the subdued room, without assigning the display an unsupported color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The dusty library shelves and books remain disordered, and an old computer is now running. Boxes and books have been arranged as makeshift bedding in one corner. 찰리: He is awake at the old computer, practicing speech after gathering data. His worn body and retained disguise are unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 모니터 화면의 불빛이 비친 채 입 모양을 크고 둥글게 벌리고 집중하는 찰리의 낡은 금속 얼굴.\n\nLOCATION (lock): At an old computer among the abandoned library's dusty bookshelves, with monitor light illuminating the otherwise dark area. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old computer monitor (In use during 찰리's speech practice) — Only a narrow side edge is visible; the display face points toward 찰리 outside the crop; used as Locates his attention without showing invented screen content; Library books and shelving (Dust-covered) — Indistinct fragments remain behind his head; used as Retains the abandoned reading-room context at shallow focus.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the monitor's established light articulate 찰리's worn metal face against the subdued room, without assigning the display an unsupported color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The dusty library shelves and books remain disordered, and an old computer is now running. Boxes and books have been arranged as makeshift bedding in one corner. 찰리: He is awake at the old computer, practicing speech after gathering data. His worn body and retained disguise are unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "찰리의 시선과 안면부가 화면 밖 우측의 모니터를 향하고 있습니다.",
    "built_space": "어두운 도서관 환경. 배경에 먼지 쌓인 책장이 흐릿하게 보이고, 우측에 낡은 모니터의 측면이 배치되어 있습니다.",
    "entities": "찰리는 낡은 금속 재질의 베이지색 장갑판과 흰색 마스크형 얼굴을 지닌 로봇으로 레퍼런스와 일치합니다.",
    "hard_violations": [],
    "physics": "찰리의 상체는 바닥에 안정적으로 위치해 있으며, 구조적으로 불가능한 포즈나 떠 있는 요소는 없습니다."
   },
   {
    "label": "B",
    "direction": "찰리의 시선이 우측 프레임 바깥의 모니터 화면을 향해 집중되어 있습니다.",
    "built_space": "배경으로 낡은 책과 책장이 아웃포커싱되어 있으며, 우측 전경에 모니터 측면이 자리 잡고 있습니다.",
    "entities": "찰리의 외형(베이지색 장갑판, 흰색 마스크, 눈의 형태)이 레퍼런스에 매우 부합하며 입을 크게 벌리고 있습니다.",
    "hard_violations": [],
    "physics": "상체의 무게감과 기계적 관절 형태가 자연스럽고 지지 기반에 문제가 없습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "프롬프트가 요구한 '입을 크고 둥글게 벌린' 표정을 더 정확하고 생동감 있게 묘사했으며, 조명과 프레이밍 모두 훌륭합니다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "지시된 클로즈업 앵글과 배경 요소는 잘 구현되었으나, 입을 크게 벌린 형태가 경쟁작에 비해 덜 둥글고 소극적입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 시선과 안면부가 화면 밖 우측의 모니터를 향하고 있습니다.",
        "built_space": "어두운 도서관 환경. 배경에 먼지 쌓인 책장이 흐릿하게 보이고, 우측에 낡은 모니터의 측면이 배치되어 있습니다.",
        "entities": "찰리는 낡은 금속 재질의 베이지색 장갑판과 흰색 마스크형 얼굴을 지닌 로봇으로 레퍼런스와 일치합니다.",
        "hard_violations": [],
        "physics": "찰리의 상체는 바닥에 안정적으로 위치해 있으며, 구조적으로 불가능한 포즈나 떠 있는 요소는 없습니다."
       },
       {
        "label": "B",
        "direction": "찰리의 시선이 우측 프레임 바깥의 모니터 화면을 향해 집중되어 있습니다.",
        "built_space": "배경으로 낡은 책과 책장이 아웃포커싱되어 있으며, 우측 전경에 모니터 측면이 자리 잡고 있습니다.",
        "entities": "찰리의 외형(베이지색 장갑판, 흰색 마스크, 눈의 형태)이 레퍼런스에 매우 부합하며 입을 크게 벌리고 있습니다.",
        "hard_violations": [],
        "physics": "상체의 무게감과 기계적 관절 형태가 자연스럽고 지지 기반에 문제가 없습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "프롬프트가 요구한 '입을 크고 둥글게 벌린' 표정을 더 정확하고 생동감 있게 묘사했으며, 조명과 프레이밍 모두 훌륭합니다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "지시된 클로즈업 앵글과 배경 요소는 잘 구현되었으나, 입을 크게 벌린 형태가 경쟁작에 비해 덜 둥글고 소극적입니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 시선과 안면부가 화면 밖 우측의 모니터를 향하고 있습니다.",
        "built_space": "어두운 도서관 환경. 배경에 먼지 쌓인 책장이 흐릿하게 보이고, 우측에 낡은 모니터의 측면이 배치되어 있습니다.",
        "entities": "찰리는 낡은 금속 재질의 베이지색 장갑판과 흰색 마스크형 얼굴을 지닌 로봇으로 레퍼런스와 일치합니다.",
        "hard_violations": [],
        "physics": "찰리의 상체는 바닥에 안정적으로 위치해 있으며, 구조적으로 불가능한 포즈나 떠 있는 요소는 없습니다."
       },
       {
        "label": "B",
        "direction": "찰리의 시선이 우측 프레임 바깥의 모니터 화면을 향해 집중되어 있습니다.",
        "built_space": "배경으로 낡은 책과 책장이 아웃포커싱되어 있으며, 우측 전경에 모니터 측면이 자리 잡고 있습니다.",
        "entities": "찰리의 외형(베이지색 장갑판, 흰색 마스크, 눈의 형태)이 레퍼런스에 매우 부합하며 입을 크게 벌리고 있습니다.",
        "hard_violations": [],
        "physics": "상체의 무게감과 기계적 관절 형태가 자연스럽고 지지 기반에 문제가 없습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "B보다 입을 크게 벌려 발음 연습 순간에 가깝지만, 얼굴보다 상체를 넓게 담고 모니터 외장을 과다 노출해 지정된 클로즈업에는 미달한다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "주황색 눈과 낡은 장갑판은 참조에 충실하지만, 입 벌림이 작고 모니터 뒷면의 큰 노출이 얼굴 중심 클로즈업 지시에서 더 멀어진다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 얼굴과 두 눈은 화면 오른쪽의 모니터를 향한다. 카메라에는 모니터 외장만 보이며 표시 면은 찰리 쪽에 숨겨져 있어 사용 방향은 맞는다.",
        "built_space": "오른쪽 전경에 구형 모니터 한 대가 있고, 머리 뒤와 왼쪽에 책이 꽂힌 선반 및 낡은 수직 측판이 보인다. 도서관의 재질은 참조와 부합하지만, 배경이 좁은 파편 이상으로 드러나고 모니터도 가느다란 측면이 아니라 넓은 외장 부분을 차지한다. 머리뿐 아니라 양어깨와 가슴까지 크게 포함한다.",
        "entities": "찰리 한 개체만 보인다. 흰색 각진 마스크, 베이지색 마모 장갑판, 원형 귀 부품과 안테나는 참조와 대체로 일치한다. 눈은 참조의 밝은 주황색보다 어두운 적갈색이다. 입은 크게 벌어져 있으나 둥근 O형보다는 각진 세로 개구부에 가깝다. 구형 모니터와 흐린 책들은 보이며 읽을 수 있는 글자는 없다. 침구와 하체는 이 구도 밖이므로 누락으로 보지 않는다.",
        "hard_violations": [],
        "physics": "머리는 노출된 기계식 목과 몸통에 연결되고, 벌어진 아래턱도 턱 주변 기구에 연결되어 있다. 떠 있는 신체 부품은 보이지 않는다. 모니터 하단과 몸의 하부 지지점은 프레임 밖이라 확인할 수 있지만, 공중에 떠 있다는 징후는 없다."
       },
       {
        "label": "B",
        "direction": "찰리는 오른쪽 앞의 모니터를 바라본다. 모니터의 통풍구가 있는 뒷면은 카메라 쪽, 보이지 않는 표시 면은 찰리 쪽을 향해 있어 시청 방향은 타당하다.",
        "built_space": "오른쪽 전경에 구형 모니터 한 대가 있으며 왼쪽과 머리 뒤에는 여러 단의 책장과 기울어진 책들이 흐리게 보인다. 낡고 어두운 도서관이라는 장소는 유지하지만 모니터 뒷면이 화면 오른쪽의 큰 영역을 차지한다. 양어깨와 가슴을 포함해 얼굴 중심 클로즈업보다 넓다.",
        "entities": "찰리만 등장하며 흰 마스크, 주황색 발광 눈, 베이지색 장갑판, 안테나와 목의 기계 부품은 참조의 정체성을 유지한다. 입은 열려 있지만 크고 둥글게 벌린 발음 연습 표정으로 보기에는 개구부가 작고 각지다. 얼굴에는 모니터 쪽의 옅은 차가운 빛이 보인다. 책과 모니터에 읽을 수 있는 문구는 없으며 프레임 밖의 침구나 하체는 평가하지 않는다.",
        "hard_violations": [],
        "physics": "머리와 아래턱은 목 및 턱의 기계 구조에 연결되어 있고 어깨도 몸통에 붙어 있다. 모니터 받침과 찰리의 하체는 잘렸으므로 실제 바닥 접점은 확인되지 않지만, 지지 없이 떠 있는 물체나 불가능한 동작은 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "B보다 입을 크게 벌려 발음 연습 순간에 가깝지만, 얼굴보다 상체를 넓게 담고 모니터 외장을 과다 노출해 지정된 클로즈업에는 미달한다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "주황색 눈과 낡은 장갑판은 참조에 충실하지만, 입 벌림이 작고 모니터 뒷면의 큰 노출이 얼굴 중심 클로즈업 지시에서 더 멀어진다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 얼굴과 두 눈은 화면 오른쪽의 모니터를 향한다. 카메라에는 모니터 외장만 보이며 표시 면은 찰리 쪽에 숨겨져 있어 사용 방향은 맞는다.",
        "built_space": "오른쪽 전경에 구형 모니터 한 대가 있고, 머리 뒤와 왼쪽에 책이 꽂힌 선반 및 낡은 수직 측판이 보인다. 도서관의 재질은 참조와 부합하지만, 배경이 좁은 파편 이상으로 드러나고 모니터도 가느다란 측면이 아니라 넓은 외장 부분을 차지한다. 머리뿐 아니라 양어깨와 가슴까지 크게 포함한다.",
        "entities": "찰리 한 개체만 보인다. 흰색 각진 마스크, 베이지색 마모 장갑판, 원형 귀 부품과 안테나는 참조와 대체로 일치한다. 눈은 참조의 밝은 주황색보다 어두운 적갈색이다. 입은 크게 벌어져 있으나 둥근 O형보다는 각진 세로 개구부에 가깝다. 구형 모니터와 흐린 책들은 보이며 읽을 수 있는 글자는 없다. 침구와 하체는 이 구도 밖이므로 누락으로 보지 않는다.",
        "hard_violations": [],
        "physics": "머리는 노출된 기계식 목과 몸통에 연결되고, 벌어진 아래턱도 턱 주변 기구에 연결되어 있다. 떠 있는 신체 부품은 보이지 않는다. 모니터 하단과 몸의 하부 지지점은 프레임 밖이라 확인할 수 있지만, 공중에 떠 있다는 징후는 없다."
       },
       {
        "label": "A",
        "direction": "찰리는 오른쪽 앞의 모니터를 바라본다. 모니터의 통풍구가 있는 뒷면은 카메라 쪽, 보이지 않는 표시 면은 찰리 쪽을 향해 있어 시청 방향은 타당하다.",
        "built_space": "오른쪽 전경에 구형 모니터 한 대가 있으며 왼쪽과 머리 뒤에는 여러 단의 책장과 기울어진 책들이 흐리게 보인다. 낡고 어두운 도서관이라는 장소는 유지하지만 모니터 뒷면이 화면 오른쪽의 큰 영역을 차지한다. 양어깨와 가슴을 포함해 얼굴 중심 클로즈업보다 넓다.",
        "entities": "찰리만 등장하며 흰 마스크, 주황색 발광 눈, 베이지색 장갑판, 안테나와 목의 기계 부품은 참조의 정체성을 유지한다. 입은 열려 있지만 크고 둥글게 벌린 발음 연습 표정으로 보기에는 개구부가 작고 각지다. 얼굴에는 모니터 쪽의 옅은 차가운 빛이 보인다. 책과 모니터에 읽을 수 있는 문구는 없으며 프레임 밖의 침구나 하체는 평가하지 않는다.",
        "hard_violations": [],
        "physics": "머리와 아래턱은 목 및 턱의 기계 구조에 연결되어 있고 어깨도 몸통에 붙어 있다. 모니터 받침과 찰리의 하체는 잘렸으므로 실제 바닥 접점은 확인되지 않지만, 지지 없이 떠 있는 물체나 불가능한 동작은 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.69,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.69,
    "B": 2.0
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1690
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "프롬프트가 요구한 '입을 크고 둥글게 벌린' 표정을 더 정확하고 생동감 있게 묘사했으며, 조명과 프레이밍 모두 훌륭합니다."
   },
   {
    "label": "A",
    "score": 1690,
    "verdict_ko": "지시된 클로즈업 앵글과 배경 요소는 잘 구현되었으나, 입을 크게 벌린 형태가 경쟁작에 비해 덜 둥글고 소극적입니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S46sh17_sel.png",
    "asset_id": "b8987440-20be-4100-9d42-63e1f1850baa",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-1d4e-7336-ac59-cff1e86dcf7b",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S46sh17"
  }
 },
 "S46sh28::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:36:14.989854+00:00",
  "fingerprint": "48f369d31181b80ad7d900289ad3e76b37e12f741715c0ab4af7f1fd51cb423f",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S46sh28_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S46sh28_sel.png",
  "source_sha256": "2b0875bc5a2097b318793644baa528f46bd03c2d83552cd31d272ecfced23ad4",
  "file": "S46sh28_cine.png",
  "staged_sha256": "7b46cffcb114b73b8f6209a24b08a51195d9b1bfb69e37c084f0f0dad1040624",
  "latency_ms": 9748
 },
 "S47sh3::signage": {
  "fp": "0fb0867d39ad412d",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S47sh3": {
  "input_fingerprint": "ed40795200ad9be6",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): early morning.\n\nSHOT TEXT (authoritative, Korean): 화려한 바다와 인물들이 그려진 거대한 벽화 앞에서 뭉툭한 크레파스를 쥔 손을 벽면에 댄 찰리의 낡은 뒷모습.\n\nLOCATION (lock): Along the mural-covered inner wall of the abandoned library's reading room, in early-morning ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Crayon mural (Extensive and still being drawn) — The illustrated wall face is seen obliquely, showing the blue sea, marine life, palms, and drawn figures of 찰리, 현우, 라울, 앰버, 페드로, their mother, and the priest around the imagined Haenam arrival; used as Colorful narrative backdrop separating the drawing hand from 찰리's silhouette; the depicted people remain drawings, not additional physical figures; Blunt crayon (Held against the wall in 찰리's hand) — Tip meets the mural beyond the outline of his torso; used as Small, clearly separated action detail establishing authorship.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained early-morning ambient light, allowing the mural's blue sea and other explicitly rich colors to provide the scene's selective chromatic warmth and tenderness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): In the early-morning library, a crayon mural already covers a wall: a blue sea leads to a colorful Haenam paradise with marine life, palms and a camper carrying Charlie, Hyunwoo, Raul and Amber, with Pedro, Miyeon and the priest also depicted. The written message about Haenam, Soyoung and Giant Charlie accompanies the pictures. 찰리: He stands at the mural holding a crayon and continuing the drawing. His worn metal body and retained disguise remain unchanged.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): early morning.\n\nSHOT TEXT (authoritative, Korean): 화려한 바다와 인물들이 그려진 거대한 벽화 앞에서 뭉툭한 크레파스를 쥔 손을 벽면에 댄 찰리의 낡은 뒷모습.\n\nLOCATION (lock): Along the mural-covered inner wall of the abandoned library's reading room, in early-morning ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Crayon mural (Extensive and still being drawn) — The illustrated wall face is seen obliquely, showing the blue sea, marine life, palms, and drawn figures of 찰리, 현우, 라울, 앰버, 페드로, their mother, and the priest around the imagined Haenam arrival; used as Colorful narrative backdrop separating the drawing hand from 찰리's silhouette; the depicted people remain drawings, not additional physical figures; Blunt crayon (Held against the wall in 찰리's hand) — Tip meets the mural beyond the outline of his torso; used as Small, clearly separated action detail establishing authorship.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained early-morning ambient light, allowing the mural's blue sea and other explicitly rich colors to provide the scene's selective chromatic warmth and tenderness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): In the early-morning library, a crayon mural already covers a wall: a blue sea leads to a colorful Haenam paradise with marine life, palms and a camper carrying Charlie, Hyunwoo, Raul and Amber, with Pedro, Miyeon and the priest also depicted. The written message about Haenam, Soyoung and Giant Charlie accompanies the pictures. 찰리: He stands at the mural holding a crayon and continuing the drawing. His worn metal body and retained disguise remain unchanged.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): early morning.\n\nSHOT TEXT (authoritative, Korean): 화려한 바다와 인물들이 그려진 거대한 벽화 앞에서 뭉툭한 크레파스를 쥔 손을 벽면에 댄 찰리의 낡은 뒷모습.\n\nLOCATION (lock): Along the mural-covered inner wall of the abandoned library's reading room, in early-morning ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Crayon mural (Extensive and still being drawn) — The illustrated wall face is seen obliquely, showing the blue sea, marine life, palms, and drawn figures of 찰리, 현우, 라울, 앰버, 페드로, their mother, and the priest around the imagined Haenam arrival; used as Colorful narrative backdrop separating the drawing hand from 찰리's silhouette; the depicted people remain drawings, not additional physical figures; Blunt crayon (Held against the wall in 찰리's hand) — Tip meets the mural beyond the outline of his torso; used as Small, clearly separated action detail establishing authorship.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained early-morning ambient light, allowing the mural's blue sea and other explicitly rich colors to provide the scene's selective chromatic warmth and tenderness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): In the early-morning library, a crayon mural already covers a wall: a blue sea leads to a colorful Haenam paradise with marine life, palms and a camper carrying Charlie, Hyunwoo, Raul and Amber, with Pedro, Miyeon and the priest also depicted. The written message about Haenam, Soyoung and Giant Charlie accompanies the pictures. 찰리: He stands at the mural holding a crayon and continuing the drawing. His worn metal body and retained disguise remain unchanged.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "찰리가 벽화 쪽을 향해 서서 크레파스를 쥐고 그려진 소년의 얼굴 쪽에 대고 있습니다.",
    "built_space": "우측에 창문과 빈 책장이 있는 실내 공간이며, 정면 벽에 거대한 종이가 붙어 있는 형태입니다.",
    "entities": "찰리의 외형은 레퍼런스와 일치하나 뒷모습(뒷모습)이 아닌 측면에 가깝습니다. 벽화에는 인물들과 바다가 있지만 한글 텍스트가 선명하게 적혀 있습니다.",
    "hard_violations": [
     "[gemini-pro] 절대 금지된 '읽을 수 있는 텍스트(해남, 소영과... 및 인물들의 이름)'가 프레임 내에 크고 선명하게 등장합니다.",
     "[gemini-pro] 벽화가 벽면에 직접 그려진 것이 아니라 거대한 종이 스티커처럼 벽 위에 덧붙여져 있어 물질적 사실성(Material Realism) 지시를 위반했습니다.",
     "[gpt-high] 벽화에 큰 한글 문구와 여러 인물 이름이 선명하게 노출되어, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 명시적 금지 조건을 위반합니다."
    ],
    "physics": "찰리가 바닥에 서서 오른팔을 들어 크레파스를 쥐고 종이 표면에 대고 있습니다."
   },
   {
    "label": "B",
    "direction": "찰리가 벽을 향해 서서 오른손에 쥔 파란색 크레파스를 벽화의 파란 선에 가져다 대고 있습니다.",
    "built_space": "폐도서관의 열람실로, 좌측에는 낡은 책장들이 있고 우측에는 거대한 벽화가 그려진 긴 벽면이 있습니다. 상단에는 창문들이 줄지어 있습니다.",
    "entities": "찰리는 레퍼런스의 외형(고릴라형 로봇, 베이지색 장갑, 하얀 마스크)과 일치하며 뒷모습으로 나타납니다. 벽화에는 바다, 야자수, 캠핑카, 사람들이 그려져 있습니다.",
    "hard_violations": [],
    "physics": "찰리는 두 발로 바닥을 딛고 안정적으로 서 있으며, 오른손은 크레파스를 쥐고 벽에 지탱하고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "프롬프트가 요구한 '찰리의 뒷모습'을 정확히 연출했으며, 읽을 수 있는 글자를 배제하라는 지시사항과 벽면에 직접 그려진 벽화의 질감을 훌륭하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "찰리의 뒷모습이 아닌 측면/정면에 가까운 앵글이며, 절대 금지된 '읽을 수 있는 글자'가 대거 포함되었고 벽화가 벽이 아닌 종이에 그려져 붙은 것처럼 보입니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "찰리가 벽을 향해 서서 오른손에 쥔 파란색 크레파스를 벽화의 파란 선에 가져다 대고 있습니다.",
        "built_space": "폐도서관의 열람실로, 좌측에는 낡은 책장들이 있고 우측에는 거대한 벽화가 그려진 긴 벽면이 있습니다. 상단에는 창문들이 줄지어 있습니다.",
        "entities": "찰리는 레퍼런스의 외형(고릴라형 로봇, 베이지색 장갑, 하얀 마스크)과 일치하며 뒷모습으로 나타납니다. 벽화에는 바다, 야자수, 캠핑카, 사람들이 그려져 있습니다.",
        "hard_violations": [],
        "physics": "찰리는 두 발로 바닥을 딛고 안정적으로 서 있으며, 오른손은 크레파스를 쥐고 벽에 지탱하고 있습니다."
       },
       {
        "label": "A",
        "direction": "찰리가 벽화 쪽을 향해 서서 크레파스를 쥐고 그려진 소년의 얼굴 쪽에 대고 있습니다.",
        "built_space": "우측에 창문과 빈 책장이 있는 실내 공간이며, 정면 벽에 거대한 종이가 붙어 있는 형태입니다.",
        "entities": "찰리의 외형은 레퍼런스와 일치하나 뒷모습(뒷모습)이 아닌 측면에 가깝습니다. 벽화에는 인물들과 바다가 있지만 한글 텍스트가 선명하게 적혀 있습니다.",
        "hard_violations": [
         "절대 금지된 '읽을 수 있는 텍스트(해남, 소영과... 및 인물들의 이름)'가 프레임 내에 크고 선명하게 등장합니다.",
         "벽화가 벽면에 직접 그려진 것이 아니라 거대한 종이 스티커처럼 벽 위에 덧붙여져 있어 물질적 사실성(Material Realism) 지시를 위반했습니다."
        ],
        "physics": "찰리가 바닥에 서서 오른팔을 들어 크레파스를 쥐고 종이 표면에 대고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "프롬프트가 요구한 '찰리의 뒷모습'을 정확히 연출했으며, 읽을 수 있는 글자를 배제하라는 지시사항과 벽면에 직접 그려진 벽화의 질감을 훌륭하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "찰리의 뒷모습이 아닌 측면/정면에 가까운 앵글이며, 절대 금지된 '읽을 수 있는 글자'가 대거 포함되었고 벽화가 벽이 아닌 종이에 그려져 붙은 것처럼 보입니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리가 벽을 향해 서서 오른손에 쥔 파란색 크레파스를 벽화의 파란 선에 가져다 대고 있습니다.",
        "built_space": "폐도서관의 열람실로, 좌측에는 낡은 책장들이 있고 우측에는 거대한 벽화가 그려진 긴 벽면이 있습니다. 상단에는 창문들이 줄지어 있습니다.",
        "entities": "찰리는 레퍼런스의 외형(고릴라형 로봇, 베이지색 장갑, 하얀 마스크)과 일치하며 뒷모습으로 나타납니다. 벽화에는 바다, 야자수, 캠핑카, 사람들이 그려져 있습니다.",
        "hard_violations": [],
        "physics": "찰리는 두 발로 바닥을 딛고 안정적으로 서 있으며, 오른손은 크레파스를 쥐고 벽에 지탱하고 있습니다."
       },
       {
        "label": "A",
        "direction": "찰리가 벽화 쪽을 향해 서서 크레파스를 쥐고 그려진 소년의 얼굴 쪽에 대고 있습니다.",
        "built_space": "우측에 창문과 빈 책장이 있는 실내 공간이며, 정면 벽에 거대한 종이가 붙어 있는 형태입니다.",
        "entities": "찰리의 외형은 레퍼런스와 일치하나 뒷모습(뒷모습)이 아닌 측면에 가깝습니다. 벽화에는 인물들과 바다가 있지만 한글 텍스트가 선명하게 적혀 있습니다.",
        "hard_violations": [
         "절대 금지된 '읽을 수 있는 텍스트(해남, 소영과... 및 인물들의 이름)'가 프레임 내에 크고 선명하게 등장합니다.",
         "벽화가 벽면에 직접 그려진 것이 아니라 거대한 종이 스티커처럼 벽 위에 덧붙여져 있어 물질적 사실성(Material Realism) 지시를 위반했습니다."
        ],
        "physics": "찰리가 바닥에 서서 오른팔을 들어 크레파스를 쥐고 종이 표면에 대고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "비스듬한 대형 벽화 앞의 낡은 뒷모습과 몸통 밖으로 분리된 크레파스 접촉 동작을 잘 구현했지만, 인물이 크게 잡혀 와이드 숏의 공간감은 다소 부족합니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "벽에 그리는 동작과 금속 몸체는 맞지만, 벽화의 큰 한글 문구와 이름들이 선명하게 읽혀 읽을 수 있는 글자 금지 조건을 결정적으로 위반합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 머리와 몸은 오른쪽 벽화를 향하며 얼굴은 대부분 가려져 있습니다. 들어 올린 오른손의 파란 크레파스 끝은 몸통 윤곽 밖에서 벽화의 파란 선에 닿아 있어 그리는 대상과 도구 방향이 일치합니다.",
        "built_space": "오른쪽의 긴 벽화 벽이 왼쪽 뒤로 멀어지고, 그 위로 높은 창들이 이어집니다. 왼쪽에는 책장 두 열, 위에는 길쭉한 천장 조명들이 보입니다. 찰리는 벽 바로 앞 통로를 차지하며 팔이 벽까지 닿는 배치입니다. 낡은 벽과 책장은 이전 사진의 도서관 재질과 대체로 이어지지만, 참조에서 보이지 않았던 창과 방 전체 구조의 정확한 일치는 확인할 수 없습니다.",
        "entities": "실제 입체 인물은 찰리 하나이며, 샌드 베이지 장갑판, 마모된 금속, 긴 기계 팔과 머리 안테나가 참조에 부합합니다. 흰 얼굴은 옆 가장자리만 보입니다. 벽에는 푸른 바다, 돌고래와 물고기, 산호, 야자수, 캠핑카와 여러 사람이 그림으로 표현되어 있습니다. 그림 속 인물 각각을 지정된 이름과 대응시키기는 어렵고, 캠핑카 탑승 관계도 뚜렷하지 않습니다. 손에는 굵고 짧은 파란 크레파스가 있으며 읽을 수 있는 글자는 보이지 않습니다.",
        "hard_violations": [],
        "physics": "하체는 화면 아래로 잘려 발의 바닥 접촉은 확인되지 않지만, 몸통과 골반은 정상적인 직립 자세로 연결되어 있습니다. 오른팔은 어깨와 굽힌 팔꿈치로 지지되고 기계 손가락이 크레파스를 쥐고 있습니다. 도구 끝의 벽 접촉과 관절 자세는 벽에 그림을 그리는 동작으로 성립하며, 지지 없이 떠 있는 물체는 보이지 않습니다."
       },
       {
        "label": "B",
        "direction": "찰리는 화면 왼쪽 벽화를 향하고 흰 얼굴의 옆면이 보입니다. 앞으로 뻗은 손의 황토색 크레파스는 벽화 속 아이 얼굴 옆을 향하며 끝이 그림 표면에 닿아 있습니다. 손과 도구는 몸통 바깥에서 구분됩니다.",
        "built_space": "벽화 위로 높은 창들이 이어지고, 오른쪽 뒤에는 낮은 창 하나와 책장 한 열, 오른쪽 위에는 천장 조명 하나가 보입니다. 찰리는 벽 앞에 서서 팔을 뻗고 있으며 공간상 접촉은 가능합니다. 다만 벽화가 벽 전체에 이어진 그림보다는 경계가 뚜렷한 밝은 직사각형 바탕처럼 보입니다. 참조의 낡은 도서관 분위기는 유지하지만 기존 사진만으로 이 창 배치를 검증할 수는 없습니다.",
        "entities": "입체 인물은 찰리 하나이며 마모된 베이지 장갑, 긴 기계 팔, 등쪽 통풍구와 안테나는 참조와 대체로 맞습니다. 벽화에는 바다, 야자수, 해양 생물, 사제 복장의 인물과 아이들 및 성인 여성이 그려져 있습니다. 보이는 벽화에는 캠핑카가 없고 지정된 일행 전체도 명확히 구분되지 않습니다. 크레파스는 손에 들려 있습니다. 벽화 상단의 큰 한글 문구와 하단의 인물 이름은 명백히 읽을 수 있습니다.",
        "hard_violations": [
         "벽화에 큰 한글 문구와 여러 인물 이름이 선명하게 노출되어, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 명시적 금지 조건을 위반합니다."
        ],
        "physics": "발은 화면 밖이지만 골반과 다리는 직립한 몸을 자연스럽게 이어 받습니다. 앞으로 뻗은 팔은 어깨와 팔꿈치 관절에 연결되어 있고, 손가락이 크레파스를 잡아 벽 쪽으로 지지합니다. 벽에 그리는 자세는 물리적으로 가능하며 공중에 무지지 상태로 떠 있는 몸이나 도구는 보이지 않습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "비스듬한 대형 벽화 앞의 낡은 뒷모습과 몸통 밖으로 분리된 크레파스 접촉 동작을 잘 구현했지만, 인물이 크게 잡혀 와이드 숏의 공간감은 다소 부족합니다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "벽에 그리는 동작과 금속 몸체는 맞지만, 벽화의 큰 한글 문구와 이름들이 선명하게 읽혀 읽을 수 있는 글자 금지 조건을 결정적으로 위반합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 머리와 몸은 오른쪽 벽화를 향하며 얼굴은 대부분 가려져 있습니다. 들어 올린 오른손의 파란 크레파스 끝은 몸통 윤곽 밖에서 벽화의 파란 선에 닿아 있어 그리는 대상과 도구 방향이 일치합니다.",
        "built_space": "오른쪽의 긴 벽화 벽이 왼쪽 뒤로 멀어지고, 그 위로 높은 창들이 이어집니다. 왼쪽에는 책장 두 열, 위에는 길쭉한 천장 조명들이 보입니다. 찰리는 벽 바로 앞 통로를 차지하며 팔이 벽까지 닿는 배치입니다. 낡은 벽과 책장은 이전 사진의 도서관 재질과 대체로 이어지지만, 참조에서 보이지 않았던 창과 방 전체 구조의 정확한 일치는 확인할 수 없습니다.",
        "entities": "실제 입체 인물은 찰리 하나이며, 샌드 베이지 장갑판, 마모된 금속, 긴 기계 팔과 머리 안테나가 참조에 부합합니다. 흰 얼굴은 옆 가장자리만 보입니다. 벽에는 푸른 바다, 돌고래와 물고기, 산호, 야자수, 캠핑카와 여러 사람이 그림으로 표현되어 있습니다. 그림 속 인물 각각을 지정된 이름과 대응시키기는 어렵고, 캠핑카 탑승 관계도 뚜렷하지 않습니다. 손에는 굵고 짧은 파란 크레파스가 있으며 읽을 수 있는 글자는 보이지 않습니다.",
        "hard_violations": [],
        "physics": "하체는 화면 아래로 잘려 발의 바닥 접촉은 확인되지 않지만, 몸통과 골반은 정상적인 직립 자세로 연결되어 있습니다. 오른팔은 어깨와 굽힌 팔꿈치로 지지되고 기계 손가락이 크레파스를 쥐고 있습니다. 도구 끝의 벽 접촉과 관절 자세는 벽에 그림을 그리는 동작으로 성립하며, 지지 없이 떠 있는 물체는 보이지 않습니다."
       },
       {
        "label": "A",
        "direction": "찰리는 화면 왼쪽 벽화를 향하고 흰 얼굴의 옆면이 보입니다. 앞으로 뻗은 손의 황토색 크레파스는 벽화 속 아이 얼굴 옆을 향하며 끝이 그림 표면에 닿아 있습니다. 손과 도구는 몸통 바깥에서 구분됩니다.",
        "built_space": "벽화 위로 높은 창들이 이어지고, 오른쪽 뒤에는 낮은 창 하나와 책장 한 열, 오른쪽 위에는 천장 조명 하나가 보입니다. 찰리는 벽 앞에 서서 팔을 뻗고 있으며 공간상 접촉은 가능합니다. 다만 벽화가 벽 전체에 이어진 그림보다는 경계가 뚜렷한 밝은 직사각형 바탕처럼 보입니다. 참조의 낡은 도서관 분위기는 유지하지만 기존 사진만으로 이 창 배치를 검증할 수는 없습니다.",
        "entities": "입체 인물은 찰리 하나이며 마모된 베이지 장갑, 긴 기계 팔, 등쪽 통풍구와 안테나는 참조와 대체로 맞습니다. 벽화에는 바다, 야자수, 해양 생물, 사제 복장의 인물과 아이들 및 성인 여성이 그려져 있습니다. 보이는 벽화에는 캠핑카가 없고 지정된 일행 전체도 명확히 구분되지 않습니다. 크레파스는 손에 들려 있습니다. 벽화 상단의 큰 한글 문구와 하단의 인물 이름은 명백히 읽을 수 있습니다.",
        "hard_violations": [
         "벽화에 큰 한글 문구와 여러 인물 이름이 선명하게 노출되어, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 명시적 금지 조건을 위반합니다."
        ],
        "physics": "발은 화면 밖이지만 골반과 다리는 직립한 몸을 자연스럽게 이어 받습니다. 앞으로 뻗은 팔은 어깨와 팔꿈치 관절에 연결되어 있고, 손가락이 크레파스를 잡아 벽 쪽으로 지지합니다. 벽에 그리는 자세는 물리적으로 가능하며 공중에 무지지 상태로 떠 있는 몸이나 도구는 보이지 않습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.583,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.333,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 절대 금지된 '읽을 수 있는 텍스트(해남, 소영과... 및 인물들의 이름)'가 프레임 내에 크고 선명하게 등장합니다.",
     "[gemini-pro] 벽화가 벽면에 직접 그려진 것이 아니라 거대한 종이 스티커처럼 벽 위에 덧붙여져 있어 물질적 사실성(Material Realism) 지시를 위반했습니다.",
     "[gpt-high] 벽화에 큰 한글 문구와 여러 인물 이름이 선명하게 노출되어, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 명시적 금지 조건을 위반합니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 333
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "프롬프트가 요구한 '찰리의 뒷모습'을 정확히 연출했으며, 읽을 수 있는 글자를 배제하라는 지시사항과 벽면에 직접 그려진 벽화의 질감을 훌륭하게 구현했습니다."
   },
   {
    "label": "A",
    "score": 333,
    "verdict_ko": "찰리의 뒷모습이 아닌 측면/정면에 가까운 앵글이며, 절대 금지된 '읽을 수 있는 글자'가 대거 포함되었고 벽화가 벽이 아닌 종이에 그려져 붙은 것처럼 보입니다.  ★위반: [gemini-pro] 절대 금지된 '읽을 수 있는 텍스트(해남, 소영과... 및 인물들의 이름)'가 프레임 내에 크고 선명하게 등장합니다. / [gemini-pro] 벽화가 벽면에 직접 그려진 것이 아니라 거대한 종이 스티커처럼 벽 위에 덧붙여져 있어 물질적 사실성(Material Realism) 지시를 위반했습니다. / [gpt-high] 벽화에 큰 한글 문구와 여러 인물 이름이 선명하게 노출되어, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 명시적 금지 조건을 위반합니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S46sh28_sel.png",
    "asset_id": "5af04930-0130-4430-9e55-10dd66608c2f",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-1f00-7fdc-a398-cbe7954c8163",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S46sh28"
  }
 },
 "S47sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:37:09.963622+00:00",
  "fingerprint": "9db2d9853185553e7f41db6c2875aafcb69c67405b864213cc7833be531571f6",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S47sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S47sh3_sel.png",
  "source_sha256": "c24397437bdf14a53e9d961ba2aa5a1c010d61aebcf12dc8aea10efd6f0863ca",
  "file": "S47sh3_cine.png",
  "staged_sha256": "aee2d577d4e88d92f819de119789785148d2ad20417bef897a5a69d8c18d06c9",
  "latency_ms": 10238
 },
 "S47sh9::signage": {
  "fp": "5c3081ff9a2933bd",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S47sh9::bgfirst_bg": {
  "input_fingerprint": "fa83a79386c6066b",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 도서관 외진 구석 바닥에 쭈그려 앉아 벗어놓은 신발 안에서 빳빳한 카드를 반쯤 뽑아 올린 자세로 멈춘 현우의 거친 손 클로즈업.\n\nLOCATION (lock): In a secluded outdoor corner beside the abandoned library, away from the reading-room group in the morning.\n\nTIME OF DAY (lock): early morning.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Removed shoe (Off 현우's foot, with the card being extracted from inside) — Opening faces diagonally upward toward his hands and the camera; used as Reveals the hiding place while remaining smaller than the hand-and-knee context; Contact card (Stiff, intact after the flooding, and halfway withdrawn) — A narrow portion of the contact-bearing face is visible at an oblique angle, without resolving invented contact details; used as Small focal evidence of his concealed means of contact; Secluded library corner floor (Beneath 현우's crouched body and removed shoe); used as Ground plane anchoring the intimate hand detail.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral early-morning ambient illumination with controlled detail in the hands and card, without introducing a special light source into the secluded corner.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 도서관 외진 구석 바닥에 쭈그려 앉아 벗어놓은 신발 안에서 빳빳한 카드를 반쯤 뽑아 올린 자세로 멈춘 현우의 거친 손 클로즈업.\n\nLOCATION (lock): In a secluded outdoor corner beside the abandoned library, away from the reading-room group in the morning.\n\nTIME OF DAY (lock): early morning.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Removed shoe (Off 현우's foot, with the card being extracted from inside) — Opening faces diagonally upward toward his hands and the camera; used as Reveals the hiding place while remaining smaller than the hand-and-knee context; Contact card (Stiff, intact after the flooding, and halfway withdrawn) — A narrow portion of the contact-bearing face is visible at an oblique angle, without resolving invented contact details; used as Small focal evidence of his concealed means of contact; Secluded library corner floor (Beneath 현우's crouched body and removed shoe); used as Ground plane anchoring the intimate hand detail.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral early-morning ambient illumination with controlled detail in the hands and card, without introducing a special light source into the secluded corner.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S47sh9__bgfirst_bg.png",
  "asset_id": "85df77f5-d63d-4165-9bbc-cafbc3fedb4e",
  "input_asset_ids": [
   "198f725a-7e21-4a39-9d9c-8249fc9fea2c",
   "f4370321-fb89-4785-899e-192be57a9cc4"
  ]
 },
 "S47sh9": {
  "input_fingerprint": "93a96bd90df08e7c",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): early morning.\n\nSHOT TEXT (authoritative, Korean): 도서관 외진 구석 바닥에 쭈그려 앉아 벗어놓은 신발 안에서 빳빳한 카드를 반쯤 뽑아 올린 자세로 멈춘 현우의 거친 손 클로즈업.\n\nLOCATION (lock): In a secluded outdoor corner beside the abandoned library, away from the reading-room group in the morning. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Removed shoe (Off 현우's foot, with the card being extracted from inside) — Opening faces diagonally upward toward his hands and the camera; used as Reveals the hiding place while remaining smaller than the hand-and-knee context; Contact card (Stiff, intact after the flooding, and halfway withdrawn) — A narrow portion of the contact-bearing face is visible at an oblique angle, without resolving invented contact details; used as Small focal evidence of his concealed means of contact; Secluded library corner floor (Beneath 현우's crouched body and removed shoe); used as Ground plane anchoring the intimate hand detail.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral early-morning ambient illumination with controlled detail in the hands and card, without introducing a special light source into the secluded corner.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The completed portions of the colorful Haenam mural and its written message remain on the library wall. The camper remains outside the abandoned library. 현우: He has removed one shoe and is extracting the intact contact card concealed inside it, with the disposable phone ready. His facial bruises and untreated leg wound remain.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): early morning.\n\nSHOT TEXT (authoritative, Korean): 도서관 외진 구석 바닥에 쭈그려 앉아 벗어놓은 신발 안에서 빳빳한 카드를 반쯤 뽑아 올린 자세로 멈춘 현우의 거친 손 클로즈업.\n\nLOCATION (lock): In a secluded outdoor corner beside the abandoned library, away from the reading-room group in the morning. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Removed shoe (Off 현우's foot, with the card being extracted from inside) — Opening faces diagonally upward toward his hands and the camera; used as Reveals the hiding place while remaining smaller than the hand-and-knee context; Contact card (Stiff, intact after the flooding, and halfway withdrawn) — A narrow portion of the contact-bearing face is visible at an oblique angle, without resolving invented contact details; used as Small focal evidence of his concealed means of contact; Secluded library corner floor (Beneath 현우's crouched body and removed shoe); used as Ground plane anchoring the intimate hand detail.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral early-morning ambient illumination with controlled detail in the hands and card, without introducing a special light source into the secluded corner.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The completed portions of the colorful Haenam mural and its written message remain on the library wall. The camper remains outside the abandoned library. 현우: He has removed one shoe and is extracting the intact contact card concealed inside it, with the disposable phone ready. His facial bruises and untreated leg wound remain.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): early morning.\n\nSHOT TEXT (authoritative, Korean): 도서관 외진 구석 바닥에 쭈그려 앉아 벗어놓은 신발 안에서 빳빳한 카드를 반쯤 뽑아 올린 자세로 멈춘 현우의 거친 손 클로즈업.\n\nLOCATION (lock): In a secluded outdoor corner beside the abandoned library, away from the reading-room group in the morning. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Removed shoe (Off 현우's foot, with the card being extracted from inside) — Opening faces diagonally upward toward his hands and the camera; used as Reveals the hiding place while remaining smaller than the hand-and-knee context; Contact card (Stiff, intact after the flooding, and halfway withdrawn) — A narrow portion of the contact-bearing face is visible at an oblique angle, without resolving invented contact details; used as Small focal evidence of his concealed means of contact; Secluded library corner floor (Beneath 현우's crouched body and removed shoe); used as Ground plane anchoring the intimate hand detail.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral early-morning ambient illumination with controlled detail in the hands and card, without introducing a special light source into the secluded corner.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The completed portions of the colorful Haenam mural and its written message remain on the library wall. The camper remains outside the abandoned library. 현우: He has removed one shoe and is extracting the intact contact card concealed inside it, with the disposable phone ready. His facial bruises and untreated leg wound remain.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S47sh9__bgfirst_bg.png",
     "asset_id": "85df77f5-d63d-4165-9bbc-cafbc3fedb4e",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S47sh9.png",
     "asset_id": "198f725a-7e21-4a39-9d9c-8249fc9fea2c",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L202B02.png",
     "asset_id": "f4370321-fb89-4785-899e-192be57a9cc4",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "시선은 바닥에 놓인 신발과 카드를 향하고 있으며, 두 손은 정확히 신발 입구와 카드를 향해 작용하고 있습니다.",
    "built_space": "레퍼런스 이미지의 야외 구석 공간(벽면, 배관, 우측 에어컨 실외기, 젖은 바닥)이 올바른 위치와 비례로 배치되어 있습니다.",
    "entities": "현우(얼굴의 상처와 앳된 외모), 거친 손, 바닥에 벗어놓은 신발, 반쯤 뽑혀 나온 빳빳한 카드 모두 텍스트와 일치합니다.",
    "hard_violations": [],
    "physics": "인물은 쪼그려 앉아 체중을 지탱하고 있으며, 신발은 바닥에 닿아 있고, 왼손이 신발을 고정한 상태에서 오른손이 카드를 뽑고 있어 물리적으로 자연스럽습니다."
   },
   {
    "label": "B",
    "direction": "시선은 아래쪽의 신발을 향하며, 오른손은 신발 내부의 카드를 만지고 있습니다.",
    "built_space": "레퍼런스의 건물 구석, 배관, 실외기 및 바닥 요소가 잘 반영되어 있습니다.",
    "entities": "현우(얼굴 멍과 다리의 상처), 운동화, 카드 등 지시된 요소들이 등장합니다.",
    "hard_violations": [
     "[gemini-pro] 공중에 떠 있는 신발 (지탱하는 손이나 바닥이 없음)"
    ],
    "physics": "오른손이 신발 안의 카드를 잡고 있으나, 무거운 신발 전체가 바닥에 닿지 않고 공중에 완전히 떠 있으며 이를 지탱하는 물리적 구조나 손이 전혀 없습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "프롬프트가 지시한 클로즈업 구도와 쪼그려 앉은 자세, 바닥에 놓인 신발에서 카드를 뽑는 손의 동작을 물리적 오류 없이 충실하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "다리의 상처 등 설정된 디테일은 나타나나, 지탱하는 것 없이 신발이 공중에 떠 있는 치명적인 물리적 오류가 있어 감점되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 바닥에 놓인 신발과 카드를 향하고 있으며, 두 손은 정확히 신발 입구와 카드를 향해 작용하고 있습니다.",
        "built_space": "레퍼런스 이미지의 야외 구석 공간(벽면, 배관, 우측 에어컨 실외기, 젖은 바닥)이 올바른 위치와 비례로 배치되어 있습니다.",
        "entities": "현우(얼굴의 상처와 앳된 외모), 거친 손, 바닥에 벗어놓은 신발, 반쯤 뽑혀 나온 빳빳한 카드 모두 텍스트와 일치합니다.",
        "hard_violations": [],
        "physics": "인물은 쪼그려 앉아 체중을 지탱하고 있으며, 신발은 바닥에 닿아 있고, 왼손이 신발을 고정한 상태에서 오른손이 카드를 뽑고 있어 물리적으로 자연스럽습니다."
       },
       {
        "label": "B",
        "direction": "시선은 아래쪽의 신발을 향하며, 오른손은 신발 내부의 카드를 만지고 있습니다.",
        "built_space": "레퍼런스의 건물 구석, 배관, 실외기 및 바닥 요소가 잘 반영되어 있습니다.",
        "entities": "현우(얼굴 멍과 다리의 상처), 운동화, 카드 등 지시된 요소들이 등장합니다.",
        "hard_violations": [
         "공중에 떠 있는 신발 (지탱하는 손이나 바닥이 없음)"
        ],
        "physics": "오른손이 신발 안의 카드를 잡고 있으나, 무거운 신발 전체가 바닥에 닿지 않고 공중에 완전히 떠 있으며 이를 지탱하는 물리적 구조나 손이 전혀 없습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "프롬프트가 지시한 클로즈업 구도와 쪼그려 앉은 자세, 바닥에 놓인 신발에서 카드를 뽑는 손의 동작을 물리적 오류 없이 충실하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "다리의 상처 등 설정된 디테일은 나타나나, 지탱하는 것 없이 신발이 공중에 떠 있는 치명적인 물리적 오류가 있어 감점되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 바닥에 놓인 신발과 카드를 향하고 있으며, 두 손은 정확히 신발 입구와 카드를 향해 작용하고 있습니다.",
        "built_space": "레퍼런스 이미지의 야외 구석 공간(벽면, 배관, 우측 에어컨 실외기, 젖은 바닥)이 올바른 위치와 비례로 배치되어 있습니다.",
        "entities": "현우(얼굴의 상처와 앳된 외모), 거친 손, 바닥에 벗어놓은 신발, 반쯤 뽑혀 나온 빳빳한 카드 모두 텍스트와 일치합니다.",
        "hard_violations": [],
        "physics": "인물은 쪼그려 앉아 체중을 지탱하고 있으며, 신발은 바닥에 닿아 있고, 왼손이 신발을 고정한 상태에서 오른손이 카드를 뽑고 있어 물리적으로 자연스럽습니다."
       },
       {
        "label": "B",
        "direction": "시선은 아래쪽의 신발을 향하며, 오른손은 신발 내부의 카드를 만지고 있습니다.",
        "built_space": "레퍼런스의 건물 구석, 배관, 실외기 및 바닥 요소가 잘 반영되어 있습니다.",
        "entities": "현우(얼굴 멍과 다리의 상처), 운동화, 카드 등 지시된 요소들이 등장합니다.",
        "hard_violations": [
         "공중에 떠 있는 신발 (지탱하는 손이나 바닥이 없음)"
        ],
        "physics": "오른손이 신발 안의 카드를 잡고 있으나, 무거운 신발 전체가 바닥에 닿지 않고 공중에 완전히 떠 있으며 이를 지탱하는 물리적 구조나 손이 전혀 없습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "장소와 인물은 잘 맞지만 얼굴과 배경까지 넓게 담겼고, 카드 면이 거의 전부 드러나 ‘손 클로즈업에서 반쯤 뽑다 멈춘 순간’이 약하다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "거친 두 손과 무릎 중심의 근접 구도, 신발을 붙잡고 카드를 비스듬히 반쯤 꺼내는 동작이 더 정확하나 소매와 카드의 연락처 면 표현은 아쉽다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 눈은 아래쪽 신발과 카드를 향한다. 벗은 신발의 입구는 손과 카메라 쪽으로 비스듬히 올라와 있다. 검지는 카드 앞면을 누르며 아래쪽으로 뻗어 있고, 카드를 위로 집어 뽑는 방향성은 분명하지 않다. 카드의 넓은 면이 카메라에 상당히 노출된다.",
        "built_space": "회색 외벽 두 면의 모서리, 모서리 배수관 한 개, 뒤쪽 상부 창호 한 벌, 오른쪽 루버 설비 한 대와 그 받침 단이 보인다. 얼룩진 젖은 콘크리트 바닥까지 장소 사진의 구성을 따른다. 현우는 왼쪽 벽 가까이 쭈그려 있고 신발은 바닥 위로 들려 있다. 고정 설비의 중복이나 불가능한 반사는 보이지 않는다.",
        "entities": "한 명의 젊은 동아시아계 남성이 보이며 검은 헝클어진 머리, 앳된 얼굴과 남색 반팔은 인물 참조에 가깝다. 손과 팔에는 때와 상처가 있고 얼굴의 멍 및 드러난 다리의 상처도 보인다. 낡은 운동화 한 짝은 발에서 벗겨져 있으며, 안쪽에 온전한 직사각형 카드가 있다. 카드에는 작은 인쇄 흔적이 있으나 내용을 확실히 읽을 수는 없다. 다만 카드 면 대부분이 보여 좁은 일부만 노출하라는 조건과 다르다. 휴대전화, 벽화와 캠핑카는 이 구도에 보이지 않는다.",
        "hard_violations": [],
        "physics": "굽힌 손가락이 신발 입구를 걸어 잡아 들어 올리는 것으로 보이므로 신발이 무지지 상태로 떠 있지는 않다. 카드는 손가락과 신발 안쪽 사이에 놓여 지지되지만, 윗모서리를 집어 반쯤 당긴 자세보다는 앞면을 누르는 자세에 가깝다. 접힌 무릎과 아래쪽으로 이어지는 다리는 쪼그린 자세와 양립하며, 실제 발바닥 접점은 화면 밖이다."
       },
       {
        "label": "B",
        "direction": "눈은 잘려 있어 시선 자체는 확인할 수 없지만 얼굴은 손의 작업 쪽으로 숙여져 있다. 한 손은 신발 입구를 벌려 고정하고 다른 손은 카드 윗부분을 집어 위로 당긴다. 신발 입구가 손과 카메라 쪽으로 비스듬히 열려 있으며, 카드 면은 카메라에 사선으로 놓여 있다.",
        "built_space": "외벽 모서리와 배수관 한 개, 상부 창호 한 벌, 오른쪽 루버 설비 한 대 및 받침 단이 장소 참조와 대응한다. 현우는 왼쪽 바닥 가까이 쭈그려 있고, 손과 굽힌 무릎이 전경 대부분을 차지한다. 벗은 신발은 두 손 아래 전경에 있으며 하단이 잘려 있다. 설비 중복이나 불가능한 반사는 없다.",
        "entities": "현우 한 명의 아래 얼굴, 팔, 손과 굽힌 다리가 보인다. 보이는 얼굴은 젊은 동아시아계 남성으로 참조와 대체로 맞고 볼의 상처도 유지된다. 거칠고 긁힌 손이 명확하지만 남색 상의는 참조의 반팔보다 긴 소매를 걷은 형태다. 벗은 낡은 운동화와 손으로 잡은 빳빳한 카드가 있으며, 카드 아래쪽은 신발 안에 가려진다. 카드에는 띠 모양 요소가 보이지만 연락처가 실린 면이라는 단서는 약하고 읽을 수 있는 글자는 없다. 다리 상처는 바지에 가려져 확인할 수 없으며 휴대전화, 벽화와 캠핑카도 보이지 않는다.",
        "hard_violations": [],
        "physics": "한 손이 신발의 혀와 입구를 확실히 붙잡고, 다른 손의 손가락이 카드 윗부분을 집어 지지한다. 카드 하단이 신발 안에 남아 있어 반쯤 꺼내다 멈춘 동작이 물리적으로 자연스럽다. 화면 왼쪽 아래에 바닥 쪽으로 내려오는 양말 신은 발 일부가 있고 두 무릎이 접혀 있어 쪼그린 자세의 하중 배치도 납득된다. 신발 하단의 바닥 접촉은 잘렸지만 손의 지지가 명확하다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "장소와 인물은 잘 맞지만 얼굴과 배경까지 넓게 담겼고, 카드 면이 거의 전부 드러나 ‘손 클로즈업에서 반쯤 뽑다 멈춘 순간’이 약하다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "거친 두 손과 무릎 중심의 근접 구도, 신발을 붙잡고 카드를 비스듬히 반쯤 꺼내는 동작이 더 정확하나 소매와 카드의 연락처 면 표현은 아쉽다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 눈은 아래쪽 신발과 카드를 향한다. 벗은 신발의 입구는 손과 카메라 쪽으로 비스듬히 올라와 있다. 검지는 카드 앞면을 누르며 아래쪽으로 뻗어 있고, 카드를 위로 집어 뽑는 방향성은 분명하지 않다. 카드의 넓은 면이 카메라에 상당히 노출된다.",
        "built_space": "회색 외벽 두 면의 모서리, 모서리 배수관 한 개, 뒤쪽 상부 창호 한 벌, 오른쪽 루버 설비 한 대와 그 받침 단이 보인다. 얼룩진 젖은 콘크리트 바닥까지 장소 사진의 구성을 따른다. 현우는 왼쪽 벽 가까이 쭈그려 있고 신발은 바닥 위로 들려 있다. 고정 설비의 중복이나 불가능한 반사는 보이지 않는다.",
        "entities": "한 명의 젊은 동아시아계 남성이 보이며 검은 헝클어진 머리, 앳된 얼굴과 남색 반팔은 인물 참조에 가깝다. 손과 팔에는 때와 상처가 있고 얼굴의 멍 및 드러난 다리의 상처도 보인다. 낡은 운동화 한 짝은 발에서 벗겨져 있으며, 안쪽에 온전한 직사각형 카드가 있다. 카드에는 작은 인쇄 흔적이 있으나 내용을 확실히 읽을 수는 없다. 다만 카드 면 대부분이 보여 좁은 일부만 노출하라는 조건과 다르다. 휴대전화, 벽화와 캠핑카는 이 구도에 보이지 않는다.",
        "hard_violations": [],
        "physics": "굽힌 손가락이 신발 입구를 걸어 잡아 들어 올리는 것으로 보이므로 신발이 무지지 상태로 떠 있지는 않다. 카드는 손가락과 신발 안쪽 사이에 놓여 지지되지만, 윗모서리를 집어 반쯤 당긴 자세보다는 앞면을 누르는 자세에 가깝다. 접힌 무릎과 아래쪽으로 이어지는 다리는 쪼그린 자세와 양립하며, 실제 발바닥 접점은 화면 밖이다."
       },
       {
        "label": "A",
        "direction": "눈은 잘려 있어 시선 자체는 확인할 수 없지만 얼굴은 손의 작업 쪽으로 숙여져 있다. 한 손은 신발 입구를 벌려 고정하고 다른 손은 카드 윗부분을 집어 위로 당긴다. 신발 입구가 손과 카메라 쪽으로 비스듬히 열려 있으며, 카드 면은 카메라에 사선으로 놓여 있다.",
        "built_space": "외벽 모서리와 배수관 한 개, 상부 창호 한 벌, 오른쪽 루버 설비 한 대 및 받침 단이 장소 참조와 대응한다. 현우는 왼쪽 바닥 가까이 쭈그려 있고, 손과 굽힌 무릎이 전경 대부분을 차지한다. 벗은 신발은 두 손 아래 전경에 있으며 하단이 잘려 있다. 설비 중복이나 불가능한 반사는 없다.",
        "entities": "현우 한 명의 아래 얼굴, 팔, 손과 굽힌 다리가 보인다. 보이는 얼굴은 젊은 동아시아계 남성으로 참조와 대체로 맞고 볼의 상처도 유지된다. 거칠고 긁힌 손이 명확하지만 남색 상의는 참조의 반팔보다 긴 소매를 걷은 형태다. 벗은 낡은 운동화와 손으로 잡은 빳빳한 카드가 있으며, 카드 아래쪽은 신발 안에 가려진다. 카드에는 띠 모양 요소가 보이지만 연락처가 실린 면이라는 단서는 약하고 읽을 수 있는 글자는 없다. 다리 상처는 바지에 가려져 확인할 수 없으며 휴대전화, 벽화와 캠핑카도 보이지 않는다.",
        "hard_violations": [],
        "physics": "한 손이 신발의 혀와 입구를 확실히 붙잡고, 다른 손의 손가락이 카드 윗부분을 집어 지지한다. 카드 하단이 신발 안에 남아 있어 반쯤 꺼내다 멈춘 동작이 물리적으로 자연스럽다. 화면 왼쪽 아래에 바닥 쪽으로 내려오는 양말 신은 발 일부가 있고 두 무릎이 접혀 있어 쪼그린 자세의 하중 배치도 납득된다. 신발 하단의 바닥 접촉은 잘렸지만 손의 지지가 명확하다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.179
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.929
   },
   "violations": {
    "B": [
     "[gemini-pro] 공중에 떠 있는 신발 (지탱하는 손이나 바닥이 없음)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 929
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "프롬프트가 지시한 클로즈업 구도와 쪼그려 앉은 자세, 바닥에 놓인 신발에서 카드를 뽑는 손의 동작을 물리적 오류 없이 충실하게 구현했습니다."
   },
   {
    "label": "B",
    "score": 929,
    "verdict_ko": "다리의 상처 등 설정된 디테일은 나타나나, 지탱하는 것 없이 신발이 공중에 떠 있는 치명적인 물리적 오류가 있어 감점되었습니다.  ★위반: [gemini-pro] 공중에 떠 있는 신발 (지탱하는 손이나 바닥이 없음)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L202B02.png",
    "asset_id": "f4370321-fb89-4785-899e-192be57a9cc4",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-20b3-7c0c-a223-896e3c9ec95d",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S47sh9__bgfirst_bg.png",
   "bg_asset_id": "85df77f5-d63d-4165-9bbc-cafbc3fedb4e",
   "bg_record_key": "S47sh9::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S47sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:47:51.060187+00:00",
  "fingerprint": "82aaa7f6513789f5058932a9ffa05754a6f58b88f48831e0221c75754be16709",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S47sh9_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S47sh9_sel.png",
  "source_sha256": "3049a2a39139ff15c44991031e7204c31ef23392e8b174c91b5b0cc4057e2460",
  "file": "S47sh9_cine.png",
  "staged_sha256": "2f5dc8861f77cb23d3dd3dd4477ed2f7f519f51e7c37344ba908564df6603c76",
  "latency_ms": 10096
 },
 "S47sh12::signage": {
  "fp": "bc58d9dd83132e37",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S47sh12": {
  "input_fingerprint": "e885130f329b8073",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): early morning.\n\nSHOT TEXT (authoritative, Korean): 뒤돌아선 현우의 등 바로 뒤에 소리 없이 묵묵히 서 있는 찰리의 육중한 전신.\n\nLOCATION (lock): In the same secluded exterior corner of the abandoned library, where the private phone call has just ended. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Continuous corner floor between foreground 현우 and 찰리's fully visible feet in the lower-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Secluded library corner (Shared by 현우 and the silently arrived 찰리) — The floor continues visibly from 현우's foreground position to 찰리's feet, with the corner behind them; used as Proves their immediate physical proximity and prevents the reveal from reading as a separate space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the corner's restrained early-morning ambient illumination, using gentle tonal separation to make the silent full-body reveal legible without a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The Haenam crayon mural and its written message remain intact inside the dusty library. The camper remains parked outside. 현우: He has finished the call and still possesses the disposable phone and the intact contact card taken from his shoe. His facial bruises and untreated leg wound remain. 찰리: He now stands close by in the secluded library corner, with his worn metal body and retained disguise unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): early morning.\n\nSHOT TEXT (authoritative, Korean): 뒤돌아선 현우의 등 바로 뒤에 소리 없이 묵묵히 서 있는 찰리의 육중한 전신.\n\nLOCATION (lock): In the same secluded exterior corner of the abandoned library, where the private phone call has just ended. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Continuous corner floor between foreground 현우 and 찰리's fully visible feet in the lower-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Secluded library corner (Shared by 현우 and the silently arrived 찰리) — The floor continues visibly from 현우's foreground position to 찰리's feet, with the corner behind them; used as Proves their immediate physical proximity and prevents the reveal from reading as a separate space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the corner's restrained early-morning ambient illumination, using gentle tonal separation to make the silent full-body reveal legible without a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The Haenam crayon mural and its written message remain intact inside the dusty library. The camper remains parked outside. 현우: He has finished the call and still possesses the disposable phone and the intact contact card taken from his shoe. His facial bruises and untreated leg wound remain. 찰리: He now stands close by in the secluded library corner, with his worn metal body and retained disguise unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): early morning.\n\nSHOT TEXT (authoritative, Korean): 뒤돌아선 현우의 등 바로 뒤에 소리 없이 묵묵히 서 있는 찰리의 육중한 전신.\n\nLOCATION (lock): In the same secluded exterior corner of the abandoned library, where the private phone call has just ended. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Continuous corner floor between foreground 현우 and 찰리's fully visible feet in the lower-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Secluded library corner (Shared by 현우 and the silently arrived 찰리) — The floor continues visibly from 현우's foreground position to 찰리's feet, with the corner behind them; used as Proves their immediate physical proximity and prevents the reveal from reading as a separate space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the corner's restrained early-morning ambient illumination, using gentle tonal separation to make the silent full-body reveal legible without a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The Haenam crayon mural and its written message remain intact inside the dusty library. The camper remains parked outside. 현우: He has finished the call and still possesses the disposable phone and the intact contact card taken from his shoe. His facial bruises and untreated leg wound remain. 찰리: He now stands close by in the secluded library corner, with his worn metal body and retained disguise unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우는 카메라를 등진 채 앞쪽의 찰리를 향해 서 있고, 찰리도 현우를 마주보고 있음.",
    "built_space": "이전 샷과 동일한 콘크리트 벽면과 배수관이 있는 야외 구석으로, 전경의 현우부터 중경의 찰리까지 바닥면이 자연스럽게 이어짐.",
    "entities": "현우(뒷모습, 이전 샷과 일치하지 않는 회색 후드티 착용), 찰리(모래색 장갑판과 흰색 마스크 등 참조 이미지의 외형과 정확히 일치함).",
    "hard_violations": [
     "[gpt-high] 찰리를 현우의 등 바로 뒤가 아니라 현우와 마주 보는 정면 위치에 배치했다."
    ],
    "physics": "현우와 찰리 모두 지면에 두 발로 체중을 싣고 안정적으로 서 있음."
   },
   {
    "label": "B",
    "direction": "현우는 쪼그려 앉아 왼쪽 벽/창문을 향하고 있으며, 찰리는 현우의 등 뒤에서 그를 내려다보고 있음.",
    "built_space": "야외 구석이 아닌 실내 공간으로 창문과 문, 크레용 벽화가 있는 구조이며 공간 설정 지시와 완전히 다름.",
    "entities": "현우(남색 셔츠를 입고 쪼그려 앉은 뒷모습), 찰리(참조 이미지와 일치하는 로봇).",
    "hard_violations": [
     "[gemini-pro] 지정된 장소(야외 구석)를 무시하고 창문과 벽화가 있는 실내 공간을 임의로 창조함(Invented space)",
     "[gemini-pro] 절대 생성하지 말라고 명시된 '읽을 수 있는 텍스트(한글 문장)'가 벽화 옆에 뚜렷하게 생성됨(Leaked text)",
     "[gpt-high] 지정된 외부 모퉁이를 천장과 창, 출입구가 있는 다른 실내 공간으로 바꾸었다.",
     "[gpt-high] 찰리를 현우의 등 바로 뒤가 아니라 현우의 정면에 배치했다.",
     "[gpt-high] 뒤 벽에 읽을 수 있는 한글 문구가 노출되어 글자 금지 조건을 위반했다."
    ],
    "physics": "현우는 바닥에 쪼그려 앉아 발로 지탱하고 있으며, 찰리는 바닥에 똑바로 서 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "지정된 야외 구석 위치와 프레이밍을 잘 구현했으나, 현우가 이전 샷과 다른 후드티를 입고 있으며 찰리와 마주보고 있어 '등 뒤'라는 위치 관계 설정에서 감점이 있습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "현우의 등 뒤에 선 찰리의 묘사는 맞으나, 지정된 야외 위치를 완전히 무시한 실내 배경과 금지된 텍스트 생성이 겹쳐 지침을 크게 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 카메라를 등진 채 앞쪽의 찰리를 향해 서 있고, 찰리도 현우를 마주보고 있음.",
        "built_space": "이전 샷과 동일한 콘크리트 벽면과 배수관이 있는 야외 구석으로, 전경의 현우부터 중경의 찰리까지 바닥면이 자연스럽게 이어짐.",
        "entities": "현우(뒷모습, 이전 샷과 일치하지 않는 회색 후드티 착용), 찰리(모래색 장갑판과 흰색 마스크 등 참조 이미지의 외형과 정확히 일치함).",
        "hard_violations": [],
        "physics": "현우와 찰리 모두 지면에 두 발로 체중을 싣고 안정적으로 서 있음."
       },
       {
        "label": "B",
        "direction": "현우는 쪼그려 앉아 왼쪽 벽/창문을 향하고 있으며, 찰리는 현우의 등 뒤에서 그를 내려다보고 있음.",
        "built_space": "야외 구석이 아닌 실내 공간으로 창문과 문, 크레용 벽화가 있는 구조이며 공간 설정 지시와 완전히 다름.",
        "entities": "현우(남색 셔츠를 입고 쪼그려 앉은 뒷모습), 찰리(참조 이미지와 일치하는 로봇).",
        "hard_violations": [
         "지정된 장소(야외 구석)를 무시하고 창문과 벽화가 있는 실내 공간을 임의로 창조함(Invented space)",
         "절대 생성하지 말라고 명시된 '읽을 수 있는 텍스트(한글 문장)'가 벽화 옆에 뚜렷하게 생성됨(Leaked text)"
        ],
        "physics": "현우는 바닥에 쪼그려 앉아 발로 지탱하고 있으며, 찰리는 바닥에 똑바로 서 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "지정된 야외 구석 위치와 프레이밍을 잘 구현했으나, 현우가 이전 샷과 다른 후드티를 입고 있으며 찰리와 마주보고 있어 '등 뒤'라는 위치 관계 설정에서 감점이 있습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "현우의 등 뒤에 선 찰리의 묘사는 맞으나, 지정된 야외 위치를 완전히 무시한 실내 배경과 금지된 텍스트 생성이 겹쳐 지침을 크게 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우는 카메라를 등진 채 앞쪽의 찰리를 향해 서 있고, 찰리도 현우를 마주보고 있음.",
        "built_space": "이전 샷과 동일한 콘크리트 벽면과 배수관이 있는 야외 구석으로, 전경의 현우부터 중경의 찰리까지 바닥면이 자연스럽게 이어짐.",
        "entities": "현우(뒷모습, 이전 샷과 일치하지 않는 회색 후드티 착용), 찰리(모래색 장갑판과 흰색 마스크 등 참조 이미지의 외형과 정확히 일치함).",
        "hard_violations": [],
        "physics": "현우와 찰리 모두 지면에 두 발로 체중을 싣고 안정적으로 서 있음."
       },
       {
        "label": "B",
        "direction": "현우는 쪼그려 앉아 왼쪽 벽/창문을 향하고 있으며, 찰리는 현우의 등 뒤에서 그를 내려다보고 있음.",
        "built_space": "야외 구석이 아닌 실내 공간으로 창문과 문, 크레용 벽화가 있는 구조이며 공간 설정 지시와 완전히 다름.",
        "entities": "현우(남색 셔츠를 입고 쪼그려 앉은 뒷모습), 찰리(참조 이미지와 일치하는 로봇).",
        "hard_violations": [
         "지정된 장소(야외 구석)를 무시하고 창문과 벽화가 있는 실내 공간을 임의로 창조함(Invented space)",
         "절대 생성하지 말라고 명시된 '읽을 수 있는 텍스트(한글 문장)'가 벽화 옆에 뚜렷하게 생성됨(Leaked text)"
        ],
        "physics": "현우는 바닥에 쪼그려 앉아 발로 지탱하고 있으며, 찰리는 바닥에 똑바로 서 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 1,
        "verdict_ko": "찰리의 전신과 연결된 바닥은 보이지만, 외부 모퉁이를 실내로 바꾸고 현우 앞에 찰리를 배치했으며 읽히는 벽 글씨까지 노출했다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "외부 콘크리트 모퉁이와 아침빛은 더 충실하지만, 찰리가 현우의 등 뒤가 아닌 정면에 있고 현우의 옷도 후드 상의로 바뀌었다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 카메라에 등을 보인 채 찰리 쪽으로 몸을 향하고 고개를 낮추고 있다. 찰리의 얼굴은 전경의 현우 쪽을 향한다. 따라서 찰리는 현우의 등 바로 뒤가 아니라 현우가 바라보는 쪽에 있다. 무기나 이동 동작은 없다.",
        "built_space": "천장이 있는 실내이며 왼쪽에 큰 창 개구부 두 곳, 오른쪽에 캠핑카가 보이는 출입구 한 곳이 있다. 뒤 벽에는 크레용 벽화와 글씨, 작은 전기 부속 하나가 보인다. 전경에서 찰리의 두 발까지 바닥이 이어지지만, 참조의 외부 콘크리트 모퉁이·배수관·오른쪽 설비 받침 공간과는 다른 장소다.",
        "entities": "현우와 찰리만 보인다. 현우는 젊은 남성의 체격, 헝클어진 검은 머리, 남색 상의와 검은 바지, 낡은 밝은 운동화로 참조와 대체로 맞지만 얼굴은 확인할 수 없다. 신발 입구에 카드 같은 직사각형 물체가 있으며 전화기는 식별되지 않는다. 찰리는 흰 마스크형 얼굴과 샌드 베이지 금속 장갑을 갖췄으나 참조보다 다리가 길고 직립형 비례가 강하다. 벽화와 캠핑카를 굳이 노출했고 벽의 한글 문구는 읽을 수 있다. 다리 상처와 별도 위장 상태는 확인되지 않는다.",
        "hard_violations": [
         "지정된 외부 모퉁이를 천장과 창, 출입구가 있는 다른 실내 공간으로 바꾸었다.",
         "찰리를 현우의 등 바로 뒤가 아니라 현우의 정면에 배치했다.",
         "뒤 벽에 읽을 수 있는 한글 문구가 노출되어 글자 금지 조건을 위반했다."
        ],
        "physics": "현우는 무릎을 굽혀 두 운동화로 바닥을 딛고 쪼그려 앉아 있으며 지지점이 보인다. 찰리도 두 발바닥을 바닥에 붙이고 서 있다. 신발 입구의 직사각형 물체는 신발에 끼워져 지지되는 것으로 보인다. 떠 있는 신체나 물체는 없지만 현우의 쪼그린 자세는 요청된 등 뒤 등장 순간을 전달하지 못한다."
       },
       {
        "label": "B",
        "direction": "현우의 등은 카메라를 향하지만 머리와 몸의 앞쪽은 찰리를 향한다. 찰리 역시 얼굴을 현우 쪽으로 돌리고 있어 두 인물이 마주 보는 배치다. 찰리가 현우의 등 뒤에 조용히 도착한 방향 관계는 성립하지 않는다.",
        "built_space": "왼쪽 콘크리트 벽, 뒤 벽의 창 한 구역, 수직 배수관 하나, 오른쪽 높은 차폐벽이 보이는 외부 공간이다. 물기와 오염이 있는 바닥은 전경 현우에서 중경 찰리까지 연속되어 같은 공간임을 보여 준다. 참조의 외벽 재질과 배수관은 유사하지만 오른쪽 설비와 받침 대신 차폐벽이 공간을 막는다. 찰리의 한쪽 발은 온전히 보이고 다른 발은 현우에게 일부 가려져, 두 발이 완전히 드러나는 하단 중앙 배치를 충족하지 못한다.",
        "entities": "현우와 찰리만 보인다. 현우의 검은 머리와 젊은 남성 체격은 부합하나 얼굴과 멍은 뒷모습 때문에 확인되지 않는다. 참조의 남색 둥근 목 상의가 짙은 회색 후드 상의로 바뀌었다. 오른손에 얇은 직사각형 소지품이 보이지만 전화기와 온전한 연락처 카드를 각각 식별하기는 어렵다. 찰리는 긴 팔, 육중한 몸통, 닳은 베이지 장갑판, 흰 기계식 얼굴 등 참조의 주요 특징을 유지한다. 실내 벽화와 캠핑카는 이 구도에서 보이지 않으며 읽히는 글자도 없다.",
        "hard_violations": [
         "찰리를 현우의 등 바로 뒤가 아니라 현우와 마주 보는 정면 위치에 배치했다."
        ],
        "physics": "찰리는 벌린 두 다리와 바닥에 닿은 발로 무게를 지탱하며 금속 팔도 관절에 자연스럽게 연결되어 있다. 현우의 발은 프레임 밖이지만 몸통과 다리는 정상적인 직립 자세로 이어져 부유의 징후가 없다. 오른손의 직사각형 물체는 손에 잡혀 있다. 지지 없이 떠 있는 물체나 불가능한 신체 자세는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "찰리의 전신과 연결된 바닥은 보이지만, 외부 모퉁이를 실내로 바꾸고 현우 앞에 찰리를 배치했으며 읽히는 벽 글씨까지 노출했다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "외부 콘크리트 모퉁이와 아침빛은 더 충실하지만, 찰리가 현우의 등 뒤가 아닌 정면에 있고 현우의 옷도 후드 상의로 바뀌었다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 카메라에 등을 보인 채 찰리 쪽으로 몸을 향하고 고개를 낮추고 있다. 찰리의 얼굴은 전경의 현우 쪽을 향한다. 따라서 찰리는 현우의 등 바로 뒤가 아니라 현우가 바라보는 쪽에 있다. 무기나 이동 동작은 없다.",
        "built_space": "천장이 있는 실내이며 왼쪽에 큰 창 개구부 두 곳, 오른쪽에 캠핑카가 보이는 출입구 한 곳이 있다. 뒤 벽에는 크레용 벽화와 글씨, 작은 전기 부속 하나가 보인다. 전경에서 찰리의 두 발까지 바닥이 이어지지만, 참조의 외부 콘크리트 모퉁이·배수관·오른쪽 설비 받침 공간과는 다른 장소다.",
        "entities": "현우와 찰리만 보인다. 현우는 젊은 남성의 체격, 헝클어진 검은 머리, 남색 상의와 검은 바지, 낡은 밝은 운동화로 참조와 대체로 맞지만 얼굴은 확인할 수 없다. 신발 입구에 카드 같은 직사각형 물체가 있으며 전화기는 식별되지 않는다. 찰리는 흰 마스크형 얼굴과 샌드 베이지 금속 장갑을 갖췄으나 참조보다 다리가 길고 직립형 비례가 강하다. 벽화와 캠핑카를 굳이 노출했고 벽의 한글 문구는 읽을 수 있다. 다리 상처와 별도 위장 상태는 확인되지 않는다.",
        "hard_violations": [
         "지정된 외부 모퉁이를 천장과 창, 출입구가 있는 다른 실내 공간으로 바꾸었다.",
         "찰리를 현우의 등 바로 뒤가 아니라 현우의 정면에 배치했다.",
         "뒤 벽에 읽을 수 있는 한글 문구가 노출되어 글자 금지 조건을 위반했다."
        ],
        "physics": "현우는 무릎을 굽혀 두 운동화로 바닥을 딛고 쪼그려 앉아 있으며 지지점이 보인다. 찰리도 두 발바닥을 바닥에 붙이고 서 있다. 신발 입구의 직사각형 물체는 신발에 끼워져 지지되는 것으로 보인다. 떠 있는 신체나 물체는 없지만 현우의 쪼그린 자세는 요청된 등 뒤 등장 순간을 전달하지 못한다."
       },
       {
        "label": "A",
        "direction": "현우의 등은 카메라를 향하지만 머리와 몸의 앞쪽은 찰리를 향한다. 찰리 역시 얼굴을 현우 쪽으로 돌리고 있어 두 인물이 마주 보는 배치다. 찰리가 현우의 등 뒤에 조용히 도착한 방향 관계는 성립하지 않는다.",
        "built_space": "왼쪽 콘크리트 벽, 뒤 벽의 창 한 구역, 수직 배수관 하나, 오른쪽 높은 차폐벽이 보이는 외부 공간이다. 물기와 오염이 있는 바닥은 전경 현우에서 중경 찰리까지 연속되어 같은 공간임을 보여 준다. 참조의 외벽 재질과 배수관은 유사하지만 오른쪽 설비와 받침 대신 차폐벽이 공간을 막는다. 찰리의 한쪽 발은 온전히 보이고 다른 발은 현우에게 일부 가려져, 두 발이 완전히 드러나는 하단 중앙 배치를 충족하지 못한다.",
        "entities": "현우와 찰리만 보인다. 현우의 검은 머리와 젊은 남성 체격은 부합하나 얼굴과 멍은 뒷모습 때문에 확인되지 않는다. 참조의 남색 둥근 목 상의가 짙은 회색 후드 상의로 바뀌었다. 오른손에 얇은 직사각형 소지품이 보이지만 전화기와 온전한 연락처 카드를 각각 식별하기는 어렵다. 찰리는 긴 팔, 육중한 몸통, 닳은 베이지 장갑판, 흰 기계식 얼굴 등 참조의 주요 특징을 유지한다. 실내 벽화와 캠핑카는 이 구도에서 보이지 않으며 읽히는 글자도 없다.",
        "hard_violations": [
         "찰리를 현우의 등 바로 뒤가 아니라 현우와 마주 보는 정면 위치에 배치했다."
        ],
        "physics": "찰리는 벌린 두 다리와 바닥에 닿은 발로 무게를 지탱하며 금속 팔도 관절에 자연스럽게 연결되어 있다. 현우의 발은 프레임 밖이지만 몸통과 다리는 정상적인 직립 자세로 이어져 부유의 징후가 없다. 오른손의 직사각형 물체는 손에 잡혀 있다. 지지 없이 떠 있는 물체나 불가능한 신체 자세는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.583
   },
   "adjusted": {
    "A": 1.75,
    "B": 0.333
   },
   "violations": {
    "B": [
     "[gemini-pro] 지정된 장소(야외 구석)를 무시하고 창문과 벽화가 있는 실내 공간을 임의로 창조함(Invented space)",
     "[gemini-pro] 절대 생성하지 말라고 명시된 '읽을 수 있는 텍스트(한글 문장)'가 벽화 옆에 뚜렷하게 생성됨(Leaked text)",
     "[gpt-high] 지정된 외부 모퉁이를 천장과 창, 출입구가 있는 다른 실내 공간으로 바꾸었다.",
     "[gpt-high] 찰리를 현우의 등 바로 뒤가 아니라 현우의 정면에 배치했다.",
     "[gpt-high] 뒤 벽에 읽을 수 있는 한글 문구가 노출되어 글자 금지 조건을 위반했다."
    ],
    "A": [
     "[gpt-high] 찰리를 현우의 등 바로 뒤가 아니라 현우와 마주 보는 정면 위치에 배치했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 333
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "지정된 야외 구석 위치와 프레이밍을 잘 구현했으나, 현우가 이전 샷과 다른 후드티를 입고 있으며 찰리와 마주보고 있어 '등 뒤'라는 위치 관계 설정에서 감점이 있습니다.  ★위반: [gpt-high] 찰리를 현우의 등 바로 뒤가 아니라 현우와 마주 보는 정면 위치에 배치했다."
   },
   {
    "label": "B",
    "score": 333,
    "verdict_ko": "현우의 등 뒤에 선 찰리의 묘사는 맞으나, 지정된 야외 위치를 완전히 무시한 실내 배경과 금지된 텍스트 생성이 겹쳐 지침을 크게 위반했습니다.  ★위반: [gemini-pro] 지정된 장소(야외 구석)를 무시하고 창문과 벽화가 있는 실내 공간을 임의로 창조함(Invented space) / [gemini-pro] 절대 생성하지 말라고 명시된 '읽을 수 있는 텍스트(한글 문장)'가 벽화 옆에 뚜렷하게 생성됨(Leaked text) / [gpt-high] 지정된 외부 모퉁이를 천장과 창, 출입구가 있는 다른 실내 공간으로 바꾸었다. / [gpt-high] 찰리를 현우의 등 바로 뒤가 아니라 현우의 정면에 배치했다. / [gpt-high] 뒤 벽에 읽을 수 있는 한글 문구가 노출되어 글자 금지 조건을 위반했다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S47sh9_sel.png",
    "asset_id": "de289937-5838-40fa-8a50-7f246deba968",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-2405-7ee5-8b5d-80a93ba206d0",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S47sh9"
  }
 },
 "S47sh12::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:49:23.639956+00:00",
  "fingerprint": "37105b65871da89231661eaca713582723b882d20290b25df81f6a846e078ed4",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S47sh12_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S47sh12_sel.png",
  "source_sha256": "9afe6c0e536ea803525d664087ec2da2862d2051997e2b0743260f93a979d045",
  "file": "S47sh12_cine.png",
  "staged_sha256": "1668f7ec5e63c4fd3a5d622eb3f6e98f35ee9f0db1ac59f55c1477f4d7bfd955",
  "latency_ms": 11297
 },
 "S48sh5::confined_fp_apt": {
  "applies": true,
  "reason_ko": "이 샷은 캠핑카 내부 조수석을 배경으로 하며, 앰버가 조수석에 앉아 앞유리를 통해 전방의 경찰 검문소를 바라보는 상황입니다. 차량 내부의 정확한 좌석 배치와 인물의 시선 방향이 어긋나면 화면의 일관성과 몰입을 해칠 수 있으므로 평면도 형태의 레이아웃 가이드가 필요합니다.",
  "input_fingerprint": "2d90f3866d2fd521"
 },
 "S48sh5::signage": {
  "fp": "3f8891c62006ae26",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "confinedfp::0261ae55cee7": {
  "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/confinedfp_base_0261ae55cee7.png",
  "place_text": "Inside the camper's front passenger seat, looking through the windshield toward a distant road checkpoint in daylight.",
  "input_fingerprint": "a8d5a55357967c68"
 },
 "S48sh16::signage": {
  "fp": "cebdca82f1f62a9c",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S48sh16": {
  "input_fingerprint": "613a884cea18e021",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 현우가 탄 캠핑카가 도로변을 향해 차체가 기울어진 채 거친 흙먼지를 뒤로 뿜어내며 멀어지는 중간 순간의 뒷모습.\n\nLOCATION (lock): On a rough mountain access track descending toward the roadside, where the camper throws up dust. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Camper (Descending toward the roadside with its body tilted) — Rear and passenger-side flank face the elevated camera; used as Provides the receding subject and a readable measure of the uneven descent; Mountain track (Used by the departing camper) — Runs from the lower foreground toward the upper-right distance; used as Carries the departure direction through the wide composition; Trailing earth dust (Thrown up behind the moving camper); used as Marks the vehicle's recent path without obscuring its rear silhouette.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight and controlled tonal separation keep the departing vehicle legible against the mountain track.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper's wheel fasteners have just been tightened, but its tire remains punctured and its fuel is nearly exhausted; the earlier roof and window leaks have not been repaired. Charlie retains his blanket covering and worn metal exterior on the mountain path.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 현우가 탄 캠핑카가 도로변을 향해 차체가 기울어진 채 거친 흙먼지를 뒤로 뿜어내며 멀어지는 중간 순간의 뒷모습.\n\nLOCATION (lock): On a rough mountain access track descending toward the roadside, where the camper throws up dust. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Camper (Descending toward the roadside with its body tilted) — Rear and passenger-side flank face the elevated camera; used as Provides the receding subject and a readable measure of the uneven descent; Mountain track (Used by the departing camper) — Runs from the lower foreground toward the upper-right distance; used as Carries the departure direction through the wide composition; Trailing earth dust (Thrown up behind the moving camper); used as Marks the vehicle's recent path without obscuring its rear silhouette.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight and controlled tonal separation keep the departing vehicle legible against the mountain track.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper's wheel fasteners have just been tightened, but its tire remains punctured and its fuel is nearly exhausted; the earlier roof and window leaks have not been repaired. Charlie retains his blanket covering and worn metal exterior on the mountain path.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 현우가 탄 캠핑카가 도로변을 향해 차체가 기울어진 채 거친 흙먼지를 뒤로 뿜어내며 멀어지는 중간 순간의 뒷모습.\n\nLOCATION (lock): On a rough mountain access track descending toward the roadside, where the camper throws up dust. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Camper (Descending toward the roadside with its body tilted) — Rear and passenger-side flank face the elevated camera; used as Provides the receding subject and a readable measure of the uneven descent; Mountain track (Used by the departing camper) — Runs from the lower foreground toward the upper-right distance; used as Carries the departure direction through the wide composition; Trailing earth dust (Thrown up behind the moving camper); used as Marks the vehicle's recent path without obscuring its rear silhouette.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight and controlled tonal separation keep the departing vehicle legible against the mountain track.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper's wheel fasteners have just been tightened, but its tire remains punctured and its fuel is nearly exhausted; the earlier roof and window leaks have not been repaired. Charlie retains his blanket covering and worn metal exterior on the mountain path.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "캠핑카가 우측 상단 원경의 포장도로를 향해 주행 중.",
    "built_space": "흙길이 전경에서 우측 상단의 도로를 향해 완만한 내리막 혹은 평행을 이루며 이어짐.",
    "entities": "캠핑카, 뒷면에 담요가 덮인 구조물, 흙먼지가 프롬프트와 일치함.",
    "hard_violations": [],
    "physics": "캠핑카 바퀴가 지면에 닿아 먼지를 일으키며 이동 중."
   },
   {
    "label": "B",
    "direction": "캠핑카가 하단의 포장도로를 등지고 우측 상단의 산 위를 향해 주행 중.",
    "built_space": "포장도로가 하단 전경에 위치하고 흙길이 우측 상단으로 오르막을 형성함.",
    "entities": "캠핑카와 흙먼지는 존재하나, 담요가 덮인 구조물이 없고 원경에 정체불명의 사람/오토바이가 있음.",
    "hard_violations": [
     "[gemini-pro] 프레임 내 인물 금지(NO PEOPLE IN THIS SHOT) 위반: 원경에 사람 형태가 존재함",
     "[gemini-pro] 수직 방향 및 이동 방향 위반: 도로를 향해 내리막(descending)을 가야 하나, 도로를 등지고 산으로 오르막을 등반함",
     "[gpt-high] 도로변으로 내려가야 하는 차량이 전경 도로를 등지고 산길 오르막으로 향해, 명시된 경로의 상하 방향을 반대로 구현했습니다.",
     "[gpt-high] 사람이나 서 있는 인물을 넣지 말라는 구도에 원경의 직립 인물이 나타납니다."
    ],
    "physics": "캠핑카 바퀴가 지면에 닿아 먼지를 내며 주행 중."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "구도, 내리막길 설정, 뒷면의 담요 덮인 물체 등 프롬프트의 요구사항을 충실히 반영했습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "원경에 사람이 등장하여 인물 배제 규칙을 위반했으며, 도로를 향하는 내리막이 아닌 산을 향하는 오르막을 주행하여 실격입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "캠핑카가 우측 상단 원경의 포장도로를 향해 주행 중.",
        "built_space": "흙길이 전경에서 우측 상단의 도로를 향해 완만한 내리막 혹은 평행을 이루며 이어짐.",
        "entities": "캠핑카, 뒷면에 담요가 덮인 구조물, 흙먼지가 프롬프트와 일치함.",
        "hard_violations": [],
        "physics": "캠핑카 바퀴가 지면에 닿아 먼지를 일으키며 이동 중."
       },
       {
        "label": "B",
        "direction": "캠핑카가 하단의 포장도로를 등지고 우측 상단의 산 위를 향해 주행 중.",
        "built_space": "포장도로가 하단 전경에 위치하고 흙길이 우측 상단으로 오르막을 형성함.",
        "entities": "캠핑카와 흙먼지는 존재하나, 담요가 덮인 구조물이 없고 원경에 정체불명의 사람/오토바이가 있음.",
        "hard_violations": [
         "프레임 내 인물 금지(NO PEOPLE IN THIS SHOT) 위반: 원경에 사람 형태가 존재함",
         "수직 방향 및 이동 방향 위반: 도로를 향해 내리막(descending)을 가야 하나, 도로를 등지고 산으로 오르막을 등반함"
        ],
        "physics": "캠핑카 바퀴가 지면에 닿아 먼지를 내며 주행 중."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "구도, 내리막길 설정, 뒷면의 담요 덮인 물체 등 프롬프트의 요구사항을 충실히 반영했습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "원경에 사람이 등장하여 인물 배제 규칙을 위반했으며, 도로를 향하는 내리막이 아닌 산을 향하는 오르막을 주행하여 실격입니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "캠핑카가 우측 상단 원경의 포장도로를 향해 주행 중.",
        "built_space": "흙길이 전경에서 우측 상단의 도로를 향해 완만한 내리막 혹은 평행을 이루며 이어짐.",
        "entities": "캠핑카, 뒷면에 담요가 덮인 구조물, 흙먼지가 프롬프트와 일치함.",
        "hard_violations": [],
        "physics": "캠핑카 바퀴가 지면에 닿아 먼지를 일으키며 이동 중."
       },
       {
        "label": "B",
        "direction": "캠핑카가 하단의 포장도로를 등지고 우측 상단의 산 위를 향해 주행 중.",
        "built_space": "포장도로가 하단 전경에 위치하고 흙길이 우측 상단으로 오르막을 형성함.",
        "entities": "캠핑카와 흙먼지는 존재하나, 담요가 덮인 구조물이 없고 원경에 정체불명의 사람/오토바이가 있음.",
        "hard_violations": [
         "프레임 내 인물 금지(NO PEOPLE IN THIS SHOT) 위반: 원경에 사람 형태가 존재함",
         "수직 방향 및 이동 방향 위반: 도로를 향해 내리막(descending)을 가야 하나, 도로를 등지고 산으로 오르막을 등반함"
        ],
        "physics": "캠핑카 바퀴가 지면에 닿아 먼지를 내며 주행 중."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 1,
        "verdict_ko": "차량이 도로를 향해 내려가는 대신 도로에서 산길 위로 올라가며, 원경 인물과 낮고 가까운 시점도 지정된 무인 와이드 구도에 어긋납니다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "높은 시점에서 후면과 조수석 측면을 보이며 우상단 도로로 내려가는 와이드 구도를 충실히 구현하지만, 차체 기울기는 약하고 찰리의 정체는 불명확합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "캠핑카 앞부분은 화면 우상단의 오르막 산길을 향하고, 먼지는 뒤쪽인 좌하단으로 뻗습니다. 포장도로가 전경과 좌측 아래에 있어 차량은 도로변으로 내려가는 것이 아니라 도로에서 멀어져 산으로 올라갑니다. 우상단 길에는 차량에 등을 돌린 듯한 작은 인물이 보이며 시선은 판별되지 않습니다.",
        "built_space": "포장도로가 화면 하단을 가로지르고 흙길 한 갈래가 우상단 산비탈로 올라갑니다. 캠핑카에는 후면 창 하나와 조수석 쪽 출입문 하나가 보입니다. 차량이 화면을 크게 차지하며 지붕 윗면이 거의 보이지 않는 낮은 시점으로, 요구한 높은 카메라의 와이드 구도와 다릅니다.",
        "entities": "낡은 흰색 캠핑카, 돌이 많은 산길, 흙먼지와 낮의 산악 환경은 보입니다. 원경에 회색 옷 또는 덮개를 두른 사람 형태 하나가 있으며 나이·성별·민족은 판별할 수 없습니다. 찰리의 마모된 금속 외장은 확인되지 않습니다. 보이는 타이어는 뚜렷하게 주저앉지 않았고, 연료량·체결 상태·누수 지속 여부는 확인할 수 없습니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "도로변으로 내려가야 하는 차량이 전경 도로를 등지고 산길 오르막으로 향해, 명시된 경로의 상하 방향을 반대로 구현했습니다.",
         "사람이나 서 있는 인물을 넣지 말라는 구도에 원경의 직립 인물이 나타납니다."
        ],
        "physics": "차체는 왼쪽으로 크게 기울고 바퀴들은 요철이 있는 지면에 닿거나 바로 위를 통과하는 모습입니다. 바퀴 주변에서 뒤로 튀는 흙과 먼지는 거친 길을 달리는 동작으로 설명됩니다. 기울기 자체를 불가능한 자세로 볼 근거는 없으며, 원경 인물도 길 위에 서 있습니다."
       },
       {
        "label": "B",
        "direction": "캠핑카는 화면 우상단의 포장도로 접속부를 향해 멀어지고 있습니다. 후면과 조수석 측면이 카메라를 향하며, 흙먼지는 지나온 길인 좌하단으로 이어져 후면 윤곽을 크게 가리지 않습니다. 노출된 사람이나 시선은 없습니다.",
        "built_space": "거친 흙길 한 갈래가 하단 전경에서 우상단 도로 접속부까지 이어지고, 도착 지점의 포장도로와 가드레일이 보입니다. 높은 카메라에서 차량 지붕·후면·조수석 측면을 함께 봅니다. 후면 창 하나와 측면 출입문 하나가 있으며, 뒤쪽에는 금속 상자 형태의 적재물이 묶여 있습니다. 넓은 주변 지형이 포함되어 요구한 와이드 구도에 부합합니다.",
        "entities": "낡은 캠핑카, 산악 비포장길, 뒤로 퍼지는 흙먼지와 주간 환경이 모두 보입니다. 후면 적재물은 담요가 덮인 낡은 금속 외장이어서 찰리의 재질·덮개 조건에는 대응하지만, 상자 같은 형태라 찰리 자체인지는 확정하기 어렵습니다. 뒤 타이어는 접지부가 눌려 보이나 펑크 여부는 단정할 수 없습니다. 연료량·볼트 체결·누수 상태는 확인되지 않습니다. 사람이나 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "차량 바퀴가 울퉁불퉁한 흙길에 접지하고 차체가 약간 기울어 있어 내려가는 주행으로 읽힙니다. 먼지는 타이어 뒤에서 발생해 지나온 경로로 퍼집니다. 후면 금속 적재물은 받침과 감싼 고정끈으로 지지되며, 담요는 그 위에 걸쳐져 있어 지지 없이 떠 있는 물체는 없습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "차량이 도로를 향해 내려가는 대신 도로에서 산길 위로 올라가며, 원경 인물과 낮고 가까운 시점도 지정된 무인 와이드 구도에 어긋납니다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "높은 시점에서 후면과 조수석 측면을 보이며 우상단 도로로 내려가는 와이드 구도를 충실히 구현하지만, 차체 기울기는 약하고 찰리의 정체는 불명확합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "캠핑카 앞부분은 화면 우상단의 오르막 산길을 향하고, 먼지는 뒤쪽인 좌하단으로 뻗습니다. 포장도로가 전경과 좌측 아래에 있어 차량은 도로변으로 내려가는 것이 아니라 도로에서 멀어져 산으로 올라갑니다. 우상단 길에는 차량에 등을 돌린 듯한 작은 인물이 보이며 시선은 판별되지 않습니다.",
        "built_space": "포장도로가 화면 하단을 가로지르고 흙길 한 갈래가 우상단 산비탈로 올라갑니다. 캠핑카에는 후면 창 하나와 조수석 쪽 출입문 하나가 보입니다. 차량이 화면을 크게 차지하며 지붕 윗면이 거의 보이지 않는 낮은 시점으로, 요구한 높은 카메라의 와이드 구도와 다릅니다.",
        "entities": "낡은 흰색 캠핑카, 돌이 많은 산길, 흙먼지와 낮의 산악 환경은 보입니다. 원경에 회색 옷 또는 덮개를 두른 사람 형태 하나가 있으며 나이·성별·민족은 판별할 수 없습니다. 찰리의 마모된 금속 외장은 확인되지 않습니다. 보이는 타이어는 뚜렷하게 주저앉지 않았고, 연료량·체결 상태·누수 지속 여부는 확인할 수 없습니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "도로변으로 내려가야 하는 차량이 전경 도로를 등지고 산길 오르막으로 향해, 명시된 경로의 상하 방향을 반대로 구현했습니다.",
         "사람이나 서 있는 인물을 넣지 말라는 구도에 원경의 직립 인물이 나타납니다."
        ],
        "physics": "차체는 왼쪽으로 크게 기울고 바퀴들은 요철이 있는 지면에 닿거나 바로 위를 통과하는 모습입니다. 바퀴 주변에서 뒤로 튀는 흙과 먼지는 거친 길을 달리는 동작으로 설명됩니다. 기울기 자체를 불가능한 자세로 볼 근거는 없으며, 원경 인물도 길 위에 서 있습니다."
       },
       {
        "label": "A",
        "direction": "캠핑카는 화면 우상단의 포장도로 접속부를 향해 멀어지고 있습니다. 후면과 조수석 측면이 카메라를 향하며, 흙먼지는 지나온 길인 좌하단으로 이어져 후면 윤곽을 크게 가리지 않습니다. 노출된 사람이나 시선은 없습니다.",
        "built_space": "거친 흙길 한 갈래가 하단 전경에서 우상단 도로 접속부까지 이어지고, 도착 지점의 포장도로와 가드레일이 보입니다. 높은 카메라에서 차량 지붕·후면·조수석 측면을 함께 봅니다. 후면 창 하나와 측면 출입문 하나가 있으며, 뒤쪽에는 금속 상자 형태의 적재물이 묶여 있습니다. 넓은 주변 지형이 포함되어 요구한 와이드 구도에 부합합니다.",
        "entities": "낡은 캠핑카, 산악 비포장길, 뒤로 퍼지는 흙먼지와 주간 환경이 모두 보입니다. 후면 적재물은 담요가 덮인 낡은 금속 외장이어서 찰리의 재질·덮개 조건에는 대응하지만, 상자 같은 형태라 찰리 자체인지는 확정하기 어렵습니다. 뒤 타이어는 접지부가 눌려 보이나 펑크 여부는 단정할 수 없습니다. 연료량·볼트 체결·누수 상태는 확인되지 않습니다. 사람이나 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "차량 바퀴가 울퉁불퉁한 흙길에 접지하고 차체가 약간 기울어 있어 내려가는 주행으로 읽힙니다. 먼지는 타이어 뒤에서 발생해 지나온 경로로 퍼집니다. 후면 금속 적재물은 받침과 감싼 고정끈으로 지지되며, 담요는 그 위에 걸쳐져 있어 지지 없이 떠 있는 물체는 없습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.411
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.161
   },
   "violations": {
    "B": [
     "[gemini-pro] 프레임 내 인물 금지(NO PEOPLE IN THIS SHOT) 위반: 원경에 사람 형태가 존재함",
     "[gemini-pro] 수직 방향 및 이동 방향 위반: 도로를 향해 내리막(descending)을 가야 하나, 도로를 등지고 산으로 오르막을 등반함",
     "[gpt-high] 도로변으로 내려가야 하는 차량이 전경 도로를 등지고 산길 오르막으로 향해, 명시된 경로의 상하 방향을 반대로 구현했습니다.",
     "[gpt-high] 사람이나 서 있는 인물을 넣지 말라는 구도에 원경의 직립 인물이 나타납니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 161
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "구도, 내리막길 설정, 뒷면의 담요 덮인 물체 등 프롬프트의 요구사항을 충실히 반영했습니다."
   },
   {
    "label": "B",
    "score": 161,
    "verdict_ko": "원경에 사람이 등장하여 인물 배제 규칙을 위반했으며, 도로를 향하는 내리막이 아닌 산을 향하는 오르막을 주행하여 실격입니다.  ★위반: [gemini-pro] 프레임 내 인물 금지(NO PEOPLE IN THIS SHOT) 위반: 원경에 사람 형태가 존재함 / [gemini-pro] 수직 방향 및 이동 방향 위반: 도로를 향해 내리막(descending)을 가야 하나, 도로를 등지고 산으로 오르막을 등반함 / [gpt-high] 도로변으로 내려가야 하는 차량이 전경 도로를 등지고 산길 오르막으로 향해, 명시된 경로의 상하 방향을 반대로 구현했습니다. / [gpt-high] 사람이나 서 있는 인물을 넣지 말라는 구도에 원경의 직립 인물이 나타납니다."
   }
  ],
  "refs": [],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-2903-7b4f-b6da-d3082ade531f",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S48sh16::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:51:20.122995+00:00",
  "fingerprint": "8fe80be0893e7edc7850aad071e6b9afa0927b1edcd0b211e9ba5c540e29b99d",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S48sh16_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S48sh16_sel.png",
  "source_sha256": "9847de5c4c836208861c95d23b0219ab5422adbb23b27f989b3873e17bdc769f",
  "file": "S48sh16_cine.png",
  "staged_sha256": "5d44b78e6e2b5b06c7581e2ed064f463f961fd8a4a2a82aed672f62a9f91881a",
  "latency_ms": 11956
 },
 "S48sh22::signage": {
  "fp": "7baa36487147b709",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S48sh22": {
  "input_fingerprint": "458a102feffa8cf5",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 거대한 금속 손을 양손으로 꽉 감싸 쥔 채 따뜻하게 올려다보는 앰버의 다정한 상체.\n\nLOCATION (lock): On the mountain track descending toward the main road, where the children and robot walk behind the camper. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Mountain path (The group has been descending along it); used as Remains a subdued strip of spatial context behind the promise.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Softly controlled daylight keeps the human hands and metal hand equally readable, allowing tenderness to come from their contact rather than a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper remains compromised by a punctured tire and very low fuel after its wheel fasteners were tightened. Charlie stands on the mountain path with his worn metal body and blanket covering, extending a large hand for the promise. 앰버: She is standing on the mountain path with her hand extended for a promise, still wearing the replacement shoes.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 거대한 금속 손을 양손으로 꽉 감싸 쥔 채 따뜻하게 올려다보는 앰버의 다정한 상체.\n\nLOCATION (lock): On the mountain track descending toward the main road, where the children and robot walk behind the camper. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Mountain path (The group has been descending along it); used as Remains a subdued strip of spatial context behind the promise.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Softly controlled daylight keeps the human hands and metal hand equally readable, allowing tenderness to come from their contact rather than a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper remains compromised by a punctured tire and very low fuel after its wheel fasteners were tightened. Charlie stands on the mountain path with his worn metal body and blanket covering, extending a large hand for the promise. 앰버: She is standing on the mountain path with her hand extended for a promise, still wearing the replacement shoes.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 거대한 금속 손을 양손으로 꽉 감싸 쥔 채 따뜻하게 올려다보는 앰버의 다정한 상체.\n\nLOCATION (lock): On the mountain track descending toward the main road, where the children and robot walk behind the camper. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Mountain path (The group has been descending along it); used as Remains a subdued strip of spatial context behind the promise.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Softly controlled daylight keeps the human hands and metal hand equally readable, allowing tenderness to come from their contact rather than a lighting change.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper remains compromised by a punctured tire and very low fuel after its wheel fasteners were tightened. Charlie stands on the mountain path with his worn metal body and blanket covering, extending a large hand for the promise. 앰버: She is standing on the mountain path with her hand extended for a promise, still wearing the replacement shoes.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "B",
    "direction": "앰버는 고개와 눈을 오른쪽 위로 들어 프레임 밖 찰리의 얼굴이 있을 방향을 따뜻하게 바라본다. 찰리의 팔은 오른쪽 위에서 앰버의 가슴 앞으로 내려오며, 앰버의 두 손은 같은 금속 손의 양쪽을 감싸고 있다.",
    "built_space": "인공 구조물 내부가 아니라 흙과 돌이 드러난 산길이다. 뒤로 산 능선과 성긴 나무가 보이고, 길은 오른쪽의 흐릿한 띠로 남는다. 도로·가드레일·전신주는 이 구도에서 확인되지 않으며, 중복된 시설이나 불가능한 반사는 없다.",
    "entities": "앰버는 약 10세의 금발 여자아이로 보이며 큰 눈, 둥근 얼굴, 남색 반소매가 인물 참고와 잘 맞는다. 외모만으로 혼혈 배경을 확정할 수는 없다. 찰리는 마모된 샌드 베이지 장갑의 거대한 손과 팔, 몸통 일부로 나타난다. 얼굴과 짧은 다리, 앰버의 신발은 프레임 밖이다. 담요는 보이지 않지만 어깨도 잘려 있어 착용 상태를 확정하기 어렵다. 추가 인물이나 읽을 수 있는 글자는 없다.",
    "hard_violations": [],
    "physics": "금속 손은 손목 관절을 통해 찰리의 팔에 연결되어 있고, 앰버의 두 손바닥과 손가락이 그 표면에 실제로 닿는다. 팔꿈치를 굽혀 가슴 앞에서 큰 손을 잡는 자세가 자연스럽다. 발은 잘려 있지만 몸통은 정상적인 직립 자세이며 공중에 떠 있다는 징후는 없다."
   },
   {
    "label": "A",
    "direction": "앰버는 왼쪽 위의 찰리 얼굴을 올려다보고, 찰리는 고개를 숙여 앰버 쪽을 바라본다. 찰리의 큰 손은 왼쪽에서 앰버의 가슴 앞으로 수평으로 뻗는다. 금속 손 위에 얹힌 앰버의 한 손은 분명하지만, 다른 손이 반대편을 꽉 감싸는 접촉은 명확하지 않다.",
    "built_space": "돌과 바퀴 자국이 있는 산길 하나가 뒤쪽 도로로 이어지고, 오른쪽 비탈과 먼 산이 보인다. 도로를 따라 가드레일 한 줄과 전신주 두 개가 식별되어 이전 장소의 특징을 잘 잇는다. 다만 산길과 하늘이 넓게 드러나 배경을 억제된 공간 맥락으로만 두라는 지시에서는 멀어진다. 중복 시설이나 불가능한 반사는 없다.",
    "entities": "금발 여자아이의 얼굴과 남색 상의는 앰버 참고와 대체로 맞지만, 참고에 없는 회색 긴소매를 안에 입고 있다. 혼혈 배경은 외모만으로 확정할 수 없다. 찰리의 흰 마스크형 얼굴, 샌드 베이지 장갑, 큰 팔은 참고의 특징을 유지하나, 넓게 노출된 어깨와 몸통에 요구된 담요가 없다. 신발은 프레임 밖이며 추가 인물과 읽을 수 있는 글자는 없다.",
    "hard_violations": [],
    "physics": "찰리의 수평 팔은 보이는 팔꿈치와 어깨 관절을 통해 몸통에 연결되고, 손도 손목에 붙어 있어 지지가 설명된다. 앰버의 보이는 손은 금속 손등에 접촉하며 팔도 몸통으로 자연스럽게 이어진다. 두 인물의 발은 보이지 않지만 직립한 상체에 부유나 불가능한 균형의 증거는 없다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": null,
     "normalized": null,
     "ok": false
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "앰버의 다정한 상체와 금속 손을 양손으로 감싼 접촉에 집중해 핵심 연출을 가장 충실히 구현하지만, 지정된 미디엄 숏보다 다소 타이트하다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "서로 바라보는 방향과 산길의 연속성은 맞지만, 찰리와 배경까지 넓힌 구도로 앰버의 상체 중심성이 약해지고 양손의 감싸 쥠도 불명확하며 의상과 담요 상태가 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버는 고개와 눈을 오른쪽 위로 들어 프레임 밖 찰리의 얼굴이 있을 방향을 따뜻하게 바라본다. 찰리의 팔은 오른쪽 위에서 앰버의 가슴 앞으로 내려오며, 앰버의 두 손은 같은 금속 손의 양쪽을 감싸고 있다.",
        "built_space": "인공 구조물 내부가 아니라 흙과 돌이 드러난 산길이다. 뒤로 산 능선과 성긴 나무가 보이고, 길은 오른쪽의 흐릿한 띠로 남는다. 도로·가드레일·전신주는 이 구도에서 확인되지 않으며, 중복된 시설이나 불가능한 반사는 없다.",
        "entities": "앰버는 약 10세의 금발 여자아이로 보이며 큰 눈, 둥근 얼굴, 남색 반소매가 인물 참고와 잘 맞는다. 외모만으로 혼혈 배경을 확정할 수는 없다. 찰리는 마모된 샌드 베이지 장갑의 거대한 손과 팔, 몸통 일부로 나타난다. 얼굴과 짧은 다리, 앰버의 신발은 프레임 밖이다. 담요는 보이지 않지만 어깨도 잘려 있어 착용 상태를 확정하기 어렵다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "금속 손은 손목 관절을 통해 찰리의 팔에 연결되어 있고, 앰버의 두 손바닥과 손가락이 그 표면에 실제로 닿는다. 팔꿈치를 굽혀 가슴 앞에서 큰 손을 잡는 자세가 자연스럽다. 발은 잘려 있지만 몸통은 정상적인 직립 자세이며 공중에 떠 있다는 징후는 없다."
       },
       {
        "label": "B",
        "direction": "앰버는 왼쪽 위의 찰리 얼굴을 올려다보고, 찰리는 고개를 숙여 앰버 쪽을 바라본다. 찰리의 큰 손은 왼쪽에서 앰버의 가슴 앞으로 수평으로 뻗는다. 금속 손 위에 얹힌 앰버의 한 손은 분명하지만, 다른 손이 반대편을 꽉 감싸는 접촉은 명확하지 않다.",
        "built_space": "돌과 바퀴 자국이 있는 산길 하나가 뒤쪽 도로로 이어지고, 오른쪽 비탈과 먼 산이 보인다. 도로를 따라 가드레일 한 줄과 전신주 두 개가 식별되어 이전 장소의 특징을 잘 잇는다. 다만 산길과 하늘이 넓게 드러나 배경을 억제된 공간 맥락으로만 두라는 지시에서는 멀어진다. 중복 시설이나 불가능한 반사는 없다.",
        "entities": "금발 여자아이의 얼굴과 남색 상의는 앰버 참고와 대체로 맞지만, 참고에 없는 회색 긴소매를 안에 입고 있다. 혼혈 배경은 외모만으로 확정할 수 없다. 찰리의 흰 마스크형 얼굴, 샌드 베이지 장갑, 큰 팔은 참고의 특징을 유지하나, 넓게 노출된 어깨와 몸통에 요구된 담요가 없다. 신발은 프레임 밖이며 추가 인물과 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "찰리의 수평 팔은 보이는 팔꿈치와 어깨 관절을 통해 몸통에 연결되고, 손도 손목에 붙어 있어 지지가 설명된다. 앰버의 보이는 손은 금속 손등에 접촉하며 팔도 몸통으로 자연스럽게 이어진다. 두 인물의 발은 보이지 않지만 직립한 상체에 부유나 불가능한 균형의 증거는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "앰버의 다정한 상체와 금속 손을 양손으로 감싼 접촉에 집중해 핵심 연출을 가장 충실히 구현하지만, 지정된 미디엄 숏보다 다소 타이트하다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "서로 바라보는 방향과 산길의 연속성은 맞지만, 찰리와 배경까지 넓힌 구도로 앰버의 상체 중심성이 약해지고 양손의 감싸 쥠도 불명확하며 의상과 담요 상태가 다르다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "앰버는 고개와 눈을 오른쪽 위로 들어 프레임 밖 찰리의 얼굴이 있을 방향을 따뜻하게 바라본다. 찰리의 팔은 오른쪽 위에서 앰버의 가슴 앞으로 내려오며, 앰버의 두 손은 같은 금속 손의 양쪽을 감싸고 있다.",
        "built_space": "인공 구조물 내부가 아니라 흙과 돌이 드러난 산길이다. 뒤로 산 능선과 성긴 나무가 보이고, 길은 오른쪽의 흐릿한 띠로 남는다. 도로·가드레일·전신주는 이 구도에서 확인되지 않으며, 중복된 시설이나 불가능한 반사는 없다.",
        "entities": "앰버는 약 10세의 금발 여자아이로 보이며 큰 눈, 둥근 얼굴, 남색 반소매가 인물 참고와 잘 맞는다. 외모만으로 혼혈 배경을 확정할 수는 없다. 찰리는 마모된 샌드 베이지 장갑의 거대한 손과 팔, 몸통 일부로 나타난다. 얼굴과 짧은 다리, 앰버의 신발은 프레임 밖이다. 담요는 보이지 않지만 어깨도 잘려 있어 착용 상태를 확정하기 어렵다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "금속 손은 손목 관절을 통해 찰리의 팔에 연결되어 있고, 앰버의 두 손바닥과 손가락이 그 표면에 실제로 닿는다. 팔꿈치를 굽혀 가슴 앞에서 큰 손을 잡는 자세가 자연스럽다. 발은 잘려 있지만 몸통은 정상적인 직립 자세이며 공중에 떠 있다는 징후는 없다."
       },
       {
        "label": "A",
        "direction": "앰버는 왼쪽 위의 찰리 얼굴을 올려다보고, 찰리는 고개를 숙여 앰버 쪽을 바라본다. 찰리의 큰 손은 왼쪽에서 앰버의 가슴 앞으로 수평으로 뻗는다. 금속 손 위에 얹힌 앰버의 한 손은 분명하지만, 다른 손이 반대편을 꽉 감싸는 접촉은 명확하지 않다.",
        "built_space": "돌과 바퀴 자국이 있는 산길 하나가 뒤쪽 도로로 이어지고, 오른쪽 비탈과 먼 산이 보인다. 도로를 따라 가드레일 한 줄과 전신주 두 개가 식별되어 이전 장소의 특징을 잘 잇는다. 다만 산길과 하늘이 넓게 드러나 배경을 억제된 공간 맥락으로만 두라는 지시에서는 멀어진다. 중복 시설이나 불가능한 반사는 없다.",
        "entities": "금발 여자아이의 얼굴과 남색 상의는 앰버 참고와 대체로 맞지만, 참고에 없는 회색 긴소매를 안에 입고 있다. 혼혈 배경은 외모만으로 확정할 수 없다. 찰리의 흰 마스크형 얼굴, 샌드 베이지 장갑, 큰 팔은 참고의 특징을 유지하나, 넓게 노출된 어깨와 몸통에 요구된 담요가 없다. 신발은 프레임 밖이며 추가 인물과 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "찰리의 수평 팔은 보이는 팔꿈치와 어깨 관절을 통해 몸통에 연결되고, 손도 손목에 붙어 있어 지지가 설명된다. 앰버의 보이는 손은 금속 손등에 접촉하며 팔도 몸통으로 자연스럽게 이어진다. 두 인물의 발은 보이지 않지만 직립한 상체에 부유나 불가능한 균형의 증거는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gemini-pro"
   ],
   "route": "single_reverse"
  },
  "totals": {
   "B": 8,
   "A": 6
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 8,
    "verdict_ko": "앰버의 다정한 상체와 금속 손을 양손으로 감싼 접촉에 집중해 핵심 연출을 가장 충실히 구현하지만, 지정된 미디엄 숏보다 다소 타이트하다."
   },
   {
    "label": "A",
    "score": 6,
    "verdict_ko": "서로 바라보는 방향과 산길의 연속성은 맞지만, 찰리와 배경까지 넓힌 구도로 앰버의 상체 중심성이 약해지고 양손의 감싸 쥠도 불명확하며 의상과 담요 상태가 다르다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S48sh16_sel.png",
    "asset_id": "4d47e73a-4dc9-4293-bf85-04b84ecae483",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-2aa3-7b25-a1a5-42490564bf01",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S48sh16"
  }
 },
 "S48sh22::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:52:26.894005+00:00",
  "fingerprint": "53b7a14ab6429e13a6d5ce5d7296f9d894842a003c5ea9f8380c8aaf793e1079",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S48sh22_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S48sh22_sel.png",
  "source_sha256": "b3daafc8b76380cb728b3d0de9ce2143ddc94dd411c957cf3ef6f83a4f13e107",
  "file": "S48sh22_cine.png",
  "staged_sha256": "ffc1c14fe29a5c2d1944f52752ce29ab037272af4a9a6bdfb91e89aee213ac17",
  "latency_ms": 13858
 },
 "S49sh14::signage": {
  "fp": "31fab31b484392d3",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::0e2a5a685247648c": {
  "subjects": [],
  "subject_text": "휴게소 주유장과 마트 앞\n주유기와 미터기가 설치된 도로변 주유 공간. 안쪽에 작은 마트의 전면과 출입문이 있고 바깥 도로와 바로 연결된다.",
  "identity": "canonical",
  "scope_id": "L205",
  "scope_role": "location_exterior",
  "scope_sha": "8fffeb10f99bc0d1"
 },
 "S49sh14::bgfirst_bg": {
  "input_fingerprint": "726f908a2a433162",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 아이들 등 뒤에서 산탄총 총구를 차갑게 겨누고 선 조일광의 위협적인 전신.\n\nLOCATION (lock): Inside the service-station convenience store, near the medicine aisle, in daytime shop light.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Store shelving and merchandise (Still in place before the later collision) — Shelf fronts and receding ends border the confrontation aisle; used as Define the shared space and frame the gunman's full-body reveal; Shotgun (Raised and aimed at the children) — Seen obliquely along its side, with the muzzle directed toward the foreground children rather than the camera; used as Connects the distant threat to the foreground shoulders without oversized perspective.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Ambient interior illumination and restrained contrast keep the confrontation plainly visible without introducing a theatrical light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 아이들 등 뒤에서 산탄총 총구를 차갑게 겨누고 선 조일광의 위협적인 전신.\n\nLOCATION (lock): Inside the service-station convenience store, near the medicine aisle, in daytime shop light.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Store shelving and merchandise (Still in place before the later collision) — Shelf fronts and receding ends border the confrontation aisle; used as Define the shared space and frame the gunman's full-body reveal; Shotgun (Raised and aimed at the children) — Seen obliquely along its side, with the muzzle directed toward the foreground children rather than the camera; used as Connects the distant threat to the foreground shoulders without oversized perspective.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Ambient interior illumination and restrained contrast keep the confrontation plainly visible without introducing a theatrical light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S49sh14__bgfirst_bg.png",
  "asset_id": "837d1f24-2726-405b-897e-9d85c53d8fca",
  "input_asset_ids": [
   "4280069e-a896-4a04-9a87-3534c1d3220b",
   "64f67bba-2b52-476e-9e96-19519e3cc7c9"
  ]
 },
 "S49sh14": {
  "input_fingerprint": "7b85b6496dafe6bb",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 아이들 등 뒤에서 산탄총 총구를 차갑게 겨누고 선 조일광의 위협적인 전신.\n\nLOCATION (lock): Inside the service-station convenience store, near the medicine aisle, in daytime shop light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Store shelving and merchandise (Still in place before the later collision) — Shelf fronts and receding ends border the confrontation aisle; used as Define the shared space and frame the gunman's full-body reveal; Shotgun (Raised and aimed at the children) — Seen obliquely along its side, with the muzzle directed toward the foreground children rather than the camera; used as Connects the distant threat to the foreground shoulders without oversized perspective.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Ambient interior illumination and restrained contrast keep the confrontation plainly visible without introducing a theatrical light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper is outside at the fuel pump with the luggage reloaded, and the shop television carries the wanted report. Charlie retains his worn metal body and blanket covering and holds the medicine bottle taken from the medicine section. 조일광: He wears a military uniform with medals and a Marine Corps cap. He holds a shotgun in an aiming position. 앰버: She is inside the mart near the medicine section, still wearing the replacement shoes. 라울: He is inside the mart near the medicine section.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 조일광 right now, so 조일광's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 조일광: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리); 조일광 (한국인 남성, 50대, 중년의 얼굴, 눈가 주름, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 아이들 등 뒤에서 산탄총 총구를 차갑게 겨누고 선 조일광의 위협적인 전신.\n\nLOCATION (lock): Inside the service-station convenience store, near the medicine aisle, in daytime shop light. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Store shelving and merchandise (Still in place before the later collision) — Shelf fronts and receding ends border the confrontation aisle; used as Define the shared space and frame the gunman's full-body reveal; Shotgun (Raised and aimed at the children) — Seen obliquely along its side, with the muzzle directed toward the foreground children rather than the camera; used as Connects the distant threat to the foreground shoulders without oversized perspective.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Ambient interior illumination and restrained contrast keep the confrontation plainly visible without introducing a theatrical light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper is outside at the fuel pump with the luggage reloaded, and the shop television carries the wanted report. Charlie retains his worn metal body and blanket covering and holds the medicine bottle taken from the medicine section. 조일광: He wears a military uniform with medals and a Marine Corps cap. He holds a shotgun in an aiming position. 앰버: She is inside the mart near the medicine section, still wearing the replacement shoes. 라울: He is inside the mart near the medicine section.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 조일광 right now, so 조일광's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 조일광: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리); 조일광 (한국인 남성, 50대, 중년의 얼굴, 눈가 주름, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 아이들 등 뒤에서 산탄총 총구를 차갑게 겨누고 선 조일광의 위협적인 전신.\n\nLOCATION (lock): Inside the service-station convenience store, near the medicine aisle, in daytime shop light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Store shelving and merchandise (Still in place before the later collision) — Shelf fronts and receding ends border the confrontation aisle; used as Define the shared space and frame the gunman's full-body reveal; Shotgun (Raised and aimed at the children) — Seen obliquely along its side, with the muzzle directed toward the foreground children rather than the camera; used as Connects the distant threat to the foreground shoulders without oversized perspective.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Ambient interior illumination and restrained contrast keep the confrontation plainly visible without introducing a theatrical light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper is outside at the fuel pump with the luggage reloaded, and the shop television carries the wanted report. Charlie retains his worn metal body and blanket covering and holds the medicine bottle taken from the medicine section. 조일광: He wears a military uniform with medals and a Marine Corps cap. He holds a shotgun in an aiming position. 앰버: She is inside the mart near the medicine section, still wearing the replacement shoes. 라울: He is inside the mart near the medicine section.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 조일광 right now, so 조일광's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 조일광: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리); 조일광 (한국인 남성, 50대, 중년의 얼굴, 눈가 주름, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S49sh14__bgfirst_bg.png",
     "asset_id": "837d1f24-2726-405b-897e-9d85c53d8fca",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S49sh14.png",
     "asset_id": "4280069e-a896-4a04-9a87-3534c1d3220b",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163202>",
     "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 조일광: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1310373>",
     "asset_id": "5e8cf758-89bf-4628-a119-36601926db97",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:842741>",
     "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
     "role": "prop_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L205B01.png",
     "asset_id": "64f67bba-2b52-476e-9e96-19519e3cc7c9",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163202>",
     "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 조일광: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1310373>",
     "asset_id": "5e8cf758-89bf-4628-a119-36601926db97",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:842741>",
     "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
     "role": "prop_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "조일광은 산탄총을 아이들을 향해 비스듬히 아래로 겨누고 있으며, 두 아이는 카메라를 등지고 조일광을 주시하고 있다.",
    "built_space": "편의점 내부 약품 코너 통로. 좌우와 뒤쪽에 약품과 상품이 채워진 진열대가 공간을 적절히 둘러싸고 있다.",
    "entities": "조일광은 훈장이 달린 군복과 해병대 모자를 착용했고, 레퍼런스와 유사한 개머리판과 권총 손잡이가 달린 총기를 들고 있다. 전경의 앰버(금발)와 라울(어두운 꽁지머리)은 뒷모습으로 정확히 묘사되었다.",
    "hard_violations": [
     "[gpt-high] 왼쪽 진열대 위 가격 표지의 숫자와 일부 상품 문구가 읽혀, 글자를 전부 판독 불가능하게 처리하라는 조건을 위반한다."
    ],
    "physics": "세 인물 모두 바닥에 체중을 싣고 안정적으로 서 있으며, 조일광이 총기를 쥐고 있는 양손의 위치도 자연스럽다."
   },
   {
    "label": "B",
    "direction": "조일광은 산탄총을 앞을 향해 쥐고 있으나, 왼쪽의 앰버가 고개를 뒤로 돌려 카메라 쪽을 바라보고 있어 시선과 긴장감이 엇갈린다.",
    "built_space": "편의점 내부 통로로, 좌우 진열대에 상품들이 진열되어 있다.",
    "entities": "조일광의 복장은 레퍼런스와 일치하나, 들고 있는 산탄총은 레퍼런스(권총 손잡이 및 특수 개머리판)와 다른 일반적인 사냥용 총기 형태다. 앰버는 금발에 얼굴이 묘사되었고, 라울은 꽁지머리 뒷모습으로 나타난다.",
    "hard_violations": [
     "[gemini-pro] 우측 상단 배경의 간판에 식별 가능한 한자('非') 및 문자가 포함된 그래픽이 노출되어 'No readable writing anywhere' 지침을 위반함.",
     "[gpt-high] 오른쪽 위 광고판의 큰 문자와 일부 상품 포장의 글자가 식별되어, 읽을 수 있는 문구를 전면 금지한 조건을 위반한다."
    ],
    "physics": "인물들의 기립 자세와 총기를 든 손의 지지 상태는 물리적으로 가능하다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "레퍼런스의 특징을 살린 무기를 들고 전경의 아이들 어깨 너머로 위협을 가하는 와이드 샷의 구도와 인물 배치를 매우 훌륭하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "식별 가능한 문자(우측 상단 간판)가 노출되어 글자 배제 지침을 위반했으며, 앰버가 뒤를 돌아보고 있어 지시된 샷의 긴장감과 구도를 훼손했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "조일광은 산탄총을 아이들을 향해 비스듬히 아래로 겨누고 있으며, 두 아이는 카메라를 등지고 조일광을 주시하고 있다.",
        "built_space": "편의점 내부 약품 코너 통로. 좌우와 뒤쪽에 약품과 상품이 채워진 진열대가 공간을 적절히 둘러싸고 있다.",
        "entities": "조일광은 훈장이 달린 군복과 해병대 모자를 착용했고, 레퍼런스와 유사한 개머리판과 권총 손잡이가 달린 총기를 들고 있다. 전경의 앰버(금발)와 라울(어두운 꽁지머리)은 뒷모습으로 정확히 묘사되었다.",
        "hard_violations": [],
        "physics": "세 인물 모두 바닥에 체중을 싣고 안정적으로 서 있으며, 조일광이 총기를 쥐고 있는 양손의 위치도 자연스럽다."
       },
       {
        "label": "B",
        "direction": "조일광은 산탄총을 앞을 향해 쥐고 있으나, 왼쪽의 앰버가 고개를 뒤로 돌려 카메라 쪽을 바라보고 있어 시선과 긴장감이 엇갈린다.",
        "built_space": "편의점 내부 통로로, 좌우 진열대에 상품들이 진열되어 있다.",
        "entities": "조일광의 복장은 레퍼런스와 일치하나, 들고 있는 산탄총은 레퍼런스(권총 손잡이 및 특수 개머리판)와 다른 일반적인 사냥용 총기 형태다. 앰버는 금발에 얼굴이 묘사되었고, 라울은 꽁지머리 뒷모습으로 나타난다.",
        "hard_violations": [
         "우측 상단 배경의 간판에 식별 가능한 한자('非') 및 문자가 포함된 그래픽이 노출되어 'No readable writing anywhere' 지침을 위반함."
        ],
        "physics": "인물들의 기립 자세와 총기를 든 손의 지지 상태는 물리적으로 가능하다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "레퍼런스의 특징을 살린 무기를 들고 전경의 아이들 어깨 너머로 위협을 가하는 와이드 샷의 구도와 인물 배치를 매우 훌륭하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "식별 가능한 문자(우측 상단 간판)가 노출되어 글자 배제 지침을 위반했으며, 앰버가 뒤를 돌아보고 있어 지시된 샷의 긴장감과 구도를 훼손했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "조일광은 산탄총을 아이들을 향해 비스듬히 아래로 겨누고 있으며, 두 아이는 카메라를 등지고 조일광을 주시하고 있다.",
        "built_space": "편의점 내부 약품 코너 통로. 좌우와 뒤쪽에 약품과 상품이 채워진 진열대가 공간을 적절히 둘러싸고 있다.",
        "entities": "조일광은 훈장이 달린 군복과 해병대 모자를 착용했고, 레퍼런스와 유사한 개머리판과 권총 손잡이가 달린 총기를 들고 있다. 전경의 앰버(금발)와 라울(어두운 꽁지머리)은 뒷모습으로 정확히 묘사되었다.",
        "hard_violations": [],
        "physics": "세 인물 모두 바닥에 체중을 싣고 안정적으로 서 있으며, 조일광이 총기를 쥐고 있는 양손의 위치도 자연스럽다."
       },
       {
        "label": "B",
        "direction": "조일광은 산탄총을 앞을 향해 쥐고 있으나, 왼쪽의 앰버가 고개를 뒤로 돌려 카메라 쪽을 바라보고 있어 시선과 긴장감이 엇갈린다.",
        "built_space": "편의점 내부 통로로, 좌우 진열대에 상품들이 진열되어 있다.",
        "entities": "조일광의 복장은 레퍼런스와 일치하나, 들고 있는 산탄총은 레퍼런스(권총 손잡이 및 특수 개머리판)와 다른 일반적인 사냥용 총기 형태다. 앰버는 금발에 얼굴이 묘사되었고, 라울은 꽁지머리 뒷모습으로 나타난다.",
        "hard_violations": [
         "우측 상단 배경의 간판에 식별 가능한 한자('非') 및 문자가 포함된 그래픽이 노출되어 'No readable writing anywhere' 지침을 위반함."
        ],
        "physics": "인물들의 기립 자세와 총기를 든 손의 지지 상태는 물리적으로 가능하다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "조일광의 신발까지 드러나는 전신 구도가 B보다 충실하지만, 총을 낮춰 든 자세와 장소 불일치, 읽히는 상품 문구가 요구를 훼손한다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "전경 아이들과 총의 사선 연결은 보이지만 아이들이 조일광의 하체를 가려 전신 공개가 약하며, 낮은 총 자세·장소 불일치·읽히는 문구도 남는다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "조일광은 전경 아이들 쪽을 바라본다. 산탄총은 화면 오른쪽 아래로 기울어 있으며, 총열 연장선은 오른쪽 라울의 상체 쪽으로 향한다. 카메라 정면을 겨누지는 않고 총 옆면이 보이지만, 어깨에 올려 차갑게 조준하기보다는 가슴 아래에서 낮춰 든 자세다. 라울은 조일광을 향하고 앰버는 카메라 쪽으로 고개를 돌려, 두 아이의 등 뒤에서 겨누는 관계는 명확하지 않다.",
        "built_space": "좌우에 독립 진열대가 하나씩 있고 뒤 벽에도 상품 진열대가 있다. 조일광은 그 사이 통로에, 아이들은 같은 통로의 전경에 서며 신체와 선반의 충돌은 없다. 조일광의 모자부터 양쪽 신발까지 보인다. 다만 녹색 벽 띠와 빽빽한 약품 통로는 참조의 출입구 옆 중앙 진열대 및 오른쪽 업무 공간으로 구성된 장소와 다르다. 참조의 문·책상·의자·벽걸이 TV·볼록거울은 보이지 않아 고정 공간의 일치를 확인할 수 없다. 반사상은 없다.",
        "entities": "성인 남성 한 명과 어린이 두 명만 보인다. 조일광은 중년 동아시아계 남성으로 보이나 참조보다 얼굴이 더 길고 나이 들어 보인다. 훈장이 달린 군복과 정모는 있으나 해병대 모자라는 식별성은 약하다. 앰버는 금발의 어린 여자아이지만 머리를 땋았고 참조의 남색 상의 대신 카키색 상의를 입었다. 라울은 뒤로 묶은 검은 머리와 남색 상의를 갖췄으며 얼굴은 대부분 가려져 정확한 정체성은 확인하기 어렵다. 총은 긴 펌프식 산탄총으로, 짧은 권총형 몸체와 신축 개머리판을 가진 소품 참조와 다르다. 아이들의 신발, 외부 캠퍼와 짐은 프레임 밖이다. 찰리는 나타나지 않으며 추가 인물은 없다.",
        "hard_violations": [
         "오른쪽 위 광고판의 큰 문자와 일부 상품 포장의 글자가 식별되어, 읽을 수 있는 문구를 전면 금지한 조건을 위반한다."
        ],
        "physics": "조일광의 두 신발은 바닥에 닿아 있고 벌린 다리가 체중을 지탱한다. 오른손은 총의 방아쇠 부근을 잡고 왼손은 앞쪽을 받쳐 총을 지지한다. 개머리판은 어깨에 단단히 견착되지 않았지만 이처럼 낮게 들고 있는 자세 자체는 가능하다. 아이들은 하체가 잘린 서 있는 자세이며 공중에 떠 있다는 징후는 없다. 상품과 진열대도 각각 선반과 바닥에 지지된다."
       },
       {
        "label": "B",
        "direction": "조일광의 시선은 전경 아이들을 향한다. 총은 오른쪽 아래로 뻗어 총열 연장선이 오른쪽 라울의 어깨와 상체 쪽으로 이어지며, 카메라를 직접 겨누지 않는다. 총 옆면은 잘 보이지만 어깨 높이로 올린 조준이 아니라 낮은 준비 자세다. 두 아이 모두 조일광을 마주 보고 있어, 아이들의 등 뒤에 선 위협이라는 관계와는 다르다.",
        "built_space": "좌우 독립 진열대 두 개와 후면 벽 진열대 하나가 통로를 만든다. 조일광은 통로 중앙, 아이들은 전경 양쪽에 있다. 아이들이 조일광의 아래쪽 다리와 신발 대부분을 가려 전신 공개가 A보다 불완전하다. 천장 중앙에는 사각 매립형 냉방기 하나가 보인다. 참조의 창 위 냉방기와 개방적인 중앙 매대·오른쪽 업무 공간 대신 다른 약품 통로를 보여 주어 장소 잠금에 맞지 않는다. TV나 볼록거울, 업무용 책상은 보이지 않으며 평가할 반사상도 없다.",
        "entities": "군복 차림 중년 동아시아계 남성 한 명과 어린이 두 명이 보인다. 조일광은 훈장과 약장을 달고 해병대 계열 문양처럼 보이는 정모를 썼지만 얼굴은 참조와 차이가 있다. 왼쪽 앰버의 금발과 남색 상의는 맞지만 긴 땋은 머리는 참조와 다르다. 오른쪽 라울은 검은 꽁지머리를 했으나 회색 상의로 바뀌었고, 두 아이 모두 뒤통수만 보여 얼굴과 혼혈 정체성의 일치는 판별할 수 없다. 총은 긴 펌프식 산탄총이며 소품 참조의 짧은 권총형 총과 일치하지 않는다. 아이들의 신발과 외부 캠퍼는 보이지 않는다. 찰리나 별도 인물은 추가되지 않았다.",
        "hard_violations": [
         "왼쪽 진열대 위 가격 표지의 숫자와 일부 상품 문구가 읽혀, 글자를 전부 판독 불가능하게 처리하라는 조건을 위반한다."
        ],
        "physics": "조일광은 다리를 벌린 채 서 있고 아래쪽 발 접촉은 아이들에게 대부분 가려져 있다. 이를 부유 자세로 볼 근거는 없다. 오른손이 총의 손잡이 쪽을 잡고 왼손이 펌프 앞부분을 받쳐 총의 무게를 지탱한다. 낮춰 든 자세는 물리적으로 가능하지만 견착 조준은 아니다. 아이들도 바닥에 선 자세로 자연스럽게 이어지며, 선반과 상품에 지지 없는 부유는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "조일광의 신발까지 드러나는 전신 구도가 B보다 충실하지만, 총을 낮춰 든 자세와 장소 불일치, 읽히는 상품 문구가 요구를 훼손한다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "전경 아이들과 총의 사선 연결은 보이지만 아이들이 조일광의 하체를 가려 전신 공개가 약하며, 낮은 총 자세·장소 불일치·읽히는 문구도 남는다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "조일광은 전경 아이들 쪽을 바라본다. 산탄총은 화면 오른쪽 아래로 기울어 있으며, 총열 연장선은 오른쪽 라울의 상체 쪽으로 향한다. 카메라 정면을 겨누지는 않고 총 옆면이 보이지만, 어깨에 올려 차갑게 조준하기보다는 가슴 아래에서 낮춰 든 자세다. 라울은 조일광을 향하고 앰버는 카메라 쪽으로 고개를 돌려, 두 아이의 등 뒤에서 겨누는 관계는 명확하지 않다.",
        "built_space": "좌우에 독립 진열대가 하나씩 있고 뒤 벽에도 상품 진열대가 있다. 조일광은 그 사이 통로에, 아이들은 같은 통로의 전경에 서며 신체와 선반의 충돌은 없다. 조일광의 모자부터 양쪽 신발까지 보인다. 다만 녹색 벽 띠와 빽빽한 약품 통로는 참조의 출입구 옆 중앙 진열대 및 오른쪽 업무 공간으로 구성된 장소와 다르다. 참조의 문·책상·의자·벽걸이 TV·볼록거울은 보이지 않아 고정 공간의 일치를 확인할 수 없다. 반사상은 없다.",
        "entities": "성인 남성 한 명과 어린이 두 명만 보인다. 조일광은 중년 동아시아계 남성으로 보이나 참조보다 얼굴이 더 길고 나이 들어 보인다. 훈장이 달린 군복과 정모는 있으나 해병대 모자라는 식별성은 약하다. 앰버는 금발의 어린 여자아이지만 머리를 땋았고 참조의 남색 상의 대신 카키색 상의를 입었다. 라울은 뒤로 묶은 검은 머리와 남색 상의를 갖췄으며 얼굴은 대부분 가려져 정확한 정체성은 확인하기 어렵다. 총은 긴 펌프식 산탄총으로, 짧은 권총형 몸체와 신축 개머리판을 가진 소품 참조와 다르다. 아이들의 신발, 외부 캠퍼와 짐은 프레임 밖이다. 찰리는 나타나지 않으며 추가 인물은 없다.",
        "hard_violations": [
         "오른쪽 위 광고판의 큰 문자와 일부 상품 포장의 글자가 식별되어, 읽을 수 있는 문구를 전면 금지한 조건을 위반한다."
        ],
        "physics": "조일광의 두 신발은 바닥에 닿아 있고 벌린 다리가 체중을 지탱한다. 오른손은 총의 방아쇠 부근을 잡고 왼손은 앞쪽을 받쳐 총을 지지한다. 개머리판은 어깨에 단단히 견착되지 않았지만 이처럼 낮게 들고 있는 자세 자체는 가능하다. 아이들은 하체가 잘린 서 있는 자세이며 공중에 떠 있다는 징후는 없다. 상품과 진열대도 각각 선반과 바닥에 지지된다."
       },
       {
        "label": "A",
        "direction": "조일광의 시선은 전경 아이들을 향한다. 총은 오른쪽 아래로 뻗어 총열 연장선이 오른쪽 라울의 어깨와 상체 쪽으로 이어지며, 카메라를 직접 겨누지 않는다. 총 옆면은 잘 보이지만 어깨 높이로 올린 조준이 아니라 낮은 준비 자세다. 두 아이 모두 조일광을 마주 보고 있어, 아이들의 등 뒤에 선 위협이라는 관계와는 다르다.",
        "built_space": "좌우 독립 진열대 두 개와 후면 벽 진열대 하나가 통로를 만든다. 조일광은 통로 중앙, 아이들은 전경 양쪽에 있다. 아이들이 조일광의 아래쪽 다리와 신발 대부분을 가려 전신 공개가 A보다 불완전하다. 천장 중앙에는 사각 매립형 냉방기 하나가 보인다. 참조의 창 위 냉방기와 개방적인 중앙 매대·오른쪽 업무 공간 대신 다른 약품 통로를 보여 주어 장소 잠금에 맞지 않는다. TV나 볼록거울, 업무용 책상은 보이지 않으며 평가할 반사상도 없다.",
        "entities": "군복 차림 중년 동아시아계 남성 한 명과 어린이 두 명이 보인다. 조일광은 훈장과 약장을 달고 해병대 계열 문양처럼 보이는 정모를 썼지만 얼굴은 참조와 차이가 있다. 왼쪽 앰버의 금발과 남색 상의는 맞지만 긴 땋은 머리는 참조와 다르다. 오른쪽 라울은 검은 꽁지머리를 했으나 회색 상의로 바뀌었고, 두 아이 모두 뒤통수만 보여 얼굴과 혼혈 정체성의 일치는 판별할 수 없다. 총은 긴 펌프식 산탄총이며 소품 참조의 짧은 권총형 총과 일치하지 않는다. 아이들의 신발과 외부 캠퍼는 보이지 않는다. 찰리나 별도 인물은 추가되지 않았다.",
        "hard_violations": [
         "왼쪽 진열대 위 가격 표지의 숫자와 일부 상품 문구가 읽혀, 글자를 전부 판독 불가능하게 처리하라는 조건을 위반한다."
        ],
        "physics": "조일광은 다리를 벌린 채 서 있고 아래쪽 발 접촉은 아이들에게 대부분 가려져 있다. 이를 부유 자세로 볼 근거는 없다. 오른손이 총의 손잡이 쪽을 잡고 왼손이 펌프 앞부분을 받쳐 총의 무게를 지탱한다. 낮춰 든 자세는 물리적으로 가능하지만 견착 조준은 아니다. 아이들도 바닥에 선 자세로 자연스럽게 이어지며, 선반과 상품에 지지 없는 부유는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.75,
    "B": 1.556
   },
   "adjusted": {
    "A": 1.5,
    "B": 1.306
   },
   "violations": {
    "B": [
     "[gemini-pro] 우측 상단 배경의 간판에 식별 가능한 한자('非') 및 문자가 포함된 그래픽이 노출되어 'No readable writing anywhere' 지침을 위반함.",
     "[gpt-high] 오른쪽 위 광고판의 큰 문자와 일부 상품 포장의 글자가 식별되어, 읽을 수 있는 문구를 전면 금지한 조건을 위반한다."
    ],
    "A": [
     "[gpt-high] 왼쪽 진열대 위 가격 표지의 숫자와 일부 상품 문구가 읽혀, 글자를 전부 판독 불가능하게 처리하라는 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1500,
   "B": 1306
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1500,
    "verdict_ko": "레퍼런스의 특징을 살린 무기를 들고 전경의 아이들 어깨 너머로 위협을 가하는 와이드 샷의 구도와 인물 배치를 매우 훌륭하게 구현했습니다.  ★위반: [gpt-high] 왼쪽 진열대 위 가격 표지의 숫자와 일부 상품 문구가 읽혀, 글자를 전부 판독 불가능하게 처리하라는 조건을 위반한다."
   },
   {
    "label": "B",
    "score": 1306,
    "verdict_ko": "식별 가능한 문자(우측 상단 간판)가 노출되어 글자 배제 지침을 위반했으며, 앰버가 뒤를 돌아보고 있어 지시된 샷의 긴장감과 구도를 훼손했습니다.  ★위반: [gemini-pro] 우측 상단 배경의 간판에 식별 가능한 한자('非') 및 문자가 포함된 그래픽이 노출되어 'No readable writing anywhere' 지침을 위반함. / [gpt-high] 오른쪽 위 광고판의 큰 문자와 일부 상품 포장의 글자가 식별되어, 읽을 수 있는 문구를 전면 금지한 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L205B01.png",
    "asset_id": "64f67bba-2b52-476e-9e96-19519e3cc7c9",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163202>",
    "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 조일광: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1310373>",
    "asset_id": "5e8cf758-89bf-4628-a119-36601926db97",
    "role": "character_ref"
   },
   {
    "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
    "path": "<bytes:842741>",
    "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
    "role": "prop_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-2c56-7cd6-b852-f35ccd74fe2f",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S49sh14__bgfirst_bg.png",
   "bg_asset_id": "837d1f24-2726-405b-897e-9d85c53d8fca",
   "bg_record_key": "S49sh14::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S49sh14::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T13:29:11.346324+00:00",
  "fingerprint": "2963d345ed2cd1d774697673396e88fc7de2da679e335f8eee0251f21e687c73",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S49sh14_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S49sh14_sel.png",
  "source_sha256": "bb03b6aea92443bae1729ac871db911ee162ab1daa3c76c526278b53148d5fd4",
  "file": "S49sh14_cine.png",
  "staged_sha256": "e0ca41d3e8ccc66874b276927d629d143d4d340dca3407270904d34cf75b6a1a",
  "latency_ms": 10953
 },
 "S49sh47::signage": {
  "fp": "43280a2465ad2a5b",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S49sh47": {
  "input_fingerprint": "4d7c63f44b030f4c",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쓰레기 둔덕과 강하게 충돌한 반동으로 공중으로 솟구쳐 거꾸로 뒤집히기 시작한 캠핑카의 폭발적인 순간.\n\nLOCATION (lock): At a rubbish embankment beside the road beyond the service station, where the pursued camper crashes and overturns. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nSTRUCTURE LOOK AUTHORITY: the attached STRUCTURE LOOK photograph is the identity of the fixed structure at this location — wherever that structure appears in the frame, its shape, proportions, openings, materials and colors are LOCKED to it. The LOCATION PHOTOGRAPH remains the authority for this shot's sub-space, surroundings, time of day and lighting. If the two conflict on the structure itself, the STRUCTURE LOOK photo wins; for everything else, the LOCATION PHOTOGRAPH wins.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Trash mound beneath the rebounding camper in the lower-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Camper (Airborne after impact and beginning to overturn) — Rear and passenger-side surfaces rotate away from their normal upright relationship to the road; used as Its rotation against the level horizon carries the impact; Trash mound (Struck by the camper); used as Remains below the airborne vehicle as the visible cause of the rebound; Road (Beside the collision site) — The road recedes diagonally while its horizon remains level; used as Provides a stable spatial reference for the overturning vehicle.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with controlled contrast makes the airborne body and ground contact readable without adding an explosion or impact glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper strikes the rubbish embankment with a shotgun-shattered window and a separately smashed passenger-side window; its interior bulbs have flared brightly. Charlie has a bullet-grazed, sparking shoulder, and the pursuing hunting drone has exploded.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쓰레기 둔덕과 강하게 충돌한 반동으로 공중으로 솟구쳐 거꾸로 뒤집히기 시작한 캠핑카의 폭발적인 순간.\n\nLOCATION (lock): At a rubbish embankment beside the road beyond the service station, where the pursued camper crashes and overturns. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nSTRUCTURE LOOK AUTHORITY: the attached STRUCTURE LOOK photograph is the identity of the fixed structure at this location — wherever that structure appears in the frame, its shape, proportions, openings, materials and colors are LOCKED to it. The LOCATION PHOTOGRAPH remains the authority for this shot's sub-space, surroundings, time of day and lighting. If the two conflict on the structure itself, the STRUCTURE LOOK photo wins; for everything else, the LOCATION PHOTOGRAPH wins.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Trash mound beneath the rebounding camper in the lower-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Camper (Airborne after impact and beginning to overturn) — Rear and passenger-side surfaces rotate away from their normal upright relationship to the road; used as Its rotation against the level horizon carries the impact; Trash mound (Struck by the camper); used as Remains below the airborne vehicle as the visible cause of the rebound; Road (Beside the collision site) — The road recedes diagonally while its horizon remains level; used as Provides a stable spatial reference for the overturning vehicle.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with controlled contrast makes the airborne body and ground contact readable without adding an explosion or impact glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper strikes the rubbish embankment with a shotgun-shattered window and a separately smashed passenger-side window; its interior bulbs have flared brightly. Charlie has a bullet-grazed, sparking shoulder, and the pursuing hunting drone has exploded.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쓰레기 둔덕과 강하게 충돌한 반동으로 공중으로 솟구쳐 거꾸로 뒤집히기 시작한 캠핑카의 폭발적인 순간.\n\nLOCATION (lock): At a rubbish embankment beside the road beyond the service station, where the pursued camper crashes and overturns. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nSTRUCTURE LOOK AUTHORITY: the attached STRUCTURE LOOK photograph is the identity of the fixed structure at this location — wherever that structure appears in the frame, its shape, proportions, openings, materials and colors are LOCKED to it. The LOCATION PHOTOGRAPH remains the authority for this shot's sub-space, surroundings, time of day and lighting. If the two conflict on the structure itself, the STRUCTURE LOOK photo wins; for everything else, the LOCATION PHOTOGRAPH wins.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Trash mound beneath the rebounding camper in the lower-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Camper (Airborne after impact and beginning to overturn) — Rear and passenger-side surfaces rotate away from their normal upright relationship to the road; used as Its rotation against the level horizon carries the impact; Trash mound (Struck by the camper); used as Remains below the airborne vehicle as the visible cause of the rebound; Road (Beside the collision site) — The road recedes diagonally while its horizon remains level; used as Provides a stable spatial reference for the overturning vehicle.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with controlled contrast makes the airborne body and ground contact readable without adding an explosion or impact glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper strikes the rubbish embankment with a shotgun-shattered window and a separately smashed passenger-side window; its interior bulbs have flared brightly. Charlie has a bullet-grazed, sparking shoulder, and the pursuing hunting drone has exploded.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "캠핑카가 쓰레기 더미 위로 앞부분이 들린 채 전진함.",
    "built_space": "위치 사진과 유사한 주유소 캐노피와 주유기가 배경에 배치됨.",
    "entities": "캠핑카, 쓰레기 둔덕, 배경에 서 있는 인물, 파손된 창문.",
    "hard_violations": [
     "[gemini-pro] 지침에서 금지한 인물(주유기 옆 사람) 등장",
     "[gemini-pro] 읽을 수 있는 텍스트('S-OIL') 포함",
     "[gpt-high] 사람이 전혀 등장하면 안 되는 장면인데 주유소 매장 앞에 사람 한 명이 보인다.",
     "[gpt-high] 캐노피의 영문 상호와 로고가 읽혀 읽을 수 있는 글자와 로고 금지 조건을 위반한다."
    ],
    "physics": "앞바퀴가 들려 도약하는 형태이나, 텍스트가 요구한 '거꾸로 뒤집히는' 회전이나 물리적 반동은 나타나지 않음."
   },
   {
    "label": "B",
    "direction": "캠핑카가 공중에서 심하게 앞으로 쏠리며 측면으로 회전 중임.",
    "built_space": "배경에 주유소 캐노피와 주유기가 위치하며 텍스트가 생략됨.",
    "entities": "전복 중인 캠핑카, 쓰레기 둔덕, 배경에 서 있는 인물, 깨진 창문과 내부 불빛.",
    "hard_violations": [
     "[gemini-pro] 지침에서 금지한 인물(주유기 옆 사람) 등장",
     "[gpt-high] 사람이 전혀 등장하면 안 되는 장면인데 주유소 매장 앞에 사람 한 명이 보인다."
    ],
    "physics": "쓰레기 더미와 충돌한 반동으로 완전히 허공에 뜬 채 거꾸로 뒤집히기 시작하는 궤적을 제대로 받쳐줌."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "캠핑카가 전복되지 않고 도약만 하여 핵심 액션에 실패했으며, 금지된 인물과 텍스트가 모두 포함되어 위반이 심함."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "캠핑카가 공중에서 뒤집히는 역동적인 순간은 정확히 구현했으나, 배경에 금지된 인물이 그대로 복사되어 하드 위반임."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "캠핑카가 쓰레기 더미 위로 앞부분이 들린 채 전진함.",
        "built_space": "위치 사진과 유사한 주유소 캐노피와 주유기가 배경에 배치됨.",
        "entities": "캠핑카, 쓰레기 둔덕, 배경에 서 있는 인물, 파손된 창문.",
        "hard_violations": [
         "지침에서 금지한 인물(주유기 옆 사람) 등장",
         "읽을 수 있는 텍스트('S-OIL') 포함"
        ],
        "physics": "앞바퀴가 들려 도약하는 형태이나, 텍스트가 요구한 '거꾸로 뒤집히는' 회전이나 물리적 반동은 나타나지 않음."
       },
       {
        "label": "B",
        "direction": "캠핑카가 공중에서 심하게 앞으로 쏠리며 측면으로 회전 중임.",
        "built_space": "배경에 주유소 캐노피와 주유기가 위치하며 텍스트가 생략됨.",
        "entities": "전복 중인 캠핑카, 쓰레기 둔덕, 배경에 서 있는 인물, 깨진 창문과 내부 불빛.",
        "hard_violations": [
         "지침에서 금지한 인물(주유기 옆 사람) 등장"
        ],
        "physics": "쓰레기 더미와 충돌한 반동으로 완전히 허공에 뜬 채 거꾸로 뒤집히기 시작하는 궤적을 제대로 받쳐줌."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "캠핑카가 전복되지 않고 도약만 하여 핵심 액션에 실패했으며, 금지된 인물과 텍스트가 모두 포함되어 위반이 심함."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "캠핑카가 공중에서 뒤집히는 역동적인 순간은 정확히 구현했으나, 배경에 금지된 인물이 그대로 복사되어 하드 위반임."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "캠핑카가 쓰레기 더미 위로 앞부분이 들린 채 전진함.",
        "built_space": "위치 사진과 유사한 주유소 캐노피와 주유기가 배경에 배치됨.",
        "entities": "캠핑카, 쓰레기 둔덕, 배경에 서 있는 인물, 파손된 창문.",
        "hard_violations": [
         "지침에서 금지한 인물(주유기 옆 사람) 등장",
         "읽을 수 있는 텍스트('S-OIL') 포함"
        ],
        "physics": "앞바퀴가 들려 도약하는 형태이나, 텍스트가 요구한 '거꾸로 뒤집히는' 회전이나 물리적 반동은 나타나지 않음."
       },
       {
        "label": "B",
        "direction": "캠핑카가 공중에서 심하게 앞으로 쏠리며 측면으로 회전 중임.",
        "built_space": "배경에 주유소 캐노피와 주유기가 위치하며 텍스트가 생략됨.",
        "entities": "전복 중인 캠핑카, 쓰레기 둔덕, 배경에 서 있는 인물, 깨진 창문과 내부 불빛.",
        "hard_violations": [
         "지침에서 금지한 인물(주유기 옆 사람) 등장"
        ],
        "physics": "쓰레기 더미와 충돌한 반동으로 완전히 허공에 뜬 채 거꾸로 뒤집히기 시작하는 궤적을 제대로 받쳐줌."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "쓰레기 둔덕 위로 튀어 올라 크게 뒤집히는 차량의 후면과 하부는 요청한 순간에 더 가깝지만, 배경 인물 등장과 구조 참조 불일치, 과도한 충돌 섬광 때문에 사용할 수 없다."
       },
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "뒤집히기보다 둔덕을 타고 앞머리가 들리는 모습이며, 배경 인물과 읽히는 주유소 로고까지 보여 핵심 동작과 명시적 금지 조건을 모두 어긴다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "캠핑카 후면이 카메라를 향하고 차체가 화면 오른쪽으로 크게 기울어 하부가 드러난다. 아래의 쓰레기 둔덕에서 위로 튕겨 나온 뒤 옆으로 전복되기 시작하는 방향이 읽힌다. 조수석 쪽 면은 좁게 보여 방향 확인이 제한적이다. 도로 경계는 대각선으로 이어지고 배경의 수평 기준은 유지된다. 인물의 시선은 식별되지 않는다.",
        "built_space": "왼쪽에 대형 캐노피 하나, 사각 지지 기둥 세 개와 주유기 세 대가 보이고, 뒤에는 파란 띠와 유리 전면이 있는 낮은 매장이 있다. 오른쪽에는 전신주 하나, 큰 금속 쓰레기통 하나, 낮은 상자 하나와 쓰레기 둔덕이 있다. 둔덕은 차량 바로 아래 오른쪽에 있으나 차량은 화면 상단에 거의 닿을 만큼 커서 중경 와이드 배치보다 가까워 보인다. 주유소는 노란 지붕 테두리만 구조 참조에 가까우며, 참조의 원형 금속 기둥·노란 주유 설비·현대적인 매장 전면을 재현하지 않았다.",
        "entities": "캠핑카 한 대, 쓰레기 둔덕, 도로, 주유소가 있다. 후면 유리에는 큰 파손 구멍이 있고 왼쪽으로 드러난 다른 창도 깨져 있지만, 별도로 파손된 조수석 창인지는 확정하기 어렵다. 내부의 밝은 전구는 보인다. 차량 바깥 왼쪽에도 강한 섬광이 있어 충돌 발광 금지와 맞지 않는다. 중앙 기둥 오른쪽 매장 앞에 검은 옷을 입은 사람이 한 명 보인다. 성별·연령·민족성은 판별할 수 없다. 찰리나 드론은 식별되지 않으며, 뚜렷하게 읽히는 문구는 보이지 않는다.",
        "hard_violations": [
         "사람이 전혀 등장하면 안 되는 장면인데 주유소 매장 앞에 사람 한 명이 보인다."
        ],
        "physics": "차량 바퀴는 지면에서 떨어져 있지만 바로 아래 둔덕에서 솟는 먼지와 파편, 둔덕 가까이에 남은 차체 하단이 충돌에 의한 이륙 원인을 보여 준다. 기울어진 차체는 반동과 회전으로 설명 가능하며, 근처 도로와 둔덕이 낙하할 지면이다. 공중 파편도 같은 충격에서 튀어나온 것으로 읽힌다. 근거 없이 정지해 떠 있는 차량은 아니다."
       },
       {
        "label": "B",
        "direction": "캠핑카 앞머리는 화면 오른쪽 쓰레기 둔덕을 향하고 전면과 조수석 쪽 측면이 카메라에 보인다. 앞바퀴가 올라가고 뒤쪽이 낮아 둔덕을 타고 오르는 방향은 분명하지만, 후면을 보이며 거꾸로 뒤집히기 시작하는 회전은 약하다. 도로는 왼쪽으로 대각선 후퇴하며 배경은 수평을 유지한다. 배경 인물의 시선은 확인되지 않는다.",
        "built_space": "대형 캐노피 하나 아래 사각 기둥 세 개와 주유기 세 대가 보인다. 뒤에는 파란 띠의 낮은 매장과 회색 상부 매스가 있다. 오른쪽에는 줄무늬 전신주 하나, 금속 쓰레기통 하나, 낮은 상자 하나가 있고 둔덕은 캠핑카 아래 오른쪽에 놓인다. 차량 전체를 담은 와이드 구도이지만 위치 참조의 시점을 상당히 그대로 따른다. 구조 참조의 원형 금속 기둥, 노란 주유기와 보조 차양, 유리 매장 구성은 재현되지 않았다.",
        "entities": "캠핑카 한 대와 도로 옆 쓰레기 둔덕이 있다. 측면 뒤 창에 파손 구멍이 있고 앞 유리와 조수석 창에도 균열이 보이며, 여러 창 너머로 내부 조명이 밝게 빛난다. 산탄총으로 깨진 창과 별도로 부서진 조수석 창이라는 손상 구분은 충분히 선명하지 않다. 주유소 중앙 기둥 오른쪽에 검은 옷을 입은 사람이 한 명 있다. 작아서 성별·연령·민족성은 확정할 수 없다. 캐노피에는 주유소 영문 상호와 로고가 선명하게 읽힌다. 찰리와 드론은 식별되지 않는다.",
        "hard_violations": [
         "사람이 전혀 등장하면 안 되는 장면인데 주유소 매장 앞에 사람 한 명이 보인다.",
         "캐노피의 영문 상호와 로고가 읽혀 읽을 수 있는 글자와 로고 금지 조건을 위반한다."
        ],
        "physics": "앞바퀴는 둔덕 위 공중에 있고 뒷바퀴는 노면에 매우 가깝다. 앞부분 아래의 쓰레기와 충돌 먼지, 튀는 파편이 차량을 들어 올린 원인을 제공한다. 따라서 무근거한 부유는 아니지만, 차체가 아직 대체로 바로 서 있어 충돌 반동으로 전복되는 순간보다는 둔덕을 타고 들리는 단계로 보인다. 파편은 충돌 지점에서 튀어나오며 지면으로 떨어질 수 있는 배치다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "쓰레기 둔덕 위로 튀어 올라 크게 뒤집히는 차량의 후면과 하부는 요청한 순간에 더 가깝지만, 배경 인물 등장과 구조 참조 불일치, 과도한 충돌 섬광 때문에 사용할 수 없다."
       },
       {
        "label": "A",
        "score": 1,
        "verdict_ko": "뒤집히기보다 둔덕을 타고 앞머리가 들리는 모습이며, 배경 인물과 읽히는 주유소 로고까지 보여 핵심 동작과 명시적 금지 조건을 모두 어긴다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "캠핑카 후면이 카메라를 향하고 차체가 화면 오른쪽으로 크게 기울어 하부가 드러난다. 아래의 쓰레기 둔덕에서 위로 튕겨 나온 뒤 옆으로 전복되기 시작하는 방향이 읽힌다. 조수석 쪽 면은 좁게 보여 방향 확인이 제한적이다. 도로 경계는 대각선으로 이어지고 배경의 수평 기준은 유지된다. 인물의 시선은 식별되지 않는다.",
        "built_space": "왼쪽에 대형 캐노피 하나, 사각 지지 기둥 세 개와 주유기 세 대가 보이고, 뒤에는 파란 띠와 유리 전면이 있는 낮은 매장이 있다. 오른쪽에는 전신주 하나, 큰 금속 쓰레기통 하나, 낮은 상자 하나와 쓰레기 둔덕이 있다. 둔덕은 차량 바로 아래 오른쪽에 있으나 차량은 화면 상단에 거의 닿을 만큼 커서 중경 와이드 배치보다 가까워 보인다. 주유소는 노란 지붕 테두리만 구조 참조에 가까우며, 참조의 원형 금속 기둥·노란 주유 설비·현대적인 매장 전면을 재현하지 않았다.",
        "entities": "캠핑카 한 대, 쓰레기 둔덕, 도로, 주유소가 있다. 후면 유리에는 큰 파손 구멍이 있고 왼쪽으로 드러난 다른 창도 깨져 있지만, 별도로 파손된 조수석 창인지는 확정하기 어렵다. 내부의 밝은 전구는 보인다. 차량 바깥 왼쪽에도 강한 섬광이 있어 충돌 발광 금지와 맞지 않는다. 중앙 기둥 오른쪽 매장 앞에 검은 옷을 입은 사람이 한 명 보인다. 성별·연령·민족성은 판별할 수 없다. 찰리나 드론은 식별되지 않으며, 뚜렷하게 읽히는 문구는 보이지 않는다.",
        "hard_violations": [
         "사람이 전혀 등장하면 안 되는 장면인데 주유소 매장 앞에 사람 한 명이 보인다."
        ],
        "physics": "차량 바퀴는 지면에서 떨어져 있지만 바로 아래 둔덕에서 솟는 먼지와 파편, 둔덕 가까이에 남은 차체 하단이 충돌에 의한 이륙 원인을 보여 준다. 기울어진 차체는 반동과 회전으로 설명 가능하며, 근처 도로와 둔덕이 낙하할 지면이다. 공중 파편도 같은 충격에서 튀어나온 것으로 읽힌다. 근거 없이 정지해 떠 있는 차량은 아니다."
       },
       {
        "label": "A",
        "direction": "캠핑카 앞머리는 화면 오른쪽 쓰레기 둔덕을 향하고 전면과 조수석 쪽 측면이 카메라에 보인다. 앞바퀴가 올라가고 뒤쪽이 낮아 둔덕을 타고 오르는 방향은 분명하지만, 후면을 보이며 거꾸로 뒤집히기 시작하는 회전은 약하다. 도로는 왼쪽으로 대각선 후퇴하며 배경은 수평을 유지한다. 배경 인물의 시선은 확인되지 않는다.",
        "built_space": "대형 캐노피 하나 아래 사각 기둥 세 개와 주유기 세 대가 보인다. 뒤에는 파란 띠의 낮은 매장과 회색 상부 매스가 있다. 오른쪽에는 줄무늬 전신주 하나, 금속 쓰레기통 하나, 낮은 상자 하나가 있고 둔덕은 캠핑카 아래 오른쪽에 놓인다. 차량 전체를 담은 와이드 구도이지만 위치 참조의 시점을 상당히 그대로 따른다. 구조 참조의 원형 금속 기둥, 노란 주유기와 보조 차양, 유리 매장 구성은 재현되지 않았다.",
        "entities": "캠핑카 한 대와 도로 옆 쓰레기 둔덕이 있다. 측면 뒤 창에 파손 구멍이 있고 앞 유리와 조수석 창에도 균열이 보이며, 여러 창 너머로 내부 조명이 밝게 빛난다. 산탄총으로 깨진 창과 별도로 부서진 조수석 창이라는 손상 구분은 충분히 선명하지 않다. 주유소 중앙 기둥 오른쪽에 검은 옷을 입은 사람이 한 명 있다. 작아서 성별·연령·민족성은 확정할 수 없다. 캐노피에는 주유소 영문 상호와 로고가 선명하게 읽힌다. 찰리와 드론은 식별되지 않는다.",
        "hard_violations": [
         "사람이 전혀 등장하면 안 되는 장면인데 주유소 매장 앞에 사람 한 명이 보인다.",
         "캐노피의 영문 상호와 로고가 읽혀 읽을 수 있는 글자와 로고 금지 조건을 위반한다."
        ],
        "physics": "앞바퀴는 둔덕 위 공중에 있고 뒷바퀴는 노면에 매우 가깝다. 앞부분 아래의 쓰레기와 충돌 먼지, 튀는 파편이 차량을 들어 올린 원인을 제공한다. 따라서 무근거한 부유는 아니지만, 차체가 아직 대체로 바로 서 있어 충돌 반동으로 전복되는 순간보다는 둔덕을 타고 들리는 단계로 보인다. 파편은 충돌 지점에서 튀어나오며 지면으로 떨어질 수 있는 배치다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.083,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.833,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 지침에서 금지한 인물(주유기 옆 사람) 등장",
     "[gemini-pro] 읽을 수 있는 텍스트('S-OIL') 포함",
     "[gpt-high] 사람이 전혀 등장하면 안 되는 장면인데 주유소 매장 앞에 사람 한 명이 보인다.",
     "[gpt-high] 캐노피의 영문 상호와 로고가 읽혀 읽을 수 있는 글자와 로고 금지 조건을 위반한다."
    ],
    "B": [
     "[gemini-pro] 지침에서 금지한 인물(주유기 옆 사람) 등장",
     "[gpt-high] 사람이 전혀 등장하면 안 되는 장면인데 주유소 매장 앞에 사람 한 명이 보인다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "A": 833,
   "B": 1750
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 833,
    "verdict_ko": "캠핑카가 전복되지 않고 도약만 하여 핵심 액션에 실패했으며, 금지된 인물과 텍스트가 모두 포함되어 위반이 심함.  ★위반: [gemini-pro] 지침에서 금지한 인물(주유기 옆 사람) 등장 / [gemini-pro] 읽을 수 있는 텍스트('S-OIL') 포함 / [gpt-high] 사람이 전혀 등장하면 안 되는 장면인데 주유소 매장 앞에 사람 한 명이 보인다. / [gpt-high] 캐노피의 영문 상호와 로고가 읽혀 읽을 수 있는 글자와 로고 금지 조건을 위반한다."
   },
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "캠핑카가 공중에서 뒤집히는 역동적인 순간은 정확히 구현했으나, 배경에 금지된 인물이 그대로 복사되어 하드 위반임.  ★위반: [gemini-pro] 지침에서 금지한 인물(주유기 옆 사람) 등장 / [gpt-high] 사람이 전혀 등장하면 안 되는 장면인데 주유소 매장 앞에 사람 한 명이 보인다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its spatial layout, surroundings, fixed features, time of day and lighting mood are spatial truth; stage the moment inside this place. If a STRUCTURE LOOK photograph is also attached, that photo wins for the fixed structure itself — this photograph wins for everything around it. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L205B02.png",
    "asset_id": "d012fbdf-5de4-474f-8f94-751d1a131761",
    "role": "location_plate"
   },
   {
    "label": "STRUCTURE LOOK — the confirmed photograph of the fixed structure at this location: wherever the structure appears in the frame, its shape, proportions, materials, colors and openings are LOCKED to this photo. Never copy its camera framing, time of day or lighting — the shot text and the LOCATION PHOTOGRAPH are the authorities for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_roadside_service_station_sel.png",
    "asset_id": "ab322096-550b-4b37-8cda-d1efdc3c511b",
    "role": "structure_seed_look"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-2fcd-70bb-9cfa-31c7c265a90f",
  "ref_mode": "플레이트+seed만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  },
  "lane_policy": "ab_select_bypass:bg_only"
 },
 "S49sh47::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T06:56:46.705788+00:00",
  "fingerprint": "d30f7a91812bc58c672ea89fbe54598fc630ca77d0415cc124648601b9a251c2",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S49sh47_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S49sh47_sel.png",
  "source_sha256": "6f9a7ba0b7af615ffc41fc50bfb71a1699a15629c4c7fc4e6acf537190644102",
  "file": "S49sh47_cine.png",
  "staged_sha256": "64123d32276da34aad45604efcfdbb4eccda811ad9118321de3b71e5e4f0bc79",
  "latency_ms": 13632
 },
 "S49sh51::signage": {
  "fp": "532422eeffb2d15a",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S49sh51::bgfirst_bg": {
  "input_fingerprint": "d1c5362b7fa4f95d",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 깨진 창문 밖에서 거꾸로 뒤집힌 차 안을 심각한 표정으로 들여다보는 은영과 태진의 상체.\n\nLOCATION (lock): Outside the overturned camper's broken window at the roadside rubbish embankment, seen from the wreck's interior.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Broken camper window opening (Broken in the overturned vehicle) — The interior-facing rim surrounds an unobstructed view of the women outside; used as Creates an irregular frame around the upper-body two-shot without treating the view as a reflection; Overturned interior edges (Inverted after the crash) — Partial interior surfaces enter the lower and side margins at their overturned angles; used as Establishes the camera's vulnerable position inside the wreck.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight outside contrasts with the dim, intermittently lit interior, while the scripted interior smoke remains subtle and both women's faces stay undistorted.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.\n\nThe THIRD attached image (STRUCTURE LOOK) is the identity source of the fixed structure at this location: its faces, openings, levels, materials and signage are truth. Where it conflicts with the LOCATION PHOTOGRAPH about the structure itself, the STRUCTURE LOOK wins; the photograph still governs the surroundings, time of day and lighting.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 깨진 창문 밖에서 거꾸로 뒤집힌 차 안을 심각한 표정으로 들여다보는 은영과 태진의 상체.\n\nLOCATION (lock): Outside the overturned camper's broken window at the roadside rubbish embankment, seen from the wreck's interior.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Broken camper window opening (Broken in the overturned vehicle) — The interior-facing rim surrounds an unobstructed view of the women outside; used as Creates an irregular frame around the upper-body two-shot without treating the view as a reflection; Overturned interior edges (Inverted after the crash) — Partial interior surfaces enter the lower and side margins at their overturned angles; used as Establishes the camera's vulnerable position inside the wreck.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight outside contrasts with the dim, intermittently lit interior, while the scripted interior smoke remains subtle and both women's faces stay undistorted.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.\n\nThe THIRD attached image (STRUCTURE LOOK) is the identity source of the fixed structure at this location: its faces, openings, levels, materials and signage are truth. Where it conflicts with the LOCATION PHOTOGRAPH about the structure itself, the STRUCTURE LOOK wins; the photograph still governs the surroundings, time of day and lighting.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S49sh51__bgfirst_bg.png",
  "asset_id": "c25c9337-c13b-46af-8135-de9acaf2fcf0",
  "input_asset_ids": [
   "f4934d4b-9116-40dd-a1b2-2a834fade0e4",
   "d012fbdf-5de4-474f-8f94-751d1a131761",
   "ab322096-550b-4b37-8cda-d1efdc3c511b"
  ]
 },
 "S49sh51": {
  "input_fingerprint": "22f63187076a5ffe",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 깨진 창문 밖에서 거꾸로 뒤집힌 차 안을 심각한 표정으로 들여다보는 은영과 태진의 상체.\n\nLOCATION (lock): Outside the overturned camper's broken window at the roadside rubbish embankment, seen from the wreck's interior. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nSTRUCTURE LOOK AUTHORITY: the attached STRUCTURE LOOK photograph is the identity of the fixed structure at this location — wherever that structure appears in the frame, its shape, proportions, openings, materials and colors are LOCKED to it. The LOCATION PHOTOGRAPH remains the authority for this shot's sub-space, surroundings, time of day and lighting. If the two conflict on the structure itself, the STRUCTURE LOOK photo wins; for everything else, the LOCATION PHOTOGRAPH wins.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Broken camper window opening (Broken in the overturned vehicle) — The interior-facing rim surrounds an unobstructed view of the women outside; used as Creates an irregular frame around the upper-body two-shot without treating the view as a reflection; Overturned interior edges (Inverted after the crash) — Partial interior surfaces enter the lower and side margins at their overturned angles; used as Establishes the camera's vulnerable position inside the wreck.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight outside contrasts with the dim, intermittently lit interior, while the scripted interior smoke remains subtle and both women's faces stay undistorted.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper lies completely overturned with shattered windows, smoke inside and an interior bulb flickering faintly. Charlie is unconscious inside with the damaged shoulder sustained during the shooting. 태진: She stands outside the wreck looking into it and has short bobbed hair. 은영: She stands outside the wreck looking into it.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 태진 (한국인 여성, 30세, 검은 머리, 짧은 단발); 은영 (한국인 여성, 31세, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 깨진 창문 밖에서 거꾸로 뒤집힌 차 안을 심각한 표정으로 들여다보는 은영과 태진의 상체.\n\nLOCATION (lock): Outside the overturned camper's broken window at the roadside rubbish embankment, seen from the wreck's interior. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Broken camper window opening (Broken in the overturned vehicle) — The interior-facing rim surrounds an unobstructed view of the women outside; used as Creates an irregular frame around the upper-body two-shot without treating the view as a reflection; Overturned interior edges (Inverted after the crash) — Partial interior surfaces enter the lower and side margins at their overturned angles; used as Establishes the camera's vulnerable position inside the wreck.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight outside contrasts with the dim, intermittently lit interior, while the scripted interior smoke remains subtle and both women's faces stay undistorted.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper lies completely overturned with shattered windows, smoke inside and an interior bulb flickering faintly. Charlie is unconscious inside with the damaged shoulder sustained during the shooting. 태진: She stands outside the wreck looking into it and has short bobbed hair. 은영: She stands outside the wreck looking into it.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 태진 (한국인 여성, 30세, 검은 머리, 짧은 단발); 은영 (한국인 여성, 31세, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 깨진 창문 밖에서 거꾸로 뒤집힌 차 안을 심각한 표정으로 들여다보는 은영과 태진의 상체.\n\nLOCATION (lock): Outside the overturned camper's broken window at the roadside rubbish embankment, seen from the wreck's interior. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nSTRUCTURE LOOK AUTHORITY: the attached STRUCTURE LOOK photograph is the identity of the fixed structure at this location — wherever that structure appears in the frame, its shape, proportions, openings, materials and colors are LOCKED to it. The LOCATION PHOTOGRAPH remains the authority for this shot's sub-space, surroundings, time of day and lighting. If the two conflict on the structure itself, the STRUCTURE LOOK photo wins; for everything else, the LOCATION PHOTOGRAPH wins.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Broken camper window opening (Broken in the overturned vehicle) — The interior-facing rim surrounds an unobstructed view of the women outside; used as Creates an irregular frame around the upper-body two-shot without treating the view as a reflection; Overturned interior edges (Inverted after the crash) — Partial interior surfaces enter the lower and side margins at their overturned angles; used as Establishes the camera's vulnerable position inside the wreck.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight outside contrasts with the dim, intermittently lit interior, while the scripted interior smoke remains subtle and both women's faces stay undistorted.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper lies completely overturned with shattered windows, smoke inside and an interior bulb flickering faintly. Charlie is unconscious inside with the damaged shoulder sustained during the shooting. 태진: She stands outside the wreck looking into it and has short bobbed hair. 은영: She stands outside the wreck looking into it.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 태진 (한국인 여성, 30세, 검은 머리, 짧은 단발); 은영 (한국인 여성, 31세, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S49sh51__bgfirst_bg.png",
     "asset_id": "c25c9337-c13b-46af-8135-de9acaf2fcf0",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S49sh51.png",
     "asset_id": "f4934d4b-9116-40dd-a1b2-2a834fade0e4",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 태진: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1170167>",
     "asset_id": "db7f72e4-1c8f-4ab0-8fe8-4f944e2278d3",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 은영: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1258866>",
     "asset_id": "d98e5c15-3084-4367-9745-556e40a4a1c8",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its spatial layout, surroundings, fixed features, time of day and lighting mood are spatial truth; stage the moment inside this place. If a STRUCTURE LOOK photograph is also attached, that photo wins for the fixed structure itself — this photograph wins for everything around it. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L205B02.png",
     "asset_id": "d012fbdf-5de4-474f-8f94-751d1a131761",
     "role": "location_plate"
    },
    {
     "label": "STRUCTURE LOOK — the confirmed photograph of the fixed structure at this location: wherever the structure appears in the frame, its shape, proportions, materials, colors and openings are LOCKED to this photo. Never copy its camera framing, time of day or lighting — the shot text and the LOCATION PHOTOGRAPH are the authorities for those.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_roadside_service_station_sel.png",
     "asset_id": "ab322096-550b-4b37-8cda-d1efdc3c511b",
     "role": "structure_seed_look"
    },
    {
     "label": "CHARACTER REFERENCE — 태진: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1170167>",
     "asset_id": "db7f72e4-1c8f-4ab0-8fe8-4f944e2278d3",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 은영: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1258866>",
     "asset_id": "d98e5c15-3084-4367-9745-556e40a4a1c8",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "두 여성(태진, 은영)이 깨진 창문 밖에서 차량 내부 아래쪽을 내려다보고 있음.",
    "built_space": "배경의 주유소 캐노피는 구조물 레퍼런스와 일치함. 전경의 차량 내부는 좌석과 헤드레스트가 위를 향하고 있어 뒤집힌 상태가 아님.",
    "entities": "태진은 단발로 레퍼런스와 일치하나, 은영은 앞머리 없이 묶은 머리라 불일치함. 화면 하단에 텍스트에 없는 사람의 머리가 나타남.",
    "hard_violations": [
     "[gemini-pro] 샷 텍스트에 없는 인물(화면 하단 머리) 등장 (PEOPLE 규정 위반)",
     "[gemini-pro] 차량 내부가 거꾸로 뒤집히지 않음 (수직 방향 구조 위반)",
     "[gpt-high] 하단 전경에 제삼자의 머리와 몸 일부를 추가했다. 찰리의 차내 상태 설정과 별개로, 이 샷은 은영과 태진만 보이도록 명시하고 다른 인물의 신체 일부도 금지한다."
    ],
    "physics": "두 인물은 지면에 안정적으로 서 있으며, 부서진 유리는 창틀에 자연스럽게 얹혀 있음."
   },
   {
    "label": "B",
    "direction": "두 여성이 깨진 창문 너머 차량 안쪽을 정면으로 심각하게 응시함.",
    "built_space": "배경 우측에 위치 레퍼런스의 쓰레기 더미가 보임. 캠핑카 내부는 상단에 수납장, 하단에 의자가 배치되어 있어 전혀 뒤집히지 않은 상태임.",
    "entities": "은영(앞머리)과 태진(단발) 모두 레퍼런스의 얼굴 및 헤어스타일과 일치하며, 텍스트 외의 불필요한 인물은 없음.",
    "hard_violations": [
     "[gemini-pro] 캠핑카 내부가 거꾸로 뒤집히지 않고 똑바로 서 있음 (수직 방향 구조 위반)"
    ],
    "physics": "외부에 선 두 인물의 자세와 시선 처리가 자연스러우며 구조적 무리가 없음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "은영과 태진의 외모 묘사는 훌륭하나, 캠핑카 내부가 전혀 뒤집히지 않은 똑바른 상태로 렌더링되어 프롬프트의 핵심 상황 설정에 실패했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "차량 내부가 뒤집히지 않은 채 똑바로 서 있으며, 샷 텍스트에 명시되지 않은 제3의 인물을 화면 하단에 추가해 치명적인 규정 위반을 범했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 여성(태진, 은영)이 깨진 창문 밖에서 차량 내부 아래쪽을 내려다보고 있음.",
        "built_space": "배경의 주유소 캐노피는 구조물 레퍼런스와 일치함. 전경의 차량 내부는 좌석과 헤드레스트가 위를 향하고 있어 뒤집힌 상태가 아님.",
        "entities": "태진은 단발로 레퍼런스와 일치하나, 은영은 앞머리 없이 묶은 머리라 불일치함. 화면 하단에 텍스트에 없는 사람의 머리가 나타남.",
        "hard_violations": [
         "샷 텍스트에 없는 인물(화면 하단 머리) 등장 (PEOPLE 규정 위반)",
         "차량 내부가 거꾸로 뒤집히지 않음 (수직 방향 구조 위반)"
        ],
        "physics": "두 인물은 지면에 안정적으로 서 있으며, 부서진 유리는 창틀에 자연스럽게 얹혀 있음."
       },
       {
        "label": "B",
        "direction": "두 여성이 깨진 창문 너머 차량 안쪽을 정면으로 심각하게 응시함.",
        "built_space": "배경 우측에 위치 레퍼런스의 쓰레기 더미가 보임. 캠핑카 내부는 상단에 수납장, 하단에 의자가 배치되어 있어 전혀 뒤집히지 않은 상태임.",
        "entities": "은영(앞머리)과 태진(단발) 모두 레퍼런스의 얼굴 및 헤어스타일과 일치하며, 텍스트 외의 불필요한 인물은 없음.",
        "hard_violations": [
         "캠핑카 내부가 거꾸로 뒤집히지 않고 똑바로 서 있음 (수직 방향 구조 위반)"
        ],
        "physics": "외부에 선 두 인물의 자세와 시선 처리가 자연스러우며 구조적 무리가 없음."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "은영과 태진의 외모 묘사는 훌륭하나, 캠핑카 내부가 전혀 뒤집히지 않은 똑바른 상태로 렌더링되어 프롬프트의 핵심 상황 설정에 실패했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "차량 내부가 뒤집히지 않은 채 똑바로 서 있으며, 샷 텍스트에 명시되지 않은 제3의 인물을 화면 하단에 추가해 치명적인 규정 위반을 범했습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "두 여성(태진, 은영)이 깨진 창문 밖에서 차량 내부 아래쪽을 내려다보고 있음.",
        "built_space": "배경의 주유소 캐노피는 구조물 레퍼런스와 일치함. 전경의 차량 내부는 좌석과 헤드레스트가 위를 향하고 있어 뒤집힌 상태가 아님.",
        "entities": "태진은 단발로 레퍼런스와 일치하나, 은영은 앞머리 없이 묶은 머리라 불일치함. 화면 하단에 텍스트에 없는 사람의 머리가 나타남.",
        "hard_violations": [
         "샷 텍스트에 없는 인물(화면 하단 머리) 등장 (PEOPLE 규정 위반)",
         "차량 내부가 거꾸로 뒤집히지 않음 (수직 방향 구조 위반)"
        ],
        "physics": "두 인물은 지면에 안정적으로 서 있으며, 부서진 유리는 창틀에 자연스럽게 얹혀 있음."
       },
       {
        "label": "B",
        "direction": "두 여성이 깨진 창문 너머 차량 안쪽을 정면으로 심각하게 응시함.",
        "built_space": "배경 우측에 위치 레퍼런스의 쓰레기 더미가 보임. 캠핑카 내부는 상단에 수납장, 하단에 의자가 배치되어 있어 전혀 뒤집히지 않은 상태임.",
        "entities": "은영(앞머리)과 태진(단발) 모두 레퍼런스의 얼굴 및 헤어스타일과 일치하며, 텍스트 외의 불필요한 인물은 없음.",
        "hard_violations": [
         "캠핑카 내부가 거꾸로 뒤집히지 않고 똑바로 서 있음 (수직 방향 구조 위반)"
        ],
        "physics": "외부에 선 두 인물의 자세와 시선 처리가 자연스러우며 구조적 무리가 없음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "창밖의 두 여성만 보여 인원 제한과 내부 시점을 지키지만, 인물이 작고 전복된 실내 표현과 의상·배경 구조물의 일치도가 부족하다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "상체 투숏과 아래쪽 실내를 살피는 시선은 정확하지만, 전경에 제삼자의 머리와 몸을 노출해 명시적인 등장인물 제한을 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 태진과 오른쪽 은영은 모두 깨진 창을 통해 카메라가 있는 실내 쪽을 바라본다. 태진의 시선은 화면 오른쪽 안쪽으로 조금 치우치며, 은영은 정면에 가까운 실내를 응시한다. 서로를 보거나 바깥을 보는 모습은 아니며, 두 표정 모두 진지하다.",
        "built_space": "중앙에 큰 깨진 창 하나, 왼쪽 가장자리에 다른 창 일부, 좌상단에 수납장 문 두 면이 보인다. 아래에는 비스듬한 판 하나와 잔해가 있고 두 여성은 중앙 창 바깥에 나란히 있다. 반사가 아닌 직접적인 시야다. 다만 창과 주변 실내가 화면 대부분을 차지해 상체 투숏이 상대적으로 작다. 수납장은 위쪽, 좌석 같은 면은 아래쪽에 있어 완전히 뒤집힌 차라는 단서는 약하다. 바깥 오른쪽 쓰레기와 덤불은 장소 사진에 부합하지만, 뒤편 건물은 구조 기준 사진의 노란 처마와 현대적인 외장을 충분히 재현하지 못한다.",
        "entities": "보이는 사람은 두 명뿐이다. 왼쪽은 짧은 단발의 태진, 오른쪽은 앞머리와 더 긴 단발의 은영으로 읽히며, 두 사람 모두 설정된 한국인 젊은 성인 여성의 외형과 대체로 맞는다. 태진의 머리색은 기준보다 갈색 기운이 강하다. 두 사람의 회갈색 겉옷은 기준의 남색 상의와 다르다. 깨진 유리, 쓰레기봉투, 옅은 실내 연기가 보이고 전구는 확인되지 않는다. 읽을 수 있는 문구나 추가 인물은 보이지 않는다.",
        "hard_violations": [],
        "physics": "두 여성의 하체와 발은 창 아래에 가려져 있지만, 바깥에 서서 약간 몸을 기울인 상체로 자연스럽게 연결된다. 공중에 떠 있다는 징후는 없다. 전경의 기울어진 판은 아래 잔해와 가느다란 지지대에 걸쳐 있으며, 유리 조각은 창틀에 남아 있다. 연기는 실내에 얇게 퍼져 있다."
       },
       {
        "label": "B",
        "direction": "왼쪽 여성과 오른쪽 단발 여성 모두 고개와 눈을 아래로 향해 창 안쪽 전경의 쓰러진 사람을 살핀다. 실내를 심각하게 들여다보는 행동과 시선의 목표가 명확하다.",
        "built_space": "깨진 창 하나가 화면 가장자리를 둘러싸고, 하단과 오른쪽에 손상된 실내 패널이 들어온다. 두 여성은 창 바깥에서 허리를 조금 숙이고 있으며 상체가 크게 잡혀 요청한 미디엄 투숏에 가깝다. 뒤에는 노란 테두리의 주유소 캐노피 하나와 유리 전면 점포 하나, 오른쪽에는 경고 띠 전신주와 쓰레기 더미가 보인다. 구조 기준의 색과 재료는 A보다 가깝지만 세부 비례는 다르다. 창은 반사면으로 처리되지 않았다.",
        "entities": "오른쪽 태진은 검은 짧은 단발과 남색 상의를 갖췄지만 기준의 긴소매 니트 대신 반소매로 보인다. 왼쪽 은영은 젊은 한국인 여성 외형이나, 기준의 앞머리 있는 단발 대신 뒤로 묶은 머리이고 밝은 셔츠를 입었다. 두 여성 외에 하단 전경에 검은 머리의 제삼자와 몸 일부가 분명히 보인다. 깨진 유리와 옅은 연기는 있으며, 전구와 읽을 수 있는 글자는 확인되지 않는다.",
        "hard_violations": [
         "하단 전경에 제삼자의 머리와 몸 일부를 추가했다. 찰리의 차내 상태 설정과 별개로, 이 샷은 은영과 태진만 보이도록 명시하고 다른 인물의 신체 일부도 금지한다."
        ],
        "physics": "두 여성은 창밖 지면에 선 상태에서 앞으로 기울이는 자연스러운 자세이며, 발은 구도 밖이다. 전경 인물은 아래쪽 실내 면과 잔해에 기대 누운 것으로 보이고 공중에 뜬 형태는 아니다. 깨진 유리는 창틀 가장자리에 걸려 있다. 얼굴과 팔에 명백한 해부학적 불가능성은 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "창밖의 두 여성만 보여 인원 제한과 내부 시점을 지키지만, 인물이 작고 전복된 실내 표현과 의상·배경 구조물의 일치도가 부족하다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "상체 투숏과 아래쪽 실내를 살피는 시선은 정확하지만, 전경에 제삼자의 머리와 몸을 노출해 명시적인 등장인물 제한을 위반한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽 태진과 오른쪽 은영은 모두 깨진 창을 통해 카메라가 있는 실내 쪽을 바라본다. 태진의 시선은 화면 오른쪽 안쪽으로 조금 치우치며, 은영은 정면에 가까운 실내를 응시한다. 서로를 보거나 바깥을 보는 모습은 아니며, 두 표정 모두 진지하다.",
        "built_space": "중앙에 큰 깨진 창 하나, 왼쪽 가장자리에 다른 창 일부, 좌상단에 수납장 문 두 면이 보인다. 아래에는 비스듬한 판 하나와 잔해가 있고 두 여성은 중앙 창 바깥에 나란히 있다. 반사가 아닌 직접적인 시야다. 다만 창과 주변 실내가 화면 대부분을 차지해 상체 투숏이 상대적으로 작다. 수납장은 위쪽, 좌석 같은 면은 아래쪽에 있어 완전히 뒤집힌 차라는 단서는 약하다. 바깥 오른쪽 쓰레기와 덤불은 장소 사진에 부합하지만, 뒤편 건물은 구조 기준 사진의 노란 처마와 현대적인 외장을 충분히 재현하지 못한다.",
        "entities": "보이는 사람은 두 명뿐이다. 왼쪽은 짧은 단발의 태진, 오른쪽은 앞머리와 더 긴 단발의 은영으로 읽히며, 두 사람 모두 설정된 한국인 젊은 성인 여성의 외형과 대체로 맞는다. 태진의 머리색은 기준보다 갈색 기운이 강하다. 두 사람의 회갈색 겉옷은 기준의 남색 상의와 다르다. 깨진 유리, 쓰레기봉투, 옅은 실내 연기가 보이고 전구는 확인되지 않는다. 읽을 수 있는 문구나 추가 인물은 보이지 않는다.",
        "hard_violations": [],
        "physics": "두 여성의 하체와 발은 창 아래에 가려져 있지만, 바깥에 서서 약간 몸을 기울인 상체로 자연스럽게 연결된다. 공중에 떠 있다는 징후는 없다. 전경의 기울어진 판은 아래 잔해와 가느다란 지지대에 걸쳐 있으며, 유리 조각은 창틀에 남아 있다. 연기는 실내에 얇게 퍼져 있다."
       },
       {
        "label": "A",
        "direction": "왼쪽 여성과 오른쪽 단발 여성 모두 고개와 눈을 아래로 향해 창 안쪽 전경의 쓰러진 사람을 살핀다. 실내를 심각하게 들여다보는 행동과 시선의 목표가 명확하다.",
        "built_space": "깨진 창 하나가 화면 가장자리를 둘러싸고, 하단과 오른쪽에 손상된 실내 패널이 들어온다. 두 여성은 창 바깥에서 허리를 조금 숙이고 있으며 상체가 크게 잡혀 요청한 미디엄 투숏에 가깝다. 뒤에는 노란 테두리의 주유소 캐노피 하나와 유리 전면 점포 하나, 오른쪽에는 경고 띠 전신주와 쓰레기 더미가 보인다. 구조 기준의 색과 재료는 A보다 가깝지만 세부 비례는 다르다. 창은 반사면으로 처리되지 않았다.",
        "entities": "오른쪽 태진은 검은 짧은 단발과 남색 상의를 갖췄지만 기준의 긴소매 니트 대신 반소매로 보인다. 왼쪽 은영은 젊은 한국인 여성 외형이나, 기준의 앞머리 있는 단발 대신 뒤로 묶은 머리이고 밝은 셔츠를 입었다. 두 여성 외에 하단 전경에 검은 머리의 제삼자와 몸 일부가 분명히 보인다. 깨진 유리와 옅은 연기는 있으며, 전구와 읽을 수 있는 글자는 확인되지 않는다.",
        "hard_violations": [
         "하단 전경에 제삼자의 머리와 몸 일부를 추가했다. 찰리의 차내 상태 설정과 별개로, 이 샷은 은영과 태진만 보이도록 명시하고 다른 인물의 신체 일부도 금지한다."
        ],
        "physics": "두 여성은 창밖 지면에 선 상태에서 앞으로 기울이는 자연스러운 자세이며, 발은 구도 밖이다. 전경 인물은 아래쪽 실내 면과 잔해에 기대 누운 것으로 보이고 공중에 뜬 형태는 아니다. 깨진 유리는 창틀 가장자리에 걸려 있다. 얼굴과 팔에 명백한 해부학적 불가능성은 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.25,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.0,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 샷 텍스트에 없는 인물(화면 하단 머리) 등장 (PEOPLE 규정 위반)",
     "[gemini-pro] 차량 내부가 거꾸로 뒤집히지 않음 (수직 방향 구조 위반)",
     "[gpt-high] 하단 전경에 제삼자의 머리와 몸 일부를 추가했다. 찰리의 차내 상태 설정과 별개로, 이 샷은 은영과 태진만 보이도록 명시하고 다른 인물의 신체 일부도 금지한다."
    ],
    "B": [
     "[gemini-pro] 캠핑카 내부가 거꾸로 뒤집히지 않고 똑바로 서 있음 (수직 방향 구조 위반)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 1750,
   "A": 1000
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "은영과 태진의 외모 묘사는 훌륭하나, 캠핑카 내부가 전혀 뒤집히지 않은 똑바른 상태로 렌더링되어 프롬프트의 핵심 상황 설정에 실패했습니다.  ★위반: [gemini-pro] 캠핑카 내부가 거꾸로 뒤집히지 않고 똑바로 서 있음 (수직 방향 구조 위반)"
   },
   {
    "label": "A",
    "score": 1000,
    "verdict_ko": "차량 내부가 뒤집히지 않은 채 똑바로 서 있으며, 샷 텍스트에 명시되지 않은 제3의 인물을 화면 하단에 추가해 치명적인 규정 위반을 범했습니다.  ★위반: [gemini-pro] 샷 텍스트에 없는 인물(화면 하단 머리) 등장 (PEOPLE 규정 위반) / [gemini-pro] 차량 내부가 거꾸로 뒤집히지 않음 (수직 방향 구조 위반) / [gpt-high] 하단 전경에 제삼자의 머리와 몸 일부를 추가했다. 찰리의 차내 상태 설정과 별개로, 이 샷은 은영과 태진만 보이도록 명시하고 다른 인물의 신체 일부도 금지한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its spatial layout, surroundings, fixed features, time of day and lighting mood are spatial truth; stage the moment inside this place. If a STRUCTURE LOOK photograph is also attached, that photo wins for the fixed structure itself — this photograph wins for everything around it. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L205B02.png",
    "asset_id": "d012fbdf-5de4-474f-8f94-751d1a131761",
    "role": "location_plate"
   },
   {
    "label": "STRUCTURE LOOK — the confirmed photograph of the fixed structure at this location: wherever the structure appears in the frame, its shape, proportions, materials, colors and openings are LOCKED to this photo. Never copy its camera framing, time of day or lighting — the shot text and the LOCATION PHOTOGRAPH are the authorities for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_roadside_service_station_sel.png",
    "asset_id": "ab322096-550b-4b37-8cda-d1efdc3c511b",
    "role": "structure_seed_look"
   },
   {
    "label": "CHARACTER REFERENCE — 태진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1170167>",
    "asset_id": "db7f72e4-1c8f-4ab0-8fe8-4f944e2278d3",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 은영: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1258866>",
    "asset_id": "d98e5c15-3084-4367-9745-556e40a4a1c8",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-3179-726c-b05c-c721096c5143",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S49sh51__bgfirst_bg.png",
   "bg_asset_id": "c25c9337-c13b-46af-8135-de9acaf2fcf0",
   "bg_record_key": "S49sh51::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate",
   "seed_attached": true
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  },
  "lane_policy": "ab_select_ready"
 },
 "S49sh51::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T13:30:27.233677+00:00",
  "fingerprint": "d1e85f0583039108f17c83850e492323dc2127ade0cad5b100e9802c257b710f",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S49sh51_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S49sh51_sel.png",
  "source_sha256": "df7787f600593fa8e153667d7ac2b54f6f155997d8ef72300e1e8dccc58e6b49",
  "file": "S49sh51_cine.png",
  "staged_sha256": "0c4bd03a1aa897c3db8049f8ab6aad6c7631b945689a5079e420a9bad51ee33c",
  "latency_ms": 11611
 },
 "S50sh4::signage": {
  "fp": "e9cc6e53e13eeb1e",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::5ab1d05fce833b16": {
  "subjects": [],
  "subject_text": "태진과 은영의 카센터 창고\n벽에 여러 록밴드 포스터가 붙은 정비 창고. 거대한 스피커와 수리 공구, 펜치, 정비용 용액 병들이 놓여 있다.",
  "identity": "canonical",
  "scope_id": "L209",
  "scope_role": "location_interior",
  "scope_sha": "fd8d31a02c054e5b"
 },
 "S50sh4::bgfirst_bg": {
  "input_fingerprint": "35650877010a2ee4",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 태진의 뒷모습을 날카롭게 노려보며 바닥의 커다란 펜치를 향해 손을 뻗은 현우의 굳은 상체.\n\nLOCATION (lock): In the resting corner inside an auto-repair warehouse, with band posters nearby and daytime ambient light from the workshop.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Large pliers on the floor beneath 현우's reaching hand in the lower-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Large pliers (On the floor, not yet grasped) — Lie obliquely below 현우's approaching hand; used as Make his defensive intention legible without enlarging the tool beyond realistic scale; Damaged camper (Under repair by 태진) — A partial side section appears beside 태진 in the background; used as Anchors her activity and the depth separating the characters; Rock-band posters (Attached to the room's walls) — Printed band imagery faces into the room and is visible obliquely, without invented readable titles; used as Gives the unfamiliar space its scripted visual identity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral ambient interior light preserves readable depth between the guarded foreground and the repair area without specifying an unsupported fixture.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 태진의 뒷모습을 날카롭게 노려보며 바닥의 커다란 펜치를 향해 손을 뻗은 현우의 굳은 상체.\n\nLOCATION (lock): In the resting corner inside an auto-repair warehouse, with band posters nearby and daytime ambient light from the workshop.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Large pliers on the floor beneath 현우's reaching hand in the lower-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Large pliers (On the floor, not yet grasped) — Lie obliquely below 현우's approaching hand; used as Make his defensive intention legible without enlarging the tool beyond realistic scale; Damaged camper (Under repair by 태진) — A partial side section appears beside 태진 in the background; used as Anchors her activity and the depth separating the characters; Rock-band posters (Attached to the room's walls) — Printed band imagery faces into the room and is visible obliquely, without invented readable titles; used as Gives the unfamiliar space its scripted visual identity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral ambient interior light preserves readable depth between the guarded foreground and the repair area without specifying an unsupported fixture.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S50sh4__bgfirst_bg.png",
  "asset_id": "7a6d7074-e7b2-4b88-bab5-afa565d34ff1",
  "input_asset_ids": [
   "36a68918-70ef-4f62-a881-1c4458c88531",
   "f0f3a2cb-dd51-4ebd-a380-2b0c867d50dd"
  ]
 },
 "S50sh4": {
  "input_fingerprint": "8ae75911d73e3eab",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 태진의 뒷모습을 날카롭게 노려보며 바닥의 커다란 펜치를 향해 손을 뻗은 현우의 굳은 상체.\n\nLOCATION (lock): In the resting corner inside an auto-repair warehouse, with band posters nearby and daytime ambient light from the workshop. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Large pliers on the floor beneath 현우's reaching hand in the lower-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Large pliers (On the floor, not yet grasped) — Lie obliquely below 현우's approaching hand; used as Make his defensive intention legible without enlarging the tool beyond realistic scale; Damaged camper (Under repair by 태진) — A partial side section appears beside 태진 in the background; used as Anchors her activity and the depth separating the characters; Rock-band posters (Attached to the room's walls) — Printed band imagery faces into the room and is visible obliquely, without invented readable titles; used as Gives the unfamiliar space its scripted visual identity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral ambient interior light preserves readable depth between the guarded foreground and the repair area without specifying an unsupported fixture.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Rock-band posters cover the room, with a large pair of pliers within reach and a large speaker in the repair area. The camper is under repair and has extensive damage to its engine and brakes. 현우: He has awakened with his injured leg bandaged and visible treatment marks elsewhere on his body. The contact card remains concealed in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 태진의 뒷모습을 날카롭게 노려보며 바닥의 커다란 펜치를 향해 손을 뻗은 현우의 굳은 상체.\n\nLOCATION (lock): In the resting corner inside an auto-repair warehouse, with band posters nearby and daytime ambient light from the workshop. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Large pliers on the floor beneath 현우's reaching hand in the lower-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Large pliers (On the floor, not yet grasped) — Lie obliquely below 현우's approaching hand; used as Make his defensive intention legible without enlarging the tool beyond realistic scale; Damaged camper (Under repair by 태진) — A partial side section appears beside 태진 in the background; used as Anchors her activity and the depth separating the characters; Rock-band posters (Attached to the room's walls) — Printed band imagery faces into the room and is visible obliquely, without invented readable titles; used as Gives the unfamiliar space its scripted visual identity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral ambient interior light preserves readable depth between the guarded foreground and the repair area without specifying an unsupported fixture.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Rock-band posters cover the room, with a large pair of pliers within reach and a large speaker in the repair area. The camper is under repair and has extensive damage to its engine and brakes. 현우: He has awakened with his injured leg bandaged and visible treatment marks elsewhere on his body. The contact card remains concealed in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 태진의 뒷모습을 날카롭게 노려보며 바닥의 커다란 펜치를 향해 손을 뻗은 현우의 굳은 상체.\n\nLOCATION (lock): In the resting corner inside an auto-repair warehouse, with band posters nearby and daytime ambient light from the workshop. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Large pliers on the floor beneath 현우's reaching hand in the lower-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Large pliers (On the floor, not yet grasped) — Lie obliquely below 현우's approaching hand; used as Make his defensive intention legible without enlarging the tool beyond realistic scale; Damaged camper (Under repair by 태진) — A partial side section appears beside 태진 in the background; used as Anchors her activity and the depth separating the characters; Rock-band posters (Attached to the room's walls) — Printed band imagery faces into the room and is visible obliquely, without invented readable titles; used as Gives the unfamiliar space its scripted visual identity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral ambient interior light preserves readable depth between the guarded foreground and the repair area without specifying an unsupported fixture.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Rock-band posters cover the room, with a large pair of pliers within reach and a large speaker in the repair area. The camper is under repair and has extensive damage to its engine and brakes. 현우: He has awakened with his injured leg bandaged and visible treatment marks elsewhere on his body. The contact card remains concealed in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S50sh4__bgfirst_bg.png",
     "asset_id": "7a6d7074-e7b2-4b88-bab5-afa565d34ff1",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S50sh4.png",
     "asset_id": "36a68918-70ef-4f62-a881-1c4458c88531",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L209B01.png",
     "asset_id": "f0f3a2cb-dd51-4ebd-a380-2b0c867d50dd",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선이 배경에서 작업 중인 태진의 뒷모습을 정확히 향하고 있으며, 오른손은 바닥에 놓인 커다란 펜치를 향해 뻗어 있음.",
    "built_space": "참조된 차고의 구조를 잘 따름. 화면 좌측의 간이 침대, 벽면의 포스터, 우측의 손상된 캠핑카 등 모든 요소가 적절한 위치에 배치됨.",
    "entities": "현우의 외모, 헤어스타일, 붕대 감은 다리가 참조와 일치함. 태진의 뒷모습, 캠핑카, 바닥의 펜치, 포스터 모두 프롬프트대로 묘사됨.",
    "hard_violations": [],
    "physics": "현우가 간이 침대에 걸터앉아 몸을 숙이는 자세가 자연스러우며, 바닥에 놓인 펜치 등 모든 사물의 물리적 상태가 안정적임."
   },
   {
    "label": "B",
    "direction": "현우의 시선이 카메라 우측을 향하고 있어 뒤에 있는 태진을 보지 않음. 손은 펜치에 닿아 있으나 쥐는 형태가 아님.",
    "built_space": "차고 내부의 침대, 포스터, 캠핑카 등 공간적 배경 요소들은 참조 이미지와 일치하게 배치됨.",
    "entities": "현우의 얼굴과 붕대 감은 다리는 일치하나, 펜치가 아직 잡히지 않은 바닥 상태가 아님.",
    "hard_violations": [
     "[gemini-pro] 지탱하는 힘 없이 허공에 떠 있는 펜치(물리적 불가능)",
     "[gemini-pro] 바닥에 놓여 있어야 할 펜치의 위치 오류",
     "[gpt-high] 왼쪽 벽 포스터의 ‘다시, 앞으로’ 등 한글 문구가 읽혀, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
    ],
    "physics": "커다란 펜치가 손에 제대로 쥐어지지 않은 채 수평으로 허공에 떠 있어 심각한 물리적 오류가 발생함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "바닥에 놓인 펜치, 태진의 등을 향한 날카로운 시선, 손을 뻗는 동작 등 프롬프트의 지시를 매우 충실하게 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "펜치가 허공에 떠 있는 심각한 물리적 오류가 있으며, 시선 방향도 프롬프트와 불일치함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선이 배경에서 작업 중인 태진의 뒷모습을 정확히 향하고 있으며, 오른손은 바닥에 놓인 커다란 펜치를 향해 뻗어 있음.",
        "built_space": "참조된 차고의 구조를 잘 따름. 화면 좌측의 간이 침대, 벽면의 포스터, 우측의 손상된 캠핑카 등 모든 요소가 적절한 위치에 배치됨.",
        "entities": "현우의 외모, 헤어스타일, 붕대 감은 다리가 참조와 일치함. 태진의 뒷모습, 캠핑카, 바닥의 펜치, 포스터 모두 프롬프트대로 묘사됨.",
        "hard_violations": [],
        "physics": "현우가 간이 침대에 걸터앉아 몸을 숙이는 자세가 자연스러우며, 바닥에 놓인 펜치 등 모든 사물의 물리적 상태가 안정적임."
       },
       {
        "label": "B",
        "direction": "현우의 시선이 카메라 우측을 향하고 있어 뒤에 있는 태진을 보지 않음. 손은 펜치에 닿아 있으나 쥐는 형태가 아님.",
        "built_space": "차고 내부의 침대, 포스터, 캠핑카 등 공간적 배경 요소들은 참조 이미지와 일치하게 배치됨.",
        "entities": "현우의 얼굴과 붕대 감은 다리는 일치하나, 펜치가 아직 잡히지 않은 바닥 상태가 아님.",
        "hard_violations": [
         "지탱하는 힘 없이 허공에 떠 있는 펜치(물리적 불가능)",
         "바닥에 놓여 있어야 할 펜치의 위치 오류"
        ],
        "physics": "커다란 펜치가 손에 제대로 쥐어지지 않은 채 수평으로 허공에 떠 있어 심각한 물리적 오류가 발생함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "바닥에 놓인 펜치, 태진의 등을 향한 날카로운 시선, 손을 뻗는 동작 등 프롬프트의 지시를 매우 충실하게 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "펜치가 허공에 떠 있는 심각한 물리적 오류가 있으며, 시선 방향도 프롬프트와 불일치함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선이 배경에서 작업 중인 태진의 뒷모습을 정확히 향하고 있으며, 오른손은 바닥에 놓인 커다란 펜치를 향해 뻗어 있음.",
        "built_space": "참조된 차고의 구조를 잘 따름. 화면 좌측의 간이 침대, 벽면의 포스터, 우측의 손상된 캠핑카 등 모든 요소가 적절한 위치에 배치됨.",
        "entities": "현우의 외모, 헤어스타일, 붕대 감은 다리가 참조와 일치함. 태진의 뒷모습, 캠핑카, 바닥의 펜치, 포스터 모두 프롬프트대로 묘사됨.",
        "hard_violations": [],
        "physics": "현우가 간이 침대에 걸터앉아 몸을 숙이는 자세가 자연스러우며, 바닥에 놓인 펜치 등 모든 사물의 물리적 상태가 안정적임."
       },
       {
        "label": "B",
        "direction": "현우의 시선이 카메라 우측을 향하고 있어 뒤에 있는 태진을 보지 않음. 손은 펜치에 닿아 있으나 쥐는 형태가 아님.",
        "built_space": "차고 내부의 침대, 포스터, 캠핑카 등 공간적 배경 요소들은 참조 이미지와 일치하게 배치됨.",
        "entities": "현우의 얼굴과 붕대 감은 다리는 일치하나, 펜치가 아직 잡히지 않은 바닥 상태가 아님.",
        "hard_violations": [
         "지탱하는 힘 없이 허공에 떠 있는 펜치(물리적 불가능)",
         "바닥에 놓여 있어야 할 펜치의 위치 오류"
        ],
        "physics": "커다란 펜치가 손에 제대로 쥐어지지 않은 채 수평으로 허공에 떠 있어 심각한 물리적 오류가 발생함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "읽히는 벽면 문구가 금지 조건을 위반하며, 펜치를 이미 들어 올려 ‘바닥의 펜치에 손을 뻗는 순간’을 놓치고 태진의 등을 노리는 시선도 불명확하다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "태진의 등으로 향한 시선과 바닥의 펜치에 닿기 전 손동작은 정확하지만, 상체 중심 미디엄 숏보다 넓고 현우의 의상 및 태진의 여성 설정은 맞지 않는다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 눈은 화면 오른쪽을 향하지만 뒤쪽 정비 인물의 등으로 향한다고 보기에는 얼굴과 시선이 카메라 쪽으로 너무 열려 있다. 정비 인물은 캠퍼를 향해 몸을 숙여 측면이 보인다. 현우의 오른손은 펜치 손잡이를 이미 잡고 있어 바닥의 도구로 접근하는 손이 아니다.",
        "built_space": "왼쪽 야전침대 한 개와 열린 출입구 한 개, 뒤쪽 창 한 개·작업대 한 개·공구판 한 개, 오른쪽 선반 위 대형 스피커 한 개와 캠퍼 한 대가 보인다. 콘크리트 벽과 바닥, 천장 보 및 배치는 장소 참조와 대체로 일치한다. 현우는 침대 앞에 서서 숙이고 정비 인물은 캠퍼 앞부분 옆에 있다. 캠퍼가 오른쪽 화면을 크게 차지하며 현우도 정강이까지 보여 상체 중심 미디엄 숏보다 넓다.",
        "entities": "두 사람이 보인다. 현우는 앳된 동아시아계 남성으로 검은 머리와 남색 티셔츠가 참조에 가깝고, 정강이 붕대와 팔의 상처가 보인다. 태진에 해당하는 배경 인물은 짧은 머리와 체격이 남성적으로 보이며 여성 설정과 부합하는지 불명확하다. 붉은 손잡이 도구는 참조의 펜치보다 긴 볼트커터 형태이고 바닥에 놓여 있지 않다. 캠퍼 외판의 심한 부식은 보이지만 엔진과 브레이크 손상은 확인되지 않는다. 벽에는 포스터가 있으나 일부 한글 문구가 뚜렷하게 읽힌다. 신발 속 연락처 카드는 노출되지 않는다.",
        "hard_violations": [
         "왼쪽 벽 포스터의 ‘다시, 앞으로’ 등 한글 문구가 읽혀, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
        ],
        "physics": "현우는 바닥에 닿은 뒤쪽 신발과 화면 아래로 이어지는 앞다리로 몸을 지탱하며 앞으로 숙인다. 오른손이 도구 손잡이를 감싸 들어 올리고 있어 도구가 무지지 상태로 떠 있는 것은 아니다. 다만 이 지지 방식 자체가 아직 잡지 않은 바닥의 펜치라는 지정 순간과 다르다. 배경 인물은 두 발로 서 있고 캠퍼와 침대도 바닥에 지지되어 있다."
       },
       {
        "label": "B",
        "direction": "현우는 고개를 오른쪽 뒤로 돌려 캠퍼를 정비하는 인물의 등을 바라본다. 배경 인물은 현우에게 등을 보이고 열린 차량 앞부분을 내려다본다. 현우의 뻗은 손 아래에는 붉은 손잡이 펜치가 대각선으로 놓여 있으며, 손가락은 도구에 아직 닿지 않았다. 경계하는 시선과 도구를 향하는 손의 대상이 모두 명확하다.",
        "built_space": "왼쪽 야전침대 한 개와 출입구 한 개, 뒤쪽 창 한 개·작업대 한 개·공구판 한 개, 오른쪽 선반 위 대형 스피커 한 개와 캠퍼 한 대가 보인다. 현우는 침대 가장자리에 앉고 태진에 해당하는 인물은 캠퍼의 열린 앞부분 옆에 서 있어 휴식 공간과 수리 공간의 깊이가 성립한다. 장소의 주요 재료와 설비 배치는 참조에 가깝다. 펜치는 전경 하단 중앙에 있고 캠퍼는 오른쪽에 부분적으로 보인다. 다만 현우의 정강이와 넓은 작업장 바닥까지 포함해 요청한 상체 중심 미디엄 숏보다 넓다.",
        "entities": "두 사람이 보인다. 현우는 검은 헝클어진 머리의 젊은 동아시아계 남성으로 보이지만 얼굴이 옆으로 돌아 정확한 동일인 여부는 제한적으로만 확인된다. 회갈색 티셔츠는 참조의 남색 티셔츠와 다르다. 정강이 붕대와 팔의 상처는 보인다. 배경 정비 인물은 짧은 머리의 남성으로 보여 태진을 여성으로 지칭한 설정과 어긋난다. 바닥에는 현실적인 크기의 붉은 손잡이 펜치 한 개가 있고, 손상된 캠퍼와 열린 엔진 구역, 대형 스피커, 밴드 공연 이미지 포스터가 보인다. 브레이크 손상은 확인되지 않으며 판독 가능한 글자나 노출된 연락처 카드는 없다.",
        "hard_violations": [],
        "physics": "현우의 엉덩이는 침대에 놓이고 왼손도 침대 쪽을 짚어, 상체를 비틀면서 오른손을 내미는 자세를 지지한다. 뻗은 팔은 어깨에서 자연스럽게 이어지고 펜치는 콘크리트 바닥에 완전히 놓여 있다. 정비 인물은 양발을 바닥에 둔 채 차량 앞부분으로 상체를 기울인다. 캠퍼의 열린 덮개는 차량에 연결되어 있으며 지지 없이 떠 있는 인체나 도구는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "읽히는 벽면 문구가 금지 조건을 위반하며, 펜치를 이미 들어 올려 ‘바닥의 펜치에 손을 뻗는 순간’을 놓치고 태진의 등을 노리는 시선도 불명확하다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "태진의 등으로 향한 시선과 바닥의 펜치에 닿기 전 손동작은 정확하지만, 상체 중심 미디엄 숏보다 넓고 현우의 의상 및 태진의 여성 설정은 맞지 않는다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 눈은 화면 오른쪽을 향하지만 뒤쪽 정비 인물의 등으로 향한다고 보기에는 얼굴과 시선이 카메라 쪽으로 너무 열려 있다. 정비 인물은 캠퍼를 향해 몸을 숙여 측면이 보인다. 현우의 오른손은 펜치 손잡이를 이미 잡고 있어 바닥의 도구로 접근하는 손이 아니다.",
        "built_space": "왼쪽 야전침대 한 개와 열린 출입구 한 개, 뒤쪽 창 한 개·작업대 한 개·공구판 한 개, 오른쪽 선반 위 대형 스피커 한 개와 캠퍼 한 대가 보인다. 콘크리트 벽과 바닥, 천장 보 및 배치는 장소 참조와 대체로 일치한다. 현우는 침대 앞에 서서 숙이고 정비 인물은 캠퍼 앞부분 옆에 있다. 캠퍼가 오른쪽 화면을 크게 차지하며 현우도 정강이까지 보여 상체 중심 미디엄 숏보다 넓다.",
        "entities": "두 사람이 보인다. 현우는 앳된 동아시아계 남성으로 검은 머리와 남색 티셔츠가 참조에 가깝고, 정강이 붕대와 팔의 상처가 보인다. 태진에 해당하는 배경 인물은 짧은 머리와 체격이 남성적으로 보이며 여성 설정과 부합하는지 불명확하다. 붉은 손잡이 도구는 참조의 펜치보다 긴 볼트커터 형태이고 바닥에 놓여 있지 않다. 캠퍼 외판의 심한 부식은 보이지만 엔진과 브레이크 손상은 확인되지 않는다. 벽에는 포스터가 있으나 일부 한글 문구가 뚜렷하게 읽힌다. 신발 속 연락처 카드는 노출되지 않는다.",
        "hard_violations": [
         "왼쪽 벽 포스터의 ‘다시, 앞으로’ 등 한글 문구가 읽혀, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
        ],
        "physics": "현우는 바닥에 닿은 뒤쪽 신발과 화면 아래로 이어지는 앞다리로 몸을 지탱하며 앞으로 숙인다. 오른손이 도구 손잡이를 감싸 들어 올리고 있어 도구가 무지지 상태로 떠 있는 것은 아니다. 다만 이 지지 방식 자체가 아직 잡지 않은 바닥의 펜치라는 지정 순간과 다르다. 배경 인물은 두 발로 서 있고 캠퍼와 침대도 바닥에 지지되어 있다."
       },
       {
        "label": "A",
        "direction": "현우는 고개를 오른쪽 뒤로 돌려 캠퍼를 정비하는 인물의 등을 바라본다. 배경 인물은 현우에게 등을 보이고 열린 차량 앞부분을 내려다본다. 현우의 뻗은 손 아래에는 붉은 손잡이 펜치가 대각선으로 놓여 있으며, 손가락은 도구에 아직 닿지 않았다. 경계하는 시선과 도구를 향하는 손의 대상이 모두 명확하다.",
        "built_space": "왼쪽 야전침대 한 개와 출입구 한 개, 뒤쪽 창 한 개·작업대 한 개·공구판 한 개, 오른쪽 선반 위 대형 스피커 한 개와 캠퍼 한 대가 보인다. 현우는 침대 가장자리에 앉고 태진에 해당하는 인물은 캠퍼의 열린 앞부분 옆에 서 있어 휴식 공간과 수리 공간의 깊이가 성립한다. 장소의 주요 재료와 설비 배치는 참조에 가깝다. 펜치는 전경 하단 중앙에 있고 캠퍼는 오른쪽에 부분적으로 보인다. 다만 현우의 정강이와 넓은 작업장 바닥까지 포함해 요청한 상체 중심 미디엄 숏보다 넓다.",
        "entities": "두 사람이 보인다. 현우는 검은 헝클어진 머리의 젊은 동아시아계 남성으로 보이지만 얼굴이 옆으로 돌아 정확한 동일인 여부는 제한적으로만 확인된다. 회갈색 티셔츠는 참조의 남색 티셔츠와 다르다. 정강이 붕대와 팔의 상처는 보인다. 배경 정비 인물은 짧은 머리의 남성으로 보여 태진을 여성으로 지칭한 설정과 어긋난다. 바닥에는 현실적인 크기의 붉은 손잡이 펜치 한 개가 있고, 손상된 캠퍼와 열린 엔진 구역, 대형 스피커, 밴드 공연 이미지 포스터가 보인다. 브레이크 손상은 확인되지 않으며 판독 가능한 글자나 노출된 연락처 카드는 없다.",
        "hard_violations": [],
        "physics": "현우의 엉덩이는 침대에 놓이고 왼손도 침대 쪽을 짚어, 상체를 비틀면서 오른손을 내미는 자세를 지지한다. 뻗은 팔은 어깨에서 자연스럽게 이어지고 펜치는 콘크리트 바닥에 완전히 놓여 있다. 정비 인물은 양발을 바닥에 둔 채 차량 앞부분으로 상체를 기울인다. 캠퍼의 열린 덮개는 차량에 연결되어 있으며 지지 없이 떠 있는 인체나 도구는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.661
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.411
   },
   "violations": {
    "B": [
     "[gemini-pro] 지탱하는 힘 없이 허공에 떠 있는 펜치(물리적 불가능)",
     "[gemini-pro] 바닥에 놓여 있어야 할 펜치의 위치 오류",
     "[gpt-high] 왼쪽 벽 포스터의 ‘다시, 앞으로’ 등 한글 문구가 읽혀, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 411
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "바닥에 놓인 펜치, 태진의 등을 향한 날카로운 시선, 손을 뻗는 동작 등 프롬프트의 지시를 매우 충실하게 구현함."
   },
   {
    "label": "B",
    "score": 411,
    "verdict_ko": "펜치가 허공에 떠 있는 심각한 물리적 오류가 있으며, 시선 방향도 프롬프트와 불일치함.  ★위반: [gemini-pro] 지탱하는 힘 없이 허공에 떠 있는 펜치(물리적 불가능) / [gemini-pro] 바닥에 놓여 있어야 할 펜치의 위치 오류 / [gpt-high] 왼쪽 벽 포스터의 ‘다시, 앞으로’ 등 한글 문구가 읽혀, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L209B01.png",
    "asset_id": "f0f3a2cb-dd51-4ebd-a380-2b0c867d50dd",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-34db-7749-af20-2fa5459c84ad",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S50sh4__bgfirst_bg.png",
   "bg_asset_id": "7a6d7074-e7b2-4b88-bab5-afa565d34ff1",
   "bg_record_key": "S50sh4::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S50sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:00:25.719041+00:00",
  "fingerprint": "f45f815ff501fd7cf504615e76cdf852645ac81c1a40489dd8b6d216705021f4",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S50sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S50sh4_sel.png",
  "source_sha256": "de4bb04c4d02fde1696897fbe722eeaeb7ceb63af64a5ed598fd3076d9761cad",
  "file": "S50sh4_cine.png",
  "staged_sha256": "cbc38c9ac4ffae7a592e00fe2e9a8e1b02fe1337877331666fb2bcaa014ab94e",
  "latency_ms": 13270
 },
 "S50sh9::signage": {
  "fp": "1a15b76936124854",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S50sh9": {
  "input_fingerprint": "b0210eef3d4642a8",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 완전히 망가진 캠핑카를 가리키며 미간을 찌푸린 태진의 상체.\n\nLOCATION (lock): Beside the damaged camper inside the auto-repair workshop bay, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Damaged camper beside 태진, receiving her pointing gesture in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Damaged camper (Under repair, with major engine and brake damage described in dialogue) — Only the side section beside 태진 is visible; hidden mechanical damage is not illustrated as a cutaway; used as Receives her pointing gesture and supplies evidence of the subject under discussion.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the warehouse's neutral ambient illumination and controlled contrast across 태진 and the adjacent camper.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper remains under repair with extensive engine and brake damage, and repair fluid has just been poured into its filler opening. Rock-band posters, the large speaker and the large pliers remain in the workshop space. 태진: She has short bobbed hair and holds the repair-fluid bottle after pouring from it at the camper.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 태진 (한국인 여성, 30세, 검은 머리, 짧은 단발) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 완전히 망가진 캠핑카를 가리키며 미간을 찌푸린 태진의 상체.\n\nLOCATION (lock): Beside the damaged camper inside the auto-repair workshop bay, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Damaged camper beside 태진, receiving her pointing gesture in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Damaged camper (Under repair, with major engine and brake damage described in dialogue) — Only the side section beside 태진 is visible; hidden mechanical damage is not illustrated as a cutaway; used as Receives her pointing gesture and supplies evidence of the subject under discussion.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the warehouse's neutral ambient illumination and controlled contrast across 태진 and the adjacent camper.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper remains under repair with extensive engine and brake damage, and repair fluid has just been poured into its filler opening. Rock-band posters, the large speaker and the large pliers remain in the workshop space. 태진: She has short bobbed hair and holds the repair-fluid bottle after pouring from it at the camper.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 태진 (한국인 여성, 30세, 검은 머리, 짧은 단발) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 완전히 망가진 캠핑카를 가리키며 미간을 찌푸린 태진의 상체.\n\nLOCATION (lock): Beside the damaged camper inside the auto-repair workshop bay, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Damaged camper beside 태진, receiving her pointing gesture in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Damaged camper (Under repair, with major engine and brake damage described in dialogue) — Only the side section beside 태진 is visible; hidden mechanical damage is not illustrated as a cutaway; used as Receives her pointing gesture and supplies evidence of the subject under discussion.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the warehouse's neutral ambient illumination and controlled contrast across 태진 and the adjacent camper.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper remains under repair with extensive engine and brake damage, and repair fluid has just been poured into its filler opening. Rock-band posters, the large speaker and the large pliers remain in the workshop space. 태진: She has short bobbed hair and holds the repair-fluid bottle after pouring from it at the camper.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 태진 (한국인 여성, 30세, 검은 머리, 짧은 단발) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "시선은 정면 카메라를 향하고, 오른팔을 뻗어 우측의 망가진 캠핑카를 가리킴.",
    "built_space": "수리소 내부. 벽의 포스터와 배경 선반의 스피커는 유지되었으나, 바닥에 있어야 할 대형 플라이어가 보이지 않음.",
    "entities": "태진(단발머리, 남색 상의)이 미간을 찌푸린 채 왼손에 수리용 액체 병을 들고 있음. 병에 읽을 수 있는 영문 글씨가 있음.",
    "hard_violations": [
     "[gpt-high] 수리액 병 라벨에 읽을 수 있는 영문이 노출되어, 이미지 어디에도 읽히는 글자를 두지 말라는 명시적 금지 조건을 위반한다."
    ],
    "physics": "서 있는 자세에서 하체와 상체가 안정적으로 지지되며, 병을 쥔 손과 뻗은 팔 모두 물리적으로 자연스러움."
   },
   {
    "label": "B",
    "direction": "시선은 정면을 향하며, 오른손으로 우측 캠핑카를 가리킴.",
    "built_space": "수리소 내부. 포스터와 스피커는 배치되었으나 플라이어는 누락됨. 캠핑카 바퀴 부분이 금지된 내부 단면(cutaway) 형태로 노출됨.",
    "entities": "태진이 쪼그려 앉은 자세로 작은 병을 들고 있으며, 오른손의 형태가 심하게 일그러짐.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 오른손 손가락 해부학 구조 (physically impossible anatomy)"
    ],
    "physics": "쪼그려 앉은 하체의 지지는 확인되나, 가리키는 오른손의 뼈대와 관절 구조가 물리적으로 불가능함. 절단된 휠 하우스 안의 전선들이 허공에 매달려 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지시된 상반신 미디엄 샷, 찌푸린 표정, 캠핑카를 가리키는 동작을 정확히 구현했으나, 바닥의 플라이어가 누락되고 병 표면에 금지된 글씨가 노출되어 감점됨."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "상반신 샷 지시를 어기고 쪼그려 앉은 자세를 취했으며, 명시적으로 금지된 차량 내부 기계 단면(cutaway)이 노출되고 심각한 손가락 해부학 오류가 발생함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 정면 카메라를 향하고, 오른팔을 뻗어 우측의 망가진 캠핑카를 가리킴.",
        "built_space": "수리소 내부. 벽의 포스터와 배경 선반의 스피커는 유지되었으나, 바닥에 있어야 할 대형 플라이어가 보이지 않음.",
        "entities": "태진(단발머리, 남색 상의)이 미간을 찌푸린 채 왼손에 수리용 액체 병을 들고 있음. 병에 읽을 수 있는 영문 글씨가 있음.",
        "hard_violations": [],
        "physics": "서 있는 자세에서 하체와 상체가 안정적으로 지지되며, 병을 쥔 손과 뻗은 팔 모두 물리적으로 자연스러움."
       },
       {
        "label": "B",
        "direction": "시선은 정면을 향하며, 오른손으로 우측 캠핑카를 가리킴.",
        "built_space": "수리소 내부. 포스터와 스피커는 배치되었으나 플라이어는 누락됨. 캠핑카 바퀴 부분이 금지된 내부 단면(cutaway) 형태로 노출됨.",
        "entities": "태진이 쪼그려 앉은 자세로 작은 병을 들고 있으며, 오른손의 형태가 심하게 일그러짐.",
        "hard_violations": [
         "물리적으로 불가능한 오른손 손가락 해부학 구조 (physically impossible anatomy)"
        ],
        "physics": "쪼그려 앉은 하체의 지지는 확인되나, 가리키는 오른손의 뼈대와 관절 구조가 물리적으로 불가능함. 절단된 휠 하우스 안의 전선들이 허공에 매달려 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지시된 상반신 미디엄 샷, 찌푸린 표정, 캠핑카를 가리키는 동작을 정확히 구현했으나, 바닥의 플라이어가 누락되고 병 표면에 금지된 글씨가 노출되어 감점됨."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "상반신 샷 지시를 어기고 쪼그려 앉은 자세를 취했으며, 명시적으로 금지된 차량 내부 기계 단면(cutaway)이 노출되고 심각한 손가락 해부학 오류가 발생함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 정면 카메라를 향하고, 오른팔을 뻗어 우측의 망가진 캠핑카를 가리킴.",
        "built_space": "수리소 내부. 벽의 포스터와 배경 선반의 스피커는 유지되었으나, 바닥에 있어야 할 대형 플라이어가 보이지 않음.",
        "entities": "태진(단발머리, 남색 상의)이 미간을 찌푸린 채 왼손에 수리용 액체 병을 들고 있음. 병에 읽을 수 있는 영문 글씨가 있음.",
        "hard_violations": [],
        "physics": "서 있는 자세에서 하체와 상체가 안정적으로 지지되며, 병을 쥔 손과 뻗은 팔 모두 물리적으로 자연스러움."
       },
       {
        "label": "B",
        "direction": "시선은 정면을 향하며, 오른손으로 우측 캠핑카를 가리킴.",
        "built_space": "수리소 내부. 포스터와 스피커는 배치되었으나 플라이어는 누락됨. 캠핑카 바퀴 부분이 금지된 내부 단면(cutaway) 형태로 노출됨.",
        "entities": "태진이 쪼그려 앉은 자세로 작은 병을 들고 있으며, 오른손의 형태가 심하게 일그러짐.",
        "hard_violations": [
         "물리적으로 불가능한 오른손 손가락 해부학 구조 (physically impossible anatomy)"
        ],
        "physics": "쪼그려 앉은 하체의 지지는 확인되나, 가리키는 오른손의 뼈대와 관절 구조가 물리적으로 불가능함. 절단된 휠 하우스 안의 전선들이 허공에 매달려 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "상체 미디엄 숏과 찌푸린 미간, 캠핑카를 가리키는 동작을 구현했으나 반팔 의상과 전경으로 크게 나온 차체는 지시와 다르다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "차체를 정확히 가리키는 동작과 참조의 긴소매 의상은 더 충실하지만, 병 라벨의 읽히는 영문이 문자 금지 조건을 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "태진의 검지는 화면 오른쪽을 향하며, 연장선은 캠핑카 측면에 닿는다. 따라서 지목 대상은 캠핑카로 읽힌다. 눈은 차체가 아니라 카메라 쪽을 보고 있으며 미간은 뚜렷하게 찌푸려져 있다.",
        "built_space": "왼쪽 벽의 여러 포스터, 뒤쪽 창 하나, 작업대 하나, 공구판 하나, 오른쪽 수납 선반 위 대형 스피커 하나가 보인다. 콘크리트 벽과 철제 천장, 낮의 확산광은 장소 참조와 대체로 이어진다. 태진은 왼쪽, 캠핑카 측면은 오른쪽에 있지만 차체와 빈 휠하우스가 중경보다 전경을 크게 점유한다. 중복된 고정 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "인물은 검은 짧은 단발의 성인 여성 한 명뿐이며, 얼굴과 체형은 태진 참조에 대체로 부합한다. 남색 상의는 맞지만 참조의 긴소매 니트와 달리 반팔이다. 흰 차체의 회색 줄무늬와 심한 부식은 참조 캠핑카를 따른다. 열린 수리액 병을 손에 들고 있다. 포스터와 대형 스피커는 보이지만 대형 플라이어는 이 구도에서 확인되지 않는다. 엔진 내부 절개도는 없으나 바퀴가 빠진 휠하우스와 늘어진 선이 추가로 드러난다. 병 라벨의 작은 인쇄는 명확히 읽히지 않는다.",
        "hard_violations": [],
        "physics": "가리키는 팔은 어깨와 팔꿈치에서 자연스럽게 이어지고, 다른 손은 병 몸통을 감싸 지지한다. 하체는 프레임 밖이므로 발의 접촉은 확인할 수 없지만 상체가 공중에 떠 있는 징후는 없다. 캠핑카에는 지면에 닿은 다른 바퀴가 보이며, 휠하우스의 선은 차체 안쪽에서 매달려 있다. 지지 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "태진은 팔을 오른쪽으로 뻗어 검지 끝으로 캠핑카 회색 줄무늬 부근의 녹슨 손상 부분을 정확히 가리킨다. 지목 대상이 A보다 직접적이고 분명하다. 미간을 찌푸린 채 시선은 캠핑카가 아닌 카메라 쪽을 향한다.",
        "built_space": "왼쪽 포스터 벽, 뒤쪽 창 하나와 작업대 하나, 공구판 하나, 수납 선반 위 대형 스피커 하나가 보인다. 참조의 콘크리트 정비소와 낮 조명이 대체로 유지된다. 태진은 화면 왼쪽에서 차량 옆에 서 있고, 캠핑카 측면은 중앙 오른쪽부터 화면 끝까지 이어진다. 차체가 중경에만 머무르지 않고 오른쪽 전경까지 크게 차지한다. 설비 중복이나 불가능한 반사는 없다.",
        "entities": "검은 단발의 성인 여성 한 명만 등장하며 얼굴, 체형, 남색 긴소매 니트가 태진 참조에 가깝다. 흰색과 회색 줄무늬의 부식된 캠핑카 측면이 보이고 숨은 엔진·브레이크 손상을 절개도로 표현하지 않았다. 태진은 뚜껑이 열린 수리액 용기의 손잡이를 잡고 있다. 포스터와 스피커는 보이며 대형 플라이어는 프레임에서 확인되지 않는다. 병의 파란 라벨에는 읽을 수 있는 영문 제품 표기가 노출되어 있다.",
        "hard_violations": [
         "수리액 병 라벨에 읽을 수 있는 영문이 노출되어, 이미지 어디에도 읽히는 글자를 두지 말라는 명시적 금지 조건을 위반한다."
        ],
        "physics": "뻗은 팔과 검지는 해부학적으로 자연스럽게 연결되어 있으며 차체 손상 부분을 가리키는 자세가 가능하다. 반대 손은 용기 손잡이를 움켜쥐어 무게를 지지한다. 발은 잘렸지만 몸통은 정상적인 서 있는 자세이고, 보이는 차량 바퀴는 바닥에 닿아 있다. 지지 없이 떠 있는 신체나 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "상체 미디엄 숏과 찌푸린 미간, 캠핑카를 가리키는 동작을 구현했으나 반팔 의상과 전경으로 크게 나온 차체는 지시와 다르다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "차체를 정확히 가리키는 동작과 참조의 긴소매 의상은 더 충실하지만, 병 라벨의 읽히는 영문이 문자 금지 조건을 위반한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "태진의 검지는 화면 오른쪽을 향하며, 연장선은 캠핑카 측면에 닿는다. 따라서 지목 대상은 캠핑카로 읽힌다. 눈은 차체가 아니라 카메라 쪽을 보고 있으며 미간은 뚜렷하게 찌푸려져 있다.",
        "built_space": "왼쪽 벽의 여러 포스터, 뒤쪽 창 하나, 작업대 하나, 공구판 하나, 오른쪽 수납 선반 위 대형 스피커 하나가 보인다. 콘크리트 벽과 철제 천장, 낮의 확산광은 장소 참조와 대체로 이어진다. 태진은 왼쪽, 캠핑카 측면은 오른쪽에 있지만 차체와 빈 휠하우스가 중경보다 전경을 크게 점유한다. 중복된 고정 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "인물은 검은 짧은 단발의 성인 여성 한 명뿐이며, 얼굴과 체형은 태진 참조에 대체로 부합한다. 남색 상의는 맞지만 참조의 긴소매 니트와 달리 반팔이다. 흰 차체의 회색 줄무늬와 심한 부식은 참조 캠핑카를 따른다. 열린 수리액 병을 손에 들고 있다. 포스터와 대형 스피커는 보이지만 대형 플라이어는 이 구도에서 확인되지 않는다. 엔진 내부 절개도는 없으나 바퀴가 빠진 휠하우스와 늘어진 선이 추가로 드러난다. 병 라벨의 작은 인쇄는 명확히 읽히지 않는다.",
        "hard_violations": [],
        "physics": "가리키는 팔은 어깨와 팔꿈치에서 자연스럽게 이어지고, 다른 손은 병 몸통을 감싸 지지한다. 하체는 프레임 밖이므로 발의 접촉은 확인할 수 없지만 상체가 공중에 떠 있는 징후는 없다. 캠핑카에는 지면에 닿은 다른 바퀴가 보이며, 휠하우스의 선은 차체 안쪽에서 매달려 있다. 지지 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "태진은 팔을 오른쪽으로 뻗어 검지 끝으로 캠핑카 회색 줄무늬 부근의 녹슨 손상 부분을 정확히 가리킨다. 지목 대상이 A보다 직접적이고 분명하다. 미간을 찌푸린 채 시선은 캠핑카가 아닌 카메라 쪽을 향한다.",
        "built_space": "왼쪽 포스터 벽, 뒤쪽 창 하나와 작업대 하나, 공구판 하나, 수납 선반 위 대형 스피커 하나가 보인다. 참조의 콘크리트 정비소와 낮 조명이 대체로 유지된다. 태진은 화면 왼쪽에서 차량 옆에 서 있고, 캠핑카 측면은 중앙 오른쪽부터 화면 끝까지 이어진다. 차체가 중경에만 머무르지 않고 오른쪽 전경까지 크게 차지한다. 설비 중복이나 불가능한 반사는 없다.",
        "entities": "검은 단발의 성인 여성 한 명만 등장하며 얼굴, 체형, 남색 긴소매 니트가 태진 참조에 가깝다. 흰색과 회색 줄무늬의 부식된 캠핑카 측면이 보이고 숨은 엔진·브레이크 손상을 절개도로 표현하지 않았다. 태진은 뚜껑이 열린 수리액 용기의 손잡이를 잡고 있다. 포스터와 스피커는 보이며 대형 플라이어는 프레임에서 확인되지 않는다. 병의 파란 라벨에는 읽을 수 있는 영문 제품 표기가 노출되어 있다.",
        "hard_violations": [
         "수리액 병 라벨에 읽을 수 있는 영문이 노출되어, 이미지 어디에도 읽히는 글자를 두지 말라는 명시적 금지 조건을 위반한다."
        ],
        "physics": "뻗은 팔과 검지는 해부학적으로 자연스럽게 연결되어 있으며 차체 손상 부분을 가리키는 자세가 가능하다. 반대 손은 용기 손잡이를 움켜쥐어 무게를 지지한다. 발은 잘렸지만 몸통은 정상적인 서 있는 자세이고, 보이는 차량 바퀴는 바닥에 닿아 있다. 지지 없이 떠 있는 신체나 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.571,
    "B": 1.429
   },
   "adjusted": {
    "A": 1.321,
    "B": 1.179
   },
   "violations": {
    "B": [
     "[gemini-pro] 물리적으로 불가능한 오른손 손가락 해부학 구조 (physically impossible anatomy)"
    ],
    "A": [
     "[gpt-high] 수리액 병 라벨에 읽을 수 있는 영문이 노출되어, 이미지 어디에도 읽히는 글자를 두지 말라는 명시적 금지 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1321,
   "B": 1179
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1321,
    "verdict_ko": "지시된 상반신 미디엄 샷, 찌푸린 표정, 캠핑카를 가리키는 동작을 정확히 구현했으나, 바닥의 플라이어가 누락되고 병 표면에 금지된 글씨가 노출되어 감점됨.  ★위반: [gpt-high] 수리액 병 라벨에 읽을 수 있는 영문이 노출되어, 이미지 어디에도 읽히는 글자를 두지 말라는 명시적 금지 조건을 위반한다."
   },
   {
    "label": "B",
    "score": 1179,
    "verdict_ko": "상반신 샷 지시를 어기고 쪼그려 앉은 자세를 취했으며, 명시적으로 금지된 차량 내부 기계 단면(cutaway)이 노출되고 심각한 손가락 해부학 오류가 발생함.  ★위반: [gemini-pro] 물리적으로 불가능한 오른손 손가락 해부학 구조 (physically impossible anatomy)"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S50sh4_sel.png",
    "asset_id": "3a73ff18-cb78-45b1-921c-e07f7bd2d984",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 태진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1170167>",
    "asset_id": "db7f72e4-1c8f-4ab0-8fe8-4f944e2278d3",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-3830-7739-bab3-11cf70f71e14",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S50sh4"
  }
 },
 "S50sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:01:56.238137+00:00",
  "fingerprint": "427f51723aedd3352977d99bb9b8dac126a01ff0b0e94c1ebf6a9a45896a07d0",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S50sh9_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S50sh9_sel.png",
  "source_sha256": "9deb17ef9553f9823138831c09bc818cf942ddf8ba339251048290fd571e6698",
  "file": "S50sh9_cine.png",
  "staged_sha256": "9c7b54e367791270b90b91d311d70074fbf5ae64e07b6e704496832516f0f0cf",
  "latency_ms": 13474
 },
 "S50sh10::signage": {
  "fp": "91250177605fb4e8",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S50sh10": {
  "input_fingerprint": "06c63504e14cd9c0",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 잔뜩 미간을 찌푸린 채 앰버의 행방을 묻듯 다급하게 입을 연 현우의 얼굴.\n\nLOCATION (lock): At the resting area adjoining the repair bay inside the auto-center warehouse, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Rock-band posters (Attached to the wall behind the seated area) — Their printed faces remain partially visible and softly resolved behind 현우; used as Maintains location continuity without competing with his face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the same neutral interior light and restrained contrast, allowing urgency to register through expression rather than a lighting shift.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the recovery area's interior surfaces, rock-band posters, and the tools already present there. Exclude the repair bay as a substitute for this resting area.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper remains under repair with extensive engine and brake damage and newly added repair fluid. The workshop retains its rock-band posters, large speaker and large pliers. 현우: His injured leg remains bandaged, with visible treatment marks elsewhere on his body. The contact card remains concealed in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 잔뜩 미간을 찌푸린 채 앰버의 행방을 묻듯 다급하게 입을 연 현우의 얼굴.\n\nLOCATION (lock): At the resting area adjoining the repair bay inside the auto-center warehouse, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Rock-band posters (Attached to the wall behind the seated area) — Their printed faces remain partially visible and softly resolved behind 현우; used as Maintains location continuity without competing with his face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the same neutral interior light and restrained contrast, allowing urgency to register through expression rather than a lighting shift.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the recovery area's interior surfaces, rock-band posters, and the tools already present there. Exclude the repair bay as a substitute for this resting area.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper remains under repair with extensive engine and brake damage and newly added repair fluid. The workshop retains its rock-band posters, large speaker and large pliers. 현우: His injured leg remains bandaged, with visible treatment marks elsewhere on his body. The contact card remains concealed in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 잔뜩 미간을 찌푸린 채 앰버의 행방을 묻듯 다급하게 입을 연 현우의 얼굴.\n\nLOCATION (lock): At the resting area adjoining the repair bay inside the auto-center warehouse, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Rock-band posters (Attached to the wall behind the seated area) — Their printed faces remain partially visible and softly resolved behind 현우; used as Maintains location continuity without competing with his face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the same neutral interior light and restrained contrast, allowing urgency to register through expression rather than a lighting shift.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the recovery area's interior surfaces, rock-band posters, and the tools already present there. Exclude the repair bay as a substitute for this resting area.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper remains under repair with extensive engine and brake damage and newly added repair fluid. The workshop retains its rock-band posters, large speaker and large pliers. 현우: His injured leg remains bandaged, with visible treatment marks elsewhere on his body. The contact card remains concealed in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "화면 오른쪽 밖을 향해 시선이 향하며, 다급하게 말하듯 입을 벌리고 있음.",
    "built_space": "정비소 휴게 공간. 왼쪽 벽에 록밴드 포스터, 배경에 공구판, 오른쪽에 스피커가 놓인 선반이 레퍼런스와 동일한 구조로 배치됨.",
    "entities": "현우(앳된 얼굴, 헝클어진 검은 머리)의 외모가 레퍼런스와 일치하며 짙은 남색 티셔츠를 착용함.",
    "hard_violations": [],
    "physics": "화면 밖의 하체에 의해 안정적으로 지지되는 자연스러운 상체 자세."
   },
   {
    "label": "B",
    "direction": "화면 왼쪽 밖을 향해 시선이 향하며 입을 벌리고 있음.",
    "built_space": "정비소 내부. 인물 바로 뒤쪽 벽면에 록밴드 포스터들이 평면적으로 부착됨.",
    "entities": "현우의 외모 특징이 일치함. 붕대를 감은 무릎이 프레임 전면에 노출됨.",
    "hard_violations": [],
    "physics": "클로즈업 샷임에도 무릎이 가슴 앞까지 비정상적으로 높게 당겨져 구도상 어색하게 위치함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "미간을 잔뜩 찌푸린 다급한 표정과 클로즈업 프레이밍을 완벽하게 구현했으며, 레퍼런스의 공간 요소(포스터, 공구판, 스피커)를 입체적으로 잘 살렸습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "요구된 찌푸린 표정 연기가 부족하고, 얼굴 클로즈업 프레임에 무릎을 억지로 포함시켜 구도가 부자연스럽습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "화면 오른쪽 밖을 향해 시선이 향하며, 다급하게 말하듯 입을 벌리고 있음.",
        "built_space": "정비소 휴게 공간. 왼쪽 벽에 록밴드 포스터, 배경에 공구판, 오른쪽에 스피커가 놓인 선반이 레퍼런스와 동일한 구조로 배치됨.",
        "entities": "현우(앳된 얼굴, 헝클어진 검은 머리)의 외모가 레퍼런스와 일치하며 짙은 남색 티셔츠를 착용함.",
        "hard_violations": [],
        "physics": "화면 밖의 하체에 의해 안정적으로 지지되는 자연스러운 상체 자세."
       },
       {
        "label": "B",
        "direction": "화면 왼쪽 밖을 향해 시선이 향하며 입을 벌리고 있음.",
        "built_space": "정비소 내부. 인물 바로 뒤쪽 벽면에 록밴드 포스터들이 평면적으로 부착됨.",
        "entities": "현우의 외모 특징이 일치함. 붕대를 감은 무릎이 프레임 전면에 노출됨.",
        "hard_violations": [],
        "physics": "클로즈업 샷임에도 무릎이 가슴 앞까지 비정상적으로 높게 당겨져 구도상 어색하게 위치함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "미간을 잔뜩 찌푸린 다급한 표정과 클로즈업 프레이밍을 완벽하게 구현했으며, 레퍼런스의 공간 요소(포스터, 공구판, 스피커)를 입체적으로 잘 살렸습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "요구된 찌푸린 표정 연기가 부족하고, 얼굴 클로즈업 프레임에 무릎을 억지로 포함시켜 구도가 부자연스럽습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "화면 오른쪽 밖을 향해 시선이 향하며, 다급하게 말하듯 입을 벌리고 있음.",
        "built_space": "정비소 휴게 공간. 왼쪽 벽에 록밴드 포스터, 배경에 공구판, 오른쪽에 스피커가 놓인 선반이 레퍼런스와 동일한 구조로 배치됨.",
        "entities": "현우(앳된 얼굴, 헝클어진 검은 머리)의 외모가 레퍼런스와 일치하며 짙은 남색 티셔츠를 착용함.",
        "hard_violations": [],
        "physics": "화면 밖의 하체에 의해 안정적으로 지지되는 자연스러운 상체 자세."
       },
       {
        "label": "B",
        "direction": "화면 왼쪽 밖을 향해 시선이 향하며 입을 벌리고 있음.",
        "built_space": "정비소 내부. 인물 바로 뒤쪽 벽면에 록밴드 포스터들이 평면적으로 부착됨.",
        "entities": "현우의 외모 특징이 일치함. 붕대를 감은 무릎이 프레임 전면에 노출됨.",
        "hard_violations": [],
        "physics": "클로즈업 샷임에도 무릎이 가슴 앞까지 비정상적으로 높게 당겨져 구도상 어색하게 위치함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "다급하게 말하는 표정과 부상 흔적은 적절하지만, 무릎과 상체까지 넓힌 구도가 얼굴 중심 클로즈업 지시에서 벗어나고 의상 색도 참조보다 밝다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "얼굴 중심 클로즈업, 깊게 찌푸린 미간과 말을 꺼내는 순간, 참조의 얼굴·남색 의상을 더 충실히 구현했으나 노출된 피부의 치료 흔적은 분명하지 않다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴과 눈은 화면 왼쪽의 카메라 밖 상대를 향하고, 입을 크게 벌려 질문하는 순간처럼 보인다. 상대는 보이지 않으므로 앰버의 행방을 묻는 구체적 내용까지 확인할 수는 없지만 대화 방향은 자연스럽다. 겨누는 물건이나 이동하는 물체는 없다.",
        "built_space": "얼룩진 벽에 가장자리와 머리 뒤의 일부를 포함해 포스터 약 7장이 보이며, 인쇄된 인물들은 흐릿하다. 오른쪽에는 수납 가구와 용기들, 인물 뒤에는 의자의 금속 테두리가 일부 보인다. 휴식 공간에 앉은 배치로 읽히고 정비 차량이 배경을 대신하지 않는다. 다만 무릎까지 들어오면서 지정된 얼굴 클로즈업보다 넓어졌다. 고정 설비의 명백한 중복이나 불가능한 반사는 없다.",
        "entities": "실제 인물은 젊은 동아시아계 남성 한 명뿐이며, 헝클어진 검은 머리와 얼굴 윤곽은 현우 참조에 대체로 부합한다. 국적은 외형만으로 확인할 수 없다. 청색 티셔츠는 참조의 짙은 남색보다 밝다. 뺨의 상처와 무릎 아래 붕대가 보인다. 포스터 속 얼굴은 인쇄물이지 추가 인물이 아니다. 캠핑카·대형 스피커·대형 펜치·신발 속 카드는 이 구도에서 확인되지 않으며, 판독 가능한 글자는 뚜렷하지 않다.",
        "hard_violations": [],
        "physics": "화면 오른쪽 의자 테두리와 몸의 배치로 앉아 있는 자세가 읽힌다. 굽힌 무릎은 아래쪽 화면 밖 다리로 이어지고 몸통과 목의 연결도 자연스럽다. 엉덩이와 발의 접촉점은 잘렸지만 공중에 떠 있다는 징후는 없다. 손에 든 물체나 지지 없이 떠 있는 물체도 없다."
       },
       {
        "label": "B",
        "direction": "얼굴은 약간 화면 오른쪽으로 돌아가 있고 눈은 그쪽 카메라 밖 대화 상대를 향한다. 미간을 깊게 모으고 입을 열어 다급하게 질문을 시작하는 순간으로 읽힌다. 상대 자체는 보이지 않는다. 겨누는 물건이나 이동하는 물체는 없다.",
        "built_space": "낡은 벽의 왼쪽에 주요 포스터 4장과 가장자리의 포스터 일부 1장이 보인다. 오른쪽에는 공구판 1개, 금속 선반 1개, 그 위 대형 스피커 1개가 있다. 이는 참조의 벽 재질과 작업장 설비를 이어가며, 정비 차량으로 휴식 공간을 대체하지 않는다. 얼굴과 어깨 위주로 잘려 좌석과 하체 위치는 확인할 수 없다. 포스터의 얼굴은 부분적으로 가려지고 부드럽게 흐려져 있으며, 설비 중복이나 반사 문제는 보이지 않는다.",
        "entities": "실제 인물은 현우에 해당하는 젊은 동아시아계 남성 한 명이며, 앳된 얼굴·검은 헝클어진 머리·체형·짙은 남색 티셔츠가 참조와 잘 맞는다. 한국계 미국인이라는 국적 배경 자체는 사진으로 확인할 수 없다. 눈은 정상적인 홍채와 동공을 지닌다. 포스터와 대형 스피커, 여러 공구가 보이지만 대형 펜치를 특정하기는 어렵다. 얼굴과 목에는 뚜렷한 치료 흔적이 보이지 않는다. 다리 붕대와 신발 속 카드, 캠핑카 상태는 올바른 클로즈업 밖이므로 평가할 수 없다. 명확히 읽히는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리는 목과 어깨에 자연스럽게 지지되고, 말하면서 미간과 입 주변 근육을 움직이는 해부학적으로 가능한 표정이다. 하체와 좌석은 화면 밖이므로 앉거나 선 상태의 접촉점은 확인되지 않지만 부유하는 몸으로 보이지 않는다. 스피커는 선반 위에 놓였고 공구는 벽 공구판에 걸려 있어 지지가 성립한다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "다급하게 말하는 표정과 부상 흔적은 적절하지만, 무릎과 상체까지 넓힌 구도가 얼굴 중심 클로즈업 지시에서 벗어나고 의상 색도 참조보다 밝다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "얼굴 중심 클로즈업, 깊게 찌푸린 미간과 말을 꺼내는 순간, 참조의 얼굴·남색 의상을 더 충실히 구현했으나 노출된 피부의 치료 흔적은 분명하지 않다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴과 눈은 화면 왼쪽의 카메라 밖 상대를 향하고, 입을 크게 벌려 질문하는 순간처럼 보인다. 상대는 보이지 않으므로 앰버의 행방을 묻는 구체적 내용까지 확인할 수는 없지만 대화 방향은 자연스럽다. 겨누는 물건이나 이동하는 물체는 없다.",
        "built_space": "얼룩진 벽에 가장자리와 머리 뒤의 일부를 포함해 포스터 약 7장이 보이며, 인쇄된 인물들은 흐릿하다. 오른쪽에는 수납 가구와 용기들, 인물 뒤에는 의자의 금속 테두리가 일부 보인다. 휴식 공간에 앉은 배치로 읽히고 정비 차량이 배경을 대신하지 않는다. 다만 무릎까지 들어오면서 지정된 얼굴 클로즈업보다 넓어졌다. 고정 설비의 명백한 중복이나 불가능한 반사는 없다.",
        "entities": "실제 인물은 젊은 동아시아계 남성 한 명뿐이며, 헝클어진 검은 머리와 얼굴 윤곽은 현우 참조에 대체로 부합한다. 국적은 외형만으로 확인할 수 없다. 청색 티셔츠는 참조의 짙은 남색보다 밝다. 뺨의 상처와 무릎 아래 붕대가 보인다. 포스터 속 얼굴은 인쇄물이지 추가 인물이 아니다. 캠핑카·대형 스피커·대형 펜치·신발 속 카드는 이 구도에서 확인되지 않으며, 판독 가능한 글자는 뚜렷하지 않다.",
        "hard_violations": [],
        "physics": "화면 오른쪽 의자 테두리와 몸의 배치로 앉아 있는 자세가 읽힌다. 굽힌 무릎은 아래쪽 화면 밖 다리로 이어지고 몸통과 목의 연결도 자연스럽다. 엉덩이와 발의 접촉점은 잘렸지만 공중에 떠 있다는 징후는 없다. 손에 든 물체나 지지 없이 떠 있는 물체도 없다."
       },
       {
        "label": "A",
        "direction": "얼굴은 약간 화면 오른쪽으로 돌아가 있고 눈은 그쪽 카메라 밖 대화 상대를 향한다. 미간을 깊게 모으고 입을 열어 다급하게 질문을 시작하는 순간으로 읽힌다. 상대 자체는 보이지 않는다. 겨누는 물건이나 이동하는 물체는 없다.",
        "built_space": "낡은 벽의 왼쪽에 주요 포스터 4장과 가장자리의 포스터 일부 1장이 보인다. 오른쪽에는 공구판 1개, 금속 선반 1개, 그 위 대형 스피커 1개가 있다. 이는 참조의 벽 재질과 작업장 설비를 이어가며, 정비 차량으로 휴식 공간을 대체하지 않는다. 얼굴과 어깨 위주로 잘려 좌석과 하체 위치는 확인할 수 없다. 포스터의 얼굴은 부분적으로 가려지고 부드럽게 흐려져 있으며, 설비 중복이나 반사 문제는 보이지 않는다.",
        "entities": "실제 인물은 현우에 해당하는 젊은 동아시아계 남성 한 명이며, 앳된 얼굴·검은 헝클어진 머리·체형·짙은 남색 티셔츠가 참조와 잘 맞는다. 한국계 미국인이라는 국적 배경 자체는 사진으로 확인할 수 없다. 눈은 정상적인 홍채와 동공을 지닌다. 포스터와 대형 스피커, 여러 공구가 보이지만 대형 펜치를 특정하기는 어렵다. 얼굴과 목에는 뚜렷한 치료 흔적이 보이지 않는다. 다리 붕대와 신발 속 카드, 캠핑카 상태는 올바른 클로즈업 밖이므로 평가할 수 없다. 명확히 읽히는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리는 목과 어깨에 자연스럽게 지지되고, 말하면서 미간과 입 주변 근육을 움직이는 해부학적으로 가능한 표정이다. 하체와 좌석은 화면 밖이므로 앉거나 선 상태의 접촉점은 확인되지 않지만 부유하는 몸으로 보이지 않는다. 스피커는 선반 위에 놓였고 공구는 벽 공구판에 걸려 있어 지지가 성립한다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.278
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.278
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1278
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "미간을 잔뜩 찌푸린 다급한 표정과 클로즈업 프레이밍을 완벽하게 구현했으며, 레퍼런스의 공간 요소(포스터, 공구판, 스피커)를 입체적으로 잘 살렸습니다."
   },
   {
    "label": "B",
    "score": 1278,
    "verdict_ko": "요구된 찌푸린 표정 연기가 부족하고, 얼굴 클로즈업 프레임에 무릎을 억지로 포함시켜 구도가 부자연스럽습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S50sh9_sel.png",
    "asset_id": "7bc8f972-2e1e-401a-b66d-c7420a3ed479",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-39dc-7960-9d53-91015eb9e007",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S50sh9"
  }
 },
 "S50sh10::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:03:02.528353+00:00",
  "fingerprint": "b5a903dfd92aa862a48f46c23f98e3d2a248317ff1ae1ea4370a329574b4b532",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S50sh10_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S50sh10_sel.png",
  "source_sha256": "7bee7554b31abae26696de1d5cbffdd7057bf5065273d7babb6b9d85dccf123e",
  "file": "S50sh10_cine.png",
  "staged_sha256": "64d64fe26dbbbcf1d4e37d8c75f107b5da642a2c983a6ac8546c1a25c9aed132",
  "latency_ms": 14271
 },
 "S51sh10::signage": {
  "fp": "40137920c01153b5",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::511bcff39f0e633e": {
  "subjects": [],
  "subject_text": "카센터 컨테이너 간이식당\n정비 시설 옆 컨테이너 안에 마련된 간이식당. 비좁은 직사각형 공간에 식탁과 좌석이 놓여 있고 한쪽에 출입문이 있다.",
  "identity": "canonical",
  "scope_id": "L211",
  "scope_role": "location_interior",
  "scope_sha": "01d05264bb880c65"
 },
 "S51sh10::bgfirst_bg": {
  "input_fingerprint": "8865703c2344ba96",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 해남이라는 말에 눈을 부릅뜬 채 앰버를 향해 고함을 치듯 입을 크게 벌린 현우의 분노한 얼굴.\n\nLOCATION (lock): At a dining table inside the converted-container canteen beside the auto-repair shop, in daytime ambient light.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container dining-space interior (Occupied during the conversation) — A narrow, softly resolved interior background remains behind 현우; used as Retains the shared location while excluding unrelated diners from the tight confrontation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Consistent ambient dining-space light and controlled contrast preserve the intimacy of the confrontation without turning the outburst into a lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 해남이라는 말에 눈을 부릅뜬 채 앰버를 향해 고함을 치듯 입을 크게 벌린 현우의 분노한 얼굴.\n\nLOCATION (lock): At a dining table inside the converted-container canteen beside the auto-repair shop, in daytime ambient light.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container dining-space interior (Occupied during the conversation) — A narrow, softly resolved interior background remains behind 현우; used as Retains the shared location while excluding unrelated diners from the tight confrontation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Consistent ambient dining-space light and controlled contrast preserve the intimacy of the confrontation without turning the outburst into a lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S51sh10__bgfirst_bg.png",
  "asset_id": "683d25c0-58aa-44f7-a776-222b7b276425",
  "input_asset_ids": [
   "30650cc7-b31b-433e-b1bc-a519c8b82177",
   "78695b5c-f8e2-4047-9612-b490195deea7"
  ]
 },
 "S51sh10": {
  "input_fingerprint": "ec761e1e6415aea7",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 해남이라는 말에 눈을 부릅뜬 채 앰버를 향해 고함을 치듯 입을 크게 벌린 현우의 분노한 얼굴.\n\nLOCATION (lock): At a dining table inside the converted-container canteen beside the auto-repair shop, in daytime ambient light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container dining-space interior (Occupied during the conversation) — A narrow, softly resolved interior background remains behind 현우; used as Retains the shared location while excluding unrelated diners from the tight confrontation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Consistent ambient dining-space light and controlled contrast preserve the intimacy of the confrontation without turning the outburst into a lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Bowls of rice soup are in the container dining area beside the repair shop. Charlie has been repaired and can rotate his arm again; the previous shoulder malfunction should no longer be shown as active. 현우: His leg remains bandaged and his body retains treatment marks. The contact card remains hidden in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 해남이라는 말에 눈을 부릅뜬 채 앰버를 향해 고함을 치듯 입을 크게 벌린 현우의 분노한 얼굴.\n\nLOCATION (lock): At a dining table inside the converted-container canteen beside the auto-repair shop, in daytime ambient light. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container dining-space interior (Occupied during the conversation) — A narrow, softly resolved interior background remains behind 현우; used as Retains the shared location while excluding unrelated diners from the tight confrontation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Consistent ambient dining-space light and controlled contrast preserve the intimacy of the confrontation without turning the outburst into a lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Bowls of rice soup are in the container dining area beside the repair shop. Charlie has been repaired and can rotate his arm again; the previous shoulder malfunction should no longer be shown as active. 현우: His leg remains bandaged and his body retains treatment marks. The contact card remains hidden in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 해남이라는 말에 눈을 부릅뜬 채 앰버를 향해 고함을 치듯 입을 크게 벌린 현우의 분노한 얼굴.\n\nLOCATION (lock): At a dining table inside the converted-container canteen beside the auto-repair shop, in daytime ambient light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Container dining-space interior (Occupied during the conversation) — A narrow, softly resolved interior background remains behind 현우; used as Retains the shared location while excluding unrelated diners from the tight confrontation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Consistent ambient dining-space light and controlled contrast preserve the intimacy of the confrontation without turning the outburst into a lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Bowls of rice soup are in the container dining area beside the repair shop. Charlie has been repaired and can rotate his arm again; the previous shoulder malfunction should no longer be shown as active. 현우: His leg remains bandaged and his body retains treatment marks. The contact card remains hidden in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S51sh10__bgfirst_bg.png",
     "asset_id": "683d25c0-58aa-44f7-a776-222b7b276425",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S51sh10.png",
     "asset_id": "30650cc7-b31b-433e-b1bc-a519c8b82177",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L211B01.png",
     "asset_id": "78695b5c-f8e2-4047-9612-b490195deea7",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우의 분노한 시선과 크게 벌린 입이 화면 앞쪽 앰버의 뒷모습을 정확히 향하고 있음.",
    "built_space": "컨테이너 식당 내부의 출입문, 포스터, 정수기 및 테이블 배치가 레퍼런스 이미지와 완벽하게 일치함.",
    "entities": "현우의 얼굴 특징 및 상처, 앰버의 어깨, 테이블 위 국밥 그릇이 요구사항에 맞게 표현됨.",
    "hard_violations": [
     "[gpt-high] 현우만 보이도록 제한한 인물 조건을 어기고 왼쪽 전경에 상대 인물의 머리와 어깨를 추가했다."
    ],
    "physics": "테이블에 체중을 싣고 상체를 앞으로 기울인 분노한 자세가 물리적으로 자연스럽게 지지됨."
   },
   {
    "label": "B",
    "direction": "현우가 눈을 부릅뜨고 화면 전경의 인물을 향해 고함을 치고 있음.",
    "built_space": "식당 내부이나, 창문과 가구 배치가 레퍼런스의 실제 구조와 다소 차이를 보임.",
    "entities": "현우와 국밥 그릇은 존재하나, 프레임에서 배제해야 할 배경의 식당 손님들이 포함되어 있음.",
    "hard_violations": [
     "[gpt-high] 현우만 보이도록 제한한 인물 조건을 어기고 왼쪽 전경에 상대 인물의 머리와 어깨를 추가했다.",
     "[gpt-high] 타 식객을 제외해야 하는 장면에 배경 식객 다섯 명을 추가했다."
    ],
    "physics": "상체를 앞으로 크게 기울였으나, 팔의 위치와 신체를 지지하는 방식이 화면상에서 다소 불분명함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "레퍼런스와 일치하는 완벽한 공간 구현을 보여주며, 관계없는 인물을 배제하라는 프레이밍 지시를 정확히 따랐습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "현우의 분노한 표정은 잘 표현되었으나, 지시와 달리 배경에 관계없는 인물들이 다수 등장하여 몰입을 방해합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 분노한 시선과 크게 벌린 입이 화면 앞쪽 앰버의 뒷모습을 정확히 향하고 있음.",
        "built_space": "컨테이너 식당 내부의 출입문, 포스터, 정수기 및 테이블 배치가 레퍼런스 이미지와 완벽하게 일치함.",
        "entities": "현우의 얼굴 특징 및 상처, 앰버의 어깨, 테이블 위 국밥 그릇이 요구사항에 맞게 표현됨.",
        "hard_violations": [],
        "physics": "테이블에 체중을 싣고 상체를 앞으로 기울인 분노한 자세가 물리적으로 자연스럽게 지지됨."
       },
       {
        "label": "B",
        "direction": "현우가 눈을 부릅뜨고 화면 전경의 인물을 향해 고함을 치고 있음.",
        "built_space": "식당 내부이나, 창문과 가구 배치가 레퍼런스의 실제 구조와 다소 차이를 보임.",
        "entities": "현우와 국밥 그릇은 존재하나, 프레임에서 배제해야 할 배경의 식당 손님들이 포함되어 있음.",
        "hard_violations": [],
        "physics": "상체를 앞으로 크게 기울였으나, 팔의 위치와 신체를 지지하는 방식이 화면상에서 다소 불분명함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "레퍼런스와 일치하는 완벽한 공간 구현을 보여주며, 관계없는 인물을 배제하라는 프레이밍 지시를 정확히 따랐습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "현우의 분노한 표정은 잘 표현되었으나, 지시와 달리 배경에 관계없는 인물들이 다수 등장하여 몰입을 방해합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 분노한 시선과 크게 벌린 입이 화면 앞쪽 앰버의 뒷모습을 정확히 향하고 있음.",
        "built_space": "컨테이너 식당 내부의 출입문, 포스터, 정수기 및 테이블 배치가 레퍼런스 이미지와 완벽하게 일치함.",
        "entities": "현우의 얼굴 특징 및 상처, 앰버의 어깨, 테이블 위 국밥 그릇이 요구사항에 맞게 표현됨.",
        "hard_violations": [],
        "physics": "테이블에 체중을 싣고 상체를 앞으로 기울인 분노한 자세가 물리적으로 자연스럽게 지지됨."
       },
       {
        "label": "B",
        "direction": "현우가 눈을 부릅뜨고 화면 전경의 인물을 향해 고함을 치고 있음.",
        "built_space": "식당 내부이나, 창문과 가구 배치가 레퍼런스의 실제 구조와 다소 차이를 보임.",
        "entities": "현우와 국밥 그릇은 존재하나, 프레임에서 배제해야 할 배경의 식당 손님들이 포함되어 있음.",
        "hard_violations": [],
        "physics": "상체를 앞으로 크게 기울였으나, 팔의 위치와 신체를 지지하는 방식이 화면상에서 다소 불분명함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "눈을 부릅뜨고 고함치는 표정은 적합하지만, 금지된 전경 인물과 배경 식객 5명을 추가하고 얼굴 클로즈업을 넓은 식탁 대화 구도로 바꿨다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "참조 식당의 구조와 현우의 외형은 더 충실하지만, 금지된 상대 인물이 들어오고 상반신까지 넓혀 얼굴 클로즈업을 놓쳤으며 눈을 부릅뜨는 연기도 약하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 눈과 얼굴은 화면 왼쪽 전경 인물의 얼굴을 향하며, 입을 크게 벌려 그 사람에게 고함치는 모습이다. 앰버를 향한 대화 방향으로는 자연스럽지만, 전경 인물의 신원은 확정할 수 없다. 배경 식객들은 각각 같은 식탁의 상대나 음식 쪽을 보고 있다.",
        "built_space": "현우 앞 식탁 하나와 뒤쪽 식탁 두 구역, 나무 벤치들이 보인다. 좌우 벽에 창 구역이 하나씩 있고, 천장 형광등 하나, 뒤 왼쪽 대형 냉방기 하나, 오른쪽 조리기구 선반이 보인다. 금속 벽과 나무 식탁은 컨테이너 식당에 부합하지만 참조의 문·창·선반 배치보다 일반적인 다른 식당처럼 읽힌다. 얼굴 뒤에 좁고 흐린 배경만 남기는 대신 실내와 식객을 넓게 노출했다.",
        "entities": "현우는 앳된 동아시아계 남성 외형, 헝클어진 검은 머리, 남색 티셔츠로 참조와 대체로 일치한다. 국적은 외형만으로 확인할 수 없다. 눈은 정상적인 해부학적 형태로 크게 떠 있고 입도 크게 열려 있다. 현우 외에 왼쪽 전경 인물 한 명과 배경 식객 다섯 명이 보인다. 앞에는 밥이 보이는 검은 국그릇, 금속 수저, 접시와 컵이 있다. 다리 붕대와 신발 속 카드는 프레임 밖이며, 읽을 수 있는 글자는 확인되지 않는다.",
        "hard_violations": [
         "현우만 보이도록 제한한 인물 조건을 어기고 왼쪽 전경에 상대 인물의 머리와 어깨를 추가했다.",
         "타 식객을 제외해야 하는 장면에 배경 식객 다섯 명을 추가했다."
        ],
        "physics": "현우는 식탁 뒤에서 상체를 앞으로 기울이고 있으며 목과 몸통의 연결은 자연스럽다. 골반과 좌석 접촉은 가려져 있지만 공중에 떠 있는 자세는 아니다. 그릇과 컵은 식탁에 놓여 있고 금속 수저는 그릇 가장자리에 기대어 지지된다. 배경 인물들도 식탁과 벤치 주변에 앉아 있는 자세로 읽히며 명백한 부유나 불가능한 관절은 없다."
       },
       {
        "label": "B",
        "direction": "현우의 눈과 얼굴은 왼쪽 전경의 갈색 머리 인물 얼굴을 향한다. 앞으로 몸을 숙여 그 상대에게 고함치는 방향은 명확하다. 다만 눈썹을 강하게 찌푸리고 눈을 다소 좁혀, 요구된 '눈을 부릅뜬' 상태보다는 노려보는 표정에 가깝다.",
        "built_space": "왼쪽 뒤 창 달린 문 하나, 중앙 왼쪽 금속 선반 하나와 밥솥 하나, 오른쪽 뒤 큰 창 하나, 천장 형광등 하나가 보인다. 회색 골이 진 벽, 노출 배관, 낡은 나무 식탁과 벤치가 참조 장소에 비교적 잘 대응한다. 앞 식탁과 오른쪽 뒤 식탁 일부가 보이고 현우는 앞 식탁 뒤에 위치한다. 그러나 상반신과 실내 설비를 상당히 보여 주어 좁고 부드러운 배경을 둔 얼굴 클로즈업은 아니다.",
        "entities": "현우의 앳된 동아시아계 남성 외형, 검은 헝클어진 머리, 체격과 남색 티셔츠는 참조에 가깝다. 뺨에 붉은 상처성 흔적이 보이고 눈과 입은 정상적인 인체 형태다. 왼쪽에는 허용되지 않은 상대 인물의 머리와 어깨가 크게 들어온다. 식탁에는 국물이 남은 흰 그릇, 숟가락, 젓가락과 금속 컵이 있으며 밥의 존재는 뚜렷하지 않다. 다리 붕대와 숨긴 카드는 화면 밖이다. 벽 게시물의 글자는 판독되지 않는다.",
        "hard_violations": [
         "현우만 보이도록 제한한 인물 조건을 어기고 왼쪽 전경에 상대 인물의 머리와 어깨를 추가했다."
        ],
        "physics": "현우는 식탁 뒤에서 상체를 앞으로 숙이고 양팔을 아래로 내린다. 손과 하체의 지지점은 프레임에 보이지 않지만, 허리 아래가 가려진 자연스러운 전경사 자세이며 부유를 나타내지는 않는다. 흰 그릇과 금속 컵은 식탁에 닿아 있고 숟가락은 그릇에 기대며 젓가락은 식탁 위에 놓여 있다. 명백한 해부학적 불가능이나 지지 없는 물체는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "눈을 부릅뜨고 고함치는 표정은 적합하지만, 금지된 전경 인물과 배경 식객 5명을 추가하고 얼굴 클로즈업을 넓은 식탁 대화 구도로 바꿨다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "참조 식당의 구조와 현우의 외형은 더 충실하지만, 금지된 상대 인물이 들어오고 상반신까지 넓혀 얼굴 클로즈업을 놓쳤으며 눈을 부릅뜨는 연기도 약하다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 눈과 얼굴은 화면 왼쪽 전경 인물의 얼굴을 향하며, 입을 크게 벌려 그 사람에게 고함치는 모습이다. 앰버를 향한 대화 방향으로는 자연스럽지만, 전경 인물의 신원은 확정할 수 없다. 배경 식객들은 각각 같은 식탁의 상대나 음식 쪽을 보고 있다.",
        "built_space": "현우 앞 식탁 하나와 뒤쪽 식탁 두 구역, 나무 벤치들이 보인다. 좌우 벽에 창 구역이 하나씩 있고, 천장 형광등 하나, 뒤 왼쪽 대형 냉방기 하나, 오른쪽 조리기구 선반이 보인다. 금속 벽과 나무 식탁은 컨테이너 식당에 부합하지만 참조의 문·창·선반 배치보다 일반적인 다른 식당처럼 읽힌다. 얼굴 뒤에 좁고 흐린 배경만 남기는 대신 실내와 식객을 넓게 노출했다.",
        "entities": "현우는 앳된 동아시아계 남성 외형, 헝클어진 검은 머리, 남색 티셔츠로 참조와 대체로 일치한다. 국적은 외형만으로 확인할 수 없다. 눈은 정상적인 해부학적 형태로 크게 떠 있고 입도 크게 열려 있다. 현우 외에 왼쪽 전경 인물 한 명과 배경 식객 다섯 명이 보인다. 앞에는 밥이 보이는 검은 국그릇, 금속 수저, 접시와 컵이 있다. 다리 붕대와 신발 속 카드는 프레임 밖이며, 읽을 수 있는 글자는 확인되지 않는다.",
        "hard_violations": [
         "현우만 보이도록 제한한 인물 조건을 어기고 왼쪽 전경에 상대 인물의 머리와 어깨를 추가했다.",
         "타 식객을 제외해야 하는 장면에 배경 식객 다섯 명을 추가했다."
        ],
        "physics": "현우는 식탁 뒤에서 상체를 앞으로 기울이고 있으며 목과 몸통의 연결은 자연스럽다. 골반과 좌석 접촉은 가려져 있지만 공중에 떠 있는 자세는 아니다. 그릇과 컵은 식탁에 놓여 있고 금속 수저는 그릇 가장자리에 기대어 지지된다. 배경 인물들도 식탁과 벤치 주변에 앉아 있는 자세로 읽히며 명백한 부유나 불가능한 관절은 없다."
       },
       {
        "label": "A",
        "direction": "현우의 눈과 얼굴은 왼쪽 전경의 갈색 머리 인물 얼굴을 향한다. 앞으로 몸을 숙여 그 상대에게 고함치는 방향은 명확하다. 다만 눈썹을 강하게 찌푸리고 눈을 다소 좁혀, 요구된 '눈을 부릅뜬' 상태보다는 노려보는 표정에 가깝다.",
        "built_space": "왼쪽 뒤 창 달린 문 하나, 중앙 왼쪽 금속 선반 하나와 밥솥 하나, 오른쪽 뒤 큰 창 하나, 천장 형광등 하나가 보인다. 회색 골이 진 벽, 노출 배관, 낡은 나무 식탁과 벤치가 참조 장소에 비교적 잘 대응한다. 앞 식탁과 오른쪽 뒤 식탁 일부가 보이고 현우는 앞 식탁 뒤에 위치한다. 그러나 상반신과 실내 설비를 상당히 보여 주어 좁고 부드러운 배경을 둔 얼굴 클로즈업은 아니다.",
        "entities": "현우의 앳된 동아시아계 남성 외형, 검은 헝클어진 머리, 체격과 남색 티셔츠는 참조에 가깝다. 뺨에 붉은 상처성 흔적이 보이고 눈과 입은 정상적인 인체 형태다. 왼쪽에는 허용되지 않은 상대 인물의 머리와 어깨가 크게 들어온다. 식탁에는 국물이 남은 흰 그릇, 숟가락, 젓가락과 금속 컵이 있으며 밥의 존재는 뚜렷하지 않다. 다리 붕대와 숨긴 카드는 화면 밖이다. 벽 게시물의 글자는 판독되지 않는다.",
        "hard_violations": [
         "현우만 보이도록 제한한 인물 조건을 어기고 왼쪽 전경에 상대 인물의 머리와 어깨를 추가했다."
        ],
        "physics": "현우는 식탁 뒤에서 상체를 앞으로 숙이고 양팔을 아래로 내린다. 손과 하체의 지지점은 프레임에 보이지 않지만, 허리 아래가 가려진 자연스러운 전경사 자세이며 부유를 나타내지는 않는다. 흰 그릇과 금속 컵은 식탁에 닿아 있고 숟가락은 그릇에 기대며 젓가락은 식탁 위에 놓여 있다. 명백한 해부학적 불가능이나 지지 없는 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.071
   },
   "adjusted": {
    "A": 1.75,
    "B": 0.821
   },
   "violations": {
    "B": [
     "[gpt-high] 현우만 보이도록 제한한 인물 조건을 어기고 왼쪽 전경에 상대 인물의 머리와 어깨를 추가했다.",
     "[gpt-high] 타 식객을 제외해야 하는 장면에 배경 식객 다섯 명을 추가했다."
    ],
    "A": [
     "[gpt-high] 현우만 보이도록 제한한 인물 조건을 어기고 왼쪽 전경에 상대 인물의 머리와 어깨를 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 821
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "레퍼런스와 일치하는 완벽한 공간 구현을 보여주며, 관계없는 인물을 배제하라는 프레이밍 지시를 정확히 따랐습니다.  ★위반: [gpt-high] 현우만 보이도록 제한한 인물 조건을 어기고 왼쪽 전경에 상대 인물의 머리와 어깨를 추가했다."
   },
   {
    "label": "B",
    "score": 821,
    "verdict_ko": "현우의 분노한 표정은 잘 표현되었으나, 지시와 달리 배경에 관계없는 인물들이 다수 등장하여 몰입을 방해합니다.  ★위반: [gpt-high] 현우만 보이도록 제한한 인물 조건을 어기고 왼쪽 전경에 상대 인물의 머리와 어깨를 추가했다. / [gpt-high] 타 식객을 제외해야 하는 장면에 배경 식객 다섯 명을 추가했다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L211B01.png",
    "asset_id": "78695b5c-f8e2-4047-9612-b490195deea7",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-3b8b-7f6e-a254-611c8192c4c4",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S51sh10__bgfirst_bg.png",
   "bg_asset_id": "683d25c0-58aa-44f7-a776-222b7b276425",
   "bg_record_key": "S51sh10::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S51sh10::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:04:41.459011+00:00",
  "fingerprint": "1e51331110221f2341d1aaef73bd5dca058dc0e769d4d9b620583680ad462cd7",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S51sh10_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S51sh10_sel.png",
  "source_sha256": "6a7bb02ec9efd02a2369da4164eafb87177d0c3ba4e6f9e51aa9c4d8456c3362",
  "file": "S51sh10_cine.png",
  "staged_sha256": "e3c73e58819710bee49e51c1bab9d5e6746674514aa495fbfda59876887469b0",
  "latency_ms": 11856
 },
 "S51sh14::signage": {
  "fp": "64e7725cce9cf68e",
  "inscriptions": [],
  "cues": [
   {
    "text_native": "",
    "source": "scene_text_implied",
    "source_quote": "삐뚤빼뚤한 글씨가 적힌 낡은 종이"
   }
  ],
  "dropped": []
 },
 "S51sh14": {
  "input_fingerprint": "52b3f5af97f867d5",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 앰버가 쥔 삐뚤빼뚤한 글씨가 적힌 낡은 종이를 뚫어져라 내려다보는 현우의 시점 쇼트.\n\nLOCATION (lock): In the temporary sleeping area at the auto-repair premises, in morning light as the farewell note is examined. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 찰리's farewell letter (Worn paper held by 앰버, bearing uneven handwriting) — The written face is visible to the camera and reads: 고마워. 해남에 오게 되면 꼭 들러. 앰버 널 기억할게 .친구 찰리; used as Central reading focus, surrounded by the holder's hands and partial torso rather than enlarged beyond natural scale.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate daytime illumination keeps the handwriting legible with restrained contrast and no perceptual distortion.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie's farewell note says that he is going to Haenam and will remember Amber as his friend. The camper is still at the repair shop; Charlie has departed alone. 앰버: She is presenting the farewell note.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 앰버, 현우 right now, so 앰버, 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 앰버, 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 앰버가 쥔 삐뚤빼뚤한 글씨가 적힌 낡은 종이를 뚫어져라 내려다보는 현우의 시점 쇼트.\n\nLOCATION (lock): In the temporary sleeping area at the auto-repair premises, in morning light as the farewell note is examined. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 찰리's farewell letter (Worn paper held by 앰버, bearing uneven handwriting) — The written face is visible to the camera and reads: 고마워. 해남에 오게 되면 꼭 들러. 앰버 널 기억할게 .친구 찰리; used as Central reading focus, surrounded by the holder's hands and partial torso rather than enlarged beyond natural scale.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate daytime illumination keeps the handwriting legible with restrained contrast and no perceptual distortion.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie's farewell note says that he is going to Haenam and will remember Amber as his friend. The camper is still at the repair shop; Charlie has departed alone. 앰버: She is presenting the farewell note.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 앰버, 현우 right now, so 앰버, 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 앰버, 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 앰버가 쥔 삐뚤빼뚤한 글씨가 적힌 낡은 종이를 뚫어져라 내려다보는 현우의 시점 쇼트.\n\nLOCATION (lock): In the temporary sleeping area at the auto-repair premises, in morning light as the farewell note is examined. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 찰리's farewell letter (Worn paper held by 앰버, bearing uneven handwriting) — The written face is visible to the camera and reads: 고마워. 해남에 오게 되면 꼭 들러. 앰버 널 기억할게 .친구 찰리; used as Central reading focus, surrounded by the holder's hands and partial torso rather than enlarged beyond natural scale.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate daytime illumination keeps the handwriting legible with restrained contrast and no perceptual distortion.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie's farewell note says that he is going to Haenam and will remember Amber as his friend. The camper is still at the repair shop; Charlie has departed alone. 앰버: She is presenting the farewell note.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 앰버, 현우 right now, so 앰버, 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 앰버, 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시점(POV)에서 앰버가 들고 있는 종이의 정면을 똑바로 내려다보고 있습니다.",
    "built_space": "배경에 이전 숏 레퍼런스에서 확인되는 공구 걸이(페그보드), 로큰롤 포스터, 스피커가 있는 금속 선반과 오일 통들이 정확히 동일하게 배치되어 있습니다.",
    "entities": "레퍼런스와 동일한 짙은 파란색 셔츠를 입은 금발의 앰버가 종이를 들고 있으며, 종이에는 '고마워. 해남에 오게 되면 꼭 들러. 앰버 널 기억할게 .친구 찰리'라는 문구가 정확히 적혀 있습니다.",
    "hard_violations": [
     "[gpt-high] 편지의 한국어 문장이 또렷하게 읽혀, 이미지 어디에도 판독 가능한 글자가 없어야 한다는 최종 금지 조건을 위반한다."
    ],
    "physics": "앰버의 양손이 종이의 양쪽 가장자리를 물리적으로 자연스럽게 쥐고 지지하고 있습니다."
   },
   {
    "label": "B",
    "direction": "현우의 시점(POV)에서 종이를 바라보고 있습니다.",
    "built_space": "창문, 커튼, 침대가 있는 방의 모습으로 묘사되어, 이전 숏 레퍼런스의 정비소 내부 배경과 일치하지 않습니다.",
    "entities": "금발의 앰버가 종이를 들고 있으나 레퍼런스와 다른 회색 긴팔 셔츠를 입고 있습니다. 종이의 텍스트는 프롬프트와 일치합니다.",
    "hard_violations": [
     "[gemini-pro] 지정된 이전 숏 레퍼런스의 장소(Location lock)를 반영하지 않고 임의의 침실 배경을 생성함.",
     "[gpt-high] 편지에 여러 줄의 한국어 문장이 선명하게 판독되어, 이미지 어디에도 읽을 수 있는 글자가 없어야 한다는 최종 금지 조건을 위반한다."
    ],
    "physics": "앰버의 양손이 종이를 안정적으로 쥐고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "이전 숏 레퍼런스의 배경(공구 벽걸이, 포스터, 선반)과 캐릭터의 의상을 완벽하게 일치시켰으며, 요구된 텍스트와 시점(POV)을 충실히 구현했습니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "텍스트는 자연스럽게 적혀 있으나, 이전 숏의 배경을 따르지 않고 임의의 침실로 렌더링했으며 앰버의 의상 색상도 레퍼런스와 다릅니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시점(POV)에서 앰버가 들고 있는 종이의 정면을 똑바로 내려다보고 있습니다.",
        "built_space": "배경에 이전 숏 레퍼런스에서 확인되는 공구 걸이(페그보드), 로큰롤 포스터, 스피커가 있는 금속 선반과 오일 통들이 정확히 동일하게 배치되어 있습니다.",
        "entities": "레퍼런스와 동일한 짙은 파란색 셔츠를 입은 금발의 앰버가 종이를 들고 있으며, 종이에는 '고마워. 해남에 오게 되면 꼭 들러. 앰버 널 기억할게 .친구 찰리'라는 문구가 정확히 적혀 있습니다.",
        "hard_violations": [],
        "physics": "앰버의 양손이 종이의 양쪽 가장자리를 물리적으로 자연스럽게 쥐고 지지하고 있습니다."
       },
       {
        "label": "B",
        "direction": "현우의 시점(POV)에서 종이를 바라보고 있습니다.",
        "built_space": "창문, 커튼, 침대가 있는 방의 모습으로 묘사되어, 이전 숏 레퍼런스의 정비소 내부 배경과 일치하지 않습니다.",
        "entities": "금발의 앰버가 종이를 들고 있으나 레퍼런스와 다른 회색 긴팔 셔츠를 입고 있습니다. 종이의 텍스트는 프롬프트와 일치합니다.",
        "hard_violations": [
         "지정된 이전 숏 레퍼런스의 장소(Location lock)를 반영하지 않고 임의의 침실 배경을 생성함."
        ],
        "physics": "앰버의 양손이 종이를 안정적으로 쥐고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "이전 숏 레퍼런스의 배경(공구 벽걸이, 포스터, 선반)과 캐릭터의 의상을 완벽하게 일치시켰으며, 요구된 텍스트와 시점(POV)을 충실히 구현했습니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "텍스트는 자연스럽게 적혀 있으나, 이전 숏의 배경을 따르지 않고 임의의 침실로 렌더링했으며 앰버의 의상 색상도 레퍼런스와 다릅니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시점(POV)에서 앰버가 들고 있는 종이의 정면을 똑바로 내려다보고 있습니다.",
        "built_space": "배경에 이전 숏 레퍼런스에서 확인되는 공구 걸이(페그보드), 로큰롤 포스터, 스피커가 있는 금속 선반과 오일 통들이 정확히 동일하게 배치되어 있습니다.",
        "entities": "레퍼런스와 동일한 짙은 파란색 셔츠를 입은 금발의 앰버가 종이를 들고 있으며, 종이에는 '고마워. 해남에 오게 되면 꼭 들러. 앰버 널 기억할게 .친구 찰리'라는 문구가 정확히 적혀 있습니다.",
        "hard_violations": [],
        "physics": "앰버의 양손이 종이의 양쪽 가장자리를 물리적으로 자연스럽게 쥐고 지지하고 있습니다."
       },
       {
        "label": "B",
        "direction": "현우의 시점(POV)에서 종이를 바라보고 있습니다.",
        "built_space": "창문, 커튼, 침대가 있는 방의 모습으로 묘사되어, 이전 숏 레퍼런스의 정비소 내부 배경과 일치하지 않습니다.",
        "entities": "금발의 앰버가 종이를 들고 있으나 레퍼런스와 다른 회색 긴팔 셔츠를 입고 있습니다. 종이의 텍스트는 프롬프트와 일치합니다.",
        "hard_violations": [
         "지정된 이전 숏 레퍼런스의 장소(Location lock)를 반영하지 않고 임의의 침실 배경을 생성함."
        ],
        "physics": "앰버의 양손이 종이를 안정적으로 쥐고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "편지 중심의 근접 촬영은 맞지만, 판독 가능한 글자가 금지 조건을 위반하고 회색 소매와 배경은 앰버 및 장소 참조와 덜 일치한다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "읽을 수 있는 글자라는 공통 위반은 있으나, 자연스러운 편지 크기와 손·부분 몸통 구성, 남색 상의 및 정비소 설비의 연속성이 더 충실하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "편지의 글씨 면은 카메라와 오른쪽 전경의 앰버 쪽을 향한다. 앰버의 눈은 보이지 않아 실제 시선은 확인할 수 없다. 카메라는 앰버의 뒤쪽 어깨 가까이에 있어, 현우에게 편지를 내보이는 장면보다는 앰버가 읽는 모습을 뒤에서 보는 구도로 읽힌다. 다만 현우가 곁에서 내려다보는 위치일 가능성은 있다.",
        "built_space": "왼쪽 위에 창 하나, 그 아래 작업대 하나, 뒤쪽에 선반 구조, 아래쪽에 침구가 놓인 잠자리 하나가 보인다. 낮빛과 낡은 작업장 재질은 어울리지만, 참조의 벽면 공구판과 포스터 배치는 이 구도에서 확인되지 않는다. 앰버는 잠자리 옆에서 편지를 들고 있으며 설비와 충돌하는 배치는 없다. 반사는 없다.",
        "entities": "낡고 얼룩지며 접힌 종이 한 장을 두 손으로 들고 있다. 작별 인사의 주요 문구는 선명하게 읽히지만, 이는 마지막의 모든 판독 가능한 글자 금지 조건과 충돌한다. 오른쪽에는 금발과 작은 손이 보이며 얼굴이 가려져 정확한 나이·혼혈 정체성·얼굴 일치는 확인할 수 없다. 회색 긴소매는 참조의 남색 반소매 상의와 다르다. 추가 인물은 보이지 않는다.",
        "hard_violations": [
         "편지에 여러 줄의 한국어 문장이 선명하게 판독되어, 이미지 어디에도 읽을 수 있는 글자가 없어야 한다는 최종 금지 조건을 위반한다."
        ],
        "physics": "두 손이 종이의 좌우 가장자리를 엄지와 나머지 손가락으로 잡아 지지한다. 종이의 접힘과 약한 휨은 이 지지 방식으로 가능하다. 손목과 팔의 연결도 자연스럽고, 침구와 작업장 물품은 각각 받침 면에 놓여 있다. 지지 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "글씨 면은 카메라와 앰버가 있는 오른쪽 뒤편을 향하고, 종이는 약간 기울어져 있다. 앰버의 얼굴 일부만 보여 눈의 목표는 확인할 수 없다. 현우가 앰버 옆에서 편지를 내려다보는 시점으로 해석할 여지는 있지만, 앰버가 현우를 향해 내미는 관계보다는 어깨 너머로 함께 읽는 관계가 더 강하다.",
        "built_space": "왼쪽 벽에 포스터 일부 하나, 중앙 위에 공구판 하나, 오른쪽에 금속 선반 하나, 아래에 긴 작업대 하나와 그 하부 수납면이 보인다. 렌치와 손공구들은 공구판에 걸려 있고 용기들은 선반에 놓여 있다. 참조의 거친 벽, 공구판, 오른쪽 선반 관계가 유지된다. 오른쪽 아래의 침구는 임시 잠자리와 부합하며 앰버는 그 옆에 있다. 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "금발, 남색 반소매 상의, 어린 인물의 손과 부분 몸통이 보여 앰버 참조의 가시적 특징에 부합한다. 얼굴 대부분이 잘려 정확한 얼굴과 혼혈 정체성은 확인할 수 없다. 종이는 손에 비례하는 자연스러운 크기이며 누렇게 닳고 구겨져 있다. 작별 편지의 주요 문구는 읽히지만 마지막의 글자 판독 금지 조건을 위반한다. 추가 인물은 없다.",
        "hard_violations": [
         "편지의 한국어 문장이 또렷하게 읽혀, 이미지 어디에도 판독 가능한 글자가 없어야 한다는 최종 금지 조건을 위반한다."
        ],
        "physics": "왼쪽 아래 가장자리와 오른쪽 가장자리를 두 손이 실제로 집고 있어 편지가 지지된다. 종이의 기울기와 휨, 손목 각도는 가벼운 종이를 들고 보여주는 행동으로 가능하다. 팔은 몸통 쪽으로 자연스럽게 이어지고 배경 용기와 공구에도 선반이나 걸이의 지지가 있다. 떠 있는 신체나 물체는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "편지 중심의 근접 촬영은 맞지만, 판독 가능한 글자가 금지 조건을 위반하고 회색 소매와 배경은 앰버 및 장소 참조와 덜 일치한다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "읽을 수 있는 글자라는 공통 위반은 있으나, 자연스러운 편지 크기와 손·부분 몸통 구성, 남색 상의 및 정비소 설비의 연속성이 더 충실하다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "편지의 글씨 면은 카메라와 오른쪽 전경의 앰버 쪽을 향한다. 앰버의 눈은 보이지 않아 실제 시선은 확인할 수 없다. 카메라는 앰버의 뒤쪽 어깨 가까이에 있어, 현우에게 편지를 내보이는 장면보다는 앰버가 읽는 모습을 뒤에서 보는 구도로 읽힌다. 다만 현우가 곁에서 내려다보는 위치일 가능성은 있다.",
        "built_space": "왼쪽 위에 창 하나, 그 아래 작업대 하나, 뒤쪽에 선반 구조, 아래쪽에 침구가 놓인 잠자리 하나가 보인다. 낮빛과 낡은 작업장 재질은 어울리지만, 참조의 벽면 공구판과 포스터 배치는 이 구도에서 확인되지 않는다. 앰버는 잠자리 옆에서 편지를 들고 있으며 설비와 충돌하는 배치는 없다. 반사는 없다.",
        "entities": "낡고 얼룩지며 접힌 종이 한 장을 두 손으로 들고 있다. 작별 인사의 주요 문구는 선명하게 읽히지만, 이는 마지막의 모든 판독 가능한 글자 금지 조건과 충돌한다. 오른쪽에는 금발과 작은 손이 보이며 얼굴이 가려져 정확한 나이·혼혈 정체성·얼굴 일치는 확인할 수 없다. 회색 긴소매는 참조의 남색 반소매 상의와 다르다. 추가 인물은 보이지 않는다.",
        "hard_violations": [
         "편지에 여러 줄의 한국어 문장이 선명하게 판독되어, 이미지 어디에도 읽을 수 있는 글자가 없어야 한다는 최종 금지 조건을 위반한다."
        ],
        "physics": "두 손이 종이의 좌우 가장자리를 엄지와 나머지 손가락으로 잡아 지지한다. 종이의 접힘과 약한 휨은 이 지지 방식으로 가능하다. 손목과 팔의 연결도 자연스럽고, 침구와 작업장 물품은 각각 받침 면에 놓여 있다. 지지 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "글씨 면은 카메라와 앰버가 있는 오른쪽 뒤편을 향하고, 종이는 약간 기울어져 있다. 앰버의 얼굴 일부만 보여 눈의 목표는 확인할 수 없다. 현우가 앰버 옆에서 편지를 내려다보는 시점으로 해석할 여지는 있지만, 앰버가 현우를 향해 내미는 관계보다는 어깨 너머로 함께 읽는 관계가 더 강하다.",
        "built_space": "왼쪽 벽에 포스터 일부 하나, 중앙 위에 공구판 하나, 오른쪽에 금속 선반 하나, 아래에 긴 작업대 하나와 그 하부 수납면이 보인다. 렌치와 손공구들은 공구판에 걸려 있고 용기들은 선반에 놓여 있다. 참조의 거친 벽, 공구판, 오른쪽 선반 관계가 유지된다. 오른쪽 아래의 침구는 임시 잠자리와 부합하며 앰버는 그 옆에 있다. 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "금발, 남색 반소매 상의, 어린 인물의 손과 부분 몸통이 보여 앰버 참조의 가시적 특징에 부합한다. 얼굴 대부분이 잘려 정확한 얼굴과 혼혈 정체성은 확인할 수 없다. 종이는 손에 비례하는 자연스러운 크기이며 누렇게 닳고 구겨져 있다. 작별 편지의 주요 문구는 읽히지만 마지막의 글자 판독 금지 조건을 위반한다. 추가 인물은 없다.",
        "hard_violations": [
         "편지의 한국어 문장이 또렷하게 읽혀, 이미지 어디에도 판독 가능한 글자가 없어야 한다는 최종 금지 조건을 위반한다."
        ],
        "physics": "왼쪽 아래 가장자리와 오른쪽 가장자리를 두 손이 실제로 집고 있어 편지가 지지된다. 종이의 기울기와 휨, 손목 각도는 가벼운 종이를 들고 보여주는 행동으로 가능하다. 팔은 몸통 쪽으로 자연스럽게 이어지고 배경 용기와 공구에도 선반이나 걸이의 지지가 있다. 떠 있는 신체나 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.306
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.056
   },
   "violations": {
    "B": [
     "[gemini-pro] 지정된 이전 숏 레퍼런스의 장소(Location lock)를 반영하지 않고 임의의 침실 배경을 생성함.",
     "[gpt-high] 편지에 여러 줄의 한국어 문장이 선명하게 판독되어, 이미지 어디에도 읽을 수 있는 글자가 없어야 한다는 최종 금지 조건을 위반한다."
    ],
    "A": [
     "[gpt-high] 편지의 한국어 문장이 또렷하게 읽혀, 이미지 어디에도 판독 가능한 글자가 없어야 한다는 최종 금지 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 1056
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "이전 숏 레퍼런스의 배경(공구 벽걸이, 포스터, 선반)과 캐릭터의 의상을 완벽하게 일치시켰으며, 요구된 텍스트와 시점(POV)을 충실히 구현했습니다.  ★위반: [gpt-high] 편지의 한국어 문장이 또렷하게 읽혀, 이미지 어디에도 판독 가능한 글자가 없어야 한다는 최종 금지 조건을 위반한다."
   },
   {
    "label": "B",
    "score": 1056,
    "verdict_ko": "텍스트는 자연스럽게 적혀 있으나, 이전 숏의 배경을 따르지 않고 임의의 침실로 렌더링했으며 앰버의 의상 색상도 레퍼런스와 다릅니다.  ★위반: [gemini-pro] 지정된 이전 숏 레퍼런스의 장소(Location lock)를 반영하지 않고 임의의 침실 배경을 생성함. / [gpt-high] 편지에 여러 줄의 한국어 문장이 선명하게 판독되어, 이미지 어디에도 읽을 수 있는 글자가 없어야 한다는 최종 금지 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S50sh10_sel.png",
    "asset_id": "b2f3f11c-b7ff-42ba-b244-e993ff76d45f",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-3ee3-79a3-a192-923b4ed110ef",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S50sh10"
  }
 },
 "S51sh14::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:05:46.405590+00:00",
  "fingerprint": "fbfafe1a1a0c9f524e87295d96e24b403522300956a10da6d7a27199904cfa55",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S51sh14_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S51sh14_sel.png",
  "source_sha256": "e4ada98a57a24f91a96bd3ca38b65476739364096cda84fe18c1613f77ea0597",
  "file": "S51sh14_cine.png",
  "staged_sha256": "c44e0c1fb3a7117d87f7bc6264c9ec6fa0a95887e8ffed4d8c2b258af907b3e1",
  "latency_ms": 11211
 },
 "S51sh19::signage": {
  "fp": "077a34b083b69dc3",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S51sh19::bgfirst_bg": {
  "input_fingerprint": "56bfb11813fc2114",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 출발하는 차 안에서 은영과 태진을 향해 양손을 높이 든 채 활짝 웃는 앰버와 라울의 환한 상체.\n\nLOCATION (lock): Inside the departing camper's passenger area, with daylight entering through the windows facing the auto-repair shop.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Camper interior (Occupied by the children as the vehicle departs) — An oblique interior view establishes the children's outward-facing position; used as A narrow contextual surround anchors the upper-body farewell without competing with their hands.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination preserves open facial detail and a gentle, restrained contrast appropriate to the affectionate farewell.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 출발하는 차 안에서 은영과 태진을 향해 양손을 높이 든 채 활짝 웃는 앰버와 라울의 환한 상체.\n\nLOCATION (lock): Inside the departing camper's passenger area, with daylight entering through the windows facing the auto-repair shop.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Camper interior (Occupied by the children as the vehicle departs) — An oblique interior view establishes the children's outward-facing position; used as A narrow contextual surround anchors the upper-body farewell without competing with their hands.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination preserves open facial detail and a gentle, restrained contrast appropriate to the affectionate farewell.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S51sh19__bgfirst_bg.png",
  "asset_id": "3517c78d-2559-4fa2-9d40-d3795bb05053",
  "input_asset_ids": [
   "9849057f-ede4-459a-804c-8659aaf33808",
   "2d5b2b8c-9d71-4f7f-acce-b7f26107156a"
  ]
 },
 "S51sh19": {
  "input_fingerprint": "ff8fb9121145384f",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 출발하는 차 안에서 은영과 태진을 향해 양손을 높이 든 채 활짝 웃는 앰버와 라울의 환한 상체.\n\nLOCATION (lock): Inside the departing camper's passenger area, with daylight entering through the windows facing the auto-repair shop. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Camper interior (Occupied by the children as the vehicle departs) — An oblique interior view establishes the children's outward-facing position; used as A narrow contextual surround anchors the upper-body farewell without competing with their hands.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination preserves open facial detail and a gentle, restrained contrast appropriate to the affectionate farewell.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper is operational again and has the luggage loaded for departure. Additional food and Amber's medicine have been placed aboard. 앰버: She is aboard the departing camper. 라울: He is aboard the departing camper, with his earlier injury treated.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 출발하는 차 안에서 은영과 태진을 향해 양손을 높이 든 채 활짝 웃는 앰버와 라울의 환한 상체.\n\nLOCATION (lock): Inside the departing camper's passenger area, with daylight entering through the windows facing the auto-repair shop. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Camper interior (Occupied by the children as the vehicle departs) — An oblique interior view establishes the children's outward-facing position; used as A narrow contextual surround anchors the upper-body farewell without competing with their hands.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination preserves open facial detail and a gentle, restrained contrast appropriate to the affectionate farewell.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper is operational again and has the luggage loaded for departure. Additional food and Amber's medicine have been placed aboard. 앰버: She is aboard the departing camper. 라울: He is aboard the departing camper, with his earlier injury treated.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 출발하는 차 안에서 은영과 태진을 향해 양손을 높이 든 채 활짝 웃는 앰버와 라울의 환한 상체.\n\nLOCATION (lock): Inside the departing camper's passenger area, with daylight entering through the windows facing the auto-repair shop. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Camper interior (Occupied by the children as the vehicle departs) — An oblique interior view establishes the children's outward-facing position; used as A narrow contextual surround anchors the upper-body farewell without competing with their hands.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination preserves open facial detail and a gentle, restrained contrast appropriate to the affectionate farewell.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper is operational again and has the luggage loaded for departure. Additional food and Amber's medicine have been placed aboard. 앰버: She is aboard the departing camper. 라울: He is aboard the departing camper, with his earlier injury treated.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S51sh19__bgfirst_bg.png",
     "asset_id": "3517c78d-2559-4fa2-9d40-d3795bb05053",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S51sh19.png",
     "asset_id": "9849057f-ede4-459a-804c-8659aaf33808",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163202>",
     "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L211B02.png",
     "asset_id": "2d5b2b8c-9d71-4f7f-acce-b7f26107156a",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163202>",
     "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "아이들이 창밖 오른쪽을 향해 시선을 두고 양손을 들고 외부를 향해 인사함.",
    "built_space": "캠핑카 내부 좌석과 창문이 보이며, 밖으로는 흐릿한 철제 건물이 보이나 간판은 식별 불가함.",
    "entities": "앰버와 라울의 인물 레퍼런스는 잘 맞으나, 짐이나 약 등의 적재 물품은 확인되지 않음.",
    "hard_violations": [
     "[gemini-pro] 신체 구조 오류: 라울의 들고 있는 두 손이 모두 왼손의 형태를 띠고 있음."
    ],
    "physics": "좌석에 앉아 팔을 위로 뻗고 있으며, 시트가 체중을 안정적으로 지탱함."
   },
   {
    "label": "B",
    "direction": "아이들이 창밖 왼쪽을 향해 시선을 두고 활짝 웃고 있으며, 앰버는 양손, 라울은 오른손을 듦.",
    "built_space": "캠핑카 내부 선반과 창문이 명확하며, 창밖으로 레퍼런스의 '식당' 건물이 정확한 각도와 비례로 보임.",
    "entities": "인물 외모가 레퍼런스와 일치하며, 선반 위에 여행 가방, 보온병(식량), 약병들이 정확히 배치됨.",
    "hard_violations": [
     "[gpt-high] 창밖 간판에 ‘식당’이라는 읽을 수 있는 글자가 노출되어, 이미지 어디에도 읽을 수 있는 문구를 두지 말라는 조건을 위반한다."
    ],
    "physics": "차체와 의자에 기대어 자세를 지탱하고 있으며 팔의 움직임과 각도가 물리적으로 자연스러움."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "레퍼런스의 식당 배경과 명시된 소품(약병, 짐, 식량)을 완벽히 구현했으나, 라울이 한 손만 들고 있는 점이 아쉽습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "라울의 양손이 모두 왼손으로 그려지는 치명적 오류가 있으며, 지정된 적재 물품이 생략되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "아이들이 창밖 오른쪽을 향해 시선을 두고 양손을 들고 외부를 향해 인사함.",
        "built_space": "캠핑카 내부 좌석과 창문이 보이며, 밖으로는 흐릿한 철제 건물이 보이나 간판은 식별 불가함.",
        "entities": "앰버와 라울의 인물 레퍼런스는 잘 맞으나, 짐이나 약 등의 적재 물품은 확인되지 않음.",
        "hard_violations": [
         "신체 구조 오류: 라울의 들고 있는 두 손이 모두 왼손의 형태를 띠고 있음."
        ],
        "physics": "좌석에 앉아 팔을 위로 뻗고 있으며, 시트가 체중을 안정적으로 지탱함."
       },
       {
        "label": "B",
        "direction": "아이들이 창밖 왼쪽을 향해 시선을 두고 활짝 웃고 있으며, 앰버는 양손, 라울은 오른손을 듦.",
        "built_space": "캠핑카 내부 선반과 창문이 명확하며, 창밖으로 레퍼런스의 '식당' 건물이 정확한 각도와 비례로 보임.",
        "entities": "인물 외모가 레퍼런스와 일치하며, 선반 위에 여행 가방, 보온병(식량), 약병들이 정확히 배치됨.",
        "hard_violations": [],
        "physics": "차체와 의자에 기대어 자세를 지탱하고 있으며 팔의 움직임과 각도가 물리적으로 자연스러움."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "레퍼런스의 식당 배경과 명시된 소품(약병, 짐, 식량)을 완벽히 구현했으나, 라울이 한 손만 들고 있는 점이 아쉽습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "라울의 양손이 모두 왼손으로 그려지는 치명적 오류가 있으며, 지정된 적재 물품이 생략되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "아이들이 창밖 오른쪽을 향해 시선을 두고 양손을 들고 외부를 향해 인사함.",
        "built_space": "캠핑카 내부 좌석과 창문이 보이며, 밖으로는 흐릿한 철제 건물이 보이나 간판은 식별 불가함.",
        "entities": "앰버와 라울의 인물 레퍼런스는 잘 맞으나, 짐이나 약 등의 적재 물품은 확인되지 않음.",
        "hard_violations": [
         "신체 구조 오류: 라울의 들고 있는 두 손이 모두 왼손의 형태를 띠고 있음."
        ],
        "physics": "좌석에 앉아 팔을 위로 뻗고 있으며, 시트가 체중을 안정적으로 지탱함."
       },
       {
        "label": "B",
        "direction": "아이들이 창밖 왼쪽을 향해 시선을 두고 활짝 웃고 있으며, 앰버는 양손, 라울은 오른손을 듦.",
        "built_space": "캠핑카 내부 선반과 창문이 명확하며, 창밖으로 레퍼런스의 '식당' 건물이 정확한 각도와 비례로 보임.",
        "entities": "인물 외모가 레퍼런스와 일치하며, 선반 위에 여행 가방, 보온병(식량), 약병들이 정확히 배치됨.",
        "hard_violations": [],
        "physics": "차체와 의자에 기대어 자세를 지탱하고 있으며 팔의 움직임과 각도가 물리적으로 자연스러움."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "두 아이의 양손을 높이 든 동작과 의상은 맞지만, 앰버가 활짝 웃지 않고 창밖에 읽을 수 있는 ‘식당’ 간판이 노출되어 명시적 금지 조건을 위반한다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "창밖을 향해 웃는 두 아이의 상체와 자연스러운 착석을 잘 구현했지만, 라울의 한 손이 낮고 앰버의 티셔츠 색이 참조와 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 아이 모두 화면 왼쪽 창밖을 바라보며 손바닥을 바깥쪽으로 향하고 있다. 은영과 태진은 보이지 않아 실제 시선 도착점은 확인할 수 없지만, 차 밖의 배웅 상대를 향하는 방향은 자연스럽다. 네 손 모두 머리 위로 올라가 있다. 차량의 이동 방향은 정지 화면에서 확인되지 않는다.",
        "built_space": "왼쪽에 큰 측면 창 하나, 그 뒤에 좁은 창 하나, 오른쪽 뒤에 커튼 달린 창 하나가 보인다. 상단 목재 수납장, 뒤쪽 선반과 작업대, 아이들 뒤의 천 소재 벤치 등받이가 있다. 아이들은 벤치 앞에 상체를 세우고 있으며 하체는 잘려 있다. 창밖의 녹슨 골판금 벽, 출입문과 창문은 장소 참조와 유사하다. 다만 실내와 건물 배경이 화면에서 차지하는 비중이 커서 좁은 주변 맥락이라는 요구에는 덜 맞는다. 참조에는 캠퍼 내부가 없어 실내 고정 설비의 정확한 수량은 대조할 수 없다.",
        "entities": "추가 인물 없이 앰버와 라울에 해당하는 두 아이만 보인다. 앰버는 금발, 둥근 얼굴, 밝은 피부와 남색 티셔츠가 참조에 부합하지만 입을 다문 차분한 표정이어서 요구된 활짝 웃음이 없다. 라울은 갈색 피부, 뒤로 묶은 곱슬머리, 어린 얼굴과 남색 티셔츠가 대체로 부합하며 환하게 웃는다. 혼혈 배경 자체는 외모만으로 확정할 수 없다. 선반의 여행가방과 약병 형태 용기들은 적재 상태에 부합하지만 음식 및 앰버의 약인지, 라울의 상처가 치료되었는지는 확인되지 않는다. 창밖 간판의 ‘식당’은 읽을 수 있다.",
        "hard_violations": [
         "창밖 간판에 ‘식당’이라는 읽을 수 있는 글자가 노출되어, 이미지 어디에도 읽을 수 있는 문구를 두지 말라는 조건을 위반한다."
        ],
        "physics": "두 아이의 팔과 손은 어깨와 팔꿈치에 자연스럽게 연결되어 있으며 양팔을 드는 동작이 가능하다. 몸통은 수직으로 이어지고 하체와 발은 프레임 밖이므로 정확한 지지 접점은 보이지 않지만, 공중에 떠 있다고 볼 근거는 없다. 가방과 용기들은 선반 및 작업대 위에 놓여 있어 지지가 보인다."
       },
       {
        "label": "B",
        "direction": "두 아이의 얼굴과 시선은 화면 오른쪽 큰 창 바깥을 향하며 손바닥도 배웅 상대 쪽으로 열려 있다. 은영과 태진은 화면 밖이므로 직접 확인할 수 없지만 인사의 방향은 창밖으로 일관된다. 앰버는 양손을 머리 위로 높이 들었으나 라울은 한 손만 높이 들고 다른 손은 어깨 부근에 두어 ‘양손을 높이’라는 동작을 완전히 충족하지 못한다. 차량 이동 방향은 확정할 수 없다.",
        "built_space": "아이들은 하나의 넓은 벤치에서 각자 나란한 자리를 차지하고 등받이를 뒤에 둔다. 그 뒤로 좌석 등받이들이 이어지며, 오른쪽 큰 측면 창과 뒤쪽 측면 창, 왼쪽 창열, 천장 채광창 하나, 뒤쪽 목재 수납장이 보인다. 사선 실내 시점과 상체 중심의 중간 크기 구도가 창밖 인사를 명확하게 보여준다. 오른쪽 창밖의 골판금 건물, 창문, 실외기와 산 배경은 장소 참조의 특징과 부합한다. 실내 설비 수량은 외관뿐인 참조로 확정할 수 없으며, 명백한 설비 복제나 불가능한 반사는 보이지 않는다.",
        "entities": "앰버와 라울에 해당하는 두 아이만 있으며 추가 사람이나 읽을 수 있는 글자는 보이지 않는다. 앰버의 금발, 둥근 얼굴과 어린 외형은 참조에 가깝지만 티셔츠는 남색이 아닌 밝은 회색이다. 라울은 갈색 피부, 뒤로 묶인 곱슬머리, 어린 얼굴과 남색 티셔츠로 참조의 주요 특징을 따른다. 두 아이 모두 입을 열고 웃는다. 혼혈 배경 자체는 외모만으로 확정할 수 없다. 안전벨트가 보이며, 짐·음식·약과 치료 부위는 이 상체 구도에서 명확히 확인되지 않는다.",
        "hard_violations": [],
        "physics": "두 아이의 골반과 허벅지는 벤치 좌면에 놓이고 등은 등받이 쪽을 향해 자연스럽게 지지된다. 안전벨트가 몸통을 가로지르며, 팔을 들어 인사하는 동안에도 착석을 유지한다. 손과 팔의 연결 및 관절 각도는 가능한 범위이며, 지지 없이 떠 있는 신체나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "두 아이의 양손을 높이 든 동작과 의상은 맞지만, 앰버가 활짝 웃지 않고 창밖에 읽을 수 있는 ‘식당’ 간판이 노출되어 명시적 금지 조건을 위반한다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "창밖을 향해 웃는 두 아이의 상체와 자연스러운 착석을 잘 구현했지만, 라울의 한 손이 낮고 앰버의 티셔츠 색이 참조와 다르다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "두 아이 모두 화면 왼쪽 창밖을 바라보며 손바닥을 바깥쪽으로 향하고 있다. 은영과 태진은 보이지 않아 실제 시선 도착점은 확인할 수 없지만, 차 밖의 배웅 상대를 향하는 방향은 자연스럽다. 네 손 모두 머리 위로 올라가 있다. 차량의 이동 방향은 정지 화면에서 확인되지 않는다.",
        "built_space": "왼쪽에 큰 측면 창 하나, 그 뒤에 좁은 창 하나, 오른쪽 뒤에 커튼 달린 창 하나가 보인다. 상단 목재 수납장, 뒤쪽 선반과 작업대, 아이들 뒤의 천 소재 벤치 등받이가 있다. 아이들은 벤치 앞에 상체를 세우고 있으며 하체는 잘려 있다. 창밖의 녹슨 골판금 벽, 출입문과 창문은 장소 참조와 유사하다. 다만 실내와 건물 배경이 화면에서 차지하는 비중이 커서 좁은 주변 맥락이라는 요구에는 덜 맞는다. 참조에는 캠퍼 내부가 없어 실내 고정 설비의 정확한 수량은 대조할 수 없다.",
        "entities": "추가 인물 없이 앰버와 라울에 해당하는 두 아이만 보인다. 앰버는 금발, 둥근 얼굴, 밝은 피부와 남색 티셔츠가 참조에 부합하지만 입을 다문 차분한 표정이어서 요구된 활짝 웃음이 없다. 라울은 갈색 피부, 뒤로 묶은 곱슬머리, 어린 얼굴과 남색 티셔츠가 대체로 부합하며 환하게 웃는다. 혼혈 배경 자체는 외모만으로 확정할 수 없다. 선반의 여행가방과 약병 형태 용기들은 적재 상태에 부합하지만 음식 및 앰버의 약인지, 라울의 상처가 치료되었는지는 확인되지 않는다. 창밖 간판의 ‘식당’은 읽을 수 있다.",
        "hard_violations": [
         "창밖 간판에 ‘식당’이라는 읽을 수 있는 글자가 노출되어, 이미지 어디에도 읽을 수 있는 문구를 두지 말라는 조건을 위반한다."
        ],
        "physics": "두 아이의 팔과 손은 어깨와 팔꿈치에 자연스럽게 연결되어 있으며 양팔을 드는 동작이 가능하다. 몸통은 수직으로 이어지고 하체와 발은 프레임 밖이므로 정확한 지지 접점은 보이지 않지만, 공중에 떠 있다고 볼 근거는 없다. 가방과 용기들은 선반 및 작업대 위에 놓여 있어 지지가 보인다."
       },
       {
        "label": "A",
        "direction": "두 아이의 얼굴과 시선은 화면 오른쪽 큰 창 바깥을 향하며 손바닥도 배웅 상대 쪽으로 열려 있다. 은영과 태진은 화면 밖이므로 직접 확인할 수 없지만 인사의 방향은 창밖으로 일관된다. 앰버는 양손을 머리 위로 높이 들었으나 라울은 한 손만 높이 들고 다른 손은 어깨 부근에 두어 ‘양손을 높이’라는 동작을 완전히 충족하지 못한다. 차량 이동 방향은 확정할 수 없다.",
        "built_space": "아이들은 하나의 넓은 벤치에서 각자 나란한 자리를 차지하고 등받이를 뒤에 둔다. 그 뒤로 좌석 등받이들이 이어지며, 오른쪽 큰 측면 창과 뒤쪽 측면 창, 왼쪽 창열, 천장 채광창 하나, 뒤쪽 목재 수납장이 보인다. 사선 실내 시점과 상체 중심의 중간 크기 구도가 창밖 인사를 명확하게 보여준다. 오른쪽 창밖의 골판금 건물, 창문, 실외기와 산 배경은 장소 참조의 특징과 부합한다. 실내 설비 수량은 외관뿐인 참조로 확정할 수 없으며, 명백한 설비 복제나 불가능한 반사는 보이지 않는다.",
        "entities": "앰버와 라울에 해당하는 두 아이만 있으며 추가 사람이나 읽을 수 있는 글자는 보이지 않는다. 앰버의 금발, 둥근 얼굴과 어린 외형은 참조에 가깝지만 티셔츠는 남색이 아닌 밝은 회색이다. 라울은 갈색 피부, 뒤로 묶인 곱슬머리, 어린 얼굴과 남색 티셔츠로 참조의 주요 특징을 따른다. 두 아이 모두 입을 열고 웃는다. 혼혈 배경 자체는 외모만으로 확정할 수 없다. 안전벨트가 보이며, 짐·음식·약과 치료 부위는 이 상체 구도에서 명확히 확인되지 않는다.",
        "hard_violations": [],
        "physics": "두 아이의 골반과 허벅지는 벤치 좌면에 놓이고 등은 등받이 쪽을 향해 자연스럽게 지지된다. 안전벨트가 몸통을 가로지르며, 팔을 들어 인사하는 동안에도 착석을 유지한다. 손과 팔의 연결 및 관절 각도는 가능한 범위이며, 지지 없이 떠 있는 신체나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.429,
    "B": 1.375
   },
   "adjusted": {
    "A": 1.179,
    "B": 1.125
   },
   "violations": {
    "A": [
     "[gemini-pro] 신체 구조 오류: 라울의 들고 있는 두 손이 모두 왼손의 형태를 띠고 있음."
    ],
    "B": [
     "[gpt-high] 창밖 간판에 ‘식당’이라는 읽을 수 있는 글자가 노출되어, 이미지 어디에도 읽을 수 있는 문구를 두지 말라는 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1125,
   "A": 1179
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1125,
    "verdict_ko": "레퍼런스의 식당 배경과 명시된 소품(약병, 짐, 식량)을 완벽히 구현했으나, 라울이 한 손만 들고 있는 점이 아쉽습니다.  ★위반: [gpt-high] 창밖 간판에 ‘식당’이라는 읽을 수 있는 글자가 노출되어, 이미지 어디에도 읽을 수 있는 문구를 두지 말라는 조건을 위반한다."
   },
   {
    "label": "A",
    "score": 1179,
    "verdict_ko": "라울의 양손이 모두 왼손으로 그려지는 치명적 오류가 있으며, 지정된 적재 물품이 생략되었습니다.  ★위반: [gemini-pro] 신체 구조 오류: 라울의 들고 있는 두 손이 모두 왼손의 형태를 띠고 있음."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L211B02.png",
    "asset_id": "2d5b2b8c-9d71-4f7f-acce-b7f26107156a",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163202>",
    "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-4095-76d1-b6aa-79595924d6c5",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S51sh19__bgfirst_bg.png",
   "bg_asset_id": "3517c78d-2559-4fa2-9d40-d3795bb05053",
   "bg_record_key": "S51sh19::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S51sh19::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:07:44.497320+00:00",
  "fingerprint": "e272eb90b2f671a0e277e6ce3051b9cfa5d094537e585cb981edbb147544dfe6",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S51sh19_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S51sh19_sel.png",
  "source_sha256": "2b9a171f7d447c209a90e730b6516904f9e268e576d3afd8c77a0844a3428dca",
  "file": "S51sh19_cine.png",
  "staged_sha256": "c68b726ed7806396ffb90a7bb576b8a991a9427a208a3a66df2c3b1d819259f9",
  "latency_ms": 13719
 },
 "S52sh1::signage": {
  "fp": "9e37ac79ee8f4c1a",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S52sh1": {
  "input_fingerprint": "29b2ab822243e2be",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 커다란 밀짚모자를 쓰고 알록달록한 우비를 걸친 채 텅 빈 잿빛 도로 위를 걷는 도중 뒷발로 바닥을 밀어내고 앞발을 든 mid-stride 자세의 찰리의 낡은 금속 전신.\n\nLOCATION (lock): On an empty provincial asphalt road, with the disguised robot walking alone along the exposed roadway. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Provincial road (Empty and gray around the walking figure) — The route extends diagonally ahead of 찰리 toward frame right; used as Open space around the full figure establishes isolation and makes the lifted-foot phase readable; Oversized straw hat, raincoat, and boots (Worn by 찰리; the raincoat is multicolored) — Seen obliquely with the hat brim and separated boots clearly readable; used as Costume silhouettes establish the comic discrepancy between his appearance and determined walk.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light maintains restrained tonal contrast while allowing the explicitly multicolored raincoat to retain its stronger color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie now wears an oversized straw hat, oversized rubber boots and a colorful raincoat over his repaired, worn metal body.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 커다란 밀짚모자를 쓰고 알록달록한 우비를 걸친 채 텅 빈 잿빛 도로 위를 걷는 도중 뒷발로 바닥을 밀어내고 앞발을 든 mid-stride 자세의 찰리의 낡은 금속 전신.\n\nLOCATION (lock): On an empty provincial asphalt road, with the disguised robot walking alone along the exposed roadway. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Provincial road (Empty and gray around the walking figure) — The route extends diagonally ahead of 찰리 toward frame right; used as Open space around the full figure establishes isolation and makes the lifted-foot phase readable; Oversized straw hat, raincoat, and boots (Worn by 찰리; the raincoat is multicolored) — Seen obliquely with the hat brim and separated boots clearly readable; used as Costume silhouettes establish the comic discrepancy between his appearance and determined walk.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light maintains restrained tonal contrast while allowing the explicitly multicolored raincoat to retain its stronger color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie now wears an oversized straw hat, oversized rubber boots and a colorful raincoat over his repaired, worn metal body.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 커다란 밀짚모자를 쓰고 알록달록한 우비를 걸친 채 텅 빈 잿빛 도로 위를 걷는 도중 뒷발로 바닥을 밀어내고 앞발을 든 mid-stride 자세의 찰리의 낡은 금속 전신.\n\nLOCATION (lock): On an empty provincial asphalt road, with the disguised robot walking alone along the exposed roadway. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Provincial road (Empty and gray around the walking figure) — The route extends diagonally ahead of 찰리 toward frame right; used as Open space around the full figure establishes isolation and makes the lifted-foot phase readable; Oversized straw hat, raincoat, and boots (Worn by 찰리; the raincoat is multicolored) — Seen obliquely with the hat brim and separated boots clearly readable; used as Costume silhouettes establish the comic discrepancy between his appearance and determined walk.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light maintains restrained tonal contrast while allowing the explicitly multicolored raincoat to retain its stronger color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie now wears an oversized straw hat, oversized rubber boots and a colorful raincoat over his repaired, worn metal body.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S52sh1__bgfirst_bg.png",
     "asset_id": "4bc8f5c0-751d-4e2e-8589-aaf2e188a7ae",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S52sh1.png",
     "asset_id": "364391b6-7633-4178-a784-c9f3499899a6",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_rural_walking_road_5c3d2f.png",
     "asset_id": "2c211078-5c11-4ab2-abaa-988d6405f566",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "캐릭터가 화면 오른쪽 도로를 향해 시선을 두고 걷고 있음.",
    "built_space": "참조 이미지와 동일한 젖은 아스팔트 도로, 우측의 전신주 및 배경의 산맥.",
    "entities": "밀짚모자, 알록달록한 우비, 고무 장화가 명확히 보임. 단, 캐릭터의 얼굴과 체형이 레퍼런스의 찰리와 전혀 다름.",
    "hard_violations": [
     "[gemini-pro] 캐릭터가 실사 배경 위에 합성된 평면적인 2D 애니메이션 스타일로 렌더링되어 물질적 사실주의(MATERIAL REALISM) 제약 위반"
    ],
    "physics": "오른발은 바닥을 딛고 왼발은 공중에 들어 올린 걷는 자세를 취함."
   },
   {
    "label": "B",
    "direction": "캐릭터가 화면 정면 약간 왼쪽을 향해 걷고 있음.",
    "built_space": "참조 사진의 장소와 일치하는 텅 빈 잿빛 도로, 전신주, 농경지 및 산 배경.",
    "entities": "찰리의 육중한 금속 체형과 마스크형 얼굴이 레퍼런스와 완벽히 일치함. 밀짚모자와 알록달록한 우비는 있으나 고무 장화가 누락됨.",
    "hard_violations": [],
    "physics": "왼발로 바닥을 단단히 지지하고 오른발을 들어 올린 mid-stride 자세를 보여줌."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "고무 장화가 누락되고 걷는 방향이 프롬프트와 완전히 일치하지는 않으나, 찰리의 외형 레퍼런스를 정확히 반영하고 실사 질감을 훌륭하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "장화와 걷는 방향은 맞췄으나, 캐릭터가 2D 그래픽처럼 렌더링되어 실사 제약을 심각하게 위반했으며 찰리의 외형도 일치하지 않습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "캐릭터가 화면 오른쪽 도로를 향해 시선을 두고 걷고 있음.",
        "built_space": "참조 이미지와 동일한 젖은 아스팔트 도로, 우측의 전신주 및 배경의 산맥.",
        "entities": "밀짚모자, 알록달록한 우비, 고무 장화가 명확히 보임. 단, 캐릭터의 얼굴과 체형이 레퍼런스의 찰리와 전혀 다름.",
        "hard_violations": [
         "캐릭터가 실사 배경 위에 합성된 평면적인 2D 애니메이션 스타일로 렌더링되어 물질적 사실주의(MATERIAL REALISM) 제약 위반"
        ],
        "physics": "오른발은 바닥을 딛고 왼발은 공중에 들어 올린 걷는 자세를 취함."
       },
       {
        "label": "B",
        "direction": "캐릭터가 화면 정면 약간 왼쪽을 향해 걷고 있음.",
        "built_space": "참조 사진의 장소와 일치하는 텅 빈 잿빛 도로, 전신주, 농경지 및 산 배경.",
        "entities": "찰리의 육중한 금속 체형과 마스크형 얼굴이 레퍼런스와 완벽히 일치함. 밀짚모자와 알록달록한 우비는 있으나 고무 장화가 누락됨.",
        "hard_violations": [],
        "physics": "왼발로 바닥을 단단히 지지하고 오른발을 들어 올린 mid-stride 자세를 보여줌."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "고무 장화가 누락되고 걷는 방향이 프롬프트와 완전히 일치하지는 않으나, 찰리의 외형 레퍼런스를 정확히 반영하고 실사 질감을 훌륭하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "장화와 걷는 방향은 맞췄으나, 캐릭터가 2D 그래픽처럼 렌더링되어 실사 제약을 심각하게 위반했으며 찰리의 외형도 일치하지 않습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "캐릭터가 화면 오른쪽 도로를 향해 시선을 두고 걷고 있음.",
        "built_space": "참조 이미지와 동일한 젖은 아스팔트 도로, 우측의 전신주 및 배경의 산맥.",
        "entities": "밀짚모자, 알록달록한 우비, 고무 장화가 명확히 보임. 단, 캐릭터의 얼굴과 체형이 레퍼런스의 찰리와 전혀 다름.",
        "hard_violations": [
         "캐릭터가 실사 배경 위에 합성된 평면적인 2D 애니메이션 스타일로 렌더링되어 물질적 사실주의(MATERIAL REALISM) 제약 위반"
        ],
        "physics": "오른발은 바닥을 딛고 왼발은 공중에 들어 올린 걷는 자세를 취함."
       },
       {
        "label": "B",
        "direction": "캐릭터가 화면 정면 약간 왼쪽을 향해 걷고 있음.",
        "built_space": "참조 사진의 장소와 일치하는 텅 빈 잿빛 도로, 전신주, 농경지 및 산 배경.",
        "entities": "찰리의 육중한 금속 체형과 마스크형 얼굴이 레퍼런스와 완벽히 일치함. 밀짚모자와 알록달록한 우비는 있으나 고무 장화가 누락됨.",
        "hard_violations": [],
        "physics": "왼발로 바닥을 단단히 지지하고 오른발을 들어 올린 mid-stride 자세를 보여줌."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "찰리의 장갑판과 체형, 장소는 잘 맞지만 오른쪽으로 뻗는 길을 등지고 관객 쪽으로 걸으며, 필수 고무장화가 없다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "오른쪽을 향한 전신 보행과 분리된 대형 장화, 밀짚모자·다색 우비를 구현했으나 찰리의 육중하고 팔이 긴 체형은 약해졌다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴과 몸은 카메라 쪽에서 약간 화면 왼쪽을 향하고, 든 앞발도 왼쪽 전경으로 나간다. 도로는 인물 뒤에서 화면 오른쪽 원경으로 이어지므로, 찰리 앞쪽으로 길이 오른쪽 대각선으로 뻗어야 한다는 진행 관계와 반대다.",
        "built_space": "젖은 아스팔트 도로 하나, 노란 점선 중앙선 한 줄, 양쪽 흰 경계선 두 줄이 보인다. 오른쪽에는 전주가 한 열로 반복되고 왼쪽에는 물찬 논과 낮은 농업 시설, 뒤에는 산지가 있어 장소 참조와 부합한다. 먼 전주의 정확한 개수는 분간하기 어렵다. 찰리는 중앙선 부근에 홀로 있으며, 젖은 노면에 아래로 이어지는 반사는 가능한 위치다. 전신과 주변의 빈 도로가 함께 잡힌다.",
        "entities": "인물은 로봇 찰리 하나뿐이다. 흰 각진 마스크형 얼굴, 주황빛 눈, 마모된 샌드 베이지 장갑판, 큰 손과 비교적 짧은 다리는 참조와 가깝다. 큰 밀짚모자 하나와 여러 색 조각으로 된 우비 하나를 착용했다. 그러나 두 발은 노출된 기계 발이며 요구된 대형 고무장화가 아니다. 추가 인물이나 읽을 수 있는 문자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "화면 오른쪽의 뒤쪽 기계 발이 노면에 닿아 몸을 지지하며, 왼쪽 앞발은 바닥에서 들려 있다. 지지 없는 부유는 아니다. 다만 지지발이 평평하게 놓여 있어 뒷발로 바닥을 밀어내는 추진 순간은 약하다. 모자는 머리에, 우비는 어깨에 지지되며 옷자락도 자연스럽게 늘어진다."
       },
       {
        "label": "B",
        "direction": "머리와 몸, 앞으로 뻗은 장화가 모두 화면 오른쪽을 향한다. 시선은 모자와 옆얼굴 때문에 정확히 확인하기 어렵지만 머리 방향은 진행 방향과 일치한다. 도로도 인물 앞에서 오른쪽 원경으로 이어져 요구한 방향 관계를 충족한다.",
        "built_space": "젖은 지방도로 하나에 노란 점선 중앙선 한 줄과 흰 경계선 두 줄이 있고, 오른쪽 전주 한 열은 원경으로 작아진다. 먼 전주들은 겹쳐 정확한 수를 판별하기 어렵다. 왼쪽 논과 낮은 시설, 배후 산지가 장소 참조와 일치한다. 찰리는 중앙선 근처에 혼자 있으며 전신과 오른쪽의 넓은 빈 도로가 보인다. 노면의 장화와 의상 반사는 카메라 위치에서 가능한 방향이다.",
        "entities": "로봇 찰리 하나가 넓은 챙의 밀짚모자, 청색·녹색·노란색·보라색 계열의 긴 우비, 큰 밝은색 고무장화 두 짝을 착용했다. 드러난 팔과 무릎에는 낡은 베이지 장갑판이 있다. 흰 얼굴은 옆면만 보여 참조의 정면 세부는 확인할 수 없다. 보이는 팔다리 비율은 참조보다 가늘고 길어 고릴라형의 육중함이 덜하다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "화면 왼쪽 뒷장화의 앞쪽 밑창이 노면에 닿아 몸을 지지하고, 오른쪽 앞장화는 발끝이 들려 밑창이 드러난다. 앞장화 뒤꿈치가 노면에 매우 가까워 완전히 든 순간의 간격은 작지만, 다리의 벌어짐과 뒤쪽 지지로 보행은 성립한다. 모자와 우비는 몸에 지지되고, 뒤로 흐르는 옷자락도 걸음으로 설명 가능한 형태다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "찰리의 장갑판과 체형, 장소는 잘 맞지만 오른쪽으로 뻗는 길을 등지고 관객 쪽으로 걸으며, 필수 고무장화가 없다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "오른쪽을 향한 전신 보행과 분리된 대형 장화, 밀짚모자·다색 우비를 구현했으나 찰리의 육중하고 팔이 긴 체형은 약해졌다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴과 몸은 카메라 쪽에서 약간 화면 왼쪽을 향하고, 든 앞발도 왼쪽 전경으로 나간다. 도로는 인물 뒤에서 화면 오른쪽 원경으로 이어지므로, 찰리 앞쪽으로 길이 오른쪽 대각선으로 뻗어야 한다는 진행 관계와 반대다.",
        "built_space": "젖은 아스팔트 도로 하나, 노란 점선 중앙선 한 줄, 양쪽 흰 경계선 두 줄이 보인다. 오른쪽에는 전주가 한 열로 반복되고 왼쪽에는 물찬 논과 낮은 농업 시설, 뒤에는 산지가 있어 장소 참조와 부합한다. 먼 전주의 정확한 개수는 분간하기 어렵다. 찰리는 중앙선 부근에 홀로 있으며, 젖은 노면에 아래로 이어지는 반사는 가능한 위치다. 전신과 주변의 빈 도로가 함께 잡힌다.",
        "entities": "인물은 로봇 찰리 하나뿐이다. 흰 각진 마스크형 얼굴, 주황빛 눈, 마모된 샌드 베이지 장갑판, 큰 손과 비교적 짧은 다리는 참조와 가깝다. 큰 밀짚모자 하나와 여러 색 조각으로 된 우비 하나를 착용했다. 그러나 두 발은 노출된 기계 발이며 요구된 대형 고무장화가 아니다. 추가 인물이나 읽을 수 있는 문자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "화면 오른쪽의 뒤쪽 기계 발이 노면에 닿아 몸을 지지하며, 왼쪽 앞발은 바닥에서 들려 있다. 지지 없는 부유는 아니다. 다만 지지발이 평평하게 놓여 있어 뒷발로 바닥을 밀어내는 추진 순간은 약하다. 모자는 머리에, 우비는 어깨에 지지되며 옷자락도 자연스럽게 늘어진다."
       },
       {
        "label": "A",
        "direction": "머리와 몸, 앞으로 뻗은 장화가 모두 화면 오른쪽을 향한다. 시선은 모자와 옆얼굴 때문에 정확히 확인하기 어렵지만 머리 방향은 진행 방향과 일치한다. 도로도 인물 앞에서 오른쪽 원경으로 이어져 요구한 방향 관계를 충족한다.",
        "built_space": "젖은 지방도로 하나에 노란 점선 중앙선 한 줄과 흰 경계선 두 줄이 있고, 오른쪽 전주 한 열은 원경으로 작아진다. 먼 전주들은 겹쳐 정확한 수를 판별하기 어렵다. 왼쪽 논과 낮은 시설, 배후 산지가 장소 참조와 일치한다. 찰리는 중앙선 근처에 혼자 있으며 전신과 오른쪽의 넓은 빈 도로가 보인다. 노면의 장화와 의상 반사는 카메라 위치에서 가능한 방향이다.",
        "entities": "로봇 찰리 하나가 넓은 챙의 밀짚모자, 청색·녹색·노란색·보라색 계열의 긴 우비, 큰 밝은색 고무장화 두 짝을 착용했다. 드러난 팔과 무릎에는 낡은 베이지 장갑판이 있다. 흰 얼굴은 옆면만 보여 참조의 정면 세부는 확인할 수 없다. 보이는 팔다리 비율은 참조보다 가늘고 길어 고릴라형의 육중함이 덜하다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "화면 왼쪽 뒷장화의 앞쪽 밑창이 노면에 닿아 몸을 지지하고, 오른쪽 앞장화는 발끝이 들려 밑창이 드러난다. 앞장화 뒤꿈치가 노면에 매우 가까워 완전히 든 순간의 간격은 작지만, 다리의 벌어짐과 뒤쪽 지지로 보행은 성립한다. 모자와 우비는 몸에 지지되고, 뒤로 흐르는 옷자락도 걸음으로 설명 가능한 형태다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.429,
    "B": 1.75
   },
   "adjusted": {
    "A": 1.179,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 캐릭터가 실사 배경 위에 합성된 평면적인 2D 애니메이션 스타일로 렌더링되어 물질적 사실주의(MATERIAL REALISM) 제약 위반"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1750,
   "A": 1179
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "고무 장화가 누락되고 걷는 방향이 프롬프트와 완전히 일치하지는 않으나, 찰리의 외형 레퍼런스를 정확히 반영하고 실사 질감을 훌륭하게 구현했습니다."
   },
   {
    "label": "A",
    "score": 1179,
    "verdict_ko": "장화와 걷는 방향은 맞췄으나, 캐릭터가 2D 그래픽처럼 렌더링되어 실사 제약을 심각하게 위반했으며 찰리의 외형도 일치하지 않습니다.  ★위반: [gemini-pro] 캐릭터가 실사 배경 위에 합성된 평면적인 2D 애니메이션 스타일로 렌더링되어 물질적 사실주의(MATERIAL REALISM) 제약 위반"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_rural_walking_road_5c3d2f.png",
    "asset_id": "2c211078-5c11-4ab2-abaa-988d6405f566",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-43f0-7e2c-9399-bc91d82c045e",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S52sh1__bgfirst_bg.png",
   "bg_asset_id": "4bc8f5c0-751d-4e2e-8589-aaf2e188a7ae",
   "bg_record_key": "S52sh1::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "rural_walking_road",
   "groupbg_asset_id": "2c211078-5c11-4ab2-abaa-988d6405f566"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S52sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:39:12.918942+00:00",
  "fingerprint": "df49dc06999b695764e7644522ece6a99de00e37c7e61bbf72d130a7bf3b750b",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S52sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S52sh1_sel.png",
  "source_sha256": "7ed26b8e732db18b52df6640955e9bc935247db8d683f08b1181f27fa2b440a8",
  "file": "S52sh1_cine.png",
  "staged_sha256": "a1d9231e5737b509e17cefe99a23fce46f5ceb231b863cc25b457c1e20675241",
  "latency_ms": 10366
 },
 "S53sh2::signage": {
  "fp": "6f3c543335c9f5c1",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::0a3f52341ebda2d0": {
  "subjects": [],
  "subject_text": "수도권 도심 민병대 사무실\n책상 위에 큰 반도 지도가 펼쳐진 사무실. 다트판과 다트가 있고 책상 주변에 업무용 좌석이 놓여 있다.",
  "identity": "canonical",
  "scope_id": "L213",
  "scope_role": "location_interior",
  "scope_sha": "c19a59f424a5380d"
 },
 "groupbg::urban_militia_office": {
  "input_fingerprint": "8b9d4ffe98043a1c",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "urban_militia_office",
    "tags": [
     "S53sh2",
     "S53sh3"
    ]
   },
   "context_sig": "6935d1a2b18c8a4c"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At the map-covered desk inside an urban militia office, under ordinary office lighting.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n수도권 도심 민병대 사무실: 거대한 지도가 펼쳐져 있고 다트 게임판이 놓여 있는 작전 회의실 겸 휴게 공간. (특징: 책상 위를 덮을 만큼 커다란 한반도 종이 지도; 벽에 걸린 다트 과녁판과 꽂힌 다트핀)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 53. 수도권 도심 민병대 사무실\n- 책상 위에 한반도 지도가 펼쳐져 있다.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At the map-covered desk inside an urban militia office, under ordinary office lighting.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n수도권 도심 민병대 사무실: 거대한 지도가 펼쳐져 있고 다트 게임판이 놓여 있는 작전 회의실 겸 휴게 공간. (특징: 책상 위를 덮을 만큼 커다란 한반도 종이 지도; 벽에 걸린 다트 과녁판과 꽂힌 다트핀)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 53. 수도권 도심 민병대 사무실\n- 책상 위에 한반도 지도가 펼쳐져 있다.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_urban_militia_office_ccf685.png",
  "asset_id": "91778f94-deb9-4780-814b-5f1f344f3b26",
  "input_asset_ids": [
   "70416dc7-78a1-45aa-80e3-0fdad91a577a"
  ],
  "origin_tag": "S53sh2",
  "place_text": "At the map-covered desk inside an urban militia office, under ordinary office lighting.",
  "origin_inputs": {
   "place_text": "At the map-covered desk inside an urban militia office, under ordinary office lighting.",
   "time_of_day_en": "day",
   "conti_asset_id": "70416dc7-78a1-45aa-80e3-0fdad91a577a"
  }
 },
 "S53sh2::bgfirst_bg": {
  "input_fingerprint": "71eef7a3fbfb7c55",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 핏발 선 눈으로 두 주먹이 책상에 강하게 충돌한 순간, 고래고래 소리치듯 입을 크게 벌린 박철진의 광기 어린 얼굴.\n\nLOCATION (lock): At the map-covered desk inside an urban militia office, under ordinary office lighting.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Desk corner (Receiving the impact of both fists) — The near corner crosses the lower edge obliquely; used as Connects the facial outburst to its physical impact without oversized foreground distortion; Map of the Korean Peninsula (Spread across the desktop) — A partial view of the upward-facing printed map is visible between and behind the fists; used as Retains the search context at the frame boundary.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient illumination holds detail in the eyes and fists with controlled contrast rather than an invented dramatic light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 핏발 선 눈으로 두 주먹이 책상에 강하게 충돌한 순간, 고래고래 소리치듯 입을 크게 벌린 박철진의 광기 어린 얼굴.\n\nLOCATION (lock): At the map-covered desk inside an urban militia office, under ordinary office lighting.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Desk corner (Receiving the impact of both fists) — The near corner crosses the lower edge obliquely; used as Connects the facial outburst to its physical impact without oversized foreground distortion; Map of the Korean Peninsula (Spread across the desktop) — A partial view of the upward-facing printed map is visible between and behind the fists; used as Retains the search context at the frame boundary.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient illumination holds detail in the eyes and fists with controlled contrast rather than an invented dramatic light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S53sh2__bgfirst_bg.png",
  "asset_id": "e67876e3-065b-4e8d-bf17-983782ebf5b0",
  "input_asset_ids": [
   "70416dc7-78a1-45aa-80e3-0fdad91a577a",
   "91778f94-deb9-4780-814b-5f1f344f3b26"
  ]
 },
 "S53sh2": {
  "input_fingerprint": "ea7618e1114e4974",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 핏발 선 눈으로 두 주먹이 책상에 강하게 충돌한 순간, 고래고래 소리치듯 입을 크게 벌린 박철진의 광기 어린 얼굴.\n\nLOCATION (lock): At the map-covered desk inside an urban militia office, under ordinary office lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Desk corner (Receiving the impact of both fists) — The near corner crosses the lower edge obliquely; used as Connects the facial outburst to its physical impact without oversized foreground distortion; Map of the Korean Peninsula (Spread across the desktop) — A partial view of the upward-facing printed map is visible between and behind the fists; used as Retains the search context at the frame boundary.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient illumination holds detail in the eyes and fists with controlled contrast rather than an invented dramatic light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A map of the Korean Peninsula lies spread across the militia office desk.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 핏발 선 눈으로 두 주먹이 책상에 강하게 충돌한 순간, 고래고래 소리치듯 입을 크게 벌린 박철진의 광기 어린 얼굴.\n\nLOCATION (lock): At the map-covered desk inside an urban militia office, under ordinary office lighting. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Desk corner (Receiving the impact of both fists) — The near corner crosses the lower edge obliquely; used as Connects the facial outburst to its physical impact without oversized foreground distortion; Map of the Korean Peninsula (Spread across the desktop) — A partial view of the upward-facing printed map is visible between and behind the fists; used as Retains the search context at the frame boundary.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient illumination holds detail in the eyes and fists with controlled contrast rather than an invented dramatic light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A map of the Korean Peninsula lies spread across the militia office desk.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 핏발 선 눈으로 두 주먹이 책상에 강하게 충돌한 순간, 고래고래 소리치듯 입을 크게 벌린 박철진의 광기 어린 얼굴.\n\nLOCATION (lock): At the map-covered desk inside an urban militia office, under ordinary office lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Desk corner (Receiving the impact of both fists) — The near corner crosses the lower edge obliquely; used as Connects the facial outburst to its physical impact without oversized foreground distortion; Map of the Korean Peninsula (Spread across the desktop) — A partial view of the upward-facing printed map is visible between and behind the fists; used as Retains the search context at the frame boundary.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient illumination holds detail in the eyes and fists with controlled contrast rather than an invented dramatic light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A map of the Korean Peninsula lies spread across the militia office desk.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S53sh2__bgfirst_bg.png",
     "asset_id": "e67876e3-065b-4e8d-bf17-983782ebf5b0",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S53sh2.png",
     "asset_id": "70416dc7-78a1-45aa-80e3-0fdad91a577a",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1401722>",
     "asset_id": "fee7383c-fb61-4b3a-ba7c-79f2555de00b",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_urban_militia_office_ccf685.png",
     "asset_id": "91778f94-deb9-4780-814b-5f1f344f3b26",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1401722>",
     "asset_id": "fee7383c-fb61-4b3a-ba7c-79f2555de00b",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "박철진의 시선은 화면 우측 전경에 나타난 인물을 향하고 있으며, 두 주먹은 책상 위 지도에 닿아 있음.",
    "built_space": "제공된 기준 사진의 사무실 배경(다트판, 캐비닛, 창문 등)을 올바르게 구현했으나, 카메라 우측 전경에 있어서는 안 될 인물이 위치함.",
    "entities": "박철진의 얼굴과 헤어스타일은 레퍼런스와 일치하지만, 의상(회색 정장)이 기준 이미지(네이비 정장)와 다름. 프롬프트에 없는 미상의 인물이 우측에 등장함.",
    "hard_violations": [
     "[gemini-pro] Invented people: 샷 텍스트에 명시되지 않은 인물(우측 전경) 임의 추가",
     "[gpt-high] 화면 오른쪽 전경에 박철진이 아닌 별도 인물의 머리와 어깨를 추가했다. 박철진만 등장하도록 한 명시적 인물 제한을 위반한다."
    ],
    "physics": "상체를 숙여 책상 위 지도에 양 주먹을 강하게 내리찍은 자세와 무게 지탱 상태는 물리적으로 자연스러움."
   },
   {
    "label": "B",
    "direction": "시선은 렌즈 정면 혹은 약간 아래를 향하고 있으며, 두 주먹은 책상 위 지도 위에 놓여 있음.",
    "built_space": "배경은 사무실을 잘 묘사하고 있으나, 마치 책상 상판이 쪼개져 접혀 올라온 듯한 거대한 V자형 나무 구조물이 전경 양쪽을 비정상적으로 가리고 있음.",
    "entities": "박철진의 얼굴은 기준과 유사하게 연출되었으나, 의상(검은색 반팔 티셔츠)이 레퍼런스의 정장과 완전히 다름.",
    "hard_violations": [
     "[gemini-pro] Invented objects 및 physically impossible staging: 공간에 존재할 수 없는 정체불명의 거대한 판자 구조물이 전경을 가림"
    ],
    "physics": "인물의 주먹이 책상을 누르는 자세는 물리적으로 성립하나, 프레임 전경을 덮고 있는 사선 형태의 구조물들은 어떠한 물리적 지지 기반이나 논리성 없이 허공에 솟아 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "지정되지 않은 인물이 전경에 추가되어 심각한 규칙 위반(Hard Violation)이 발생했으며, 의상 색상도 기준과 다릅니다."
       },
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "전경 양쪽에 물리적으로 설명할 수 없는 거대한 구조물이 생성되었으며, 캐릭터의 의상 또한 기준과 전혀 일치하지 않아 사용할 수 없습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진의 시선은 화면 우측 전경에 나타난 인물을 향하고 있으며, 두 주먹은 책상 위 지도에 닿아 있음.",
        "built_space": "제공된 기준 사진의 사무실 배경(다트판, 캐비닛, 창문 등)을 올바르게 구현했으나, 카메라 우측 전경에 있어서는 안 될 인물이 위치함.",
        "entities": "박철진의 얼굴과 헤어스타일은 레퍼런스와 일치하지만, 의상(회색 정장)이 기준 이미지(네이비 정장)와 다름. 프롬프트에 없는 미상의 인물이 우측에 등장함.",
        "hard_violations": [
         "Invented people: 샷 텍스트에 명시되지 않은 인물(우측 전경) 임의 추가"
        ],
        "physics": "상체를 숙여 책상 위 지도에 양 주먹을 강하게 내리찍은 자세와 무게 지탱 상태는 물리적으로 자연스러움."
       },
       {
        "label": "B",
        "direction": "시선은 렌즈 정면 혹은 약간 아래를 향하고 있으며, 두 주먹은 책상 위 지도 위에 놓여 있음.",
        "built_space": "배경은 사무실을 잘 묘사하고 있으나, 마치 책상 상판이 쪼개져 접혀 올라온 듯한 거대한 V자형 나무 구조물이 전경 양쪽을 비정상적으로 가리고 있음.",
        "entities": "박철진의 얼굴은 기준과 유사하게 연출되었으나, 의상(검은색 반팔 티셔츠)이 레퍼런스의 정장과 완전히 다름.",
        "hard_violations": [
         "Invented objects 및 physically impossible staging: 공간에 존재할 수 없는 정체불명의 거대한 판자 구조물이 전경을 가림"
        ],
        "physics": "인물의 주먹이 책상을 누르는 자세는 물리적으로 성립하나, 프레임 전경을 덮고 있는 사선 형태의 구조물들은 어떠한 물리적 지지 기반이나 논리성 없이 허공에 솟아 있음."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "지정되지 않은 인물이 전경에 추가되어 심각한 규칙 위반(Hard Violation)이 발생했으며, 의상 색상도 기준과 다릅니다."
       },
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "전경 양쪽에 물리적으로 설명할 수 없는 거대한 구조물이 생성되었으며, 캐릭터의 의상 또한 기준과 전혀 일치하지 않아 사용할 수 없습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "박철진의 시선은 화면 우측 전경에 나타난 인물을 향하고 있으며, 두 주먹은 책상 위 지도에 닿아 있음.",
        "built_space": "제공된 기준 사진의 사무실 배경(다트판, 캐비닛, 창문 등)을 올바르게 구현했으나, 카메라 우측 전경에 있어서는 안 될 인물이 위치함.",
        "entities": "박철진의 얼굴과 헤어스타일은 레퍼런스와 일치하지만, 의상(회색 정장)이 기준 이미지(네이비 정장)와 다름. 프롬프트에 없는 미상의 인물이 우측에 등장함.",
        "hard_violations": [
         "Invented people: 샷 텍스트에 명시되지 않은 인물(우측 전경) 임의 추가"
        ],
        "physics": "상체를 숙여 책상 위 지도에 양 주먹을 강하게 내리찍은 자세와 무게 지탱 상태는 물리적으로 자연스러움."
       },
       {
        "label": "B",
        "direction": "시선은 렌즈 정면 혹은 약간 아래를 향하고 있으며, 두 주먹은 책상 위 지도 위에 놓여 있음.",
        "built_space": "배경은 사무실을 잘 묘사하고 있으나, 마치 책상 상판이 쪼개져 접혀 올라온 듯한 거대한 V자형 나무 구조물이 전경 양쪽을 비정상적으로 가리고 있음.",
        "entities": "박철진의 얼굴은 기준과 유사하게 연출되었으나, 의상(검은색 반팔 티셔츠)이 레퍼런스의 정장과 완전히 다름.",
        "hard_violations": [
         "Invented objects 및 physically impossible staging: 공간에 존재할 수 없는 정체불명의 거대한 판자 구조물이 전경을 가림"
        ],
        "physics": "인물의 주먹이 책상을 누르는 자세는 물리적으로 성립하나, 프레임 전경을 덮고 있는 사선 형태의 구조물들은 어떠한 물리적 지지 기반이나 논리성 없이 허공에 솟아 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "추가 인물 없이 포효와 양 주먹의 접촉은 구현했지만, 얼굴 클로즈업보다 넓은 구도와 전경을 가리는 지도 말림, 검은 티셔츠가 지시에서 벗어난다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "양 주먹의 책상 접촉과 고함치는 표정은 명확하지만, 금지된 상대 인물을 추가해 단독 얼굴 클로즈업을 어깨너머 구도로 바꾼 것이 결정적이다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진의 시선과 벌린 입은 카메라 쪽을 향한다. 화면 안에 별도의 시선 대상은 없다. 두 주먹은 아래쪽 지도와 책상 면에 닿아 있어 타격 대상은 맞는다.",
        "built_space": "왼쪽 블라인드 창 하나, 뒤쪽 격자 게시판 하나, 중앙 오른쪽 금속 선반, 열린 출입문 하나, 오른쪽 다트판 하나가 보여 참고 장소의 배치를 대체로 유지한다. 인물은 책상 뒤에서 앞으로 숙였다. 다만 책상 앞 모서리가 화면 하단을 비스듬히 가로지르는 대신 넓은 지도와 솟아오른 양옆 종이가 전경을 차지한다. 얼굴뿐 아니라 상체와 주변 사무실이 상당히 포함되어 지정된 클로즈업보다 넓다.",
        "entities": "보이는 인물은 중년 한국인 남성으로 묘사된 한 명이며, 짧은 검은 가르마 머리와 얼굴 윤곽은 참고 인물에 가깝다. 참고의 정장·흰 셔츠·줄무늬 넥타이 대신 검은 반소매 티셔츠를 입었다. 입을 크게 벌리고 눈 주위를 긴장시켰으나 흰자위의 핏발은 선명하지 않다. 책상 위에는 한반도 인쇄 지도가 위를 향해 놓여 있지만 양옆이 크게 들려 있어 펼쳐진 상태와 다르다. 명확히 판독되는 문구나 화면 자막은 보이지 않는다.",
        "hard_violations": [],
        "physics": "양 주먹의 아래쪽이 지도와 책상에 접촉하고, 팔과 상체가 자연스럽게 이어져 있다. 하체는 프레임 밖이므로 발의 지지는 확인할 수 없지만 공중에 떠 있는 몸으로 보이지 않는다. 양옆으로 치솟은 종이는 하단의 굽은 부분을 통해 책상 위 지도와 이어져 있어 무지지 부유물로 단정할 수 없다. 다만 두 주먹이 부딪치는 순간에 생긴 변형으로 보기에는 말림의 규모가 지나치게 크고, 충돌보다 전경 종이 구조가 강조된다."
       },
       {
        "label": "B",
        "direction": "박철진의 눈과 얼굴은 화면 오른쪽 전경 인물을 향한다. 두 주먹은 각각 아래쪽 책상 위 지도에 닿아 있다. 타격 방향은 맞지만, 고함의 대상으로 별도 인물을 시각화한 것은 허용된 인물 구성에서 벗어난다.",
        "built_space": "왼쪽 블라인드 창 하나, 격자 게시판 하나, 뒤쪽 금속 선반, 열린 문 하나, 오른쪽 다트판 하나가 보여 참고 사무실과 대체로 일치한다. 박철진은 책상 뒤에서 앞으로 숙이고 전경 인물은 책상 반대쪽에 배치되어 있다. 넓은 상체와 책상 면을 보여주는 어깨너머 구도이며, 요청된 얼굴 클로즈업과 하단의 비스듬한 책상 모서리 구성은 구현되지 않았다.",
        "entities": "중심 인물은 참고와 유사한 중년 남성의 얼굴, 짧은 검은 머리, 귀의 작은 장식을 갖췄다. 흰 셔츠와 사선 줄무늬 넥타이는 맞지만 재킷은 참고의 짙은 남색보다 회갈색에 가깝다. 눈 주위가 붉고 입을 크게 벌려 고함치는 표정이 명확하다. 위를 향한 한반도 지도가 책상에 펼쳐져 있다. 오른쪽에는 별도 남성의 뒤통수·귀·어깨가 크게 보여 허용된 한 명 외의 인물이 존재한다. 명확한 자막이나 판독 가능한 문구는 보이지 않는다.",
        "hard_violations": [
         "화면 오른쪽 전경에 박철진이 아닌 별도 인물의 머리와 어깨를 추가했다. 박철진만 등장하도록 한 명시적 인물 제한을 위반한다."
        ],
        "physics": "두 주먹 모두 책상 위 지도에 안정적으로 접촉하며 손목과 팔의 연결도 자연스럽다. 앞으로 기울어진 몸은 양팔을 통해 책상에 힘을 싣는 자세로 읽힌다. 지도는 책상 면에 지지되어 있고 부유하는 물체는 없다. 하체가 잘린 인물들의 발 지지는 확인할 수 없지만, 보이는 부분에 물리적으로 불가능한 부유나 관절 배치는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "추가 인물 없이 포효와 양 주먹의 접촉은 구현했지만, 얼굴 클로즈업보다 넓은 구도와 전경을 가리는 지도 말림, 검은 티셔츠가 지시에서 벗어난다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "양 주먹의 책상 접촉과 고함치는 표정은 명확하지만, 금지된 상대 인물을 추가해 단독 얼굴 클로즈업을 어깨너머 구도로 바꾼 것이 결정적이다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "박철진의 시선과 벌린 입은 카메라 쪽을 향한다. 화면 안에 별도의 시선 대상은 없다. 두 주먹은 아래쪽 지도와 책상 면에 닿아 있어 타격 대상은 맞는다.",
        "built_space": "왼쪽 블라인드 창 하나, 뒤쪽 격자 게시판 하나, 중앙 오른쪽 금속 선반, 열린 출입문 하나, 오른쪽 다트판 하나가 보여 참고 장소의 배치를 대체로 유지한다. 인물은 책상 뒤에서 앞으로 숙였다. 다만 책상 앞 모서리가 화면 하단을 비스듬히 가로지르는 대신 넓은 지도와 솟아오른 양옆 종이가 전경을 차지한다. 얼굴뿐 아니라 상체와 주변 사무실이 상당히 포함되어 지정된 클로즈업보다 넓다.",
        "entities": "보이는 인물은 중년 한국인 남성으로 묘사된 한 명이며, 짧은 검은 가르마 머리와 얼굴 윤곽은 참고 인물에 가깝다. 참고의 정장·흰 셔츠·줄무늬 넥타이 대신 검은 반소매 티셔츠를 입었다. 입을 크게 벌리고 눈 주위를 긴장시켰으나 흰자위의 핏발은 선명하지 않다. 책상 위에는 한반도 인쇄 지도가 위를 향해 놓여 있지만 양옆이 크게 들려 있어 펼쳐진 상태와 다르다. 명확히 판독되는 문구나 화면 자막은 보이지 않는다.",
        "hard_violations": [],
        "physics": "양 주먹의 아래쪽이 지도와 책상에 접촉하고, 팔과 상체가 자연스럽게 이어져 있다. 하체는 프레임 밖이므로 발의 지지는 확인할 수 없지만 공중에 떠 있는 몸으로 보이지 않는다. 양옆으로 치솟은 종이는 하단의 굽은 부분을 통해 책상 위 지도와 이어져 있어 무지지 부유물로 단정할 수 없다. 다만 두 주먹이 부딪치는 순간에 생긴 변형으로 보기에는 말림의 규모가 지나치게 크고, 충돌보다 전경 종이 구조가 강조된다."
       },
       {
        "label": "A",
        "direction": "박철진의 눈과 얼굴은 화면 오른쪽 전경 인물을 향한다. 두 주먹은 각각 아래쪽 책상 위 지도에 닿아 있다. 타격 방향은 맞지만, 고함의 대상으로 별도 인물을 시각화한 것은 허용된 인물 구성에서 벗어난다.",
        "built_space": "왼쪽 블라인드 창 하나, 격자 게시판 하나, 뒤쪽 금속 선반, 열린 문 하나, 오른쪽 다트판 하나가 보여 참고 사무실과 대체로 일치한다. 박철진은 책상 뒤에서 앞으로 숙이고 전경 인물은 책상 반대쪽에 배치되어 있다. 넓은 상체와 책상 면을 보여주는 어깨너머 구도이며, 요청된 얼굴 클로즈업과 하단의 비스듬한 책상 모서리 구성은 구현되지 않았다.",
        "entities": "중심 인물은 참고와 유사한 중년 남성의 얼굴, 짧은 검은 머리, 귀의 작은 장식을 갖췄다. 흰 셔츠와 사선 줄무늬 넥타이는 맞지만 재킷은 참고의 짙은 남색보다 회갈색에 가깝다. 눈 주위가 붉고 입을 크게 벌려 고함치는 표정이 명확하다. 위를 향한 한반도 지도가 책상에 펼쳐져 있다. 오른쪽에는 별도 남성의 뒤통수·귀·어깨가 크게 보여 허용된 한 명 외의 인물이 존재한다. 명확한 자막이나 판독 가능한 문구는 보이지 않는다.",
        "hard_violations": [
         "화면 오른쪽 전경에 박철진이 아닌 별도 인물의 머리와 어깨를 추가했다. 박철진만 등장하도록 한 명시적 인물 제한을 위반한다."
        ],
        "physics": "두 주먹 모두 책상 위 지도에 안정적으로 접촉하며 손목과 팔의 연결도 자연스럽다. 앞으로 기울어진 몸은 양팔을 통해 책상에 힘을 싣는 자세로 읽힌다. 지도는 책상 면에 지지되어 있고 부유하는 물체는 없다. 하체가 잘린 인물들의 발 지지는 확인할 수 없지만, 보이는 부분에 물리적으로 불가능한 부유나 관절 배치는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.5,
    "B": 1.5
   },
   "adjusted": {
    "A": 1.25,
    "B": 1.25
   },
   "violations": {
    "A": [
     "[gemini-pro] Invented people: 샷 텍스트에 명시되지 않은 인물(우측 전경) 임의 추가",
     "[gpt-high] 화면 오른쪽 전경에 박철진이 아닌 별도 인물의 머리와 어깨를 추가했다. 박철진만 등장하도록 한 명시적 인물 제한을 위반한다."
    ],
    "B": [
     "[gemini-pro] Invented objects 및 physically impossible staging: 공간에 존재할 수 없는 정체불명의 거대한 판자 구조물이 전경을 가림"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1250,
   "B": 1250
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1250,
    "verdict_ko": "지정되지 않은 인물이 전경에 추가되어 심각한 규칙 위반(Hard Violation)이 발생했으며, 의상 색상도 기준과 다릅니다.  ★위반: [gemini-pro] Invented people: 샷 텍스트에 명시되지 않은 인물(우측 전경) 임의 추가 / [gpt-high] 화면 오른쪽 전경에 박철진이 아닌 별도 인물의 머리와 어깨를 추가했다. 박철진만 등장하도록 한 명시적 인물 제한을 위반한다."
   },
   {
    "label": "B",
    "score": 1250,
    "verdict_ko": "전경 양쪽에 물리적으로 설명할 수 없는 거대한 구조물이 생성되었으며, 캐릭터의 의상 또한 기준과 전혀 일치하지 않아 사용할 수 없습니다.  ★위반: [gemini-pro] Invented objects 및 physically impossible staging: 공간에 존재할 수 없는 정체불명의 거대한 판자 구조물이 전경을 가림"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_urban_militia_office_ccf685.png",
    "asset_id": "91778f94-deb9-4780-814b-5f1f344f3b26",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1401722>",
    "asset_id": "fee7383c-fb61-4b3a-ba7c-79f2555de00b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-48bd-7923-b69f-014d0b570921",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S53sh2__bgfirst_bg.png",
   "bg_asset_id": "e67876e3-065b-4e8d-bf17-983782ebf5b0",
   "bg_record_key": "S53sh2::bgfirst_bg",
   "chain_winner": true,
   "authority": "groupbg",
   "group_key": "urban_militia_office",
   "groupbg_asset_id": "91778f94-deb9-4780-814b-5f1f344f3b26"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S53sh2::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:10:55.462391+00:00",
  "fingerprint": "208a9cfef59a8fe62a75f073d951a07c0f69f790592fe4f17d188bcc72cd9aab",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S53sh2_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S53sh2_sel.png",
  "source_sha256": "e206adf0e28c7977e54a61aae179849680a48817f5d04cfdb86c77ded2d63012",
  "file": "S53sh2_cine.png",
  "staged_sha256": "76f585a9ba7183d9a24309d5323f6b274f5d964cb3501274f20f576a002c698a",
  "latency_ms": 11243
 },
 "S53sh3::signage": {
  "fp": "a6d7c9c683d47ee0",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S53sh3": {
  "input_fingerprint": "f70dfaf31b3aec04",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 박철진의 귓가에 입을 바짝 댄 채 은밀하게 속삭이는 수하 1의 조심스러운 상체.\n\nLOCATION (lock): Beside the commander's desk inside the urban militia office, under ordinary office lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Desk (Remaining beneath the pair during the whispered report) — The same side edge recedes beneath their upper bodies; used as Maintains spatial continuity with the fist impact and anchors their close positioning.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding office illumination and contrast so the lowered voices, not a lighting shift, create the secrecy.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The map of the Korean Peninsula remains spread across the militia office desk.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리); 수하 1 (한국인, 성인 남성, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 박철진의 귓가에 입을 바짝 댄 채 은밀하게 속삭이는 수하 1의 조심스러운 상체.\n\nLOCATION (lock): Beside the commander's desk inside the urban militia office, under ordinary office lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Desk (Remaining beneath the pair during the whispered report) — The same side edge recedes beneath their upper bodies; used as Maintains spatial continuity with the fist impact and anchors their close positioning.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding office illumination and contrast so the lowered voices, not a lighting shift, create the secrecy.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The map of the Korean Peninsula remains spread across the militia office desk.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리); 수하 1 (한국인, 성인 남성, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 박철진의 귓가에 입을 바짝 댄 채 은밀하게 속삭이는 수하 1의 조심스러운 상체.\n\nLOCATION (lock): Beside the commander's desk inside the urban militia office, under ordinary office lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Desk (Remaining beneath the pair during the whispered report) — The same side edge recedes beneath their upper bodies; used as Maintains spatial continuity with the fist impact and anchors their close positioning.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding office illumination and contrast so the lowered voices, not a lighting shift, create the secrecy.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The map of the Korean Peninsula remains spread across the militia office desk.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리); 수하 1 (한국인, 성인 남성, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "수하 1은 박철진의 귀가 아닌 얼굴 정면을 바라보며 대화하고 있음.",
    "built_space": "사무실의 블라인드, 출입문, 다트판 등 배경 요소와 책상이 이전 컷과 일관성 있게 배치됨.",
    "entities": "두 인물의 외양과 복장이 레퍼런스와 일치하며, 책상 위 한반도 지도도 존재함.",
    "hard_violations": [
     "[gpt-high] 배경 다트판의 둘레 숫자가 읽혀, 어디에도 읽을 수 있는 문자를 두지 말라는 조건을 위반한다."
    ],
    "physics": "두 사람 모두 책상에 양손이나 팔을 짚고 상체를 숙인 자연스러운 자세를 취하고 있음."
   },
   {
    "label": "B",
    "direction": "수하 1의 시선과 고개가 박철진의 귀를 향해 있으며, 박철진은 정면 아래를 응시하고 있음.",
    "built_space": "이전 컷의 사무실 배경(좌측 블라인드, 우측 출입문 및 다트판)과 책상이 정확한 위치와 비례로 묘사됨.",
    "entities": "박철진과 수하 1의 얼굴, 헤어스타일, 복장이 레퍼런스와 일치함. 책상 위에 한반도 지도가 올바르게 놓여 있음.",
    "hard_violations": [
     "[gpt-high] 배경 다트판의 둘레 숫자가 읽혀, 어디에도 읽을 수 있는 문자를 두지 말라는 조건을 위반한다."
    ],
    "physics": "수하 1은 상체를 숙여 박철진에게 다가가며 오른손으로 입을 가리고 있고, 박철진은 책상에 양 주먹을 짚고 안정적으로 체중을 지지하고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "수하 1이 박철진의 귀에 입을 바짝 대고 속삭이는 핵심 행동을 매우 정확하게 구현하였으며, 이전 컷과의 연속성도 훌륭합니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "수하 1이 귓가에 속삭이지 않고 거리를 둔 채 얼굴을 마주보고 있어 프롬프트의 핵심 행동 지시를 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "수하 1의 시선과 고개가 박철진의 귀를 향해 있으며, 박철진은 정면 아래를 응시하고 있음.",
        "built_space": "이전 컷의 사무실 배경(좌측 블라인드, 우측 출입문 및 다트판)과 책상이 정확한 위치와 비례로 묘사됨.",
        "entities": "박철진과 수하 1의 얼굴, 헤어스타일, 복장이 레퍼런스와 일치함. 책상 위에 한반도 지도가 올바르게 놓여 있음.",
        "hard_violations": [],
        "physics": "수하 1은 상체를 숙여 박철진에게 다가가며 오른손으로 입을 가리고 있고, 박철진은 책상에 양 주먹을 짚고 안정적으로 체중을 지지하고 있음."
       },
       {
        "label": "A",
        "direction": "수하 1은 박철진의 귀가 아닌 얼굴 정면을 바라보며 대화하고 있음.",
        "built_space": "사무실의 블라인드, 출입문, 다트판 등 배경 요소와 책상이 이전 컷과 일관성 있게 배치됨.",
        "entities": "두 인물의 외양과 복장이 레퍼런스와 일치하며, 책상 위 한반도 지도도 존재함.",
        "hard_violations": [],
        "physics": "두 사람 모두 책상에 양손이나 팔을 짚고 상체를 숙인 자연스러운 자세를 취하고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "수하 1이 박철진의 귀에 입을 바짝 대고 속삭이는 핵심 행동을 매우 정확하게 구현하였으며, 이전 컷과의 연속성도 훌륭합니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "수하 1이 귓가에 속삭이지 않고 거리를 둔 채 얼굴을 마주보고 있어 프롬프트의 핵심 행동 지시를 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "수하 1의 시선과 고개가 박철진의 귀를 향해 있으며, 박철진은 정면 아래를 응시하고 있음.",
        "built_space": "이전 컷의 사무실 배경(좌측 블라인드, 우측 출입문 및 다트판)과 책상이 정확한 위치와 비례로 묘사됨.",
        "entities": "박철진과 수하 1의 얼굴, 헤어스타일, 복장이 레퍼런스와 일치함. 책상 위에 한반도 지도가 올바르게 놓여 있음.",
        "hard_violations": [],
        "physics": "수하 1은 상체를 숙여 박철진에게 다가가며 오른손으로 입을 가리고 있고, 박철진은 책상에 양 주먹을 짚고 안정적으로 체중을 지지하고 있음."
       },
       {
        "label": "A",
        "direction": "수하 1은 박철진의 귀가 아닌 얼굴 정면을 바라보며 대화하고 있음.",
        "built_space": "사무실의 블라인드, 출입문, 다트판 등 배경 요소와 책상이 이전 컷과 일관성 있게 배치됨.",
        "entities": "두 인물의 외양과 복장이 레퍼런스와 일치하며, 책상 위 한반도 지도도 존재함.",
        "hard_violations": [],
        "physics": "두 사람 모두 책상에 양손이나 팔을 짚고 상체를 숙인 자연스러운 자세를 취하고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "귓가에 입을 붙이고 조심스럽게 속삭이는 상체와 책상 위 공간 관계가 더 정확하지만, 읽히는 다트판 숫자는 문자 금지 조건을 위반한다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "박철진의 외모와 책상 연속성은 좋지만 수하의 입이 귀보다 뺨 쪽을 향하며 떨어져 있고, 다트판 숫자도 문자 금지 조건을 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 수하가 오른쪽 박철진의 가까운 귀 부위로 얼굴과 입을 밀착하고 손으로 입 주변을 가린다. 귀 자체는 가려졌지만 속삭임의 대상과 방향은 분명하다. 수하는 박철진의 옆얼굴 쪽을 내려다보고, 박철진은 화면 왼쪽 아래를 경계하듯 본다.",
        "built_space": "두 사람의 허리 부근까지 담은 미디엄 숏이며 목재 책상 한 개가 상체 아래로 이어진다. 왼쪽 블라인드 창 한 구역, 중앙 뒤 수납 선반 한 개, 오른쪽 열린 출입문 한 개, 우측 벽 다트판 한 개와 그 아래 서류함·금속 용기들이 보인다. 천장 형광등과 낮의 창빛도 앞 장면에 부합한다. 박철진은 책상 뒤에서 몸을 기울이고 수하는 바로 옆에 붙어 있으며, 불가능한 반사나 중복된 고정 설비는 보이지 않는다.",
        "entities": "등장인물은 짧은 검은 머리의 한국인 성인 남성 두 명뿐이다. 박철진의 회갈색 정장, 흰 셔츠, 사선 넥타이와 귀걸이는 이전 장면에 맞지만 얼굴 윤곽은 참조보다 다소 젊고 갸름하다. 수하의 남색 정장과 넥타이 없는 흰 셔츠, 머리와 옆얼굴은 인물 참조에 대체로 맞는다. 한반도 지도는 책상에 펼쳐져 있으며 종이 더미와 보조 지도도 남아 있다. 배경 다트판 둘레에는 판독 가능한 숫자가 있다.",
        "hard_violations": [
         "배경 다트판의 둘레 숫자가 읽혀, 어디에도 읽을 수 있는 문자를 두지 말라는 조건을 위반한다."
        ],
        "physics": "수하의 굽힌 팔꿈치가 책상에 닿아 기울어진 상체를 받치며, 들어 올린 손은 팔과 자연스럽게 연결되어 입을 가린다. 박철진의 오른쪽에 보이는 주먹은 지도 위에 놓이고 반대 손도 책상 가장자리에 닿아 있다. 지도와 서류는 책상 위에 놓여 있다. 하체는 프레임 밖이지만 보이는 접촉과 상체 자세에 부유나 불가능한 관절은 없다."
       },
       {
        "label": "B",
        "direction": "오른쪽 수하가 왼쪽 박철진을 향해 몸을 숙이지만, 입은 귀에 붙기보다 옆눈과 뺨 가까운 앞쪽 공간을 향하고 사이에 간격이 남는다. 수하의 시선도 박철진의 옆얼굴을 향한다. 박철진은 앞쪽 아래를 응시한다. 은밀한 대화는 보이나 '귓가에 입을 바짝 댄' 정확한 동작은 약하다.",
        "built_space": "책상 뒤 박철진과 옆에서 숙인 수하를 허리 부근까지 담은 미디엄 숏이다. 목재 책상 한 개, 왼쪽 블라인드 창 한 구역, 박철진 뒤 의자 한 개, 벽면 철망 게시판 한 개, 중앙 수납 선반 한 개, 뒤쪽 열린 문 한 개, 오른쪽 다트판 한 개가 보인다. 책상 가장자리와 한반도 지도가 두 상체 아래를 연결하고 기존 사무실의 재료와 낮 조명을 유지한다. 고정 설비 중복이나 불가능한 반사는 없다.",
        "entities": "두 명 모두 한국인 성인 남성으로 보이고 추가 인물은 없다. 박철진의 중년 얼굴, 짧게 넘긴 검은 머리, 귀걸이, 회갈색 정장과 사선 넥타이는 이전 장면 및 인물 참조에 잘 맞는다. 수하의 검은 머리, 남색 정장과 열린 흰 셔츠도 참조와 부합한다. 한반도 지도는 책상 위에 펼쳐져 있고 보조 지도와 종이 더미도 보인다. 다트판의 숫자는 판독 가능하다.",
        "hard_violations": [
         "배경 다트판의 둘레 숫자가 읽혀, 어디에도 읽을 수 있는 문자를 두지 말라는 조건을 위반한다."
        ],
        "physics": "박철진은 양손을 지도와 책상에 짚어 앞으로 기울어진 몸을 받친다. 수하도 앞쪽 손바닥과 뒤쪽 손을 책상에 대고 상체를 숙여 지지 관계가 자연스럽다. 손가락과 팔의 연결에 명백한 불가능성은 없고 지도와 서류도 책상에 지지된다. 프레임 밖 다리를 확인할 수 없다는 이유만으로 부유라고 볼 근거는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "귓가에 입을 붙이고 조심스럽게 속삭이는 상체와 책상 위 공간 관계가 더 정확하지만, 읽히는 다트판 숫자는 문자 금지 조건을 위반한다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "박철진의 외모와 책상 연속성은 좋지만 수하의 입이 귀보다 뺨 쪽을 향하며 떨어져 있고, 다트판 숫자도 문자 금지 조건을 위반한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽 수하가 오른쪽 박철진의 가까운 귀 부위로 얼굴과 입을 밀착하고 손으로 입 주변을 가린다. 귀 자체는 가려졌지만 속삭임의 대상과 방향은 분명하다. 수하는 박철진의 옆얼굴 쪽을 내려다보고, 박철진은 화면 왼쪽 아래를 경계하듯 본다.",
        "built_space": "두 사람의 허리 부근까지 담은 미디엄 숏이며 목재 책상 한 개가 상체 아래로 이어진다. 왼쪽 블라인드 창 한 구역, 중앙 뒤 수납 선반 한 개, 오른쪽 열린 출입문 한 개, 우측 벽 다트판 한 개와 그 아래 서류함·금속 용기들이 보인다. 천장 형광등과 낮의 창빛도 앞 장면에 부합한다. 박철진은 책상 뒤에서 몸을 기울이고 수하는 바로 옆에 붙어 있으며, 불가능한 반사나 중복된 고정 설비는 보이지 않는다.",
        "entities": "등장인물은 짧은 검은 머리의 한국인 성인 남성 두 명뿐이다. 박철진의 회갈색 정장, 흰 셔츠, 사선 넥타이와 귀걸이는 이전 장면에 맞지만 얼굴 윤곽은 참조보다 다소 젊고 갸름하다. 수하의 남색 정장과 넥타이 없는 흰 셔츠, 머리와 옆얼굴은 인물 참조에 대체로 맞는다. 한반도 지도는 책상에 펼쳐져 있으며 종이 더미와 보조 지도도 남아 있다. 배경 다트판 둘레에는 판독 가능한 숫자가 있다.",
        "hard_violations": [
         "배경 다트판의 둘레 숫자가 읽혀, 어디에도 읽을 수 있는 문자를 두지 말라는 조건을 위반한다."
        ],
        "physics": "수하의 굽힌 팔꿈치가 책상에 닿아 기울어진 상체를 받치며, 들어 올린 손은 팔과 자연스럽게 연결되어 입을 가린다. 박철진의 오른쪽에 보이는 주먹은 지도 위에 놓이고 반대 손도 책상 가장자리에 닿아 있다. 지도와 서류는 책상 위에 놓여 있다. 하체는 프레임 밖이지만 보이는 접촉과 상체 자세에 부유나 불가능한 관절은 없다."
       },
       {
        "label": "A",
        "direction": "오른쪽 수하가 왼쪽 박철진을 향해 몸을 숙이지만, 입은 귀에 붙기보다 옆눈과 뺨 가까운 앞쪽 공간을 향하고 사이에 간격이 남는다. 수하의 시선도 박철진의 옆얼굴을 향한다. 박철진은 앞쪽 아래를 응시한다. 은밀한 대화는 보이나 '귓가에 입을 바짝 댄' 정확한 동작은 약하다.",
        "built_space": "책상 뒤 박철진과 옆에서 숙인 수하를 허리 부근까지 담은 미디엄 숏이다. 목재 책상 한 개, 왼쪽 블라인드 창 한 구역, 박철진 뒤 의자 한 개, 벽면 철망 게시판 한 개, 중앙 수납 선반 한 개, 뒤쪽 열린 문 한 개, 오른쪽 다트판 한 개가 보인다. 책상 가장자리와 한반도 지도가 두 상체 아래를 연결하고 기존 사무실의 재료와 낮 조명을 유지한다. 고정 설비 중복이나 불가능한 반사는 없다.",
        "entities": "두 명 모두 한국인 성인 남성으로 보이고 추가 인물은 없다. 박철진의 중년 얼굴, 짧게 넘긴 검은 머리, 귀걸이, 회갈색 정장과 사선 넥타이는 이전 장면 및 인물 참조에 잘 맞는다. 수하의 검은 머리, 남색 정장과 열린 흰 셔츠도 참조와 부합한다. 한반도 지도는 책상 위에 펼쳐져 있고 보조 지도와 종이 더미도 보인다. 다트판의 숫자는 판독 가능하다.",
        "hard_violations": [
         "배경 다트판의 둘레 숫자가 읽혀, 어디에도 읽을 수 있는 문자를 두지 말라는 조건을 위반한다."
        ],
        "physics": "박철진은 양손을 지도와 책상에 짚어 앞으로 기울어진 몸을 받친다. 수하도 앞쪽 손바닥과 뒤쪽 손을 책상에 대고 상체를 숙여 지지 관계가 자연스럽다. 손가락과 팔의 연결에 명백한 불가능성은 없고 지도와 서류도 책상에 지지된다. 프레임 밖 다리를 확인할 수 없다는 이유만으로 부유라고 볼 근거는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.3,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.05,
    "B": 1.75
   },
   "violations": {
    "B": [
     "[gpt-high] 배경 다트판의 둘레 숫자가 읽혀, 어디에도 읽을 수 있는 문자를 두지 말라는 조건을 위반한다."
    ],
    "A": [
     "[gpt-high] 배경 다트판의 둘레 숫자가 읽혀, 어디에도 읽을 수 있는 문자를 두지 말라는 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 1750,
   "A": 1050
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "수하 1이 박철진의 귀에 입을 바짝 대고 속삭이는 핵심 행동을 매우 정확하게 구현하였으며, 이전 컷과의 연속성도 훌륭합니다.  ★위반: [gpt-high] 배경 다트판의 둘레 숫자가 읽혀, 어디에도 읽을 수 있는 문자를 두지 말라는 조건을 위반한다."
   },
   {
    "label": "A",
    "score": 1050,
    "verdict_ko": "수하 1이 귓가에 속삭이지 않고 거리를 둔 채 얼굴을 마주보고 있어 프롬프트의 핵심 행동 지시를 위반했습니다.  ★위반: [gpt-high] 배경 다트판의 둘레 숫자가 읽혀, 어디에도 읽을 수 있는 문자를 두지 말라는 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S53sh2_sel.png",
    "asset_id": "c2f2e1ce-4a99-4a64-bd1f-0a21d2090144",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1401722>",
    "asset_id": "fee7383c-fb61-4b3a-ba7c-79f2555de00b",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 수하 1: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1192369>",
    "asset_id": "a969dd53-900f-4bf4-b80a-e38b77b03fbf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-4d97-728a-bc5c-666e144ce42e",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S53sh2"
  }
 },
 "S53sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:12:12.205723+00:00",
  "fingerprint": "5dda2cf29bace7dcfb52f9e9763526c92d702e1f648c757e011450a5d648ca9b",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S53sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S53sh3_sel.png",
  "source_sha256": "2ee99ac0018d89892427168a92ae0d544132322d8cf3b7404eeed96e6d54fa97",
  "file": "S53sh3_cine.png",
  "staged_sha256": "b0dbd181c072b47375699cccbf449d80ca66b4ff2ef321ebe49b5ecd60842011",
  "latency_ms": 11093
 },
 "S54sh2::signage": {
  "fp": "9e517e5410cb2101",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S54sh2": {
  "input_fingerprint": "645eb72edd6de713",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 국방장관을 향해 몸을 쑥 내민 채 두 눈을 번뜩이며 속삭이듯 입을 반쯤 벌린 윤성찬의 굳은 상체.\n\nLOCATION (lock): In the visitor conversation area inside the defense minister's office, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate daytime illumination gives the faces controlled tonal separation without adding a specific fixture or colored light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 국방장관을 향해 몸을 쑥 내민 채 두 눈을 번뜩이며 속삭이듯 입을 반쯤 벌린 윤성찬의 굳은 상체.\n\nLOCATION (lock): In the visitor conversation area inside the defense minister's office, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate daytime illumination gives the faces controlled tonal separation without adding a specific fixture or colored light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 국방장관을 향해 몸을 쑥 내민 채 두 눈을 번뜩이며 속삭이듯 입을 반쯤 벌린 윤성찬의 굳은 상체.\n\nLOCATION (lock): In the visitor conversation area inside the defense minister's office, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate daytime illumination gives the faces controlled tonal separation without adding a specific fixture or colored light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "윤성찬이 화면 전경의 인물(국방장관)을 향해 상체를 기울이고 강렬한 시선을 보냄.",
    "built_space": "사무실 내부 공간. 벽에 산 풍경 액자가 있으나, 레퍼런스의 검은색 가죽 가구와 달리 갈색 가죽 소파와 의자가 배치됨.",
    "entities": "윤성찬의 얼굴 특징은 일치하나 레퍼런스의 안경이 누락됨. 샷 명단에 없는 인물이 전경에 등장하며, 이전 샷 인물의 회색 셔츠를 입고 있음.",
    "hard_violations": [
     "[gemini-pro] 명단에 없는 인물 등장",
     "[gemini-pro] 이전 샷 인물의 의상(회색 셔츠) 복사",
     "[gemini-pro] 지정된 공간의 가구 재질 위반 (갈색 가죽)",
     "[gpt-high] 등장 가능한 인물이 윤성찬 한 명으로 제한되어 있는데, 오른쪽 전경에 상대 남성의 머리와 상체를 추가했다.",
     "[gpt-high] 장소가 고정된 접견 구역에 참조의 좌석과 다른 갈색 가죽 안락의자를 추가했다."
    ],
    "physics": "몸을 앞으로 숙인 자세가 자연스러우며 하체에 의해 지탱되고 있음."
   },
   {
    "label": "B",
    "direction": "윤성찬이 화면 전경의 인물(국방장관)에게 몸을 내밀며 시선을 맞추고 있음.",
    "built_space": "사무실 내부 공간. 창문, 책상, 작은 식물이 놓인 나무 테이블 등 레퍼런스의 요소들이 잘 유지되었고 검은색 가죽 의자 일부가 보임.",
    "entities": "윤성찬의 얼굴 특징은 일치하나 안경이 누락됨. 샷 명단에 없는 인물이 전경에 등장함.",
    "hard_violations": [
     "[gemini-pro] 명단에 없는 인물 등장",
     "[gpt-high] 등장 가능한 인물이 윤성찬 한 명으로 제한되어 있는데, 오른쪽 전경에 상대 남성의 머리와 상체를 추가했다."
    ],
    "physics": "상체를 쑥 내민 자세가 물리적으로 어색함 없이 지탱되고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "명단에 없는 인물이 화면 앞쪽에 등장하는 위반이 있으나, 배경 공간의 디테일을 유지하고 금지된 의상 복사를 피하여 상대적으로 우수함."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "명단에 없는 인물이 등장할 뿐만 아니라 이전 샷의 의상을 복사하였고, 공간의 가구 재질(갈색 가죽)을 임의로 변경하여 지침을 다수 위반함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "윤성찬이 화면 전경의 인물(국방장관)을 향해 상체를 기울이고 강렬한 시선을 보냄.",
        "built_space": "사무실 내부 공간. 벽에 산 풍경 액자가 있으나, 레퍼런스의 검은색 가죽 가구와 달리 갈색 가죽 소파와 의자가 배치됨.",
        "entities": "윤성찬의 얼굴 특징은 일치하나 레퍼런스의 안경이 누락됨. 샷 명단에 없는 인물이 전경에 등장하며, 이전 샷 인물의 회색 셔츠를 입고 있음.",
        "hard_violations": [
         "명단에 없는 인물 등장",
         "이전 샷 인물의 의상(회색 셔츠) 복사",
         "지정된 공간의 가구 재질 위반 (갈색 가죽)"
        ],
        "physics": "몸을 앞으로 숙인 자세가 자연스러우며 하체에 의해 지탱되고 있음."
       },
       {
        "label": "B",
        "direction": "윤성찬이 화면 전경의 인물(국방장관)에게 몸을 내밀며 시선을 맞추고 있음.",
        "built_space": "사무실 내부 공간. 창문, 책상, 작은 식물이 놓인 나무 테이블 등 레퍼런스의 요소들이 잘 유지되었고 검은색 가죽 의자 일부가 보임.",
        "entities": "윤성찬의 얼굴 특징은 일치하나 안경이 누락됨. 샷 명단에 없는 인물이 전경에 등장함.",
        "hard_violations": [
         "명단에 없는 인물 등장"
        ],
        "physics": "상체를 쑥 내민 자세가 물리적으로 어색함 없이 지탱되고 있음."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "명단에 없는 인물이 화면 앞쪽에 등장하는 위반이 있으나, 배경 공간의 디테일을 유지하고 금지된 의상 복사를 피하여 상대적으로 우수함."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "명단에 없는 인물이 등장할 뿐만 아니라 이전 샷의 의상을 복사하였고, 공간의 가구 재질(갈색 가죽)을 임의로 변경하여 지침을 다수 위반함."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "윤성찬이 화면 전경의 인물(국방장관)을 향해 상체를 기울이고 강렬한 시선을 보냄.",
        "built_space": "사무실 내부 공간. 벽에 산 풍경 액자가 있으나, 레퍼런스의 검은색 가죽 가구와 달리 갈색 가죽 소파와 의자가 배치됨.",
        "entities": "윤성찬의 얼굴 특징은 일치하나 레퍼런스의 안경이 누락됨. 샷 명단에 없는 인물이 전경에 등장하며, 이전 샷 인물의 회색 셔츠를 입고 있음.",
        "hard_violations": [
         "명단에 없는 인물 등장",
         "이전 샷 인물의 의상(회색 셔츠) 복사",
         "지정된 공간의 가구 재질 위반 (갈색 가죽)"
        ],
        "physics": "몸을 앞으로 숙인 자세가 자연스러우며 하체에 의해 지탱되고 있음."
       },
       {
        "label": "B",
        "direction": "윤성찬이 화면 전경의 인물(국방장관)에게 몸을 내밀며 시선을 맞추고 있음.",
        "built_space": "사무실 내부 공간. 창문, 책상, 작은 식물이 놓인 나무 테이블 등 레퍼런스의 요소들이 잘 유지되었고 검은색 가죽 의자 일부가 보임.",
        "entities": "윤성찬의 얼굴 특징은 일치하나 안경이 누락됨. 샷 명단에 없는 인물이 전경에 등장함.",
        "hard_violations": [
         "명단에 없는 인물 등장"
        ],
        "physics": "상체를 쑥 내민 자세가 물리적으로 어색함 없이 지탱되고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "앞으로 내민 굳은 상체와 상대를 겨눈 눈빛, 기존 접견실의 공간 연속성은 더 충실하지만, 허용되지 않은 상대 인물을 추가해 실격이며 참조의 안경도 빠졌다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "미디엄 숏의 전경사 상체와 반쯤 열린 입은 부합하지만, 허용되지 않은 상대 인물과 참조에 없던 갈색 안락의자를 추가했고 접견실의 연속성도 떨어진다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "윤성찬의 상체와 얼굴이 화면 오른쪽 전경의 검은 머리 남자를 향하며, 두 눈도 그 남자의 얼굴 쪽에 고정되어 있다. 국방장관에게 가까이 다가가 속삭이는 방향 관계는 읽히지만, 그 상대를 실제 인물로 화면에 넣은 것은 등장인물 제한에 어긋난다. 무기나 방향성 소품은 없다.",
        "built_space": "뒤쪽에 창 구역 두 곳과 그 사이 벽의 액자 일부 하나, 목재 수납장, 검은 가죽 좌석이 보인다. 전경에는 낮은 탁자 하나와 작은 화분 하나, 쟁반 하나, 책 더미가 있다. 창가의 별도 화분 하나와 장식패들도 보이며, 참조의 창가 수납장·가죽 좌석·목재 탁자로 이루어진 접견 구역을 비교적 잘 유지한다. 윤성찬은 왼쪽 좌석 앞에서 탁자 너머 상대 쪽으로 기울어 있다. 불가능한 반사나 명백한 고정 설비 중복은 보이지 않는다.",
        "entities": "주인공은 회색 옆가르마 머리, 깊은 이마와 눈가 주름, 검회색 재킷과 짙은 남색 셔츠를 가진 고령의 한국인 남성으로 표현되어 참조와 대체로 맞는다. 다만 참조의 얇은 금속테 안경이 없다. 입은 조금 벌어져 있고 눈은 정상적인 홍채와 동공을 유지한다. 오른쪽에는 별개의 젊은 남성 머리·귀·목·어깨가 뚜렷하게 추가되어 있다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "등장 가능한 인물이 윤성찬 한 명으로 제한되어 있는데, 오른쪽 전경에 상대 남성의 머리와 상체를 추가했다."
        ],
        "physics": "윤성찬은 허리에서 앞으로 숙인 자세이며 하체와 손의 접촉점은 화면 밖이다. 왼쪽 뒤의 좌석과 하단으로 이어지는 몸통은 자연스러운 착석 전경사 또는 선 자세의 전경사와 양립하며, 허공에 떠 있다는 증거는 없다. 탁자는 보이는 다리로 지지되고 화분·쟁반·책은 탁자 위에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "윤성찬의 몸과 얼굴은 오른쪽 전경 남자에게 향하고, 시선도 그 남자의 얼굴에 닿는다. 입을 벌리고 앞으로 깊이 숙여 말을 건네는 동작이 분명하다. 상대 남자 역시 윤성찬 쪽을 향하지만 눈은 보이지 않는다. 겨누는 무기나 이동 물체는 없다.",
        "built_space": "왼쪽 벽에 산 풍경 액자 하나, 뒤쪽에 검은 가죽 소파 하나, 중앙에 갈색 가죽 안락의자 하나, 왼쪽 하단에 낮은 탁자 일부가 보인다. 뒤에는 넓은 목재 벽면 또는 문짝이 이어진다. 윤성찬은 갈색 의자 앞에서 몸을 숙이고 있다. 참조의 검은 가죽과 목재 팔걸이로 된 접견 좌석 대신 별개의 갈색 안락의자가 중심에 추가되어 장소 고정 조건을 깨뜨린다. 반사 문제는 보이지 않는다.",
        "entities": "회색 머리와 깊은 얼굴 주름, 검회색 재킷, 짙은 남색 셔츠는 윤성찬 참조와 대체로 일치하지만 금속테 안경은 없다. 고령 남성의 눈과 입은 정상적인 해부학적 형태로 표현되었다. 오른쪽 전경의 검은 머리 남자는 윤성찬과 별개인 추가 인물이다. 산 풍경 액자는 참조와 소재가 유사하지만 구체적인 그림은 다르며, 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "등장 가능한 인물이 윤성찬 한 명으로 제한되어 있는데, 오른쪽 전경에 상대 남성의 머리와 상체를 추가했다.",
         "장소가 고정된 접견 구역에 참조의 좌석과 다른 갈색 가죽 안락의자를 추가했다."
        ],
        "physics": "윤성찬의 몸통은 허리에서 앞으로 기울고 두 팔은 아래로 뻗는다. 손과 발이 잘려 정확한 지지 접촉은 확인할 수 없지만, 몸은 화면 아래로 자연스럽게 이어져 있으며 공중 부양이나 불가능한 관절은 보이지 않는다. 갈색 의자와 뒤쪽 소파는 바닥에 놓인 가구로 읽힌다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "앞으로 내민 굳은 상체와 상대를 겨눈 눈빛, 기존 접견실의 공간 연속성은 더 충실하지만, 허용되지 않은 상대 인물을 추가해 실격이며 참조의 안경도 빠졌다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "미디엄 숏의 전경사 상체와 반쯤 열린 입은 부합하지만, 허용되지 않은 상대 인물과 참조에 없던 갈색 안락의자를 추가했고 접견실의 연속성도 떨어진다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "윤성찬의 상체와 얼굴이 화면 오른쪽 전경의 검은 머리 남자를 향하며, 두 눈도 그 남자의 얼굴 쪽에 고정되어 있다. 국방장관에게 가까이 다가가 속삭이는 방향 관계는 읽히지만, 그 상대를 실제 인물로 화면에 넣은 것은 등장인물 제한에 어긋난다. 무기나 방향성 소품은 없다.",
        "built_space": "뒤쪽에 창 구역 두 곳과 그 사이 벽의 액자 일부 하나, 목재 수납장, 검은 가죽 좌석이 보인다. 전경에는 낮은 탁자 하나와 작은 화분 하나, 쟁반 하나, 책 더미가 있다. 창가의 별도 화분 하나와 장식패들도 보이며, 참조의 창가 수납장·가죽 좌석·목재 탁자로 이루어진 접견 구역을 비교적 잘 유지한다. 윤성찬은 왼쪽 좌석 앞에서 탁자 너머 상대 쪽으로 기울어 있다. 불가능한 반사나 명백한 고정 설비 중복은 보이지 않는다.",
        "entities": "주인공은 회색 옆가르마 머리, 깊은 이마와 눈가 주름, 검회색 재킷과 짙은 남색 셔츠를 가진 고령의 한국인 남성으로 표현되어 참조와 대체로 맞는다. 다만 참조의 얇은 금속테 안경이 없다. 입은 조금 벌어져 있고 눈은 정상적인 홍채와 동공을 유지한다. 오른쪽에는 별개의 젊은 남성 머리·귀·목·어깨가 뚜렷하게 추가되어 있다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "등장 가능한 인물이 윤성찬 한 명으로 제한되어 있는데, 오른쪽 전경에 상대 남성의 머리와 상체를 추가했다."
        ],
        "physics": "윤성찬은 허리에서 앞으로 숙인 자세이며 하체와 손의 접촉점은 화면 밖이다. 왼쪽 뒤의 좌석과 하단으로 이어지는 몸통은 자연스러운 착석 전경사 또는 선 자세의 전경사와 양립하며, 허공에 떠 있다는 증거는 없다. 탁자는 보이는 다리로 지지되고 화분·쟁반·책은 탁자 위에 놓여 있다."
       },
       {
        "label": "A",
        "direction": "윤성찬의 몸과 얼굴은 오른쪽 전경 남자에게 향하고, 시선도 그 남자의 얼굴에 닿는다. 입을 벌리고 앞으로 깊이 숙여 말을 건네는 동작이 분명하다. 상대 남자 역시 윤성찬 쪽을 향하지만 눈은 보이지 않는다. 겨누는 무기나 이동 물체는 없다.",
        "built_space": "왼쪽 벽에 산 풍경 액자 하나, 뒤쪽에 검은 가죽 소파 하나, 중앙에 갈색 가죽 안락의자 하나, 왼쪽 하단에 낮은 탁자 일부가 보인다. 뒤에는 넓은 목재 벽면 또는 문짝이 이어진다. 윤성찬은 갈색 의자 앞에서 몸을 숙이고 있다. 참조의 검은 가죽과 목재 팔걸이로 된 접견 좌석 대신 별개의 갈색 안락의자가 중심에 추가되어 장소 고정 조건을 깨뜨린다. 반사 문제는 보이지 않는다.",
        "entities": "회색 머리와 깊은 얼굴 주름, 검회색 재킷, 짙은 남색 셔츠는 윤성찬 참조와 대체로 일치하지만 금속테 안경은 없다. 고령 남성의 눈과 입은 정상적인 해부학적 형태로 표현되었다. 오른쪽 전경의 검은 머리 남자는 윤성찬과 별개인 추가 인물이다. 산 풍경 액자는 참조와 소재가 유사하지만 구체적인 그림은 다르며, 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "등장 가능한 인물이 윤성찬 한 명으로 제한되어 있는데, 오른쪽 전경에 상대 남성의 머리와 상체를 추가했다.",
         "장소가 고정된 접견 구역에 참조의 좌석과 다른 갈색 가죽 안락의자를 추가했다."
        ],
        "physics": "윤성찬의 몸통은 허리에서 앞으로 기울고 두 팔은 아래로 뻗는다. 손과 발이 잘려 정확한 지지 접촉은 확인할 수 없지만, 몸은 화면 아래로 자연스럽게 이어져 있으며 공중 부양이나 불가능한 관절은 보이지 않는다. 갈색 의자와 뒤쪽 소파는 바닥에 놓인 가구로 읽힌다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.267,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.017,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 명단에 없는 인물 등장",
     "[gemini-pro] 이전 샷 인물의 의상(회색 셔츠) 복사",
     "[gemini-pro] 지정된 공간의 가구 재질 위반 (갈색 가죽)",
     "[gpt-high] 등장 가능한 인물이 윤성찬 한 명으로 제한되어 있는데, 오른쪽 전경에 상대 남성의 머리와 상체를 추가했다.",
     "[gpt-high] 장소가 고정된 접견 구역에 참조의 좌석과 다른 갈색 가죽 안락의자를 추가했다."
    ],
    "B": [
     "[gemini-pro] 명단에 없는 인물 등장",
     "[gpt-high] 등장 가능한 인물이 윤성찬 한 명으로 제한되어 있는데, 오른쪽 전경에 상대 남성의 머리와 상체를 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 1750,
   "A": 1017
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "명단에 없는 인물이 화면 앞쪽에 등장하는 위반이 있으나, 배경 공간의 디테일을 유지하고 금지된 의상 복사를 피하여 상대적으로 우수함.  ★위반: [gemini-pro] 명단에 없는 인물 등장 / [gpt-high] 등장 가능한 인물이 윤성찬 한 명으로 제한되어 있는데, 오른쪽 전경에 상대 남성의 머리와 상체를 추가했다."
   },
   {
    "label": "A",
    "score": 1017,
    "verdict_ko": "명단에 없는 인물이 등장할 뿐만 아니라 이전 샷의 의상을 복사하였고, 공간의 가구 재질(갈색 가죽)을 임의로 변경하여 지침을 다수 위반함.  ★위반: [gemini-pro] 명단에 없는 인물 등장 / [gemini-pro] 이전 샷 인물의 의상(회색 셔츠) 복사 / [gemini-pro] 지정된 공간의 가구 재질 위반 (갈색 가죽) / [gpt-high] 등장 가능한 인물이 윤성찬 한 명으로 제한되어 있는데, 오른쪽 전경에 상대 남성의 머리와 상체를 추가했다. / [gpt-high] 장소가 고정된 접견 구역에 참조의 좌석과 다른 갈색 가죽 안락의자를 추가했다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S43sh2_sel.png",
    "asset_id": "a07c4c0c-1b6b-4bb9-9a88-091eab67664b",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 윤성찬: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1453233>",
    "asset_id": "04d34665-3829-49fb-a6d2-25b9fd051d63",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-4f51-7cd1-aa7d-0e687f01c951",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S43sh2"
  }
 },
 "S54sh2::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:13:38.959955+00:00",
  "fingerprint": "aa7a59ea22ee5f95dfaf898ff231942008fb719f01b4d6f86998ed1da8bb0dc8",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S54sh2_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S54sh2_sel.png",
  "source_sha256": "d5a83a4302b200e631b50d107dfd6ab3f4d0fc706e78c3bb469e901f31b9031e",
  "file": "S54sh2_cine.png",
  "staged_sha256": "5513facd20aa1f7f3e2f14811e31b9dce1d29e08e69a6e4f0b9ec497eaf7cad4",
  "latency_ms": 13377
 },
 "S54sh5::signage": {
  "fp": "9b9ae43c0e77a6db",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S54sh5": {
  "input_fingerprint": "3bdd2df09c4feded",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 미소를 지운 채 목적지를 안다며 날카로운 눈빛으로 허공을 꿰뚫어 보는 윤성찬의 차가운 얼굴 클로즈업.\n\nLOCATION (lock): In the meeting area inside the defense minister's executive office, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the office's established ambient illumination and restrained contrast, conveying coldness through expression rather than a new color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 윤성찬 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the minister's office surfaces, nearby furnishings, and daytime interior illumination. Exclude the militia commander's office furnishings and telephone-conversation setting.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 미소를 지운 채 목적지를 안다며 날카로운 눈빛으로 허공을 꿰뚫어 보는 윤성찬의 차가운 얼굴 클로즈업.\n\nLOCATION (lock): In the meeting area inside the defense minister's executive office, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the office's established ambient illumination and restrained contrast, conveying coldness through expression rather than a new color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 윤성찬 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the minister's office surfaces, nearby furnishings, and daytime interior illumination. Exclude the militia commander's office furnishings and telephone-conversation setting.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 미소를 지운 채 목적지를 안다며 날카로운 눈빛으로 허공을 꿰뚫어 보는 윤성찬의 차가운 얼굴 클로즈업.\n\nLOCATION (lock): In the meeting area inside the defense minister's executive office, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the office's established ambient illumination and restrained contrast, conveying coldness through expression rather than a new color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 윤성찬 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the minister's office surfaces, nearby furnishings, and daytime interior illumination. Exclude the militia commander's office furnishings and telephone-conversation setting.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "시선이 카메라 렌즈를 정면으로 향하고 있습니다.",
    "built_space": "창문과 수납장이 보이나, 우측 배경에 이전 샷에서 없던 태극기가 등장했으며 화분의 위치도 좌측으로 변경되었습니다.",
    "entities": "윤성찬의 고령의 얼굴 특징과 깊은 주름, 의상이 레퍼런스와 일치합니다.",
    "hard_violations": [
     "[gemini-pro] 이전 샷에 없던 임의의 사물(태극기) 생성"
    ],
    "physics": "화면에 보이는 상체는 안정적으로 지탱되고 있습니다."
   },
   {
    "label": "B",
    "direction": "시선이 카메라 렌즈를 정면으로 향하고 있습니다.",
    "built_space": "배경의 창문과 수납장, 우측 전경의 나무 받침 화분, 좌측의 가죽 소파 팔걸이까지 이전 샷의 미장센을 완벽하게 재현했습니다.",
    "entities": "윤성찬의 얼굴 특징, 주름, 의상이 레퍼런스와 정확히 일치합니다.",
    "hard_violations": [],
    "physics": "이전 샷의 앞으로 기울어진 자세를 자연스럽게 유지하며 안정적입니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "이전 샷의 공간 배치와 특정 소품(화분, 소파)을 완벽히 유지하며 요구된 클로즈업 척도와 차가운 분위기를 잘 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정된 인물의 외형은 잘 묘사했으나, 이전 샷에 없던 태극기가 임의로 추가되는 치명적인 오류가 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선이 카메라 렌즈를 정면으로 향하고 있습니다.",
        "built_space": "창문과 수납장이 보이나, 우측 배경에 이전 샷에서 없던 태극기가 등장했으며 화분의 위치도 좌측으로 변경되었습니다.",
        "entities": "윤성찬의 고령의 얼굴 특징과 깊은 주름, 의상이 레퍼런스와 일치합니다.",
        "hard_violations": [
         "이전 샷에 없던 임의의 사물(태극기) 생성"
        ],
        "physics": "화면에 보이는 상체는 안정적으로 지탱되고 있습니다."
       },
       {
        "label": "B",
        "direction": "시선이 카메라 렌즈를 정면으로 향하고 있습니다.",
        "built_space": "배경의 창문과 수납장, 우측 전경의 나무 받침 화분, 좌측의 가죽 소파 팔걸이까지 이전 샷의 미장센을 완벽하게 재현했습니다.",
        "entities": "윤성찬의 얼굴 특징, 주름, 의상이 레퍼런스와 정확히 일치합니다.",
        "hard_violations": [],
        "physics": "이전 샷의 앞으로 기울어진 자세를 자연스럽게 유지하며 안정적입니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "이전 샷의 공간 배치와 특정 소품(화분, 소파)을 완벽히 유지하며 요구된 클로즈업 척도와 차가운 분위기를 잘 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정된 인물의 외형은 잘 묘사했으나, 이전 샷에 없던 태극기가 임의로 추가되는 치명적인 오류가 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선이 카메라 렌즈를 정면으로 향하고 있습니다.",
        "built_space": "창문과 수납장이 보이나, 우측 배경에 이전 샷에서 없던 태극기가 등장했으며 화분의 위치도 좌측으로 변경되었습니다.",
        "entities": "윤성찬의 고령의 얼굴 특징과 깊은 주름, 의상이 레퍼런스와 일치합니다.",
        "hard_violations": [
         "이전 샷에 없던 임의의 사물(태극기) 생성"
        ],
        "physics": "화면에 보이는 상체는 안정적으로 지탱되고 있습니다."
       },
       {
        "label": "B",
        "direction": "시선이 카메라 렌즈를 정면으로 향하고 있습니다.",
        "built_space": "배경의 창문과 수납장, 우측 전경의 나무 받침 화분, 좌측의 가죽 소파 팔걸이까지 이전 샷의 미장센을 완벽하게 재현했습니다.",
        "entities": "윤성찬의 얼굴 특징, 주름, 의상이 레퍼런스와 정확히 일치합니다.",
        "hard_violations": [],
        "physics": "이전 샷의 앞으로 기울어진 자세를 자연스럽게 유지하며 안정적입니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "얼굴이 화면을 지배하는 클로즈업과 기존 접견 공간의 연속성이 더 정확하지만, 눈빛은 요구된 날카로움보다 다소 침잠해 보입니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "미소 없는 입과 좁힌 눈매는 차가운 표정에 잘 맞지만, 어깨와 배경을 더 드러낸 구도가 얼굴 클로즈업의 집중도를 낮춥니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴은 거의 정면이며 두 눈은 렌즈 부근을 향합니다. 화면 안의 사람이나 물건을 응시하지는 않습니다. 허공을 보는 시선으로 해석할 여지는 있지만, 정면 관객을 바라보는 인상도 강하며 눈매의 날카로움은 약합니다. 무기나 방향성 있는 동작은 없습니다.",
        "built_space": "뒤쪽에 창틀로 나뉜 밝은 창 영역 두 곳과 목재 수납장이 보이고, 왼쪽에는 검은 소파 일부, 오른쪽 뒤에는 팔걸이의자 한 개와 큰 목재 책상 한 개가 보입니다. 오른쪽 아래 접견 테이블 한 개 위에는 작은 화분 한 개, 사각 트레이 한 개, 어두운 서류철류가 놓여 있습니다. 인물은 테이블 왼쪽의 소파 쪽에서 상체를 앞으로 기울여 이전 장면의 배치와 자연스럽게 이어집니다. 반사상이나 명백한 고정 설비 중복은 없습니다.",
        "entities": "고령의 한국인 남성으로 제시된 윤성찬 한 명만 보입니다. 회색으로 빗어 넘긴 머리, 깊은 이마·눈가 주름, 얼굴의 반점과 윤곽이 참조 인물에 가깝습니다. 짙은 회색 재킷과 남색 셔츠가 이전 장면과 일치하고, 안경을 쓰지 않은 것도 이전 장면을 따릅니다. 입은 닫혀 있고 미소는 없습니다. 눈은 정상적인 홍채와 동공을 유지합니다. 읽을 수 있는 글자나 다른 사람은 없습니다.",
        "hard_violations": [],
        "physics": "머리와 목, 앞으로 기울어진 상체의 연결이 자연스럽습니다. 하체와 좌면 접촉은 클로즈업 밖이라 확인할 수 없으나, 공중에 떠 있다는 징후는 없습니다. 화분·트레이·서류철은 테이블에, 뒤쪽 장식물은 수납장 상판에 놓여 있습니다. 지지 없이 떠 있는 물체는 보이지 않습니다."
       },
       {
        "label": "B",
        "direction": "얼굴은 정면에 가깝고 눈은 렌즈 부근에서 약간 화면 왼쪽을 향합니다. 화면 안에 특정 응시 대상은 없습니다. 좁힌 눈꺼풀과 단단히 다문 입 때문에 A보다 날카롭고 냉정한 응시로 읽히지만, 시선은 여전히 카메라에 가까워 정면 초상 같은 인상도 있습니다. 무기나 이동 동작은 없습니다.",
        "built_space": "뒤쪽의 밝은 창 영역 두 곳, 그 사이 벽의 액자 한 개, 목재 수납장, 오른쪽의 큰 목재 책상 한 개가 보입니다. 왼쪽 뒤에는 검은 팔걸이의자 한 개와 작은 보조 탁자 일부가 있고, 접견 테이블 한 개와 작은 화분 한 개가 인물 뒤쪽으로 펼쳐집니다. 오른쪽 끝에는 흐릿한 태극기 일부도 보이나 이전 사진에서는 확인되지 않습니다. 인물은 테이블 가까운 끝 앞에 정면으로 자리하여 A보다 공간 관계가 덜 직접적으로 이어집니다. 불가능한 반사나 명백한 설비 중복은 없습니다.",
        "entities": "윤성찬으로 보이는 고령 남성 한 명만 있으며, 회색 머리, 깊은 주름, 피부 반점과 얼굴 형태가 참조에 부합합니다. 짙은 회색 재킷과 남색 셔츠, 안경 없는 모습도 이전 장면과 맞습니다. 눈의 해부학은 정상이고 미소는 없습니다. 머리 전체와 양쪽 어깨, 가슴 윗부분까지 보여 A보다 얼굴의 화면 점유율이 작습니다. 읽을 수 있는 글자나 다른 사람은 없습니다.",
        "hard_violations": [],
        "physics": "머리가 목과 상체 위에 정상적으로 지지되며, 곧게 세운 상체에도 해부학적 모순은 없습니다. 좌면과 발은 프레임 밖이므로 앉거나 선 상태의 정확한 접촉점은 확인할 수 없습니다. 화분은 테이블 위에, 배경 소품은 가구 상판 위에 놓여 있고 공중 부양이나 부자연스러운 운동은 없습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "얼굴이 화면을 지배하는 클로즈업과 기존 접견 공간의 연속성이 더 정확하지만, 눈빛은 요구된 날카로움보다 다소 침잠해 보입니다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "미소 없는 입과 좁힌 눈매는 차가운 표정에 잘 맞지만, 어깨와 배경을 더 드러낸 구도가 얼굴 클로즈업의 집중도를 낮춥니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴은 거의 정면이며 두 눈은 렌즈 부근을 향합니다. 화면 안의 사람이나 물건을 응시하지는 않습니다. 허공을 보는 시선으로 해석할 여지는 있지만, 정면 관객을 바라보는 인상도 강하며 눈매의 날카로움은 약합니다. 무기나 방향성 있는 동작은 없습니다.",
        "built_space": "뒤쪽에 창틀로 나뉜 밝은 창 영역 두 곳과 목재 수납장이 보이고, 왼쪽에는 검은 소파 일부, 오른쪽 뒤에는 팔걸이의자 한 개와 큰 목재 책상 한 개가 보입니다. 오른쪽 아래 접견 테이블 한 개 위에는 작은 화분 한 개, 사각 트레이 한 개, 어두운 서류철류가 놓여 있습니다. 인물은 테이블 왼쪽의 소파 쪽에서 상체를 앞으로 기울여 이전 장면의 배치와 자연스럽게 이어집니다. 반사상이나 명백한 고정 설비 중복은 없습니다.",
        "entities": "고령의 한국인 남성으로 제시된 윤성찬 한 명만 보입니다. 회색으로 빗어 넘긴 머리, 깊은 이마·눈가 주름, 얼굴의 반점과 윤곽이 참조 인물에 가깝습니다. 짙은 회색 재킷과 남색 셔츠가 이전 장면과 일치하고, 안경을 쓰지 않은 것도 이전 장면을 따릅니다. 입은 닫혀 있고 미소는 없습니다. 눈은 정상적인 홍채와 동공을 유지합니다. 읽을 수 있는 글자나 다른 사람은 없습니다.",
        "hard_violations": [],
        "physics": "머리와 목, 앞으로 기울어진 상체의 연결이 자연스럽습니다. 하체와 좌면 접촉은 클로즈업 밖이라 확인할 수 없으나, 공중에 떠 있다는 징후는 없습니다. 화분·트레이·서류철은 테이블에, 뒤쪽 장식물은 수납장 상판에 놓여 있습니다. 지지 없이 떠 있는 물체는 보이지 않습니다."
       },
       {
        "label": "A",
        "direction": "얼굴은 정면에 가깝고 눈은 렌즈 부근에서 약간 화면 왼쪽을 향합니다. 화면 안에 특정 응시 대상은 없습니다. 좁힌 눈꺼풀과 단단히 다문 입 때문에 A보다 날카롭고 냉정한 응시로 읽히지만, 시선은 여전히 카메라에 가까워 정면 초상 같은 인상도 있습니다. 무기나 이동 동작은 없습니다.",
        "built_space": "뒤쪽의 밝은 창 영역 두 곳, 그 사이 벽의 액자 한 개, 목재 수납장, 오른쪽의 큰 목재 책상 한 개가 보입니다. 왼쪽 뒤에는 검은 팔걸이의자 한 개와 작은 보조 탁자 일부가 있고, 접견 테이블 한 개와 작은 화분 한 개가 인물 뒤쪽으로 펼쳐집니다. 오른쪽 끝에는 흐릿한 태극기 일부도 보이나 이전 사진에서는 확인되지 않습니다. 인물은 테이블 가까운 끝 앞에 정면으로 자리하여 A보다 공간 관계가 덜 직접적으로 이어집니다. 불가능한 반사나 명백한 설비 중복은 없습니다.",
        "entities": "윤성찬으로 보이는 고령 남성 한 명만 있으며, 회색 머리, 깊은 주름, 피부 반점과 얼굴 형태가 참조에 부합합니다. 짙은 회색 재킷과 남색 셔츠, 안경 없는 모습도 이전 장면과 맞습니다. 눈의 해부학은 정상이고 미소는 없습니다. 머리 전체와 양쪽 어깨, 가슴 윗부분까지 보여 A보다 얼굴의 화면 점유율이 작습니다. 읽을 수 있는 글자나 다른 사람은 없습니다.",
        "hard_violations": [],
        "physics": "머리가 목과 상체 위에 정상적으로 지지되며, 곧게 세운 상체에도 해부학적 모순은 없습니다. 좌면과 발은 프레임 밖이므로 앉거나 선 상태의 정확한 접촉점은 확인할 수 없습니다. 화분은 테이블 위에, 배경 소품은 가구 상판 위에 놓여 있고 공중 부양이나 부자연스러운 운동은 없습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.304,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.054,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 이전 샷에 없던 임의의 사물(태극기) 생성"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1054
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "이전 샷의 공간 배치와 특정 소품(화분, 소파)을 완벽히 유지하며 요구된 클로즈업 척도와 차가운 분위기를 잘 구현했습니다."
   },
   {
    "label": "A",
    "score": 1054,
    "verdict_ko": "지정된 인물의 외형은 잘 묘사했으나, 이전 샷에 없던 태극기가 임의로 추가되는 치명적인 오류가 발생했습니다.  ★위반: [gemini-pro] 이전 샷에 없던 임의의 사물(태극기) 생성"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 윤성찬 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S54sh2_sel.png",
    "asset_id": "1cfe2b0d-20f4-4a13-a573-622fa2931b07",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 윤성찬: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1453233>",
    "asset_id": "04d34665-3829-49fb-a6d2-25b9fd051d63",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-5102-7eca-8f47-7ef8e2e0e171",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S54sh2"
  }
 },
 "S54sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:14:43.039126+00:00",
  "fingerprint": "b382c37eb734b9c76ea22163391957eb47682dc72f3c210d063ef1d9a95ed1d8",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S54sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S54sh5_sel.png",
  "source_sha256": "3c4ca6e2ad7114380623070aae03fb6ca98119a52e6ef99486facef6b61c6ce6",
  "file": "S54sh5_cine.png",
  "staged_sha256": "2862b823c3b912b459fdd60d2387904f3b64ae1d7350f1fa688d8c623f63a359",
  "latency_ms": 13912
 },
 "S54sh6::signage": {
  "fp": "769f5b928f054d73",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S54sh6": {
  "input_fingerprint": "8f0046ab418ca255",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 윤성찬의 말에 두 눈을 동그랗게 뜬 채 놀란 표정으로 입을 벌린 국방장관의 어리둥절한 얼굴.\n\nLOCATION (lock): At the minister's position in the executive office's conversation area, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Unchanged office ambient illumination keeps the minister's startled eyes and parted mouth clearly readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same ministerial office surfaces, furnishings, and daylight conditions. Exclude the militia office's wall knife and speakerphone setup.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 국방장관 (한국인, 성인 남성, 단정한 짧은 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 윤성찬의 말에 두 눈을 동그랗게 뜬 채 놀란 표정으로 입을 벌린 국방장관의 어리둥절한 얼굴.\n\nLOCATION (lock): At the minister's position in the executive office's conversation area, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Unchanged office ambient illumination keeps the minister's startled eyes and parted mouth clearly readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same ministerial office surfaces, furnishings, and daylight conditions. Exclude the militia office's wall knife and speakerphone setup.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 국방장관 (한국인, 성인 남성, 단정한 짧은 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 윤성찬의 말에 두 눈을 동그랗게 뜬 채 놀란 표정으로 입을 벌린 국방장관의 어리둥절한 얼굴.\n\nLOCATION (lock): At the minister's position in the executive office's conversation area, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Unchanged office ambient illumination keeps the minister's startled eyes and parted mouth clearly readable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same ministerial office surfaces, furnishings, and daylight conditions. Exclude the militia office's wall knife and speakerphone setup.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 국방장관 (한국인, 성인 남성, 단정한 짧은 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "인물의 시선이 카메라 렌즈를 정면으로 응시하고 있습니다.",
    "built_space": "이전 샷과 동일한 사무실 배경(창문, 가죽 의자, 탁자, 화분 등)이 올바른 위치와 비율로 나타납니다.",
    "entities": "인물의 얼굴과 헤어스타일, 놀란 표정은 프롬프트에 부합하지만, 캐릭터 레퍼런스의 복장 대신 금지된 이전 샷 인물의 어두운 재킷을 입고 있습니다.",
    "hard_violations": [],
    "physics": "자연스러운 자세로 떠 있는 부분 없이 공간에 잘 안착해 있습니다."
   },
   {
    "label": "B",
    "direction": "인물의 시선이 카메라 밖 우측을 향하고 있어 누군가의 말을 듣고 반응하는 맥락에 잘 맞습니다.",
    "built_space": "이전 샷에서 제시된 창문, 검은색 가죽 소파, 화분이 놓인 탁자 등의 사무실 배경 요소가 동일한 위치와 조명 아래 정확히 구현되었습니다.",
    "entities": "캐릭터 레퍼런스와 얼굴, 머리 모양이 일치하며, 지시된 남색 셔츠 복장도 정확히 착용하고 있습니다. 두 눈을 동그랗게 뜨고 입을 벌린 놀란 표정도 잘 표현되었습니다.",
    "hard_violations": [],
    "physics": "신체가 안정적으로 지지된 상태에서 자연스러운 자세를 취하고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "놀란 표정의 클로즈업, 이전 샷의 배경, 자연스러운 시선 처리 및 캐릭터 레퍼런스의 복장까지 모든 지시사항을 훌륭하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "지정된 표정과 공간은 잘 표현했으나, 캐릭터 레퍼런스의 복장을 따르지 않고 이전 샷 인물의 복장(어두운 재킷)을 입힌 점이 감점 요인입니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "인물의 시선이 카메라 밖 우측을 향하고 있어 누군가의 말을 듣고 반응하는 맥락에 잘 맞습니다.",
        "built_space": "이전 샷에서 제시된 창문, 검은색 가죽 소파, 화분이 놓인 탁자 등의 사무실 배경 요소가 동일한 위치와 조명 아래 정확히 구현되었습니다.",
        "entities": "캐릭터 레퍼런스와 얼굴, 머리 모양이 일치하며, 지시된 남색 셔츠 복장도 정확히 착용하고 있습니다. 두 눈을 동그랗게 뜨고 입을 벌린 놀란 표정도 잘 표현되었습니다.",
        "hard_violations": [],
        "physics": "신체가 안정적으로 지지된 상태에서 자연스러운 자세를 취하고 있습니다."
       },
       {
        "label": "A",
        "direction": "인물의 시선이 카메라 렌즈를 정면으로 응시하고 있습니다.",
        "built_space": "이전 샷과 동일한 사무실 배경(창문, 가죽 의자, 탁자, 화분 등)이 올바른 위치와 비율로 나타납니다.",
        "entities": "인물의 얼굴과 헤어스타일, 놀란 표정은 프롬프트에 부합하지만, 캐릭터 레퍼런스의 복장 대신 금지된 이전 샷 인물의 어두운 재킷을 입고 있습니다.",
        "hard_violations": [],
        "physics": "자연스러운 자세로 떠 있는 부분 없이 공간에 잘 안착해 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "놀란 표정의 클로즈업, 이전 샷의 배경, 자연스러운 시선 처리 및 캐릭터 레퍼런스의 복장까지 모든 지시사항을 훌륭하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "지정된 표정과 공간은 잘 표현했으나, 캐릭터 레퍼런스의 복장을 따르지 않고 이전 샷 인물의 복장(어두운 재킷)을 입힌 점이 감점 요인입니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "인물의 시선이 카메라 밖 우측을 향하고 있어 누군가의 말을 듣고 반응하는 맥락에 잘 맞습니다.",
        "built_space": "이전 샷에서 제시된 창문, 검은색 가죽 소파, 화분이 놓인 탁자 등의 사무실 배경 요소가 동일한 위치와 조명 아래 정확히 구현되었습니다.",
        "entities": "캐릭터 레퍼런스와 얼굴, 머리 모양이 일치하며, 지시된 남색 셔츠 복장도 정확히 착용하고 있습니다. 두 눈을 동그랗게 뜨고 입을 벌린 놀란 표정도 잘 표현되었습니다.",
        "hard_violations": [],
        "physics": "신체가 안정적으로 지지된 상태에서 자연스러운 자세를 취하고 있습니다."
       },
       {
        "label": "A",
        "direction": "인물의 시선이 카메라 렌즈를 정면으로 응시하고 있습니다.",
        "built_space": "이전 샷과 동일한 사무실 배경(창문, 가죽 의자, 탁자, 화분 등)이 올바른 위치와 비율로 나타납니다.",
        "entities": "인물의 얼굴과 헤어스타일, 놀란 표정은 프롬프트에 부합하지만, 캐릭터 레퍼런스의 복장 대신 금지된 이전 샷 인물의 어두운 재킷을 입고 있습니다.",
        "hard_violations": [],
        "physics": "자연스러운 자세로 떠 있는 부분 없이 공간에 잘 안착해 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "화면 밖 상대를 향해 눈을 크게 뜨고 입을 벌린 클로즈업으로, 장관의 반응과 참조 인물의 남색 셔츠 및 사무실 연속성을 충실히 유지한다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "놀란 얼굴의 클로즈업은 적절하지만 시선이 렌즈를 정면으로 향하고, 참조의 남색 셔츠를 밝은 셔츠와 재킷으로 바꿔 인물 연속성이 떨어진다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴과 두 눈이 화면 오른쪽의 프레임 밖 상대를 향한다. 상대는 보이지 않지만 윤성찬의 말을 듣고 반응하는 시선으로 자연스럽게 읽힌다. 무기나 방향을 확인할 휴대 소품은 없다.",
        "built_space": "뒤쪽에 격자창과 낮은 목재 수납장이 있고, 왼쪽에는 검은 가죽 소파의 등받이 구간들이, 오른쪽 뒤에는 팔걸이 의자 한 개와 책상 한 개가 보인다. 전경 오른쪽 탁자 한 개에는 화분 한 개, 사각 트레이 한 개, 검은 서류철 한 개가 놓여 있다. 창 아래에는 작은 액자·패 모양 물건 여러 개가 있으나 흐려 정확한 총수는 판별하기 어렵다. 참조의 가구 배치와 재질, 낮빛이 잘 이어지며 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "성인 남성 한 명만 등장한다. 한국인으로 설정된 참조 인물과 검은 가르마 머리, 얼굴 윤곽, 옅은 수염, 남색 셔츠가 잘 맞는다. 양눈을 크게 뜨고 입술을 벌려 놀람과 어리둥절함을 표현하며 안구는 정상적인 사람의 형태다. 이전 사진의 노인이나 추가 인물은 없고, 벽의 칼·스피커폰·읽을 수 있는 글자도 보이지 않는다.",
        "hard_violations": [],
        "physics": "머리는 목과 어깨 위에서 자연스럽게 지지되고 상체는 약간 앞으로 기울어 있다. 골반과 좌면 접촉은 클로즈업 밖이어서 확인할 수 없지만 공중에 떠 있다는 징후는 없다. 화분, 트레이, 서류철은 탁자 위에 놓여 지지되며 움직이거나 떠 있는 물체는 없다."
       },
       {
        "label": "B",
        "direction": "얼굴은 거의 정면이고 두 눈은 렌즈 쪽을 바라본다. 눈과 입의 놀란 반응은 분명하지만 화면 밖 윤성찬을 향한 시선보다는 관객을 직접 보는 모습으로 읽힌다. 방향성 있는 휴대 소품은 없다.",
        "built_space": "뒤쪽 격자창과 목재 수납장, 왼쪽 가죽 소파, 오른쪽 뒤 팔걸이 의자 한 개와 책상 한 개, 오른쪽 전경 탁자 한 개가 보인다. 탁자 끝에는 트레이 일부가 있으며 창 아래 작은 액자·패들은 흐릿하다. 왼쪽에는 큰 화분 한 개와 원형 벽 부착물 일부가 추가로 드러나지만 이전 참조의 좁은 화면으로는 그 위치를 확인할 수 없다. 사무실의 주된 재질과 낮빛은 유사하며 불가능한 반사나 명백한 설비 중복은 없다.",
        "entities": "성인 남성 한 명이며 검은 가르마 머리와 얼굴은 장관 참조에 대체로 부합한다. 그러나 남색 셔츠 대신 밝은 셔츠와 짙은 재킷을 입었다. 눈을 크게 뜨고 입을 벌린 표정은 요구와 맞고 눈의 해부학적 형태도 정상이다. 다른 사람, 칼, 스피커폰, 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "앞으로 기울어진 상체 위에 목과 머리가 자연스럽게 연결된다. 하체와 실제 좌면 접촉은 프레임 밖이므로 확인할 수 없으며, 보이는 부분에 무지지 부유나 불가능한 관절은 없다. 배경 의자는 바닥에 놓이고 트레이는 탁자에 얹혀 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "화면 밖 상대를 향해 눈을 크게 뜨고 입을 벌린 클로즈업으로, 장관의 반응과 참조 인물의 남색 셔츠 및 사무실 연속성을 충실히 유지한다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "놀란 얼굴의 클로즈업은 적절하지만 시선이 렌즈를 정면으로 향하고, 참조의 남색 셔츠를 밝은 셔츠와 재킷으로 바꿔 인물 연속성이 떨어진다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴과 두 눈이 화면 오른쪽의 프레임 밖 상대를 향한다. 상대는 보이지 않지만 윤성찬의 말을 듣고 반응하는 시선으로 자연스럽게 읽힌다. 무기나 방향을 확인할 휴대 소품은 없다.",
        "built_space": "뒤쪽에 격자창과 낮은 목재 수납장이 있고, 왼쪽에는 검은 가죽 소파의 등받이 구간들이, 오른쪽 뒤에는 팔걸이 의자 한 개와 책상 한 개가 보인다. 전경 오른쪽 탁자 한 개에는 화분 한 개, 사각 트레이 한 개, 검은 서류철 한 개가 놓여 있다. 창 아래에는 작은 액자·패 모양 물건 여러 개가 있으나 흐려 정확한 총수는 판별하기 어렵다. 참조의 가구 배치와 재질, 낮빛이 잘 이어지며 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "성인 남성 한 명만 등장한다. 한국인으로 설정된 참조 인물과 검은 가르마 머리, 얼굴 윤곽, 옅은 수염, 남색 셔츠가 잘 맞는다. 양눈을 크게 뜨고 입술을 벌려 놀람과 어리둥절함을 표현하며 안구는 정상적인 사람의 형태다. 이전 사진의 노인이나 추가 인물은 없고, 벽의 칼·스피커폰·읽을 수 있는 글자도 보이지 않는다.",
        "hard_violations": [],
        "physics": "머리는 목과 어깨 위에서 자연스럽게 지지되고 상체는 약간 앞으로 기울어 있다. 골반과 좌면 접촉은 클로즈업 밖이어서 확인할 수 없지만 공중에 떠 있다는 징후는 없다. 화분, 트레이, 서류철은 탁자 위에 놓여 지지되며 움직이거나 떠 있는 물체는 없다."
       },
       {
        "label": "A",
        "direction": "얼굴은 거의 정면이고 두 눈은 렌즈 쪽을 바라본다. 눈과 입의 놀란 반응은 분명하지만 화면 밖 윤성찬을 향한 시선보다는 관객을 직접 보는 모습으로 읽힌다. 방향성 있는 휴대 소품은 없다.",
        "built_space": "뒤쪽 격자창과 목재 수납장, 왼쪽 가죽 소파, 오른쪽 뒤 팔걸이 의자 한 개와 책상 한 개, 오른쪽 전경 탁자 한 개가 보인다. 탁자 끝에는 트레이 일부가 있으며 창 아래 작은 액자·패들은 흐릿하다. 왼쪽에는 큰 화분 한 개와 원형 벽 부착물 일부가 추가로 드러나지만 이전 참조의 좁은 화면으로는 그 위치를 확인할 수 없다. 사무실의 주된 재질과 낮빛은 유사하며 불가능한 반사나 명백한 설비 중복은 없다.",
        "entities": "성인 남성 한 명이며 검은 가르마 머리와 얼굴은 장관 참조에 대체로 부합한다. 그러나 남색 셔츠 대신 밝은 셔츠와 짙은 재킷을 입었다. 눈을 크게 뜨고 입을 벌린 표정은 요구와 맞고 눈의 해부학적 형태도 정상이다. 다른 사람, 칼, 스피커폰, 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "앞으로 기울어진 상체 위에 목과 머리가 자연스럽게 연결된다. 하체와 실제 좌면 접촉은 프레임 밖이므로 확인할 수 없으며, 보이는 부분에 무지지 부유나 불가능한 관절은 없다. 배경 의자는 바닥에 놓이고 트레이는 탁자에 얹혀 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.417,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.417,
    "B": 2.0
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1417
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "놀란 표정의 클로즈업, 이전 샷의 배경, 자연스러운 시선 처리 및 캐릭터 레퍼런스의 복장까지 모든 지시사항을 훌륭하게 구현했습니다."
   },
   {
    "label": "A",
    "score": 1417,
    "verdict_ko": "지정된 표정과 공간은 잘 표현했으나, 캐릭터 레퍼런스의 복장을 따르지 않고 이전 샷 인물의 복장(어두운 재킷)을 입힌 점이 감점 요인입니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S54sh5_sel.png",
    "asset_id": "68599ba3-931a-4bb8-9b5b-a7b8b16b9c72",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 국방장관: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1183281>",
    "asset_id": "65f118fd-34f5-434a-8c3f-1378f0e8a758",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-52b5-7903-b1dd-38e196c11c4d",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S54sh5"
  }
 },
 "S54sh6::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:15:39.267819+00:00",
  "fingerprint": "81617589b7bc1681445b288511d7dc193d8ad8e44ecce3a80f5f495fd5c6975c",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S54sh6_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S54sh6_sel.png",
  "source_sha256": "e6485f2abb8dbc4e4c399b3ee4bc15183ba10c8cd561061cb036c83201f999a4",
  "file": "S54sh6_cine.png",
  "staged_sha256": "2cc94a30140ba00253f1902be5b03b066870cc0b43419d5a46c67e592c223ee3",
  "latency_ms": 12110
 },
 "S55sh3::signage": {
  "fp": "729584fef3dbf4c4",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S55sh3::bgfirst_bg": {
  "input_fingerprint": "cf09efd0cb5df35a",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 창밖의 숲 쪽을 향해 검지손가락을 길게 뻗은 라울의 상체.\n\nLOCATION (lock): Inside the camper's passenger area, beside a window overlooking roadside woodland in daylight after rain.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Forest visible beyond the camper side window in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Camper side window (Providing a view of the forest outside) — Seen obliquely from the aisle, with the forest visible beyond the pointing hand; used as Connects the interior gesture to its exterior destination without requiring the camera to leave the camper; Forest beside the route (Visible outside as a possible route taken by 찰리) — Appears beyond the window toward the upper right and continues outside the crop; used as Provides an intelligible destination for the pointing line; Rear seating (Occupied by 라울) — A partial seat edge remains beneath his leaning torso; used as Anchors his seated posture and prevents the pointing gesture from reading as a standing pose.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination preserves the exhausted, rain-soaked appearance without introducing fresh rainfall or an unsupported window effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 창밖의 숲 쪽을 향해 검지손가락을 길게 뻗은 라울의 상체.\n\nLOCATION (lock): Inside the camper's passenger area, beside a window overlooking roadside woodland in daylight after rain.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Forest visible beyond the camper side window in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Camper side window (Providing a view of the forest outside) — Seen obliquely from the aisle, with the forest visible beyond the pointing hand; used as Connects the interior gesture to its exterior destination without requiring the camera to leave the camper; Forest beside the route (Visible outside as a possible route taken by 찰리) — Appears beyond the window toward the upper right and continues outside the crop; used as Provides an intelligible destination for the pointing line; Rear seating (Occupied by 라울) — A partial seat edge remains beneath his leaning torso; used as Anchors his seated posture and prevents the pointing gesture from reading as a standing pose.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination preserves the exhausted, rain-soaked appearance without introducing fresh rainfall or an unsupported window effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S55sh3__bgfirst_bg.png",
  "asset_id": "f40aa134-4d1d-43ab-80e4-59ce7d1bc555",
  "input_asset_ids": [
   "30ef55b8-86fa-49ef-81a1-ac257a19f52a",
   "4be906f8-5377-4122-9528-ee6aaccbd804"
  ]
 },
 "S55sh3": {
  "input_fingerprint": "f99f550d83825124",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 창밖의 숲 쪽을 향해 검지손가락을 길게 뻗은 라울의 상체.\n\nLOCATION (lock): Inside the camper's passenger area, beside a window overlooking roadside woodland in daylight after rain. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Forest visible beyond the camper side window in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Camper side window (Providing a view of the forest outside) — Seen obliquely from the aisle, with the forest visible beyond the pointing hand; used as Connects the interior gesture to its exterior destination without requiring the camera to leave the camper; Forest beside the route (Visible outside as a possible route taken by 찰리) — Appears beyond the window toward the upper right and continues outside the crop; used as Provides an intelligible destination for the pointing line; Rear seating (Occupied by 라울) — A partial seat edge remains beneath his leaning torso; used as Anchors his seated posture and prevents the pointing gesture from reading as a standing pose.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination preserves the exhausted, rain-soaked appearance without introducing fresh rainfall or an unsupported window effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A rain shower has recently passed, and the operational camper carries the luggage, added food and medicine supplied at the repair shop. 라울: He is riding in the camper, soaked and exhausted, with his earlier injury treated.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 창밖의 숲 쪽을 향해 검지손가락을 길게 뻗은 라울의 상체.\n\nLOCATION (lock): Inside the camper's passenger area, beside a window overlooking roadside woodland in daylight after rain. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Forest visible beyond the camper side window in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Camper side window (Providing a view of the forest outside) — Seen obliquely from the aisle, with the forest visible beyond the pointing hand; used as Connects the interior gesture to its exterior destination without requiring the camera to leave the camper; Forest beside the route (Visible outside as a possible route taken by 찰리) — Appears beyond the window toward the upper right and continues outside the crop; used as Provides an intelligible destination for the pointing line; Rear seating (Occupied by 라울) — A partial seat edge remains beneath his leaning torso; used as Anchors his seated posture and prevents the pointing gesture from reading as a standing pose.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination preserves the exhausted, rain-soaked appearance without introducing fresh rainfall or an unsupported window effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A rain shower has recently passed, and the operational camper carries the luggage, added food and medicine supplied at the repair shop. 라울: He is riding in the camper, soaked and exhausted, with his earlier injury treated.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 창밖의 숲 쪽을 향해 검지손가락을 길게 뻗은 라울의 상체.\n\nLOCATION (lock): Inside the camper's passenger area, beside a window overlooking roadside woodland in daylight after rain. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Forest visible beyond the camper side window in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Camper side window (Providing a view of the forest outside) — Seen obliquely from the aisle, with the forest visible beyond the pointing hand; used as Connects the interior gesture to its exterior destination without requiring the camera to leave the camper; Forest beside the route (Visible outside as a possible route taken by 찰리) — Appears beyond the window toward the upper right and continues outside the crop; used as Provides an intelligible destination for the pointing line; Rear seating (Occupied by 라울) — A partial seat edge remains beneath his leaning torso; used as Anchors his seated posture and prevents the pointing gesture from reading as a standing pose.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient illumination preserves the exhausted, rain-soaked appearance without introducing fresh rainfall or an unsupported window effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A rain shower has recently passed, and the operational camper carries the luggage, added food and medicine supplied at the repair shop. 라울: He is riding in the camper, soaked and exhausted, with his earlier injury treated.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S55sh3__bgfirst_bg.png",
     "asset_id": "f40aa134-4d1d-43ab-80e4-59ce7d1bc555",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S55sh3.png",
     "asset_id": "30ef55b8-86fa-49ef-81a1-ac257a19f52a",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163202>",
     "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L199B03.png",
     "asset_id": "4be906f8-5377-4122-9528-ee6aaccbd804",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163202>",
     "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "시선과 뻗은 오른손 검지가 우측 상단 창밖의 숲을 정확히 향함.",
    "built_space": "캠퍼 내부 구조와 가구는 참조와 완벽히 일치하나, 인물이 지시된 뒷좌석이 아닌 전경에 위치하여 뒷좌석이 비어 있음.",
    "entities": "인종, 나이, 꽁지머리 등 라울의 캐릭터 참조와 완전히 일치함.",
    "hard_violations": [
     "[gemini-pro] 지정된 후면 좌석(Rear seating)이 아닌 전경에 인물을 잘못 배치함",
     "[gpt-high] 라울이 앉아 있어야 할 후방 벤치를 비워 두고 인물을 통로 앞쪽 좌석에 배치하여, 명시된 인물·좌석 관계를 위반했다."
    ],
    "physics": "무릎과 왼손으로 바닥을 짚어 몸을 안정적으로 지지하며, 오른팔을 자연스럽게 뻗고 있음."
   },
   {
    "label": "B",
    "direction": "시선과 뻗은 손이 우측 창밖 숲을 향하고 있음.",
    "built_space": "위치 참조와 달리 뒷좌석 소파가 복제되고 없는 복도가 생성되는 등 내부 구조가 완전히 변경됨.",
    "entities": "라울의 전반적인 외형과 의상 참조와 일치함.",
    "hard_violations": [
     "[gemini-pro] 위치 구조 임의 변경 및 가구 복제 (소파 2개)",
     "[gemini-pro] 왼팔에 오른손이 연결된 물리적으로 불가능한 해부학적 오류",
     "[gpt-high] 참조에서 서로 마주 보는 측벽의 창들을 같은 쪽 벽에 연속 배치하고 별도 좌석을 추가하여, 고정된 촬영 장소의 공간 구조를 변경했다.",
     "[gpt-high] 라울이 점유해야 하는 후방 좌석 대신 별도의 앞쪽 의자에 앉아 있다."
    ],
    "physics": "소파에 앉아 체중을 지탱하고 있으나 뻗은 팔의 손 방향이 신체 구조에 맞지 않음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "공간의 디테일과 조명은 완벽에 가깝게 구현되었으나, 인물이 지정된 후면 좌석이 아닌 전경에 배치되어 공간 연출 지침을 위반했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "캠퍼 내부 구조가 심하게 왜곡되어 위치 참조를 위반했으며, 왼팔에 오른손이 달린 치명적인 해부학적 오류가 있습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선과 뻗은 오른손 검지가 우측 상단 창밖의 숲을 정확히 향함.",
        "built_space": "캠퍼 내부 구조와 가구는 참조와 완벽히 일치하나, 인물이 지시된 뒷좌석이 아닌 전경에 위치하여 뒷좌석이 비어 있음.",
        "entities": "인종, 나이, 꽁지머리 등 라울의 캐릭터 참조와 완전히 일치함.",
        "hard_violations": [
         "지정된 후면 좌석(Rear seating)이 아닌 전경에 인물을 잘못 배치함"
        ],
        "physics": "무릎과 왼손으로 바닥을 짚어 몸을 안정적으로 지지하며, 오른팔을 자연스럽게 뻗고 있음."
       },
       {
        "label": "B",
        "direction": "시선과 뻗은 손이 우측 창밖 숲을 향하고 있음.",
        "built_space": "위치 참조와 달리 뒷좌석 소파가 복제되고 없는 복도가 생성되는 등 내부 구조가 완전히 변경됨.",
        "entities": "라울의 전반적인 외형과 의상 참조와 일치함.",
        "hard_violations": [
         "위치 구조 임의 변경 및 가구 복제 (소파 2개)",
         "왼팔에 오른손이 연결된 물리적으로 불가능한 해부학적 오류"
        ],
        "physics": "소파에 앉아 체중을 지탱하고 있으나 뻗은 팔의 손 방향이 신체 구조에 맞지 않음."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "공간의 디테일과 조명은 완벽에 가깝게 구현되었으나, 인물이 지정된 후면 좌석이 아닌 전경에 배치되어 공간 연출 지침을 위반했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "캠퍼 내부 구조가 심하게 왜곡되어 위치 참조를 위반했으며, 왼팔에 오른손이 달린 치명적인 해부학적 오류가 있습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "시선과 뻗은 오른손 검지가 우측 상단 창밖의 숲을 정확히 향함.",
        "built_space": "캠퍼 내부 구조와 가구는 참조와 완벽히 일치하나, 인물이 지시된 뒷좌석이 아닌 전경에 위치하여 뒷좌석이 비어 있음.",
        "entities": "인종, 나이, 꽁지머리 등 라울의 캐릭터 참조와 완전히 일치함.",
        "hard_violations": [
         "지정된 후면 좌석(Rear seating)이 아닌 전경에 인물을 잘못 배치함"
        ],
        "physics": "무릎과 왼손으로 바닥을 짚어 몸을 안정적으로 지지하며, 오른팔을 자연스럽게 뻗고 있음."
       },
       {
        "label": "B",
        "direction": "시선과 뻗은 손이 우측 창밖 숲을 향하고 있음.",
        "built_space": "위치 참조와 달리 뒷좌석 소파가 복제되고 없는 복도가 생성되는 등 내부 구조가 완전히 변경됨.",
        "entities": "라울의 전반적인 외형과 의상 참조와 일치함.",
        "hard_violations": [
         "위치 구조 임의 변경 및 가구 복제 (소파 2개)",
         "왼팔에 오른손이 연결된 물리적으로 불가능한 해부학적 오류"
        ],
        "physics": "소파에 앉아 체중을 지탱하고 있으나 뻗은 팔의 손 방향이 신체 구조에 맞지 않음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "숲을 향한 손가락과 착석 지지는 명확하지만, 창과 좌석의 구조를 바꾸고 라울을 후방 벤치가 아닌 별도 의자에 앉혀 장소·배치 조건을 위반한다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "참조 공간과 우상단 숲을 향해 길게 뻗은 동작은 더 충실하지만, 후방 벤치를 비워 둔 채 라울을 통로 앞쪽에 앉혀 필수 착석 배치를 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "라울의 얼굴과 시선은 오른쪽 창밖 나무들을 향한다. 팔꿈치를 굽힌 오른팔의 검지는 우상향하며 창 너머 숲을 가리킨다. 손끝의 연장선에 실제 수목이 있어 지시 대상은 명확하다.",
        "built_space": "오른쪽으로 이어지는 벽에 큰 창 하나와 그 뒤 작은 창 하나가 보인다. 왼쪽 먼 곳에는 벤치, 그 앞에는 작은 스툴, 라울 아래에는 별도의 등받이·팔걸이 의자가 있다. 참조의 후방 벤치 양옆에 창이 하나씩 있는 구조와 다르다. 라울은 후방 벤치가 아니라 앞쪽 별도 의자에 앉아 있다. 상체뿐 아니라 허벅지와 넓은 실내까지 포함한다.",
        "entities": "보이는 사람은 어린 남자아이 한 명이다. 갈색 피부, 어린 얼굴, 뒤로 묶은 곱슬머리와 남색 반팔은 라울의 참조에 대체로 부합한다. 혼혈 배경 자체는 외형만으로 확정할 수 없다. 젖은 피부와 옷, 창밖 숲, 짐가방과 옷가지가 보이며 읽을 수 있는 글자는 없다. 팔에 상처 같은 흔적은 있지만 치료 상태는 확인되지 않는다.",
        "hard_violations": [
         "참조에서 서로 마주 보는 측벽의 창들을 같은 쪽 벽에 연속 배치하고 별도 좌석을 추가하여, 고정된 촬영 장소의 공간 구조를 변경했다.",
         "라울이 점유해야 하는 후방 좌석 대신 별도의 앞쪽 의자에 앉아 있다."
        ],
        "physics": "엉덩이와 허벅지는 의자 방석에 놓이고 왼손은 팔걸이에 닿아 기울어진 몸을 지지한다. 오른팔과 손가락의 거상은 어깨와 팔꿈치로 가능한 자세다. 몸이나 물체가 지지 없이 떠 있는 현상은 없다. 창의 물방울은 지나간 비의 잔류 수분으로 읽힌다."
       },
       {
        "label": "B",
        "direction": "라울은 오른쪽 창밖을 바라보며 오른팔과 검지를 길게 우상향으로 뻗는다. 손끝 너머 우상단 창 안에 숲이 이어져 지시선의 목적지가 분명하다. 카메라를 보는 시선은 아니다.",
        "built_space": "좌우 측벽 창이 각각 하나, 뒤쪽 벤치 하나, 벤치 아래 수납부, 오른쪽 벽등 하나, 천장 개구부 하나와 걸린 천, 왼쪽 커튼이 보인다. 참조의 구조와 재질을 잘 유지한다. 그러나 후방 벤치는 비어 있고 라울은 그보다 훨씬 앞쪽의 화면 하단 좌석에 앉아 있다. 통로에서 창을 비스듬히 보는 시점과 우상단 숲 배치는 맞지만, 허벅지와 실내를 상당히 포함해 요청한 상체 중심보다 넓다.",
        "entities": "어린 남자아이 한 명만 보이며 얼굴 윤곽, 갈색 피부, 뒤로 묶은 곱슬머리, 남색 반팔이 라울 참조와 가깝다. 혼혈 배경 자체는 외형만으로 확정할 수 없다. 젖어 달라붙은 셔츠와 피부의 물기가 비에 젖은 상태를 분명히 표현한다. 숲과 젖은 도로가 창밖에 보이고 읽을 수 있는 글자는 없다. 치료된 부상은 보이는 범위에서 확인되지 않는다.",
        "hard_violations": [
         "라울이 앉아 있어야 할 후방 벤치를 비워 두고 인물을 통로 앞쪽 좌석에 배치하여, 명시된 인물·좌석 관계를 위반했다."
        ],
        "physics": "화면 하단에 걸친 좌석 가장자리와 엉덩이·허벅지의 관계가 착석을 나타내며, 왼손도 아래쪽 지지면에 닿아 있다. 후방 벤치에 앉은 것은 아니지만 공중에 떠 있는 자세는 아니다. 뻗은 오른팔은 어깨에 자연스럽게 연결되고 손가락 동작도 물리적으로 가능하다. 창과 천에 남은 물기는 최근 비가 지난 상황과 양립한다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "숲을 향한 손가락과 착석 지지는 명확하지만, 창과 좌석의 구조를 바꾸고 라울을 후방 벤치가 아닌 별도 의자에 앉혀 장소·배치 조건을 위반한다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "참조 공간과 우상단 숲을 향해 길게 뻗은 동작은 더 충실하지만, 후방 벤치를 비워 둔 채 라울을 통로 앞쪽에 앉혀 필수 착석 배치를 위반한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "라울의 얼굴과 시선은 오른쪽 창밖 나무들을 향한다. 팔꿈치를 굽힌 오른팔의 검지는 우상향하며 창 너머 숲을 가리킨다. 손끝의 연장선에 실제 수목이 있어 지시 대상은 명확하다.",
        "built_space": "오른쪽으로 이어지는 벽에 큰 창 하나와 그 뒤 작은 창 하나가 보인다. 왼쪽 먼 곳에는 벤치, 그 앞에는 작은 스툴, 라울 아래에는 별도의 등받이·팔걸이 의자가 있다. 참조의 후방 벤치 양옆에 창이 하나씩 있는 구조와 다르다. 라울은 후방 벤치가 아니라 앞쪽 별도 의자에 앉아 있다. 상체뿐 아니라 허벅지와 넓은 실내까지 포함한다.",
        "entities": "보이는 사람은 어린 남자아이 한 명이다. 갈색 피부, 어린 얼굴, 뒤로 묶은 곱슬머리와 남색 반팔은 라울의 참조에 대체로 부합한다. 혼혈 배경 자체는 외형만으로 확정할 수 없다. 젖은 피부와 옷, 창밖 숲, 짐가방과 옷가지가 보이며 읽을 수 있는 글자는 없다. 팔에 상처 같은 흔적은 있지만 치료 상태는 확인되지 않는다.",
        "hard_violations": [
         "참조에서 서로 마주 보는 측벽의 창들을 같은 쪽 벽에 연속 배치하고 별도 좌석을 추가하여, 고정된 촬영 장소의 공간 구조를 변경했다.",
         "라울이 점유해야 하는 후방 좌석 대신 별도의 앞쪽 의자에 앉아 있다."
        ],
        "physics": "엉덩이와 허벅지는 의자 방석에 놓이고 왼손은 팔걸이에 닿아 기울어진 몸을 지지한다. 오른팔과 손가락의 거상은 어깨와 팔꿈치로 가능한 자세다. 몸이나 물체가 지지 없이 떠 있는 현상은 없다. 창의 물방울은 지나간 비의 잔류 수분으로 읽힌다."
       },
       {
        "label": "A",
        "direction": "라울은 오른쪽 창밖을 바라보며 오른팔과 검지를 길게 우상향으로 뻗는다. 손끝 너머 우상단 창 안에 숲이 이어져 지시선의 목적지가 분명하다. 카메라를 보는 시선은 아니다.",
        "built_space": "좌우 측벽 창이 각각 하나, 뒤쪽 벤치 하나, 벤치 아래 수납부, 오른쪽 벽등 하나, 천장 개구부 하나와 걸린 천, 왼쪽 커튼이 보인다. 참조의 구조와 재질을 잘 유지한다. 그러나 후방 벤치는 비어 있고 라울은 그보다 훨씬 앞쪽의 화면 하단 좌석에 앉아 있다. 통로에서 창을 비스듬히 보는 시점과 우상단 숲 배치는 맞지만, 허벅지와 실내를 상당히 포함해 요청한 상체 중심보다 넓다.",
        "entities": "어린 남자아이 한 명만 보이며 얼굴 윤곽, 갈색 피부, 뒤로 묶은 곱슬머리, 남색 반팔이 라울 참조와 가깝다. 혼혈 배경 자체는 외형만으로 확정할 수 없다. 젖어 달라붙은 셔츠와 피부의 물기가 비에 젖은 상태를 분명히 표현한다. 숲과 젖은 도로가 창밖에 보이고 읽을 수 있는 글자는 없다. 치료된 부상은 보이는 범위에서 확인되지 않는다.",
        "hard_violations": [
         "라울이 앉아 있어야 할 후방 벤치를 비워 두고 인물을 통로 앞쪽 좌석에 배치하여, 명시된 인물·좌석 관계를 위반했다."
        ],
        "physics": "화면 하단에 걸친 좌석 가장자리와 엉덩이·허벅지의 관계가 착석을 나타내며, 왼손도 아래쪽 지지면에 닿아 있다. 후방 벤치에 앉은 것은 아니지만 공중에 떠 있는 자세는 아니다. 뻗은 오른팔은 어깨에 자연스럽게 연결되고 손가락 동작도 물리적으로 가능하다. 창과 천에 남은 물기는 최근 비가 지난 상황과 양립한다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.5
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.25
   },
   "violations": {
    "A": [
     "[gemini-pro] 지정된 후면 좌석(Rear seating)이 아닌 전경에 인물을 잘못 배치함",
     "[gpt-high] 라울이 앉아 있어야 할 후방 벤치를 비워 두고 인물을 통로 앞쪽 좌석에 배치하여, 명시된 인물·좌석 관계를 위반했다."
    ],
    "B": [
     "[gemini-pro] 위치 구조 임의 변경 및 가구 복제 (소파 2개)",
     "[gemini-pro] 왼팔에 오른손이 연결된 물리적으로 불가능한 해부학적 오류",
     "[gpt-high] 참조에서 서로 마주 보는 측벽의 창들을 같은 쪽 벽에 연속 배치하고 별도 좌석을 추가하여, 고정된 촬영 장소의 공간 구조를 변경했다.",
     "[gpt-high] 라울이 점유해야 하는 후방 좌석 대신 별도의 앞쪽 의자에 앉아 있다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 1250
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "공간의 디테일과 조명은 완벽에 가깝게 구현되었으나, 인물이 지정된 후면 좌석이 아닌 전경에 배치되어 공간 연출 지침을 위반했습니다.  ★위반: [gemini-pro] 지정된 후면 좌석(Rear seating)이 아닌 전경에 인물을 잘못 배치함 / [gpt-high] 라울이 앉아 있어야 할 후방 벤치를 비워 두고 인물을 통로 앞쪽 좌석에 배치하여, 명시된 인물·좌석 관계를 위반했다."
   },
   {
    "label": "B",
    "score": 1250,
    "verdict_ko": "캠퍼 내부 구조가 심하게 왜곡되어 위치 참조를 위반했으며, 왼팔에 오른손이 달린 치명적인 해부학적 오류가 있습니다.  ★위반: [gemini-pro] 위치 구조 임의 변경 및 가구 복제 (소파 2개) / [gemini-pro] 왼팔에 오른손이 연결된 물리적으로 불가능한 해부학적 오류 / [gpt-high] 참조에서 서로 마주 보는 측벽의 창들을 같은 쪽 벽에 연속 배치하고 별도 좌석을 추가하여, 고정된 촬영 장소의 공간 구조를 변경했다. / [gpt-high] 라울이 점유해야 하는 후방 좌석 대신 별도의 앞쪽 의자에 앉아 있다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L199B03.png",
    "asset_id": "4be906f8-5377-4122-9528-ee6aaccbd804",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163202>",
    "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-5461-79da-8bc7-1ededba6b741",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S55sh3__bgfirst_bg.png",
   "bg_asset_id": "f40aa134-4d1d-43ab-80e4-59ce7d1bc555",
   "bg_record_key": "S55sh3::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S55sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:17:43.101706+00:00",
  "fingerprint": "b921c9af6a3a14e44ca7b13341cd6978c0a9d8a57119746bc3a6e4da3f8ad53f",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S55sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S55sh3_sel.png",
  "source_sha256": "e6318391d893a89d8af2f7e281c6b96324aa99211d0f49e39795fd5953e02fd0",
  "file": "S55sh3_cine.png",
  "staged_sha256": "969f76366b8e7cba437b54c0343329b5651ce0e447a519c596a5281d42041a48",
  "latency_ms": 11124
 },
 "S55sh5::signage": {
  "fp": "f6a6e7660f37d1bf",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S55sh5": {
  "input_fingerprint": "48766cae24ec19f2",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 거칠게 솟구친 흙먼지를 뒤로한 채 숲길을 향해 급격히 꺾인 캠핑카의 바퀴 찰나.\n\nLOCATION (lock): At the turnoff from the provincial road onto a forest track, beside the camper's sharply turning wheels. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Forest route opening ahead of the turning wheel in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Camper front wheel (Sharply turned toward the forest route) — Seen from slightly behind the axle, exposing the changed relationship between its side and forward-facing tread; used as Primary motion detail, kept at natural scale against the adjoining vehicle body; Adjacent camper bodywork (Moving with the turning wheel) — A partial side section recedes above and behind the wheel; used as Provides a scale reference and establishes the wheel's steering angle relative to the vehicle; Forest route entrance (Being entered by the turning camper) — Opens ahead of the wheel toward the upper right; used as Makes the destination of the turn legible; Kicked-up earth (Thrown backward during the abrupt turn) — Trails behind the wheel toward the lower left; used as Reinforces the turn's force without obscuring wheel contact.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate daytime illumination keeps the wheel angle, road contact, and kicked-up earth distinct without stylized lighting changes.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper turns toward the forest after a passing shower, still carrying the loaded luggage, food and medicine.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 거칠게 솟구친 흙먼지를 뒤로한 채 숲길을 향해 급격히 꺾인 캠핑카의 바퀴 찰나.\n\nLOCATION (lock): At the turnoff from the provincial road onto a forest track, beside the camper's sharply turning wheels. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Forest route opening ahead of the turning wheel in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Camper front wheel (Sharply turned toward the forest route) — Seen from slightly behind the axle, exposing the changed relationship between its side and forward-facing tread; used as Primary motion detail, kept at natural scale against the adjoining vehicle body; Adjacent camper bodywork (Moving with the turning wheel) — A partial side section recedes above and behind the wheel; used as Provides a scale reference and establishes the wheel's steering angle relative to the vehicle; Forest route entrance (Being entered by the turning camper) — Opens ahead of the wheel toward the upper right; used as Makes the destination of the turn legible; Kicked-up earth (Thrown backward during the abrupt turn) — Trails behind the wheel toward the lower left; used as Reinforces the turn's force without obscuring wheel contact.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate daytime illumination keeps the wheel angle, road contact, and kicked-up earth distinct without stylized lighting changes.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper turns toward the forest after a passing shower, still carrying the loaded luggage, food and medicine.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 거칠게 솟구친 흙먼지를 뒤로한 채 숲길을 향해 급격히 꺾인 캠핑카의 바퀴 찰나.\n\nLOCATION (lock): At the turnoff from the provincial road onto a forest track, beside the camper's sharply turning wheels. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Forest route opening ahead of the turning wheel in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Camper front wheel (Sharply turned toward the forest route) — Seen from slightly behind the axle, exposing the changed relationship between its side and forward-facing tread; used as Primary motion detail, kept at natural scale against the adjoining vehicle body; Adjacent camper bodywork (Moving with the turning wheel) — A partial side section recedes above and behind the wheel; used as Provides a scale reference and establishes the wheel's steering angle relative to the vehicle; Forest route entrance (Being entered by the turning camper) — Opens ahead of the wheel toward the upper right; used as Makes the destination of the turn legible; Kicked-up earth (Thrown backward during the abrupt turn) — Trails behind the wheel toward the lower left; used as Reinforces the turn's force without obscuring wheel contact.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate daytime illumination keeps the wheel angle, road contact, and kicked-up earth distinct without stylized lighting changes.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper turns toward the forest after a passing shower, still carrying the loaded luggage, food and medicine.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "B",
    "direction": "차체 앞부분은 화면 오른쪽에 있고 측면은 왼쪽으로 이어진다. 앞바퀴는 측면과 넓은 트레드가 함께 보이도록 꺾여 있으며 진행 방향은 우상단 흙길 입구 쪽으로 읽힌다. 흙먼지는 접지점에서 좌하단 포장도로 쪽으로 뻗는다. 사람이나 시선 대상은 없다.",
    "built_space": "앞바퀴 한 개와 휠하우스 한 개가 크게 보이고, 왼쪽 차체 아래에는 다른 바퀴 일부가 보인다. 아래쪽 포장도로와 오른쪽 흙길이 맞닿아 있어 진입 관계는 성립한다. 다만 참조의 깊게 갈라진 포장면과 바위·낙엽이 쌓인 오른쪽 둔덕 대신 비교적 평탄한 길 가장자리와 흰 도로선이 보인다. 건물이나 반사면은 없다.",
    "entities": "은색 차량의 앞바퀴, 인접 차체, 숲길, 자갈과 흙먼지가 보이며 인물이나 읽을 수 있는 글자는 없다. 차체는 참조의 짙은 하부 패널과 수납함이 있는 캠핑카보다 매끈한 승합차 외장에 가깝다. 짐·식량·의약품은 이 외부 클로즈업에서 보이지 않아 적재 상태를 판단할 수 없다. 낮 장면이지만 소나기 직후의 젖은 표면은 뚜렷하지 않다.",
    "hard_violations": [],
    "physics": "앞타이어 하단이 흙과 포장 경계에 닿아 있고 바퀴는 휠하우스와 차체 아래 장치에 연결되어 있다. 차체와 바퀴의 크기 관계도 자연스럽다. 공중의 흙과 작은 돌은 접지점 부근에서 좌하단으로 튀어나온 것으로 보여 급회전 중 지면을 긁는 작용으로 설명된다. 근거 없이 떠 있는 물체는 없다."
   },
   {
    "label": "A",
    "direction": "차체 측면은 앞바퀴에서 왼쪽 뒤로 물러나고 앞범퍼 일부는 오른쪽에 보인다. 바퀴의 측면과 트레드가 동시에 드러나 차체에 대한 조향각이 읽히며, 바퀴 앞쪽에는 우상단으로 이어지는 숲길이 열린다. 흙먼지와 자갈은 바퀴 뒤쪽인 좌하단으로 튄다. 사람이나 시선 대상은 없다.",
    "built_space": "앞바퀴와 휠하우스가 각각 하나씩 보이고, 차체 아래 왼쪽에는 다른 바퀴 하나가 부분적으로 보인다. 측면에는 큰 하부 수납 패널과 그 위쪽 패널이 이어진다. 전경의 갈라진 포장도로, 우상단의 좁은 흙길, 오른쪽 나무줄기와 돌·낙엽 둔덕이 참조 장소의 구성과 가깝다. 낮은 차축 뒤쪽 시점에서 앞바퀴와 진입로가 함께 보이는 공간 관계도 성립한다.",
    "entities": "참조와 유사한 회색 캠핑카 외장과 어두운 하부 패널, 조향된 앞바퀴, 숲길 입구, 튀는 흙과 자갈이 보인다. 인물과 읽을 수 있는 글자는 없다. 내부 적재물은 프레임 밖이므로 판단 대상이 아니다. 숲의 낮 조명은 맞지만 소나기 직후라는 단서는 표면에서 강하게 드러나지 않는다.",
    "hard_violations": [],
    "physics": "앞바퀴는 자갈 섞인 지면에 확실히 접촉하고 휠하우스 안쪽으로 연결되며, 다른 바퀴도 차체 아래 지면에 놓여 있다. 차량을 지탱하는 관계와 바퀴 크기는 자연스럽다. 떠 있는 자갈과 흙은 접지점 뒤에 집중되어 타이어가 흙을 밀어내며 튀긴 궤적으로 설명되고, 먼지가 접지부를 완전히 가리지 않는다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": null,
     "normalized": null,
     "ok": false
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "바퀴 클로즈업과 우상단 진입·좌하단 흙먼지 배치는 충실하지만, 매끈한 승합차형 차체와 단순한 길 가장자리가 참조의 캠핑카 및 장소와 다르다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "차축 뒤쪽에서 본 꺾인 앞바퀴 클로즈업을 유지하면서 참조의 캠핑카 외장, 갈라진 도로와 돌 많은 숲길, 뒤로 튀는 흙을 가장 충실하게 재현했다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "차체 앞부분은 화면 오른쪽에 있고 측면은 왼쪽으로 이어진다. 앞바퀴는 측면과 넓은 트레드가 함께 보이도록 꺾여 있으며 진행 방향은 우상단 흙길 입구 쪽으로 읽힌다. 흙먼지는 접지점에서 좌하단 포장도로 쪽으로 뻗는다. 사람이나 시선 대상은 없다.",
        "built_space": "앞바퀴 한 개와 휠하우스 한 개가 크게 보이고, 왼쪽 차체 아래에는 다른 바퀴 일부가 보인다. 아래쪽 포장도로와 오른쪽 흙길이 맞닿아 있어 진입 관계는 성립한다. 다만 참조의 깊게 갈라진 포장면과 바위·낙엽이 쌓인 오른쪽 둔덕 대신 비교적 평탄한 길 가장자리와 흰 도로선이 보인다. 건물이나 반사면은 없다.",
        "entities": "은색 차량의 앞바퀴, 인접 차체, 숲길, 자갈과 흙먼지가 보이며 인물이나 읽을 수 있는 글자는 없다. 차체는 참조의 짙은 하부 패널과 수납함이 있는 캠핑카보다 매끈한 승합차 외장에 가깝다. 짐·식량·의약품은 이 외부 클로즈업에서 보이지 않아 적재 상태를 판단할 수 없다. 낮 장면이지만 소나기 직후의 젖은 표면은 뚜렷하지 않다.",
        "hard_violations": [],
        "physics": "앞타이어 하단이 흙과 포장 경계에 닿아 있고 바퀴는 휠하우스와 차체 아래 장치에 연결되어 있다. 차체와 바퀴의 크기 관계도 자연스럽다. 공중의 흙과 작은 돌은 접지점 부근에서 좌하단으로 튀어나온 것으로 보여 급회전 중 지면을 긁는 작용으로 설명된다. 근거 없이 떠 있는 물체는 없다."
       },
       {
        "label": "B",
        "direction": "차체 측면은 앞바퀴에서 왼쪽 뒤로 물러나고 앞범퍼 일부는 오른쪽에 보인다. 바퀴의 측면과 트레드가 동시에 드러나 차체에 대한 조향각이 읽히며, 바퀴 앞쪽에는 우상단으로 이어지는 숲길이 열린다. 흙먼지와 자갈은 바퀴 뒤쪽인 좌하단으로 튄다. 사람이나 시선 대상은 없다.",
        "built_space": "앞바퀴와 휠하우스가 각각 하나씩 보이고, 차체 아래 왼쪽에는 다른 바퀴 하나가 부분적으로 보인다. 측면에는 큰 하부 수납 패널과 그 위쪽 패널이 이어진다. 전경의 갈라진 포장도로, 우상단의 좁은 흙길, 오른쪽 나무줄기와 돌·낙엽 둔덕이 참조 장소의 구성과 가깝다. 낮은 차축 뒤쪽 시점에서 앞바퀴와 진입로가 함께 보이는 공간 관계도 성립한다.",
        "entities": "참조와 유사한 회색 캠핑카 외장과 어두운 하부 패널, 조향된 앞바퀴, 숲길 입구, 튀는 흙과 자갈이 보인다. 인물과 읽을 수 있는 글자는 없다. 내부 적재물은 프레임 밖이므로 판단 대상이 아니다. 숲의 낮 조명은 맞지만 소나기 직후라는 단서는 표면에서 강하게 드러나지 않는다.",
        "hard_violations": [],
        "physics": "앞바퀴는 자갈 섞인 지면에 확실히 접촉하고 휠하우스 안쪽으로 연결되며, 다른 바퀴도 차체 아래 지면에 놓여 있다. 차량을 지탱하는 관계와 바퀴 크기는 자연스럽다. 떠 있는 자갈과 흙은 접지점 뒤에 집중되어 타이어가 흙을 밀어내며 튀긴 궤적으로 설명되고, 먼지가 접지부를 완전히 가리지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "바퀴 클로즈업과 우상단 진입·좌하단 흙먼지 배치는 충실하지만, 매끈한 승합차형 차체와 단순한 길 가장자리가 참조의 캠핑카 및 장소와 다르다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "차축 뒤쪽에서 본 꺾인 앞바퀴 클로즈업을 유지하면서 참조의 캠핑카 외장, 갈라진 도로와 돌 많은 숲길, 뒤로 튀는 흙을 가장 충실하게 재현했다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "차체 앞부분은 화면 오른쪽에 있고 측면은 왼쪽으로 이어진다. 앞바퀴는 측면과 넓은 트레드가 함께 보이도록 꺾여 있으며 진행 방향은 우상단 흙길 입구 쪽으로 읽힌다. 흙먼지는 접지점에서 좌하단 포장도로 쪽으로 뻗는다. 사람이나 시선 대상은 없다.",
        "built_space": "앞바퀴 한 개와 휠하우스 한 개가 크게 보이고, 왼쪽 차체 아래에는 다른 바퀴 일부가 보인다. 아래쪽 포장도로와 오른쪽 흙길이 맞닿아 있어 진입 관계는 성립한다. 다만 참조의 깊게 갈라진 포장면과 바위·낙엽이 쌓인 오른쪽 둔덕 대신 비교적 평탄한 길 가장자리와 흰 도로선이 보인다. 건물이나 반사면은 없다.",
        "entities": "은색 차량의 앞바퀴, 인접 차체, 숲길, 자갈과 흙먼지가 보이며 인물이나 읽을 수 있는 글자는 없다. 차체는 참조의 짙은 하부 패널과 수납함이 있는 캠핑카보다 매끈한 승합차 외장에 가깝다. 짐·식량·의약품은 이 외부 클로즈업에서 보이지 않아 적재 상태를 판단할 수 없다. 낮 장면이지만 소나기 직후의 젖은 표면은 뚜렷하지 않다.",
        "hard_violations": [],
        "physics": "앞타이어 하단이 흙과 포장 경계에 닿아 있고 바퀴는 휠하우스와 차체 아래 장치에 연결되어 있다. 차체와 바퀴의 크기 관계도 자연스럽다. 공중의 흙과 작은 돌은 접지점 부근에서 좌하단으로 튀어나온 것으로 보여 급회전 중 지면을 긁는 작용으로 설명된다. 근거 없이 떠 있는 물체는 없다."
       },
       {
        "label": "A",
        "direction": "차체 측면은 앞바퀴에서 왼쪽 뒤로 물러나고 앞범퍼 일부는 오른쪽에 보인다. 바퀴의 측면과 트레드가 동시에 드러나 차체에 대한 조향각이 읽히며, 바퀴 앞쪽에는 우상단으로 이어지는 숲길이 열린다. 흙먼지와 자갈은 바퀴 뒤쪽인 좌하단으로 튄다. 사람이나 시선 대상은 없다.",
        "built_space": "앞바퀴와 휠하우스가 각각 하나씩 보이고, 차체 아래 왼쪽에는 다른 바퀴 하나가 부분적으로 보인다. 측면에는 큰 하부 수납 패널과 그 위쪽 패널이 이어진다. 전경의 갈라진 포장도로, 우상단의 좁은 흙길, 오른쪽 나무줄기와 돌·낙엽 둔덕이 참조 장소의 구성과 가깝다. 낮은 차축 뒤쪽 시점에서 앞바퀴와 진입로가 함께 보이는 공간 관계도 성립한다.",
        "entities": "참조와 유사한 회색 캠핑카 외장과 어두운 하부 패널, 조향된 앞바퀴, 숲길 입구, 튀는 흙과 자갈이 보인다. 인물과 읽을 수 있는 글자는 없다. 내부 적재물은 프레임 밖이므로 판단 대상이 아니다. 숲의 낮 조명은 맞지만 소나기 직후라는 단서는 표면에서 강하게 드러나지 않는다.",
        "hard_violations": [],
        "physics": "앞바퀴는 자갈 섞인 지면에 확실히 접촉하고 휠하우스 안쪽으로 연결되며, 다른 바퀴도 차체 아래 지면에 놓여 있다. 차량을 지탱하는 관계와 바퀴 크기는 자연스럽다. 떠 있는 자갈과 흙은 접지점 뒤에 집중되어 타이어가 흙을 밀어내며 튀긴 궤적으로 설명되고, 먼지가 접지부를 완전히 가리지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gemini-pro"
   ],
   "route": "single_reverse"
  },
  "totals": {
   "B": 7,
   "A": 9
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 7,
    "verdict_ko": "바퀴 클로즈업과 우상단 진입·좌하단 흙먼지 배치는 충실하지만, 매끈한 승합차형 차체와 단순한 길 가장자리가 참조의 캠핑카 및 장소와 다르다."
   },
   {
    "label": "A",
    "score": 9,
    "verdict_ko": "차축 뒤쪽에서 본 꺾인 앞바퀴 클로즈업을 유지하면서 참조의 캠핑카 외장, 갈라진 도로와 돌 많은 숲길, 뒤로 튀는 흙을 가장 충실하게 재현했다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L199B01.png",
    "asset_id": "ffcb0b14-6163-46a6-a2d3-38e545216e88",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-57a2-7ca9-8a00-1193479a2c8c",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S55sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:18:39.099938+00:00",
  "fingerprint": "5b35a120db932620a374842422a1c958524683aa823751bcd8240327f438cfa3",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S55sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S55sh5_sel.png",
  "source_sha256": "09159aa142b7a96071f17257237e4b8dcc3c937f5c0b2eb4801f15ee3086ae7a",
  "file": "S55sh5_cine.png",
  "staged_sha256": "9fe50594028551319fc3b509145cee82a368c94089c6f32ef63c4fd79a69e4ed",
  "latency_ms": 12362
 },
 "S56sh5::signage": {
  "fp": "822597267d87434b",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::8860149eacdd8c67": {
  "subjects": [],
  "subject_text": "익산 늪지대의 캠핑카 고립 지점\n숲길 끝에 펼쳐진 황폐한 늪지대. 물기 많은 진흙 바닥 옆에 출입 금지 팻말과 심하게 부식된 방사능 구역 표지가 서 있다.",
  "identity": "canonical",
  "scope_id": "L215",
  "scope_role": "location_exterior",
  "scope_sha": "a6042d9cf76b5b12"
 },
 "groupbg::swamp_vehicle_edge": {
  "input_fingerprint": "7404e75b60cc272c",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "swamp_vehicle_edge",
    "tags": [
     "S56sh5",
     "S61sh1",
     "S61sh5",
     "S61sh7"
    ]
   },
   "context_sig": "b16126b42fff63ca"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On muddy ground near the stranded camper in the marsh, along the search route for boards or timber.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n익산 늪지대의 캠핑카 고립 지점: 차량 바퀴가 푹 빠지는 질척한 뻘과 낡은 경고판이 꽂혀 있는 지대. (특징: 차체 하부가 빠져 움직이지 못하는 진흙 구덩이; 주변의 메마른 늪지대 환경; 녹이 슬어 글자가 지워진 구형 방사능 표지판과 출입 금지 팻말)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 현우의 차가 늪지대 앞에서 멈추려다 풍덩 빠진다.\n- 부근에 출입 금지 팻말과 함께 부식돼 식별이 안 되는 방사능 구역 표시\n- 늪지대. 박철진과 민병대원들이 만지고 둘러보고 있는 건... 현우 일행이 탔던 자동차다. 녹슨 방사능 구역 팻말.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On muddy ground near the stranded camper in the marsh, along the search route for boards or timber.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n익산 늪지대의 캠핑카 고립 지점: 차량 바퀴가 푹 빠지는 질척한 뻘과 낡은 경고판이 꽂혀 있는 지대. (특징: 차체 하부가 빠져 움직이지 못하는 진흙 구덩이; 주변의 메마른 늪지대 환경; 녹이 슬어 글자가 지워진 구형 방사능 표지판과 출입 금지 팻말)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 현우의 차가 늪지대 앞에서 멈추려다 풍덩 빠진다.\n- 부근에 출입 금지 팻말과 함께 부식돼 식별이 안 되는 방사능 구역 표시\n- 늪지대. 박철진과 민병대원들이 만지고 둘러보고 있는 건... 현우 일행이 탔던 자동차다. 녹슨 방사능 구역 팻말.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_swamp_vehicle_edge_78578e.png",
  "asset_id": "bf15ced8-25a7-46c4-b3de-7db0c008c53e",
  "input_asset_ids": [
   "e0803289-5e5b-4591-a71c-27038d182ed9"
  ],
  "origin_tag": "S56sh5",
  "place_text": "On muddy ground near the stranded camper in the marsh, along the search route for boards or timber.",
  "origin_inputs": {
   "place_text": "On muddy ground near the stranded camper in the marsh, along the search route for boards or timber.",
   "time_of_day_en": "day",
   "conti_asset_id": "e0803289-5e5b-4591-a71c-27038d182ed9"
  }
 },
 "S56sh5::bgfirst_bg": {
  "input_fingerprint": "78af173715b24daf",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 늪지대 바닥의 진흙 속에 선명하게 찍힌 거대한 신발 자국을 향해 쪼그려 앉은 앰버의 자세.\n\nLOCATION (lock): On muddy ground near the stranded camper in the marsh, along the search route for boards or timber.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Large shoeprint in mud (Clearly impressed in the swamp ground) — The complete impression is visible obliquely from above; used as Lower-right focal evidence, kept smaller than the crouching figure; Swamp ground (Mud surrounding the shoeprint); used as Continuous ground plane connects 앰버 to the impression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with restrained saturation and controlled contrast preserves the legibility of the mud impression without theatrical emphasis.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 늪지대 바닥의 진흙 속에 선명하게 찍힌 거대한 신발 자국을 향해 쪼그려 앉은 앰버의 자세.\n\nLOCATION (lock): On muddy ground near the stranded camper in the marsh, along the search route for boards or timber.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Large shoeprint in mud (Clearly impressed in the swamp ground) — The complete impression is visible obliquely from above; used as Lower-right focal evidence, kept smaller than the crouching figure; Swamp ground (Mud surrounding the shoeprint); used as Continuous ground plane connects 앰버 to the impression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with restrained saturation and controlled contrast preserves the legibility of the mud impression without theatrical emphasis.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S56sh5__bgfirst_bg.png",
  "asset_id": "ec65b360-e758-4e20-aaa5-dfe753a8386e",
  "input_asset_ids": [
   "e0803289-5e5b-4591-a71c-27038d182ed9",
   "bf15ced8-25a7-46c4-b3de-7db0c008c53e"
  ]
 },
 "S56sh5": {
  "input_fingerprint": "53cafaa5f16a121d",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 늪지대 바닥의 진흙 속에 선명하게 찍힌 거대한 신발 자국을 향해 쪼그려 앉은 앰버의 자세.\n\nLOCATION (lock): On muddy ground near the stranded camper in the marsh, along the search route for boards or timber. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Large shoeprint in mud (Clearly impressed in the swamp ground) — The complete impression is visible obliquely from above; used as Lower-right focal evidence, kept smaller than the crouching figure; Swamp ground (Mud surrounding the shoeprint); used as Continuous ground plane connects 앰버 to the impression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with restrained saturation and controlled contrast preserves the legibility of the mud impression without theatrical emphasis.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper is stuck in the swamp, and large shoeprints mark the nearby mud beside a no-entry sign and a badly corroded radiation warning sign. Charlie is already in the forest wearing the oversized straw hat, rubber boots and colorful raincoat. 앰버: She is searching for wood near the large shoeprints and remains damp and tired from the preceding journey.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 늪지대 바닥의 진흙 속에 선명하게 찍힌 거대한 신발 자국을 향해 쪼그려 앉은 앰버의 자세.\n\nLOCATION (lock): On muddy ground near the stranded camper in the marsh, along the search route for boards or timber. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Large shoeprint in mud (Clearly impressed in the swamp ground) — The complete impression is visible obliquely from above; used as Lower-right focal evidence, kept smaller than the crouching figure; Swamp ground (Mud surrounding the shoeprint); used as Continuous ground plane connects 앰버 to the impression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with restrained saturation and controlled contrast preserves the legibility of the mud impression without theatrical emphasis.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper is stuck in the swamp, and large shoeprints mark the nearby mud beside a no-entry sign and a badly corroded radiation warning sign. Charlie is already in the forest wearing the oversized straw hat, rubber boots and colorful raincoat. 앰버: She is searching for wood near the large shoeprints and remains damp and tired from the preceding journey.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 늪지대 바닥의 진흙 속에 선명하게 찍힌 거대한 신발 자국을 향해 쪼그려 앉은 앰버의 자세.\n\nLOCATION (lock): On muddy ground near the stranded camper in the marsh, along the search route for boards or timber. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Large shoeprint in mud (Clearly impressed in the swamp ground) — The complete impression is visible obliquely from above; used as Lower-right focal evidence, kept smaller than the crouching figure; Swamp ground (Mud surrounding the shoeprint); used as Continuous ground plane connects 앰버 to the impression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with restrained saturation and controlled contrast preserves the legibility of the mud impression without theatrical emphasis.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camper is stuck in the swamp, and large shoeprints mark the nearby mud beside a no-entry sign and a badly corroded radiation warning sign. Charlie is already in the forest wearing the oversized straw hat, rubber boots and colorful raincoat. 앰버: She is searching for wood near the large shoeprints and remains damp and tired from the preceding journey.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S56sh5__bgfirst_bg.png",
     "asset_id": "ec65b360-e758-4e20-aaa5-dfe753a8386e",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S56sh5.png",
     "asset_id": "e0803289-5e5b-4591-a71c-27038d182ed9",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_swamp_vehicle_edge_78578e.png",
     "asset_id": "bf15ced8-25a7-46c4-b3de-7db0c008c53e",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "B",
    "direction": "앰버는 몸을 오른쪽으로 약간 틀고 고개와 눈을 오른쪽 아래로 내려 신발 자국이 있는 진흙을 살핀다. 두 손은 무릎 부근에 있으며 다른 대상을 가리키지 않는다. 우하단 신발 자국은 화면 안쪽 왼쪽 위에서 오른쪽 아래로 길게 놓여 전체 윤곽과 밑창 무늬가 비스듬히 내려다보인다.",
    "built_space": "중앙 왼쪽의 앰버와 우하단 자국 사이에 진흙 지면이 연속되어 있다. 오른쪽 후경에는 늪에 빠진 캠핑카 한 대, 그 오른쪽에는 원형 진입금지 표지 하나와 사각 방사능 표지 하나가 있다. 왼쪽 전경의 목재 조각, 갈대, 물웅덩이와 먼 수목 지대도 장소 참조와 부합한다. 구조물이나 표지가 중복되지 않았으며, 물웅덩이의 반사에도 명백한 모순은 없다.",
    "entities": "등장인물은 금발의 어린 여자아이 한 명뿐이며, 약 10세의 체격과 둥근 얼굴이 참조의 앰버에 대체로 부합한다. 눈은 아래를 향해 세부 비교가 제한되고 혼혈 배경 자체는 외관만으로 확정할 수 없다. 참조의 남색 티셔츠 대신 회녹색 겉옷과 바지, 부츠가 보인다. 거대한 자국은 진흙에 패인 신발 밑창 흔적으로 읽히며, 인물보다 세로 길이가 작고 전체가 보인다. 낡은 캠핑카와 부식된 두 표지도 존재한다. 읽을 수 있는 글자나 추가 인물은 없다. 옷의 오염은 보이지만 젖음과 피로는 강하게 드러나지 않는다.",
    "hard_violations": [],
    "physics": "앰버는 양쪽 무릎과 발목을 굽혀 쪼그려 있고 두 부츠가 진흙에 닿아 체중을 받친다. 손은 무릎 가까이에 자연스럽게 놓여 있으며 공중에 떠 있는 신체나 소품은 없다. 캠핑카는 진흙에 잠긴 바퀴로 지지되고 표지판 기둥은 지면에 박혀 있다. 목재는 땅에 놓여 있고 자국 내부의 물은 낮은 홈에 고여 있어 물리적으로 타당하다."
   },
   {
    "label": "A",
    "direction": "앰버는 오른쪽을 향한 측면 자세로 고개를 숙이며, 시선은 앞쪽 진흙과 우하단 신발 자국 방향으로 향한다. 카메라를 보지 않는다. 한쪽 팔은 굽힌 무릎 위에 놓이고 다른 손은 아래로 내려와 있다. 신발 자국의 긴 축과 비스듬한 관찰 방향은 A와 유사하다.",
    "built_space": "앰버가 왼쪽 전경 대부분을 차지하며 신발 자국은 오른쪽 아래에 온전히 놓여 있다. 둘 사이의 진흙 지면은 이어진다. 오른쪽 후경의 캠핑카 한 대, 원형 진입금지 표지 하나, 사각 방사능 표지 하나의 배치와 주변 늪·갈대·목재는 장소 참조를 유지한다. 다만 인물이 화면 높이의 대부분을 차지하여 A보다 공간을 관찰하는 와이드 숏의 성격이 약하다. 불가능한 반사나 중복된 고정 시설은 보이지 않는다.",
    "entities": "금발의 어린 여자아이 한 명만 보이며 나이와 체격은 앰버 설정에 부합한다. 얼굴이 측면으로 숙여져 있어 참조의 큰 눈과 정면 얼굴 윤곽을 정확하게 비교하기는 어렵고, 혼혈 배경은 외관만으로 확정할 수 없다. 회녹색 재킷, 회색 후드, 바지와 진흙 묻은 장화는 참조의 남색 티셔츠와 다르다. 우하단에는 인물보다 확실히 작게 보이는 거대한 신발 자국의 전체 인상이 있다. 캠핑카와 두 경고 표지는 요구된 종류로 읽히며, 추가 인물이나 읽을 수 있는 문자는 없다. 진흙 오염은 분명하지만 피로와 젖은 상태는 제한적으로만 표현된다.",
    "hard_violations": [],
    "physics": "앞쪽 장화 밑창과 뒤쪽 장화가 지면에 닿고, 굽힌 다리 위로 몸통이 낮아져 실제로 가능한 쪼그림 자세를 이룬다. 무릎에 얹힌 팔도 자연스러운 지지를 갖는다. 몸이나 물건이 이유 없이 떠 있지 않다. 캠핑카 바퀴는 늪에 잠겨 있고 표지판은 기둥으로 지지되며, 목재와 신발 자국 안의 물도 지면과 중력에 맞게 놓여 있다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": null,
     "normalized": null,
     "ok": false
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "넓은 늪지 공간 안에 쪼그린 앰버와 우하단의 온전한 신발 자국을 배치해 와이드 숏 지시를 더 충실히 구현하지만, 의상은 인물 참조와 다르다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "신발 자국을 향한 자세와 지면 접촉은 타당하지만, 앰버를 전경에 크게 당겨 배치하여 요구된 와이드 숏보다 인물 중심의 가까운 구도로 읽힌다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버는 몸을 오른쪽으로 약간 틀고 고개와 눈을 오른쪽 아래로 내려 신발 자국이 있는 진흙을 살핀다. 두 손은 무릎 부근에 있으며 다른 대상을 가리키지 않는다. 우하단 신발 자국은 화면 안쪽 왼쪽 위에서 오른쪽 아래로 길게 놓여 전체 윤곽과 밑창 무늬가 비스듬히 내려다보인다.",
        "built_space": "중앙 왼쪽의 앰버와 우하단 자국 사이에 진흙 지면이 연속되어 있다. 오른쪽 후경에는 늪에 빠진 캠핑카 한 대, 그 오른쪽에는 원형 진입금지 표지 하나와 사각 방사능 표지 하나가 있다. 왼쪽 전경의 목재 조각, 갈대, 물웅덩이와 먼 수목 지대도 장소 참조와 부합한다. 구조물이나 표지가 중복되지 않았으며, 물웅덩이의 반사에도 명백한 모순은 없다.",
        "entities": "등장인물은 금발의 어린 여자아이 한 명뿐이며, 약 10세의 체격과 둥근 얼굴이 참조의 앰버에 대체로 부합한다. 눈은 아래를 향해 세부 비교가 제한되고 혼혈 배경 자체는 외관만으로 확정할 수 없다. 참조의 남색 티셔츠 대신 회녹색 겉옷과 바지, 부츠가 보인다. 거대한 자국은 진흙에 패인 신발 밑창 흔적으로 읽히며, 인물보다 세로 길이가 작고 전체가 보인다. 낡은 캠핑카와 부식된 두 표지도 존재한다. 읽을 수 있는 글자나 추가 인물은 없다. 옷의 오염은 보이지만 젖음과 피로는 강하게 드러나지 않는다.",
        "hard_violations": [],
        "physics": "앰버는 양쪽 무릎과 발목을 굽혀 쪼그려 있고 두 부츠가 진흙에 닿아 체중을 받친다. 손은 무릎 가까이에 자연스럽게 놓여 있으며 공중에 떠 있는 신체나 소품은 없다. 캠핑카는 진흙에 잠긴 바퀴로 지지되고 표지판 기둥은 지면에 박혀 있다. 목재는 땅에 놓여 있고 자국 내부의 물은 낮은 홈에 고여 있어 물리적으로 타당하다."
       },
       {
        "label": "B",
        "direction": "앰버는 오른쪽을 향한 측면 자세로 고개를 숙이며, 시선은 앞쪽 진흙과 우하단 신발 자국 방향으로 향한다. 카메라를 보지 않는다. 한쪽 팔은 굽힌 무릎 위에 놓이고 다른 손은 아래로 내려와 있다. 신발 자국의 긴 축과 비스듬한 관찰 방향은 A와 유사하다.",
        "built_space": "앰버가 왼쪽 전경 대부분을 차지하며 신발 자국은 오른쪽 아래에 온전히 놓여 있다. 둘 사이의 진흙 지면은 이어진다. 오른쪽 후경의 캠핑카 한 대, 원형 진입금지 표지 하나, 사각 방사능 표지 하나의 배치와 주변 늪·갈대·목재는 장소 참조를 유지한다. 다만 인물이 화면 높이의 대부분을 차지하여 A보다 공간을 관찰하는 와이드 숏의 성격이 약하다. 불가능한 반사나 중복된 고정 시설은 보이지 않는다.",
        "entities": "금발의 어린 여자아이 한 명만 보이며 나이와 체격은 앰버 설정에 부합한다. 얼굴이 측면으로 숙여져 있어 참조의 큰 눈과 정면 얼굴 윤곽을 정확하게 비교하기는 어렵고, 혼혈 배경은 외관만으로 확정할 수 없다. 회녹색 재킷, 회색 후드, 바지와 진흙 묻은 장화는 참조의 남색 티셔츠와 다르다. 우하단에는 인물보다 확실히 작게 보이는 거대한 신발 자국의 전체 인상이 있다. 캠핑카와 두 경고 표지는 요구된 종류로 읽히며, 추가 인물이나 읽을 수 있는 문자는 없다. 진흙 오염은 분명하지만 피로와 젖은 상태는 제한적으로만 표현된다.",
        "hard_violations": [],
        "physics": "앞쪽 장화 밑창과 뒤쪽 장화가 지면에 닿고, 굽힌 다리 위로 몸통이 낮아져 실제로 가능한 쪼그림 자세를 이룬다. 무릎에 얹힌 팔도 자연스러운 지지를 갖는다. 몸이나 물건이 이유 없이 떠 있지 않다. 캠핑카 바퀴는 늪에 잠겨 있고 표지판은 기둥으로 지지되며, 목재와 신발 자국 안의 물도 지면과 중력에 맞게 놓여 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "넓은 늪지 공간 안에 쪼그린 앰버와 우하단의 온전한 신발 자국을 배치해 와이드 숏 지시를 더 충실히 구현하지만, 의상은 인물 참조와 다르다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "신발 자국을 향한 자세와 지면 접촉은 타당하지만, 앰버를 전경에 크게 당겨 배치하여 요구된 와이드 숏보다 인물 중심의 가까운 구도로 읽힌다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "앰버는 몸을 오른쪽으로 약간 틀고 고개와 눈을 오른쪽 아래로 내려 신발 자국이 있는 진흙을 살핀다. 두 손은 무릎 부근에 있으며 다른 대상을 가리키지 않는다. 우하단 신발 자국은 화면 안쪽 왼쪽 위에서 오른쪽 아래로 길게 놓여 전체 윤곽과 밑창 무늬가 비스듬히 내려다보인다.",
        "built_space": "중앙 왼쪽의 앰버와 우하단 자국 사이에 진흙 지면이 연속되어 있다. 오른쪽 후경에는 늪에 빠진 캠핑카 한 대, 그 오른쪽에는 원형 진입금지 표지 하나와 사각 방사능 표지 하나가 있다. 왼쪽 전경의 목재 조각, 갈대, 물웅덩이와 먼 수목 지대도 장소 참조와 부합한다. 구조물이나 표지가 중복되지 않았으며, 물웅덩이의 반사에도 명백한 모순은 없다.",
        "entities": "등장인물은 금발의 어린 여자아이 한 명뿐이며, 약 10세의 체격과 둥근 얼굴이 참조의 앰버에 대체로 부합한다. 눈은 아래를 향해 세부 비교가 제한되고 혼혈 배경 자체는 외관만으로 확정할 수 없다. 참조의 남색 티셔츠 대신 회녹색 겉옷과 바지, 부츠가 보인다. 거대한 자국은 진흙에 패인 신발 밑창 흔적으로 읽히며, 인물보다 세로 길이가 작고 전체가 보인다. 낡은 캠핑카와 부식된 두 표지도 존재한다. 읽을 수 있는 글자나 추가 인물은 없다. 옷의 오염은 보이지만 젖음과 피로는 강하게 드러나지 않는다.",
        "hard_violations": [],
        "physics": "앰버는 양쪽 무릎과 발목을 굽혀 쪼그려 있고 두 부츠가 진흙에 닿아 체중을 받친다. 손은 무릎 가까이에 자연스럽게 놓여 있으며 공중에 떠 있는 신체나 소품은 없다. 캠핑카는 진흙에 잠긴 바퀴로 지지되고 표지판 기둥은 지면에 박혀 있다. 목재는 땅에 놓여 있고 자국 내부의 물은 낮은 홈에 고여 있어 물리적으로 타당하다."
       },
       {
        "label": "A",
        "direction": "앰버는 오른쪽을 향한 측면 자세로 고개를 숙이며, 시선은 앞쪽 진흙과 우하단 신발 자국 방향으로 향한다. 카메라를 보지 않는다. 한쪽 팔은 굽힌 무릎 위에 놓이고 다른 손은 아래로 내려와 있다. 신발 자국의 긴 축과 비스듬한 관찰 방향은 A와 유사하다.",
        "built_space": "앰버가 왼쪽 전경 대부분을 차지하며 신발 자국은 오른쪽 아래에 온전히 놓여 있다. 둘 사이의 진흙 지면은 이어진다. 오른쪽 후경의 캠핑카 한 대, 원형 진입금지 표지 하나, 사각 방사능 표지 하나의 배치와 주변 늪·갈대·목재는 장소 참조를 유지한다. 다만 인물이 화면 높이의 대부분을 차지하여 A보다 공간을 관찰하는 와이드 숏의 성격이 약하다. 불가능한 반사나 중복된 고정 시설은 보이지 않는다.",
        "entities": "금발의 어린 여자아이 한 명만 보이며 나이와 체격은 앰버 설정에 부합한다. 얼굴이 측면으로 숙여져 있어 참조의 큰 눈과 정면 얼굴 윤곽을 정확하게 비교하기는 어렵고, 혼혈 배경은 외관만으로 확정할 수 없다. 회녹색 재킷, 회색 후드, 바지와 진흙 묻은 장화는 참조의 남색 티셔츠와 다르다. 우하단에는 인물보다 확실히 작게 보이는 거대한 신발 자국의 전체 인상이 있다. 캠핑카와 두 경고 표지는 요구된 종류로 읽히며, 추가 인물이나 읽을 수 있는 문자는 없다. 진흙 오염은 분명하지만 피로와 젖은 상태는 제한적으로만 표현된다.",
        "hard_violations": [],
        "physics": "앞쪽 장화 밑창과 뒤쪽 장화가 지면에 닿고, 굽힌 다리 위로 몸통이 낮아져 실제로 가능한 쪼그림 자세를 이룬다. 무릎에 얹힌 팔도 자연스러운 지지를 갖는다. 몸이나 물건이 이유 없이 떠 있지 않다. 캠핑카 바퀴는 늪에 잠겨 있고 표지판은 기둥으로 지지되며, 목재와 신발 자국 안의 물도 지면과 중력에 맞게 놓여 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gemini-pro"
   ],
   "route": "single_reverse"
  },
  "totals": {
   "B": 8,
   "A": 7
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 8,
    "verdict_ko": "넓은 늪지 공간 안에 쪼그린 앰버와 우하단의 온전한 신발 자국을 배치해 와이드 숏 지시를 더 충실히 구현하지만, 의상은 인물 참조와 다르다."
   },
   {
    "label": "A",
    "score": 7,
    "verdict_ko": "신발 자국을 향한 자세와 지면 접촉은 타당하지만, 앰버를 전경에 크게 당겨 배치하여 요구된 와이드 숏보다 인물 중심의 가까운 구도로 읽힌다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_swamp_vehicle_edge_78578e.png",
    "asset_id": "bf15ced8-25a7-46c4-b3de-7db0c008c53e",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-5958-71ad-b81e-aa4796b36ed9",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S56sh5__bgfirst_bg.png",
   "bg_asset_id": "ec65b360-e758-4e20-aaa5-dfe753a8386e",
   "bg_record_key": "S56sh5::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "swamp_vehicle_edge",
   "groupbg_asset_id": "bf15ced8-25a7-46c4-b3de-7db0c008c53e"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S56sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:20:51.934252+00:00",
  "fingerprint": "a8def084ed2697142e43bc91d6c66863e858d85961a19a17045ecc3a318d473a",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S56sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S56sh5_sel.png",
  "source_sha256": "426ed6bcbf842d7b39186a97361f518c43d8989e9f048ab2a9371592e303e8a0",
  "file": "S56sh5_cine.png",
  "staged_sha256": "58f6debbe59b2c0a3c1f7406b5c5d7a4f55a28fcebd176835018ba96f4f24d30",
  "latency_ms": 10868
 },
 "S56sh8::signage": {
  "fp": "8cb7151b94050d87",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S56sh8": {
  "input_fingerprint": "be22d32426a0ce4a",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 허공에서 거대한 그물망이 찰리의 육중한 금속 몸통을 덮치는 찰나.\n\nLOCATION (lock): On a flower-lined forest path near the marsh, at the clearing where the robot is caught in a falling net. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Descending net (Just contacting 찰리's torso) — Its falling edge enters from above and folds across the near side of his body; used as Interrupts the open space above his reaching gesture; Forest trees (Surrounding the capture location); used as Peripheral depth and body-scale reference behind the net.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight and controlled highlights preserve the distinction between 찰리's metal body, the net, and the forest without added atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie retains the oversized straw hat, rubber boots and colorful raincoat as a capture net descends over him in the forest. The camper remains stuck in the swamp, with the large muddy shoeprints and corroded warning signs nearby.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 허공에서 거대한 그물망이 찰리의 육중한 금속 몸통을 덮치는 찰나.\n\nLOCATION (lock): On a flower-lined forest path near the marsh, at the clearing where the robot is caught in a falling net. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Descending net (Just contacting 찰리's torso) — Its falling edge enters from above and folds across the near side of his body; used as Interrupts the open space above his reaching gesture; Forest trees (Surrounding the capture location); used as Peripheral depth and body-scale reference behind the net.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight and controlled highlights preserve the distinction between 찰리's metal body, the net, and the forest without added atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie retains the oversized straw hat, rubber boots and colorful raincoat as a capture net descends over him in the forest. The camper remains stuck in the swamp, with the large muddy shoeprints and corroded warning signs nearby.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 허공에서 거대한 그물망이 찰리의 육중한 금속 몸통을 덮치는 찰나.\n\nLOCATION (lock): On a flower-lined forest path near the marsh, at the clearing where the robot is caught in a falling net. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Descending net (Just contacting 찰리's torso) — Its falling edge enters from above and folds across the near side of his body; used as Interrupts the open space above his reaching gesture; Forest trees (Surrounding the capture location); used as Peripheral depth and body-scale reference behind the net.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight and controlled highlights preserve the distinction between 찰리's metal body, the net, and the forest without added atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie retains the oversized straw hat, rubber boots and colorful raincoat as a capture net descends over him in the forest. The camper remains stuck in the swamp, with the large muddy shoeprints and corroded warning signs nearby.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S56sh8__bgfirst_bg.png",
     "asset_id": "7e75322c-4c67-4f9e-95d0-69b421b60a39",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S56sh8.png",
     "asset_id": "b891d386-6990-4b27-8c39-13f03b4127da",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_forest_trap_site_08ebbc.png",
     "asset_id": "0948ec1a-4562-45ba-965f-9a74745e6847",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "찰리의 시선은 정면을 향하고, 왼팔을 대각선 위로 뻗고 있습니다. 그물망은 왼쪽 상단에서 떨어져 몸의 우측을 덮치고 있습니다.",
    "built_space": "참조 이미지와 동일한 늪지대 배경으로, 진흙길, 뒤편의 버려진 캠핑카, 우측의 방사능 경고 표지판이 정확한 위치에 배치되어 있습니다.",
    "entities": "마스크형 얼굴을 한 로봇 찰리입니다. 밀짚모자, 화려한 색상의 우비, 고무 장화를 모두 착용하고 있습니다.",
    "hard_violations": [],
    "physics": "두 발은 진흙 바닥을 딛고 서 있으며, 위로 뻗은 팔과 자연스럽게 몸에 감기는 그물망의 물리적 상태가 타당합니다."
   },
   {
    "label": "B",
    "direction": "찰리의 시선은 정면을 향하며, 두 팔은 동작 없이 아래로 늘어뜨리고 있습니다. 그물은 머리 위 정중앙에서 아래로 떨어집니다.",
    "built_space": "진흙길, 늪지대, 뒤편의 캠핑카와 우측의 경고 표지판 등 지정된 공간 요소들이 올바르게 배치되어 있습니다.",
    "entities": "장갑판이 드러난 형태의 로봇 찰리이며 밀짚모자, 장화, 어깨에 걸쳐진 우비를 착용하고 있습니다.",
    "hard_violations": [
     "[gemini-pro] 지탱하는 프레임이 없는 그물망의 하단 테두리가 완벽한 원형의 단단한 훌라후프처럼 허공에 떠 있는 물리적 오류"
    ],
    "physics": "찰리는 바닥에 서 있으나, 떨어지는 유연한 그물망이 뼈대 없이 완벽한 원형을 유지하며 몸통 주위에 떠 있어 물리 법칙에 어긋납니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "찰리가 손을 뻗는 동작(reaching gesture)과 위에서 떨어지는 그물망의 위치 관계를 지시문대로 가장 잘 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "찰리가 손을 뻗는 핵심 동작이 완전히 누락되었고, 그물망의 형태가 물리적으로 매우 부자연스럽게 허공에 떠 있습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 시선은 정면을 향하고, 왼팔을 대각선 위로 뻗고 있습니다. 그물망은 왼쪽 상단에서 떨어져 몸의 우측을 덮치고 있습니다.",
        "built_space": "참조 이미지와 동일한 늪지대 배경으로, 진흙길, 뒤편의 버려진 캠핑카, 우측의 방사능 경고 표지판이 정확한 위치에 배치되어 있습니다.",
        "entities": "마스크형 얼굴을 한 로봇 찰리입니다. 밀짚모자, 화려한 색상의 우비, 고무 장화를 모두 착용하고 있습니다.",
        "hard_violations": [],
        "physics": "두 발은 진흙 바닥을 딛고 서 있으며, 위로 뻗은 팔과 자연스럽게 몸에 감기는 그물망의 물리적 상태가 타당합니다."
       },
       {
        "label": "B",
        "direction": "찰리의 시선은 정면을 향하며, 두 팔은 동작 없이 아래로 늘어뜨리고 있습니다. 그물은 머리 위 정중앙에서 아래로 떨어집니다.",
        "built_space": "진흙길, 늪지대, 뒤편의 캠핑카와 우측의 경고 표지판 등 지정된 공간 요소들이 올바르게 배치되어 있습니다.",
        "entities": "장갑판이 드러난 형태의 로봇 찰리이며 밀짚모자, 장화, 어깨에 걸쳐진 우비를 착용하고 있습니다.",
        "hard_violations": [
         "지탱하는 프레임이 없는 그물망의 하단 테두리가 완벽한 원형의 단단한 훌라후프처럼 허공에 떠 있는 물리적 오류"
        ],
        "physics": "찰리는 바닥에 서 있으나, 떨어지는 유연한 그물망이 뼈대 없이 완벽한 원형을 유지하며 몸통 주위에 떠 있어 물리 법칙에 어긋납니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "찰리가 손을 뻗는 동작(reaching gesture)과 위에서 떨어지는 그물망의 위치 관계를 지시문대로 가장 잘 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "찰리가 손을 뻗는 핵심 동작이 완전히 누락되었고, 그물망의 형태가 물리적으로 매우 부자연스럽게 허공에 떠 있습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 시선은 정면을 향하고, 왼팔을 대각선 위로 뻗고 있습니다. 그물망은 왼쪽 상단에서 떨어져 몸의 우측을 덮치고 있습니다.",
        "built_space": "참조 이미지와 동일한 늪지대 배경으로, 진흙길, 뒤편의 버려진 캠핑카, 우측의 방사능 경고 표지판이 정확한 위치에 배치되어 있습니다.",
        "entities": "마스크형 얼굴을 한 로봇 찰리입니다. 밀짚모자, 화려한 색상의 우비, 고무 장화를 모두 착용하고 있습니다.",
        "hard_violations": [],
        "physics": "두 발은 진흙 바닥을 딛고 서 있으며, 위로 뻗은 팔과 자연스럽게 몸에 감기는 그물망의 물리적 상태가 타당합니다."
       },
       {
        "label": "B",
        "direction": "찰리의 시선은 정면을 향하며, 두 팔은 동작 없이 아래로 늘어뜨리고 있습니다. 그물은 머리 위 정중앙에서 아래로 떨어집니다.",
        "built_space": "진흙길, 늪지대, 뒤편의 캠핑카와 우측의 경고 표지판 등 지정된 공간 요소들이 올바르게 배치되어 있습니다.",
        "entities": "장갑판이 드러난 형태의 로봇 찰리이며 밀짚모자, 장화, 어깨에 걸쳐진 우비를 착용하고 있습니다.",
        "hard_violations": [
         "지탱하는 프레임이 없는 그물망의 하단 테두리가 완벽한 원형의 단단한 훌라후프처럼 허공에 떠 있는 물리적 오류"
        ],
        "physics": "찰리는 바닥에 서 있으나, 떨어지는 유연한 그물망이 뼈대 없이 완벽한 원형을 유지하며 몸통 주위에 떠 있어 물리 법칙에 어긋납니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "찰리의 장갑판과 얼굴은 참조에 가깝지만, 전신 구도와 내려놓은 양팔 때문에 미디엄 숏 및 뻗는 동작을 끊으며 몸통을 덮치는 순간을 놓쳤습니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "그물이 뻗은 팔 아래로 가슴을 감싸는 접촉 순간은 더 충실하지만, 거의 전신인 구도와 참조보다 가늘어진 체형·달라진 얼굴이 아쉽습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 얼굴과 눈은 대체로 카메라 쪽 정면을 향하고 양손은 지면 쪽으로 내려와 있어, 위로 뻗는 동작이 없다. 그물은 화면 상단에서 어깨 높이로 내려오며 양옆으로 벌어지지만, 몸통 앞면을 접어 덮기보다는 어깨 주변을 둘러싼 상태로 보인다.",
        "built_space": "인공 건축물 내부가 아닌 꽃과 나무로 둘러싸인 진흙길이다. 오른쪽 뒤 습지에 기울어진 캠핑카 한 대, 오른쪽 나무 앞에 원형 표지 하나와 직사각형 표지 하나가 있다. 왼쪽의 굵은 나무와 흰 꽃, 웅덩이 및 먼 숲이 장소 참조와 잘 대응한다. 다만 모자부터 장화 밑창까지 모두 보여 지정된 미디엄 숏보다 넓다.",
        "entities": "등장 개체는 찰리 한 대뿐이다. 흰 각진 마스크형 얼굴, 주황색 눈, 샌드 베이지 장갑판, 육중한 상체와 긴 팔이 참조와 가깝다. 커다란 밀짚모자와 고무장화는 보이나, 알록달록한 우비는 팔을 감싸는 외투보다 열린 민소매 옷처럼 보인다. 거대한 밧줄 그물, 캠핑카, 부식된 경고 표지는 있으며 큰 신발자국은 뚜렷하게 구별되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 장화가 진흙 지면에 닿아 몸을 지탱하고, 팔은 어깨와 팔꿈치 관절에 연결되어 있다. 모자는 머리 위에 얹혀 있고 옷은 어깨에서 늘어진다. 그물은 상단 밖에서 내려와 어깨에 일부 닿으며 가장자리가 벌어진 낙하 순간으로 해석 가능하다. 지지 없는 정지 부유나 불가능한 관절은 보이지 않지만, 몸의 자세는 충돌에 반응하기보다 정적으로 보인다."
       },
       {
        "label": "B",
        "direction": "찰리는 화면 오른쪽 위로 한 손을 뻗고 얼굴도 그 손이 향하는 상방을 향한다. 반대 손은 왼쪽 아래로 벌어져 있다. 그물은 화면 왼쪽 상단에서 내려와 어깨와 가슴 앞쪽을 가로지르며 접혀 있어, 뻗는 동작 도중 몸통에 접촉한다는 관계가 A보다 분명하다.",
        "built_space": "꽃이 핀 진흙길에서 찰리가 전경에 서 있고, 오른쪽 뒤 습지에는 캠핑카 한 대가 기울어져 있다. 오른쪽 가장자리에는 원형 경고 표지 하나와 직사각형 표지 하나가 있으며, 왼쪽 큰 나무와 꽃밭 및 먼 습지 숲도 참조와 대응한다. 장화까지 거의 전신을 담고 배경을 넓게 보여 미디엄 숏 요구에는 맞지 않는다.",
        "entities": "찰리 한 대만 등장하며 밀짚모자, 소매가 있는 다색 우비, 고무장화, 베이지 금속 사지와 흰 얼굴이 보인다. 다만 얼굴은 참조의 각진 장갑형 마스크보다 작고 둥글며 입 주변 무늬도 다르다. 몸통과 팔은 참조의 고릴라형 중량감보다 가늘게 읽힌다. 포획 그물과 캠핑카, 부식된 표지는 있고, 큰 진흙 신발자국은 명확하게 식별되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "앞으로 내민 장화와 뒤쪽 장화가 지면에 닿고 무릎이 굽혀져 있어, 팔을 뻗으며 균형을 잡는 자세가 가능하다. 그물은 어깨와 가슴에 실제로 걸쳐지고 닿지 않은 부분은 위쪽으로 이어져, 낙하 중 몸에 걸리기 시작한 형태로 읽힌다. 모자와 우비도 각각 머리와 몸에 지지되어 있으며 설명되지 않는 부유는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "찰리의 장갑판과 얼굴은 참조에 가깝지만, 전신 구도와 내려놓은 양팔 때문에 미디엄 숏 및 뻗는 동작을 끊으며 몸통을 덮치는 순간을 놓쳤습니다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "그물이 뻗은 팔 아래로 가슴을 감싸는 접촉 순간은 더 충실하지만, 거의 전신인 구도와 참조보다 가늘어진 체형·달라진 얼굴이 아쉽습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 얼굴과 눈은 대체로 카메라 쪽 정면을 향하고 양손은 지면 쪽으로 내려와 있어, 위로 뻗는 동작이 없다. 그물은 화면 상단에서 어깨 높이로 내려오며 양옆으로 벌어지지만, 몸통 앞면을 접어 덮기보다는 어깨 주변을 둘러싼 상태로 보인다.",
        "built_space": "인공 건축물 내부가 아닌 꽃과 나무로 둘러싸인 진흙길이다. 오른쪽 뒤 습지에 기울어진 캠핑카 한 대, 오른쪽 나무 앞에 원형 표지 하나와 직사각형 표지 하나가 있다. 왼쪽의 굵은 나무와 흰 꽃, 웅덩이 및 먼 숲이 장소 참조와 잘 대응한다. 다만 모자부터 장화 밑창까지 모두 보여 지정된 미디엄 숏보다 넓다.",
        "entities": "등장 개체는 찰리 한 대뿐이다. 흰 각진 마스크형 얼굴, 주황색 눈, 샌드 베이지 장갑판, 육중한 상체와 긴 팔이 참조와 가깝다. 커다란 밀짚모자와 고무장화는 보이나, 알록달록한 우비는 팔을 감싸는 외투보다 열린 민소매 옷처럼 보인다. 거대한 밧줄 그물, 캠핑카, 부식된 경고 표지는 있으며 큰 신발자국은 뚜렷하게 구별되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 장화가 진흙 지면에 닿아 몸을 지탱하고, 팔은 어깨와 팔꿈치 관절에 연결되어 있다. 모자는 머리 위에 얹혀 있고 옷은 어깨에서 늘어진다. 그물은 상단 밖에서 내려와 어깨에 일부 닿으며 가장자리가 벌어진 낙하 순간으로 해석 가능하다. 지지 없는 정지 부유나 불가능한 관절은 보이지 않지만, 몸의 자세는 충돌에 반응하기보다 정적으로 보인다."
       },
       {
        "label": "A",
        "direction": "찰리는 화면 오른쪽 위로 한 손을 뻗고 얼굴도 그 손이 향하는 상방을 향한다. 반대 손은 왼쪽 아래로 벌어져 있다. 그물은 화면 왼쪽 상단에서 내려와 어깨와 가슴 앞쪽을 가로지르며 접혀 있어, 뻗는 동작 도중 몸통에 접촉한다는 관계가 A보다 분명하다.",
        "built_space": "꽃이 핀 진흙길에서 찰리가 전경에 서 있고, 오른쪽 뒤 습지에는 캠핑카 한 대가 기울어져 있다. 오른쪽 가장자리에는 원형 경고 표지 하나와 직사각형 표지 하나가 있으며, 왼쪽 큰 나무와 꽃밭 및 먼 습지 숲도 참조와 대응한다. 장화까지 거의 전신을 담고 배경을 넓게 보여 미디엄 숏 요구에는 맞지 않는다.",
        "entities": "찰리 한 대만 등장하며 밀짚모자, 소매가 있는 다색 우비, 고무장화, 베이지 금속 사지와 흰 얼굴이 보인다. 다만 얼굴은 참조의 각진 장갑형 마스크보다 작고 둥글며 입 주변 무늬도 다르다. 몸통과 팔은 참조의 고릴라형 중량감보다 가늘게 읽힌다. 포획 그물과 캠핑카, 부식된 표지는 있고, 큰 진흙 신발자국은 명확하게 식별되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "앞으로 내민 장화와 뒤쪽 장화가 지면에 닿고 무릎이 굽혀져 있어, 팔을 뻗으며 균형을 잡는 자세가 가능하다. 그물은 어깨와 가슴에 실제로 걸쳐지고 닿지 않은 부분은 위쪽으로 이어져, 낙하 중 몸에 걸리기 시작한 형태로 읽힌다. 모자와 우비도 각각 머리와 몸에 지지되어 있으며 설명되지 않는 부유는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.095
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.845
   },
   "violations": {
    "B": [
     "[gemini-pro] 지탱하는 프레임이 없는 그물망의 하단 테두리가 완벽한 원형의 단단한 훌라후프처럼 허공에 떠 있는 물리적 오류"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 845
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "찰리가 손을 뻗는 동작(reaching gesture)과 위에서 떨어지는 그물망의 위치 관계를 지시문대로 가장 잘 구현했습니다."
   },
   {
    "label": "B",
    "score": 845,
    "verdict_ko": "찰리가 손을 뻗는 핵심 동작이 완전히 누락되었고, 그물망의 형태가 물리적으로 매우 부자연스럽게 허공에 떠 있습니다.  ★위반: [gemini-pro] 지탱하는 프레임이 없는 그물망의 하단 테두리가 완벽한 원형의 단단한 훌라후프처럼 허공에 떠 있는 물리적 오류"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_forest_trap_site_08ebbc.png",
    "asset_id": "0948ec1a-4562-45ba-965f-9a74745e6847",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-5e34-7e01-8229-f5e5189035a2",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S56sh8__bgfirst_bg.png",
   "bg_asset_id": "7e75322c-4c67-4f9e-95d0-69b421b60a39",
   "bg_record_key": "S56sh8::bgfirst_bg",
   "chain_winner": true,
   "authority": "groupbg",
   "group_key": "forest_trap_site",
   "groupbg_asset_id": "0948ec1a-4562-45ba-965f-9a74745e6847"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S56sh8::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:41:32.944295+00:00",
  "fingerprint": "2c5634e9bd6206fc109446929df810ee3bd034179c5d7219fdc31480ce9b4778",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S56sh8_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S56sh8_sel.png",
  "source_sha256": "5de3a0c56f6e3d307a48ac0340929fa791bae12be9c58e390a7c48c0f9c51f1b",
  "file": "S56sh8_cine.png",
  "staged_sha256": "8379f77451839ea3d538c14de190b255060f3191097a69197aa7a5d1f933b8e8",
  "latency_ms": 12068
 },
 "S56sh14::signage": {
  "fp": "b86d6b02e078402c",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S56sh14": {
  "input_fingerprint": "c93d516ad22c64c4",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 구덩이 바닥에 엎어진 현우와 앰버를 둥글게 둘러싼 채 내려다보는 하회탈 병사들의 전경.\n\nLOCATION (lock): At an open trap pit in the forest near the marsh, with armed masked figures gathered around its rim. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Pit rim, walls, and floor (Open trap containing 현우 and 앰버) — The near rim, inner walls, and bottom are visible in one oblique overhead view; used as Establishes the vertical separation between captives and captors; Hahoe-shaped leather masks (Worn by the surrounding men in varied colors) — Near-side masks turn away and downward; far-side masks expose downward-angled fronts; used as Repeating but nonidentical details around the enclosing perimeter; Spears, axes, and bows (Held by the men around the pit) — Different oblique angles follow individual grips rather than a uniform radial pattern; used as Breaks the perimeter into distinct threatening silhouettes.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with controlled tonal separation keeps the pit floor and the encircling figures readable without inventing a separate light source below.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): An arrow remains embedded in a tree beside the approach to the pit trap, and the camper remains stranded in the swamp. Charlie has been captured in the forest while wearing the oversized hat, boots and colorful raincoat. 현우: He is at the bottom of the pit trap with his earlier treated injuries retained. The contact card remains concealed in his shoe. 앰버: She is at the bottom of the pit trap.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 구덩이 바닥에 엎어진 현우와 앰버를 둥글게 둘러싼 채 내려다보는 하회탈 병사들의 전경.\n\nLOCATION (lock): At an open trap pit in the forest near the marsh, with armed masked figures gathered around its rim. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Pit rim, walls, and floor (Open trap containing 현우 and 앰버) — The near rim, inner walls, and bottom are visible in one oblique overhead view; used as Establishes the vertical separation between captives and captors; Hahoe-shaped leather masks (Worn by the surrounding men in varied colors) — Near-side masks turn away and downward; far-side masks expose downward-angled fronts; used as Repeating but nonidentical details around the enclosing perimeter; Spears, axes, and bows (Held by the men around the pit) — Different oblique angles follow individual grips rather than a uniform radial pattern; used as Breaks the perimeter into distinct threatening silhouettes.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with controlled tonal separation keeps the pit floor and the encircling figures readable without inventing a separate light source below.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): An arrow remains embedded in a tree beside the approach to the pit trap, and the camper remains stranded in the swamp. Charlie has been captured in the forest while wearing the oversized hat, boots and colorful raincoat. 현우: He is at the bottom of the pit trap with his earlier treated injuries retained. The contact card remains concealed in his shoe. 앰버: She is at the bottom of the pit trap.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 구덩이 바닥에 엎어진 현우와 앰버를 둥글게 둘러싼 채 내려다보는 하회탈 병사들의 전경.\n\nLOCATION (lock): At an open trap pit in the forest near the marsh, with armed masked figures gathered around its rim. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Pit rim, walls, and floor (Open trap containing 현우 and 앰버) — The near rim, inner walls, and bottom are visible in one oblique overhead view; used as Establishes the vertical separation between captives and captors; Hahoe-shaped leather masks (Worn by the surrounding men in varied colors) — Near-side masks turn away and downward; far-side masks expose downward-angled fronts; used as Repeating but nonidentical details around the enclosing perimeter; Spears, axes, and bows (Held by the men around the pit) — Different oblique angles follow individual grips rather than a uniform radial pattern; used as Breaks the perimeter into distinct threatening silhouettes.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with controlled tonal separation keeps the pit floor and the encircling figures readable without inventing a separate light source below.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): An arrow remains embedded in a tree beside the approach to the pit trap, and the camper remains stranded in the swamp. Charlie has been captured in the forest while wearing the oversized hat, boots and colorful raincoat. 현우: He is at the bottom of the pit trap with his earlier treated injuries retained. The contact card remains concealed in his shoe. 앰버: She is at the bottom of the pit trap.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "병사들의 시선과 무기가 구덩이 바닥의 피사체들을 향하고 있습니다.",
    "built_space": "직사각형 구덩이 가장자리에 병사들이 서 있으나, 레퍼런스의 늪지대와 캠핑카 배경이 누락되었습니다.",
    "entities": "현우와 앰버는 엎드려 있으나, 맞은편 병사들의 맨얼굴이 노출된 채 하회탈이 머리 옆면에 융합되어 있습니다.",
    "hard_violations": [
     "[gemini-pro] 신체 및 복장 오류: 맞은편 병사들의 맨얼굴 옆면에 하회탈이 기괴하게 융합되어 있음"
    ],
    "physics": "바닥에 누운 인물들과 무기를 든 손의 물리적 지지가 정상적입니다."
   },
   {
    "label": "B",
    "direction": "병사들의 시선과 무기가 구덩이 중앙의 현우와 앰버를 향해 집중되어 있습니다.",
    "built_space": "원형 구덩이 가장자리에 병사들이 배치되었고, 배경에 레퍼런스의 늪지대와 캠핑카가 정확히 위치합니다.",
    "entities": "인물들의 외모와 하회탈 착용 상태가 프롬프트와 일치하나, 현우와 앰버가 엎드리지 않고 앉아 있습니다.",
    "hard_violations": [],
    "physics": "모든 인물이 바닥에 안정적으로 지지되어 있으며 무기들도 손에 올바르게 쥐어져 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "지정된 늪지대 배경을 정확히 구현하고 탈 착용이 올바르나, 엎드려 있어야 할 피사체들이 앉아 있는 포즈 오류가 아쉽습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "피사체의 엎드린 자세는 맞으나, 늪지대 배경 누락과 탈이 맨얼굴 옆에 융합되는 치명적인 생성 오류가 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "병사들의 시선과 무기가 구덩이 바닥의 피사체들을 향하고 있습니다.",
        "built_space": "직사각형 구덩이 가장자리에 병사들이 서 있으나, 레퍼런스의 늪지대와 캠핑카 배경이 누락되었습니다.",
        "entities": "현우와 앰버는 엎드려 있으나, 맞은편 병사들의 맨얼굴이 노출된 채 하회탈이 머리 옆면에 융합되어 있습니다.",
        "hard_violations": [
         "신체 및 복장 오류: 맞은편 병사들의 맨얼굴 옆면에 하회탈이 기괴하게 융합되어 있음"
        ],
        "physics": "바닥에 누운 인물들과 무기를 든 손의 물리적 지지가 정상적입니다."
       },
       {
        "label": "B",
        "direction": "병사들의 시선과 무기가 구덩이 중앙의 현우와 앰버를 향해 집중되어 있습니다.",
        "built_space": "원형 구덩이 가장자리에 병사들이 배치되었고, 배경에 레퍼런스의 늪지대와 캠핑카가 정확히 위치합니다.",
        "entities": "인물들의 외모와 하회탈 착용 상태가 프롬프트와 일치하나, 현우와 앰버가 엎드리지 않고 앉아 있습니다.",
        "hard_violations": [],
        "physics": "모든 인물이 바닥에 안정적으로 지지되어 있으며 무기들도 손에 올바르게 쥐어져 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "지정된 늪지대 배경을 정확히 구현하고 탈 착용이 올바르나, 엎드려 있어야 할 피사체들이 앉아 있는 포즈 오류가 아쉽습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "피사체의 엎드린 자세는 맞으나, 늪지대 배경 누락과 탈이 맨얼굴 옆에 융합되는 치명적인 생성 오류가 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "병사들의 시선과 무기가 구덩이 바닥의 피사체들을 향하고 있습니다.",
        "built_space": "직사각형 구덩이 가장자리에 병사들이 서 있으나, 레퍼런스의 늪지대와 캠핑카 배경이 누락되었습니다.",
        "entities": "현우와 앰버는 엎드려 있으나, 맞은편 병사들의 맨얼굴이 노출된 채 하회탈이 머리 옆면에 융합되어 있습니다.",
        "hard_violations": [
         "신체 및 복장 오류: 맞은편 병사들의 맨얼굴 옆면에 하회탈이 기괴하게 융합되어 있음"
        ],
        "physics": "바닥에 누운 인물들과 무기를 든 손의 물리적 지지가 정상적입니다."
       },
       {
        "label": "B",
        "direction": "병사들의 시선과 무기가 구덩이 중앙의 현우와 앰버를 향해 집중되어 있습니다.",
        "built_space": "원형 구덩이 가장자리에 병사들이 배치되었고, 배경에 레퍼런스의 늪지대와 캠핑카가 정확히 위치합니다.",
        "entities": "인물들의 외모와 하회탈 착용 상태가 프롬프트와 일치하나, 현우와 앰버가 엎드리지 않고 앉아 있습니다.",
        "hard_violations": [],
        "physics": "모든 인물이 바닥에 안정적으로 지지되어 있으며 무기들도 손에 올바르게 쥐어져 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "구덩이를 에워싼 병사들의 와이드 구도는 맞지만, 현우와 앰버가 엎어져 있지 않고 상체를 세워 병사들을 올려다보므로 핵심 순간을 놓쳤다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "바닥에 엎어진 두 사람과 테두리를 둘러싼 하회탈 병사들을 비스듬한 부감 와이드로 담아 핵심 배치와 행동을 충실히 구현했다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "병사들은 대체로 구덩이 중심을 향하고, 가까운 병사들은 머리 뒤쪽이나 가면 옆면을 보인다. 먼 쪽 가면은 정면이 보이지만 일부는 아래보다 수평을 향한다. 창끝 대부분은 구덩이 입구나 내벽 쪽을 향하며 바닥의 두 사람을 직접 겨누지는 않는다. 현우와 앰버는 고개를 들어 위쪽 병사들을 바라본다.",
        "built_space": "흙 구덩이 하나의 가까운 테두리, 안쪽 벽과 바닥이 한 부감 화면에 보인다. 병사들은 가장자리 지면에 둥글게 서 있고 두 인물은 아래 바닥에 있다. 뒤에는 숲과 습지, 기울어진 캠핑카 한 대가 보여 장소의 주요 재료와 배경이 이어진다. 구덩이 자체의 형태는 이전 사진에 드러나지 않아 직접 대조할 수 없다.",
        "entities": "현우는 헝클어진 검은 머리의 젊은 동아시아계 남성으로, 앰버는 금발의 어린 여자아이로 보이며 둘 다 참조처럼 남색 상의를 입었다. 작은 얼굴 크기 때문에 정확한 얼굴 일치와 앰버의 혼혈 정체성은 확정하기 어렵다. 현우 상의에는 손상 또는 얼룩이 있으나 치료된 부상의 지속 여부는 분명하지 않다. 병사들은 갈색·황토색·검은색 계열의 하회탈 형태 가면과 낡은 옷을 착용했다. 창과 활은 명확하지만 도끼의 구분은 약하다. 가면은 가죽보다 단단한 목재 같은 인상이 강하다. 나무에 박힌 화살과 신발 속 카드는 확인되지 않으며 읽을 수 있는 글씨는 없다.",
        "hard_violations": [],
        "physics": "현우는 엉덩이와 다리, 바닥을 짚은 손으로 상체를 지탱하고 앰버도 접힌 다리와 손으로 몸을 받친다. 지지 자체는 가능하지만 두 사람 모두 요구된 엎어진 자세가 아니다. 병사들의 보이는 발은 테두리 지면에 놓여 있고 창과 활은 손으로 잡고 있다. 전경 인물의 하체와 일부 무기 손잡이는 화면 밖으로 이어지며, 명백히 공중에 뜬 몸이나 물체는 없다."
       },
       {
        "label": "B",
        "direction": "병사들의 몸과 무기는 구덩이 안쪽을 향하고 여러 병사가 고개를 숙여 바닥의 현우와 앰버를 내려다본다. 다만 먼 쪽 일부 가면은 아래로 숙인 각도가 얕다. 가까운 병사들은 뒤통수와 가면의 측면을 보여 카메라 위치와 맞는다. 창끝은 구덩이 내부와 내벽을 향하고 도끼와 활은 각자의 손 위치에 따라 다른 각도로 놓여 있다. 두 포로의 얼굴은 바닥이나 옆쪽을 향한다.",
        "built_space": "구덩이 하나의 가까운 테두리, 양쪽 내벽과 바닥을 비스듬한 부감 와이드로 함께 보여 준다. 병사들은 지상 테두리를 원형으로 둘러싸고 두 포로는 깊이 아래 바닥 중앙에 놓여 수직 분리가 분명하다. 배경에는 숲, 습지, 일부 보이는 캠핑카 한 대와 오른쪽 나무 옆 표지판이 있어 이전 장소의 특징을 이어 간다. 별도의 바닥 광원이나 불가능한 반사는 보이지 않는다.",
        "entities": "바닥의 현우는 검은 머리와 남색 상의를 입은 젊은 남성이고, 앰버는 금발에 남색 상의를 입은 더 작은 여자아이로 보인다. 얼굴이 아래로 돌아가 정확한 얼굴 및 민족적 특징의 대조는 제한되지만 머리색과 연령별 체격 차이는 맞는다. 현우 상의의 찢김은 보이나 치료 흔적은 판별하기 어렵다. 주변 병사들은 서로 다른 황토색·갈색 계열의 하회탈 형태 가면을 쓰고 창, 도끼, 활을 들고 있다. 가면의 가죽 질감과 색상 다양성은 요구보다 약하다. 나무의 화살과 숨긴 카드는 확인되지 않고, 이전 사진의 인물이나 읽을 수 있는 글씨는 보이지 않는다.",
        "hard_violations": [],
        "physics": "현우의 몸통과 뻗은 다리, 팔이 흙바닥에 닿아 있고 앰버도 몸통과 굽힌 다리, 팔을 바닥에 대고 엎어져 있다. 두 몸 모두 바닥의 지지를 받아 낙하 후 쓰러진 상태로 성립한다. 병사들의 보이는 발은 구덩이 밖 지면을 딛고 있으며 무기들은 손으로 지지된다. 아래쪽에 잘린 도끼 자루는 화면 밖으로 이어져 손 접촉을 직접 확인할 수 없지만, 허공에 독립적으로 떠 있는 물체로 보이지는 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "구덩이를 에워싼 병사들의 와이드 구도는 맞지만, 현우와 앰버가 엎어져 있지 않고 상체를 세워 병사들을 올려다보므로 핵심 순간을 놓쳤다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "바닥에 엎어진 두 사람과 테두리를 둘러싼 하회탈 병사들을 비스듬한 부감 와이드로 담아 핵심 배치와 행동을 충실히 구현했다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "병사들은 대체로 구덩이 중심을 향하고, 가까운 병사들은 머리 뒤쪽이나 가면 옆면을 보인다. 먼 쪽 가면은 정면이 보이지만 일부는 아래보다 수평을 향한다. 창끝 대부분은 구덩이 입구나 내벽 쪽을 향하며 바닥의 두 사람을 직접 겨누지는 않는다. 현우와 앰버는 고개를 들어 위쪽 병사들을 바라본다.",
        "built_space": "흙 구덩이 하나의 가까운 테두리, 안쪽 벽과 바닥이 한 부감 화면에 보인다. 병사들은 가장자리 지면에 둥글게 서 있고 두 인물은 아래 바닥에 있다. 뒤에는 숲과 습지, 기울어진 캠핑카 한 대가 보여 장소의 주요 재료와 배경이 이어진다. 구덩이 자체의 형태는 이전 사진에 드러나지 않아 직접 대조할 수 없다.",
        "entities": "현우는 헝클어진 검은 머리의 젊은 동아시아계 남성으로, 앰버는 금발의 어린 여자아이로 보이며 둘 다 참조처럼 남색 상의를 입었다. 작은 얼굴 크기 때문에 정확한 얼굴 일치와 앰버의 혼혈 정체성은 확정하기 어렵다. 현우 상의에는 손상 또는 얼룩이 있으나 치료된 부상의 지속 여부는 분명하지 않다. 병사들은 갈색·황토색·검은색 계열의 하회탈 형태 가면과 낡은 옷을 착용했다. 창과 활은 명확하지만 도끼의 구분은 약하다. 가면은 가죽보다 단단한 목재 같은 인상이 강하다. 나무에 박힌 화살과 신발 속 카드는 확인되지 않으며 읽을 수 있는 글씨는 없다.",
        "hard_violations": [],
        "physics": "현우는 엉덩이와 다리, 바닥을 짚은 손으로 상체를 지탱하고 앰버도 접힌 다리와 손으로 몸을 받친다. 지지 자체는 가능하지만 두 사람 모두 요구된 엎어진 자세가 아니다. 병사들의 보이는 발은 테두리 지면에 놓여 있고 창과 활은 손으로 잡고 있다. 전경 인물의 하체와 일부 무기 손잡이는 화면 밖으로 이어지며, 명백히 공중에 뜬 몸이나 물체는 없다."
       },
       {
        "label": "A",
        "direction": "병사들의 몸과 무기는 구덩이 안쪽을 향하고 여러 병사가 고개를 숙여 바닥의 현우와 앰버를 내려다본다. 다만 먼 쪽 일부 가면은 아래로 숙인 각도가 얕다. 가까운 병사들은 뒤통수와 가면의 측면을 보여 카메라 위치와 맞는다. 창끝은 구덩이 내부와 내벽을 향하고 도끼와 활은 각자의 손 위치에 따라 다른 각도로 놓여 있다. 두 포로의 얼굴은 바닥이나 옆쪽을 향한다.",
        "built_space": "구덩이 하나의 가까운 테두리, 양쪽 내벽과 바닥을 비스듬한 부감 와이드로 함께 보여 준다. 병사들은 지상 테두리를 원형으로 둘러싸고 두 포로는 깊이 아래 바닥 중앙에 놓여 수직 분리가 분명하다. 배경에는 숲, 습지, 일부 보이는 캠핑카 한 대와 오른쪽 나무 옆 표지판이 있어 이전 장소의 특징을 이어 간다. 별도의 바닥 광원이나 불가능한 반사는 보이지 않는다.",
        "entities": "바닥의 현우는 검은 머리와 남색 상의를 입은 젊은 남성이고, 앰버는 금발에 남색 상의를 입은 더 작은 여자아이로 보인다. 얼굴이 아래로 돌아가 정확한 얼굴 및 민족적 특징의 대조는 제한되지만 머리색과 연령별 체격 차이는 맞는다. 현우 상의의 찢김은 보이나 치료 흔적은 판별하기 어렵다. 주변 병사들은 서로 다른 황토색·갈색 계열의 하회탈 형태 가면을 쓰고 창, 도끼, 활을 들고 있다. 가면의 가죽 질감과 색상 다양성은 요구보다 약하다. 나무의 화살과 숨긴 카드는 확인되지 않고, 이전 사진의 인물이나 읽을 수 있는 글씨는 보이지 않는다.",
        "hard_violations": [],
        "physics": "현우의 몸통과 뻗은 다리, 팔이 흙바닥에 닿아 있고 앰버도 몸통과 굽힌 다리, 팔을 바닥에 대고 엎어져 있다. 두 몸 모두 바닥의 지지를 받아 낙하 후 쓰러진 상태로 성립한다. 병사들의 보이는 발은 구덩이 밖 지면을 딛고 있으며 무기들은 손으로 지지된다. 아래쪽에 잘린 도끼 자루는 화면 밖으로 이어져 손 접촉을 직접 확인할 수 없지만, 허공에 독립적으로 떠 있는 물체로 보이지는 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.667,
    "B": 1.556
   },
   "adjusted": {
    "A": 1.417,
    "B": 1.556
   },
   "violations": {
    "A": [
     "[gemini-pro] 신체 및 복장 오류: 맞은편 병사들의 맨얼굴 옆면에 하회탈이 기괴하게 융합되어 있음"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1556,
   "A": 1417
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1556,
    "verdict_ko": "지정된 늪지대 배경을 정확히 구현하고 탈 착용이 올바르나, 엎드려 있어야 할 피사체들이 앉아 있는 포즈 오류가 아쉽습니다."
   },
   {
    "label": "A",
    "score": 1417,
    "verdict_ko": "피사체의 엎드린 자세는 맞으나, 늪지대 배경 누락과 탈이 맨얼굴 옆에 융합되는 치명적인 생성 오류가 발생했습니다.  ★위반: [gemini-pro] 신체 및 복장 오류: 맞은편 병사들의 맨얼굴 옆면에 하회탈이 기괴하게 융합되어 있음"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S56sh8_sel.png",
    "asset_id": "b0758c43-6a93-44b2-bf49-1386d7875ba2",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-630c-761d-a20c-24892335981d",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S56sh8"
  },
  "lane_policy": "ab_select_bypass:prev"
 },
 "S56sh14::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:43:26.414267+00:00",
  "fingerprint": "567fcb4445a3b7a5ec949672e1533667734c086d660e3ce8817a805f4054919b",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S56sh14_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S56sh14_sel.png",
  "source_sha256": "aab40977331b71d3c167e06408b6a18adfa7c3b7cfe769b7c4b3ab36b3178d74",
  "file": "S56sh14_cine.png",
  "staged_sha256": "60c2c2dc225e95a8dd1083974a6c945caebc3f6df536ea0dee835d1634427e85",
  "latency_ms": 10826
 },
 "S57sh3::signage": {
  "fp": "8b9583166e8fa63b",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::416cbd07cc13fbe5": {
  "subjects": [],
  "subject_text": "익산 한옥마을 거리와 공터\n낡은 한옥들이 이어진 거리와 넓게 트인 마을 공터. 기와지붕이 겹쳐 보이며, 건물 사이로 높은 물탱크 탑이 솟아 있다.",
  "identity": "canonical",
  "scope_id": "L218",
  "scope_role": "location_exterior",
  "scope_sha": "212f1a87b313475d"
 },
 "S57sh3::bgfirst_bg": {
  "input_fingerprint": "3c72f26b33c2e5bf",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 창과 몽둥이를 높이 치켜든 채 트럭을 향해 함성을 지르는 하회탈 병사들과 사람들의 광기 어린 전경.\n\nLOCATION (lock): Along the village street beside the passing prisoner truck, amid a crowd of masked guards and residents.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Cage bars (Enclosing the captives on the moving truck) — Seen obliquely from inside, with the crowd visible between them; used as A thin edge obstruction establishes the observer's confined position; Raised spears and clubs (Held aloft by the shouting crowd) — Their tips and shafts form unequal diagonals rather than matching verticals; used as Connects the lower crowd to the space beside the elevated truck; Hahoe masks (Worn throughout the visible crowd) — Shown in varied upward-facing three-quarter and profile views; used as Creates a collective threat without duplicating individual poses.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with restrained contrast separates raised weapons and masked faces while keeping the crowd's agitation grounded rather than spectacularly lit.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 창과 몽둥이를 높이 치켜든 채 트럭을 향해 함성을 지르는 하회탈 병사들과 사람들의 광기 어린 전경.\n\nLOCATION (lock): Along the village street beside the passing prisoner truck, amid a crowd of masked guards and residents.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Cage bars (Enclosing the captives on the moving truck) — Seen obliquely from inside, with the crowd visible between them; used as A thin edge obstruction establishes the observer's confined position; Raised spears and clubs (Held aloft by the shouting crowd) — Their tips and shafts form unequal diagonals rather than matching verticals; used as Connects the lower crowd to the space beside the elevated truck; Hahoe masks (Worn throughout the visible crowd) — Shown in varied upward-facing three-quarter and profile views; used as Creates a collective threat without duplicating individual poses.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with restrained contrast separates raised weapons and masked faces while keeping the crowd's agitation grounded rather than spectacularly lit.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S57sh3__bgfirst_bg.png",
  "asset_id": "24a0e20b-6878-4179-adef-b500d62cac48",
  "input_asset_ids": [
   "eea69cbc-ad3c-4b30-99fa-fe8f252efa26",
   "e716b949-d46e-4479-90bb-be052601519b"
  ]
 },
 "S57sh3": {
  "input_fingerprint": "636c8cfbb33c267f",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 창과 몽둥이를 높이 치켜든 채 트럭을 향해 함성을 지르는 하회탈 병사들과 사람들의 광기 어린 전경.\n\nLOCATION (lock): Along the village street beside the passing prisoner truck, amid a crowd of masked guards and residents. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Cage bars (Enclosing the captives on the moving truck) — Seen obliquely from inside, with the crowd visible between them; used as A thin edge obstruction establishes the observer's confined position; Raised spears and clubs (Held aloft by the shouting crowd) — Their tips and shafts form unequal diagonals rather than matching verticals; used as Connects the lower crowd to the space beside the elevated truck; Hahoe masks (Worn throughout the visible crowd) — Shown in varied upward-facing three-quarter and profile views; used as Creates a collective threat without duplicating individual poses.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with restrained contrast separates raised weapons and masked faces while keeping the crowd's agitation grounded rather than spectacularly lit.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A barred cage is mounted on the transporting truck, and a tall water-tank-like tower stands in the village. Water bottles are being distributed nearby.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 하회탈 병사들과 사람들 right now, so 하회탈 병사들과 사람들's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 하회탈 병사들과 사람들: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 창과 몽둥이를 높이 치켜든 채 트럭을 향해 함성을 지르는 하회탈 병사들과 사람들의 광기 어린 전경.\n\nLOCATION (lock): Along the village street beside the passing prisoner truck, amid a crowd of masked guards and residents. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Cage bars (Enclosing the captives on the moving truck) — Seen obliquely from inside, with the crowd visible between them; used as A thin edge obstruction establishes the observer's confined position; Raised spears and clubs (Held aloft by the shouting crowd) — Their tips and shafts form unequal diagonals rather than matching verticals; used as Connects the lower crowd to the space beside the elevated truck; Hahoe masks (Worn throughout the visible crowd) — Shown in varied upward-facing three-quarter and profile views; used as Creates a collective threat without duplicating individual poses.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with restrained contrast separates raised weapons and masked faces while keeping the crowd's agitation grounded rather than spectacularly lit.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A barred cage is mounted on the transporting truck, and a tall water-tank-like tower stands in the village. Water bottles are being distributed nearby.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 하회탈 병사들과 사람들 right now, so 하회탈 병사들과 사람들's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 하회탈 병사들과 사람들: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 창과 몽둥이를 높이 치켜든 채 트럭을 향해 함성을 지르는 하회탈 병사들과 사람들의 광기 어린 전경.\n\nLOCATION (lock): Along the village street beside the passing prisoner truck, amid a crowd of masked guards and residents. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Cage bars (Enclosing the captives on the moving truck) — Seen obliquely from inside, with the crowd visible between them; used as A thin edge obstruction establishes the observer's confined position; Raised spears and clubs (Held aloft by the shouting crowd) — Their tips and shafts form unequal diagonals rather than matching verticals; used as Connects the lower crowd to the space beside the elevated truck; Hahoe masks (Worn throughout the visible crowd) — Shown in varied upward-facing three-quarter and profile views; used as Creates a collective threat without duplicating individual poses.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with restrained contrast separates raised weapons and masked faces while keeping the crowd's agitation grounded rather than spectacularly lit.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A barred cage is mounted on the transporting truck, and a tall water-tank-like tower stands in the village. Water bottles are being distributed nearby.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 하회탈 병사들과 사람들 right now, so 하회탈 병사들과 사람들's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 하회탈 병사들과 사람들: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nNo people appear unless the shot text itself says so.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S57sh3__bgfirst_bg.png",
     "asset_id": "24a0e20b-6878-4179-adef-b500d62cac48",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S57sh3.png",
     "asset_id": "eea69cbc-ad3c-4b30-99fa-fe8f252efa26",
     "role": "conti_light"
    }
   ],
   "B": [
    {
     "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_ruined_village_square_sel.png",
     "asset_id": "e716b949-d46e-4479-90bb-be052601519b",
     "role": "location_seed_bg"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "군중의 시선과 무기가 화면 왼쪽 전경을 향하고 있으나, 대상이 되는 트럭의 실체가 명확하지 않음.",
    "built_space": "마을 배경은 전반적으로 일치하나, 교회 앞 흰색 조각상이 있어야 할 자리에 회색 옷을 입은 사람이 서 있음.",
    "entities": "하회탈을 쓴 군중과 무기는 존재하나, 트럭의 창살은 화면 좌측 가장자리에 흐릿하게만 나타남.",
    "hard_violations": [
     "[gemini-pro] 지정된 장소의 고정된 조각상을 살아있는 사람으로 대체하여 위치 구조물을 훼손함"
    ],
    "physics": "인물들은 지면에 안정적으로 서 있으며, 무기를 쥐고 있는 손의 지탱 상태는 정상적임."
   },
   {
    "label": "B",
    "direction": "군중 전체가 화면 우측의 트럭과 카메라를 향해 뚜렷하게 시선과 무기를 겨누고 있음.",
    "built_space": "교회, 예수상, 급수탑, 건물 등 참조 이미지의 고정 구조물들이 정확한 위치와 형태로 재현됨.",
    "entities": "트럭의 창살과 갇힌 사람들, 무기를 든 군중이 잘 나타나나, 다수의 인물이 하회탈을 쓰지 않고 손이나 머리 위에 들고 있음.",
    "hard_violations": [
     "[gpt-high] 지정된 철창 내부 시점 대신 철창 바깥에서 트럭 측면과 내부를 보는 카메라 위치를 사용했다."
    ],
    "physics": "땅에 서 있는 인물들의 무게 중심과 사물을 쥐고 있는 손의 형태 등 물리적 지탱이 자연스러움."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정된 위치의 조각상을 실제 사람으로 변형한 치명적 오류가 있으며, 트럭 내부에서 바라보는 구도 지시를 거의 구현하지 못했습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "트럭의 쇠창살 구도와 배경 장소를 정확히 재현했으나, 일부 군중이 지시와 다르게 탈을 얼굴에 쓰지 않고 손에 들고 있는 점이 아쉽습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "군중의 시선과 무기가 화면 왼쪽 전경을 향하고 있으나, 대상이 되는 트럭의 실체가 명확하지 않음.",
        "built_space": "마을 배경은 전반적으로 일치하나, 교회 앞 흰색 조각상이 있어야 할 자리에 회색 옷을 입은 사람이 서 있음.",
        "entities": "하회탈을 쓴 군중과 무기는 존재하나, 트럭의 창살은 화면 좌측 가장자리에 흐릿하게만 나타남.",
        "hard_violations": [
         "지정된 장소의 고정된 조각상을 살아있는 사람으로 대체하여 위치 구조물을 훼손함"
        ],
        "physics": "인물들은 지면에 안정적으로 서 있으며, 무기를 쥐고 있는 손의 지탱 상태는 정상적임."
       },
       {
        "label": "B",
        "direction": "군중 전체가 화면 우측의 트럭과 카메라를 향해 뚜렷하게 시선과 무기를 겨누고 있음.",
        "built_space": "교회, 예수상, 급수탑, 건물 등 참조 이미지의 고정 구조물들이 정확한 위치와 형태로 재현됨.",
        "entities": "트럭의 창살과 갇힌 사람들, 무기를 든 군중이 잘 나타나나, 다수의 인물이 하회탈을 쓰지 않고 손이나 머리 위에 들고 있음.",
        "hard_violations": [],
        "physics": "땅에 서 있는 인물들의 무게 중심과 사물을 쥐고 있는 손의 형태 등 물리적 지탱이 자연스러움."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지정된 위치의 조각상을 실제 사람으로 변형한 치명적 오류가 있으며, 트럭 내부에서 바라보는 구도 지시를 거의 구현하지 못했습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "트럭의 쇠창살 구도와 배경 장소를 정확히 재현했으나, 일부 군중이 지시와 다르게 탈을 얼굴에 쓰지 않고 손에 들고 있는 점이 아쉽습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "군중의 시선과 무기가 화면 왼쪽 전경을 향하고 있으나, 대상이 되는 트럭의 실체가 명확하지 않음.",
        "built_space": "마을 배경은 전반적으로 일치하나, 교회 앞 흰색 조각상이 있어야 할 자리에 회색 옷을 입은 사람이 서 있음.",
        "entities": "하회탈을 쓴 군중과 무기는 존재하나, 트럭의 창살은 화면 좌측 가장자리에 흐릿하게만 나타남.",
        "hard_violations": [
         "지정된 장소의 고정된 조각상을 살아있는 사람으로 대체하여 위치 구조물을 훼손함"
        ],
        "physics": "인물들은 지면에 안정적으로 서 있으며, 무기를 쥐고 있는 손의 지탱 상태는 정상적임."
       },
       {
        "label": "B",
        "direction": "군중 전체가 화면 우측의 트럭과 카메라를 향해 뚜렷하게 시선과 무기를 겨누고 있음.",
        "built_space": "교회, 예수상, 급수탑, 건물 등 참조 이미지의 고정 구조물들이 정확한 위치와 형태로 재현됨.",
        "entities": "트럭의 창살과 갇힌 사람들, 무기를 든 군중이 잘 나타나나, 다수의 인물이 하회탈을 쓰지 않고 손이나 머리 위에 들고 있음.",
        "hard_violations": [],
        "physics": "땅에 서 있는 인물들의 무게 중심과 사물을 쥐고 있는 손의 형태 등 물리적 지탱이 자연스러움."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "트럭을 향한 함성과 무기 동작, 마을 구조는 잘 드러나지만 카메라가 철창 밖에서 차량 측면을 보는 구도여서 지정된 철창 내부 관찰 시점을 위반한다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "철창 안에서 내려다보는 와이드 구도와 트럭 쪽으로 무기를 치켜든 군중을 충실히 구현하지만, 일부 정면형 가면과 병사·주민의 불분명한 구별은 아쉽다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽과 중앙의 군중은 오른쪽 트럭 철창과 그 안의 사람들을 향해 얼굴을 들고 함성을 지른다. 앞사람의 몽둥이는 왼쪽 위로, 여러 창은 서로 다른 각도로 위를 향한다. 무기를 높이 드는 행동과 함성의 대상인 트럭의 관계는 명확하다.",
        "built_space": "왼쪽에 파란 지붕 교회 한 채와 십자가 하나, 입구 앞 조각상 하나가 있고, 뒤에는 녹슨 원통형 물탱크 탑 하나가 보인다. 중앙 비탈길과 낡은 저층 건물들은 참조 장소에 부합한다. 그러나 차량 철창이 오른쪽 화면의 거의 절반을 차지하고 측면 거울과 외벽이 함께 보이며, 카메라는 철창 밖에서 차량 옆면과 내부를 들여다보는 위치로 읽힌다. 철창 안에서 가느다란 가장자리 장애물 너머 군중을 보는 구도가 아니다.",
        "entities": "군중은 동아시아계 외양의 성인 남성이 주를 이루며, 하회탈 형태의 갈색 가면과 낡은 작업복을 착용한다. 일부는 탈을 얼굴에 쓰지 않고 이마 위로 올리거나 손에 들어 보인다. 창과 나무 몽둥이, 철창 차량, 물탱크 탑이 식별된다. 철창 안에는 어두운 사람 형상이 보인다. 병사와 주민의 복장 차이는 뚜렷하지 않으며 물병 배급은 보이지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "지정된 철창 내부 시점 대신 철창 바깥에서 트럭 측면과 내부를 보는 카메라 위치를 사용했다."
        ],
        "physics": "앞쪽 몽둥이는 올라간 손이 손잡이를 감싸 지지하고, 창들도 손으로 자루를 붙잡고 있다. 들어 올린 가면 역시 손이 받친다. 뒤쪽 인물의 발은 길에 닿아 있으며, 앞쪽 인물들은 하체가 잘렸지만 상체의 기울기와 팔 동작에 명백한 부유나 불가능한 관절은 없다. 철창은 차량 차체에 고정되어 있다."
       },
       {
        "label": "B",
        "direction": "전경 군중은 대체로 얼굴을 위로 들어 카메라가 있는 높은 트럭 쪽에 함성을 보낸다. 일부는 옆 사람이나 트럭의 다른 지점을 보는 듯해 시선이 한 점에 복제되어 있지 않다. 몽둥이는 좌우 위쪽으로 벌어지고 창은 수직과 사선이 섞여 있어, 무기를 높이 치켜드는 동작을 보여준다.",
        "built_space": "왼쪽 가장자리의 굵은 세로 철창과 왼쪽 아래를 비스듬히 가로지르는 가로대 너머로 군중이 펼쳐져, 차량 철창 안의 높은 관찰 위치가 성립한다. 왼쪽 교회 한 채, 십자가 하나, 입구 앞 팔을 벌린 조각상 하나, 그 옆 원통형 물탱크 탑 하나가 보인다. 중앙 오르막길, 오른쪽 큰 창고와 가까운 기와지붕도 참조의 주요 배치를 따른다. 탑의 폭과 일부 건물 형태·간격은 참조와 차이가 있지만 동일 장소의 핵심 구조는 유지된다.",
        "entities": "동아시아계 외양의 성인 남녀 군중이 낡은 회색·갈색·남색 옷을 입고 있으며, 다수는 갈색 하회탈 형태의 가면을 얼굴에 착용한다. 맨얼굴인 주민들도 섞여 있다. 창날 달린 장대와 나무 몽둥이가 구별되며, 앞쪽 철창과 뒤쪽 물탱크 탑이 보인다. 병사를 특정할 복장 표지는 약하다. 트럭 차체와 포로, 물병 배급은 이 구도에서 확인되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "전경의 창과 몽둥이에는 각각 자루를 움켜쥔 손과 연결된 팔이 보인다. 군중의 들린 팔, 젖혀진 목, 기울어진 몸통은 지상에서 환호하는 동작으로 가능하다. 중·후경 인물들은 발로 길을 딛고 있고, 전경 인물들의 잘린 하체를 부유로 볼 근거는 없다. 앞쪽 철창의 세로대와 가로대는 서로 연결되어 있으며 지지 없는 물체는 확인되지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "트럭을 향한 함성과 무기 동작, 마을 구조는 잘 드러나지만 카메라가 철창 밖에서 차량 측면을 보는 구도여서 지정된 철창 내부 관찰 시점을 위반한다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "철창 안에서 내려다보는 와이드 구도와 트럭 쪽으로 무기를 치켜든 군중을 충실히 구현하지만, 일부 정면형 가면과 병사·주민의 불분명한 구별은 아쉽다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽과 중앙의 군중은 오른쪽 트럭 철창과 그 안의 사람들을 향해 얼굴을 들고 함성을 지른다. 앞사람의 몽둥이는 왼쪽 위로, 여러 창은 서로 다른 각도로 위를 향한다. 무기를 높이 드는 행동과 함성의 대상인 트럭의 관계는 명확하다.",
        "built_space": "왼쪽에 파란 지붕 교회 한 채와 십자가 하나, 입구 앞 조각상 하나가 있고, 뒤에는 녹슨 원통형 물탱크 탑 하나가 보인다. 중앙 비탈길과 낡은 저층 건물들은 참조 장소에 부합한다. 그러나 차량 철창이 오른쪽 화면의 거의 절반을 차지하고 측면 거울과 외벽이 함께 보이며, 카메라는 철창 밖에서 차량 옆면과 내부를 들여다보는 위치로 읽힌다. 철창 안에서 가느다란 가장자리 장애물 너머 군중을 보는 구도가 아니다.",
        "entities": "군중은 동아시아계 외양의 성인 남성이 주를 이루며, 하회탈 형태의 갈색 가면과 낡은 작업복을 착용한다. 일부는 탈을 얼굴에 쓰지 않고 이마 위로 올리거나 손에 들어 보인다. 창과 나무 몽둥이, 철창 차량, 물탱크 탑이 식별된다. 철창 안에는 어두운 사람 형상이 보인다. 병사와 주민의 복장 차이는 뚜렷하지 않으며 물병 배급은 보이지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "지정된 철창 내부 시점 대신 철창 바깥에서 트럭 측면과 내부를 보는 카메라 위치를 사용했다."
        ],
        "physics": "앞쪽 몽둥이는 올라간 손이 손잡이를 감싸 지지하고, 창들도 손으로 자루를 붙잡고 있다. 들어 올린 가면 역시 손이 받친다. 뒤쪽 인물의 발은 길에 닿아 있으며, 앞쪽 인물들은 하체가 잘렸지만 상체의 기울기와 팔 동작에 명백한 부유나 불가능한 관절은 없다. 철창은 차량 차체에 고정되어 있다."
       },
       {
        "label": "A",
        "direction": "전경 군중은 대체로 얼굴을 위로 들어 카메라가 있는 높은 트럭 쪽에 함성을 보낸다. 일부는 옆 사람이나 트럭의 다른 지점을 보는 듯해 시선이 한 점에 복제되어 있지 않다. 몽둥이는 좌우 위쪽으로 벌어지고 창은 수직과 사선이 섞여 있어, 무기를 높이 치켜드는 동작을 보여준다.",
        "built_space": "왼쪽 가장자리의 굵은 세로 철창과 왼쪽 아래를 비스듬히 가로지르는 가로대 너머로 군중이 펼쳐져, 차량 철창 안의 높은 관찰 위치가 성립한다. 왼쪽 교회 한 채, 십자가 하나, 입구 앞 팔을 벌린 조각상 하나, 그 옆 원통형 물탱크 탑 하나가 보인다. 중앙 오르막길, 오른쪽 큰 창고와 가까운 기와지붕도 참조의 주요 배치를 따른다. 탑의 폭과 일부 건물 형태·간격은 참조와 차이가 있지만 동일 장소의 핵심 구조는 유지된다.",
        "entities": "동아시아계 외양의 성인 남녀 군중이 낡은 회색·갈색·남색 옷을 입고 있으며, 다수는 갈색 하회탈 형태의 가면을 얼굴에 착용한다. 맨얼굴인 주민들도 섞여 있다. 창날 달린 장대와 나무 몽둥이가 구별되며, 앞쪽 철창과 뒤쪽 물탱크 탑이 보인다. 병사를 특정할 복장 표지는 약하다. 트럭 차체와 포로, 물병 배급은 이 구도에서 확인되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "전경의 창과 몽둥이에는 각각 자루를 움켜쥔 손과 연결된 팔이 보인다. 군중의 들린 팔, 젖혀진 목, 기울어진 몸통은 지상에서 환호하는 동작으로 가능하다. 중·후경 인물들은 발로 길을 딛고 있고, 전경 인물들의 잘린 하체를 부유로 볼 근거는 없다. 앞쪽 철창의 세로대와 가로대는 서로 연결되어 있으며 지지 없는 물체는 확인되지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.429,
    "B": 1.5
   },
   "adjusted": {
    "A": 1.179,
    "B": 1.25
   },
   "violations": {
    "A": [
     "[gemini-pro] 지정된 장소의 고정된 조각상을 살아있는 사람으로 대체하여 위치 구조물을 훼손함"
    ],
    "B": [
     "[gpt-high] 지정된 철창 내부 시점 대신 철창 바깥에서 트럭 측면과 내부를 보는 카메라 위치를 사용했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "A": 1179,
   "B": 1250
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1179,
    "verdict_ko": "지정된 위치의 조각상을 실제 사람으로 변형한 치명적 오류가 있으며, 트럭 내부에서 바라보는 구도 지시를 거의 구현하지 못했습니다.  ★위반: [gemini-pro] 지정된 장소의 고정된 조각상을 살아있는 사람으로 대체하여 위치 구조물을 훼손함"
   },
   {
    "label": "B",
    "score": 1250,
    "verdict_ko": "트럭의 쇠창살 구도와 배경 장소를 정확히 재현했으나, 일부 군중이 지시와 다르게 탈을 얼굴에 쓰지 않고 손에 들고 있는 점이 아쉽습니다.  ★위반: [gpt-high] 지정된 철창 내부 시점 대신 철창 바깥에서 트럭 측면과 내부를 보는 카메라 위치를 사용했다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_ruined_village_square_sel.png",
    "asset_id": "e716b949-d46e-4479-90bb-be052601519b",
    "role": "location_seed_bg"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-64ca-7e31-928d-61d4c01c0942",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S57sh3__bgfirst_bg.png",
   "bg_asset_id": "24a0e20b-6878-4179-adef-b500d62cac48",
   "bg_record_key": "S57sh3::bgfirst_bg",
   "chain_winner": false,
   "authority": "seed_bg"
  },
  "ref_mode": "seed-bg+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  },
  "lane_policy": "ab_select_ready"
 },
 "S57sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:25:38.622807+00:00",
  "fingerprint": "63dfc1dcaa1ab310f5272800f76bb0887fde779f627167ec870a093876519e97",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S57sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S57sh3_sel.png",
  "source_sha256": "e5bd12d9985e7bc18dc655de2f32c2a16c454b386f85056cdc0700e742a968d2",
  "file": "S57sh3_cine.png",
  "staged_sha256": "42b2402c2732d9106a0e95f3b77f1d8db77a83979dfb6b44a322f8038becd1dc",
  "latency_ms": 59883
 },
 "S57sh5::signage": {
  "fp": "4be797c62f935eee",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "groupbg::church_arena_forecourt": {
  "input_fingerprint": "4d1a3d8873f88228",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "church_arena_forecourt",
    "tags": [
     "S57sh5",
     "S57sh7",
     "S60sh4"
    ]
   },
   "context_sig": "075c46da6f3309d3"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At the truck's rear loading edge in front of the village church, where bound prisoners are pulled down.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n익산 마을 성당 앞 격투장과 관중석: 원형의 넓은 흙바닥 투기장과 이를 내려다보는 폭발 현장. (특징: 모래와 흙이 깔린 평평한 원형 공터; 테두리를 밝히는 횃불 조명들; 육중하고 기계 장갑이 덧대어진 개조형 전투 병기(B-200); 팔에 부착된 회전형 고사포 총구; 관중석에 피어오르는 박격포 폭발 화염과 뽀얀 흙먼지; 방진복을 입은 최신 용병들) / 익산 한옥마을 거리와 공터: 황폐화된 전통 건축물 사이로 불길이 일고 낡은 공을 차는 흙바닥 넓은 터. (특징: 부서진 기와와 낡은 목조 한옥 잔해들; 밤을 밝히는 드럼통 모닥불과 횃불; 창, 도끼, 몽둥이를 든 하회탈 무리; 흙먼지 날리는 공터 바닥과 낡은 축구공)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 성당인 듯 보이는 건물 앞에 트럭이 서자\n- 두 팔을 벌린 예수 동상에도 하회탈 가면이 씌워진.\n- 성당 앞 마을 한가운데 만들어진 격투장.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At the truck's rear loading edge in front of the village church, where bound prisoners are pulled down.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n익산 마을 성당 앞 격투장과 관중석: 원형의 넓은 흙바닥 투기장과 이를 내려다보는 폭발 현장. (특징: 모래와 흙이 깔린 평평한 원형 공터; 테두리를 밝히는 횃불 조명들; 육중하고 기계 장갑이 덧대어진 개조형 전투 병기(B-200); 팔에 부착된 회전형 고사포 총구; 관중석에 피어오르는 박격포 폭발 화염과 뽀얀 흙먼지; 방진복을 입은 최신 용병들) / 익산 한옥마을 거리와 공터: 황폐화된 전통 건축물 사이로 불길이 일고 낡은 공을 차는 흙바닥 넓은 터. (특징: 부서진 기와와 낡은 목조 한옥 잔해들; 밤을 밝히는 드럼통 모닥불과 횃불; 창, 도끼, 몽둥이를 든 하회탈 무리; 흙먼지 날리는 공터 바닥과 낡은 축구공)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 성당인 듯 보이는 건물 앞에 트럭이 서자\n- 두 팔을 벌린 예수 동상에도 하회탈 가면이 씌워진.\n- 성당 앞 마을 한가운데 만들어진 격투장.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_church_arena_forecourt_18a654.png",
  "asset_id": "1ffdd21b-ada1-4f0a-bec6-264c2a6c4350",
  "input_asset_ids": [
   "6666cfb4-3bf5-43d4-ac6f-3942d2bd5e0a"
  ],
  "origin_tag": "S57sh5",
  "place_text": "At the truck's rear loading edge in front of the village church, where bound prisoners are pulled down.",
  "origin_inputs": {
   "place_text": "At the truck's rear loading edge in front of the village church, where bound prisoners are pulled down.",
   "time_of_day_en": "day",
   "conti_asset_id": "6666cfb4-3bf5-43d4-ac6f-3942d2bd5e0a"
  }
 },
 "S57sh5::bgfirst_bg": {
  "input_fingerprint": "28fd29c710e31d3e",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 성당 건물 앞, 트럭 적재함 밖으로 현우를 거칠게 끌어당기는 하회탈 병사들의 굳은 상체, 그 힘에 의해 현우의 몸이 트럭 밖으로 막 쏠려 나온 mid-action 순간.\n\nLOCATION (lock): At the truck's rear loading edge in front of the village church, where bound prisoners are pulled down.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: truck bed edge in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Truck bed edge (Stationary during the captives' removal) — The side edge recedes diagonally from the right of frame; used as Fixed reference proving that 현우 is being pulled out rather than pushed aboard; Church-like building (Behind the stopped truck) — A partial exterior view remains behind the extraction; used as Maintains the destination context without competing with the bodies.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with consistent exterior contrast preserves the physical strain in the bodies without changing the lighting for the extraction.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 성당 건물 앞, 트럭 적재함 밖으로 현우를 거칠게 끌어당기는 하회탈 병사들의 굳은 상체, 그 힘에 의해 현우의 몸이 트럭 밖으로 막 쏠려 나온 mid-action 순간.\n\nLOCATION (lock): At the truck's rear loading edge in front of the village church, where bound prisoners are pulled down.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: truck bed edge in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Truck bed edge (Stationary during the captives' removal) — The side edge recedes diagonally from the right of frame; used as Fixed reference proving that 현우 is being pulled out rather than pushed aboard; Church-like building (Behind the stopped truck) — A partial exterior view remains behind the extraction; used as Maintains the destination context without competing with the bodies.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with consistent exterior contrast preserves the physical strain in the bodies without changing the lighting for the extraction.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S57sh5__bgfirst_bg.png",
  "asset_id": "5771d1f0-905b-4ad6-80b0-1c684d75f479",
  "input_asset_ids": [
   "6666cfb4-3bf5-43d4-ac6f-3942d2bd5e0a",
   "1ffdd21b-ada1-4f0a-bec6-264c2a6c4350"
  ]
 },
 "S57sh5": {
  "input_fingerprint": "e7b45c2357719f1b",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 성당 건물 앞, 트럭 적재함 밖으로 현우를 거칠게 끌어당기는 하회탈 병사들의 굳은 상체, 그 힘에 의해 현우의 몸이 트럭 밖으로 막 쏠려 나온 mid-action 순간.\n\nLOCATION (lock): At the truck's rear loading edge in front of the village church, where bound prisoners are pulled down. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: truck bed edge in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Truck bed edge (Stationary during the captives' removal) — The side edge recedes diagonally from the right of frame; used as Fixed reference proving that 현우 is being pulled out rather than pushed aboard; Church-like building (Behind the stopped truck) — A partial exterior view remains behind the extraction; used as Maintains the destination context without competing with the bodies.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with consistent exterior contrast preserves the physical strain in the bodies without changing the lighting for the extraction.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The truck and its barred cage are stopped outside the church-like building. A Hahoe mask covers the face of the outstretched-armed Jesus statue. 현우: He is bound while being unloaded from the truck, retaining his treated injuries. The contact card remains hidden in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 성당 건물 앞, 트럭 적재함 밖으로 현우를 거칠게 끌어당기는 하회탈 병사들의 굳은 상체, 그 힘에 의해 현우의 몸이 트럭 밖으로 막 쏠려 나온 mid-action 순간.\n\nLOCATION (lock): At the truck's rear loading edge in front of the village church, where bound prisoners are pulled down. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: truck bed edge in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Truck bed edge (Stationary during the captives' removal) — The side edge recedes diagonally from the right of frame; used as Fixed reference proving that 현우 is being pulled out rather than pushed aboard; Church-like building (Behind the stopped truck) — A partial exterior view remains behind the extraction; used as Maintains the destination context without competing with the bodies.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with consistent exterior contrast preserves the physical strain in the bodies without changing the lighting for the extraction.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The truck and its barred cage are stopped outside the church-like building. A Hahoe mask covers the face of the outstretched-armed Jesus statue. 현우: He is bound while being unloaded from the truck, retaining his treated injuries. The contact card remains hidden in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 성당 건물 앞, 트럭 적재함 밖으로 현우를 거칠게 끌어당기는 하회탈 병사들의 굳은 상체, 그 힘에 의해 현우의 몸이 트럭 밖으로 막 쏠려 나온 mid-action 순간.\n\nLOCATION (lock): At the truck's rear loading edge in front of the village church, where bound prisoners are pulled down. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: truck bed edge in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Truck bed edge (Stationary during the captives' removal) — The side edge recedes diagonally from the right of frame; used as Fixed reference proving that 현우 is being pulled out rather than pushed aboard; Church-like building (Behind the stopped truck) — A partial exterior view remains behind the extraction; used as Maintains the destination context without competing with the bodies.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight with consistent exterior contrast preserves the physical strain in the bodies without changing the lighting for the extraction.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The truck and its barred cage are stopped outside the church-like building. A Hahoe mask covers the face of the outstretched-armed Jesus statue. 현우: He is bound while being unloaded from the truck, retaining his treated injuries. The contact card remains hidden in his shoe.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S57sh5__bgfirst_bg.png",
     "asset_id": "5771d1f0-905b-4ad6-80b0-1c684d75f479",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S57sh5.png",
     "asset_id": "6666cfb4-3bf5-43d4-ac6f-3942d2bd5e0a",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_church_arena_forecourt_18a654.png",
     "asset_id": "1ffdd21b-ada1-4f0a-bec6-264c2a6c4350",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "두 병사가 현우를 향해 시선을 두고 밖으로 거칠게 당기고 있음.",
    "built_space": "배경 중앙에 교회, 우측에 트럭 적재함 배치. 예수상 얼굴에 하회탈이 씌워지지 않음.",
    "entities": "현우(레퍼런스와 일치)와 하회탈 병사 2명만 정확하게 묘사됨.",
    "hard_violations": [
     "[gemini-pro] physically impossible staging (현우의 왼쪽 다리가 트럭의 측면 철판 구조물을 물리적으로 관통하여 렌더링됨)"
    ],
    "physics": "병사들의 물리적 억압으로 상체가 지탱되고 있으나, 하반신이 금속 재질을 뚫고 나오는 형태임."
   },
   {
    "label": "B",
    "direction": "병사들이 현우를 향해 시선을 고정하고 당김.",
    "built_space": "배경의 교회와 우측 트럭 배치됨. 예수상의 하회탈 묘사 누락.",
    "entities": "현우, 하회탈 병사와 함께 프롬프트에 없는 결박된 포로 2명이 추가됨.",
    "hard_violations": [
     "[gemini-pro] invented people (화면 좌측 하단에 프롬프트가 지시하지 않은 인물 2명 임의 추가)",
     "[gemini-pro] leaked text (우측 병사의 가슴 패치에 읽을 수 있는 문자 노출)",
     "[gpt-high] 숏 텍스트에 없는 포로 두 명을 왼쪽 전경에 추가했다.",
     "[gpt-high] 오른쪽 병사의 가슴 패치에 판독 가능한 문자가 있어 문자 금지 조건을 위반한다."
    ],
    "physics": "병사들이 양팔을 잡아 현우를 공중에 띄운 상태로 지탱함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "다리가 트럭 구조물을 관통하는 오류와 예수상 하회탈 누락이 있으나, 인물 구성과 미디엄 샷 프레이밍 지시를 더 잘 따름."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "프롬프트에 명시되지 않은 추가 인물(포로들) 등장 및 병사 의상의 텍스트 누출이라는 치명적 위반으로 실패함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 병사가 현우를 향해 시선을 두고 밖으로 거칠게 당기고 있음.",
        "built_space": "배경 중앙에 교회, 우측에 트럭 적재함 배치. 예수상 얼굴에 하회탈이 씌워지지 않음.",
        "entities": "현우(레퍼런스와 일치)와 하회탈 병사 2명만 정확하게 묘사됨.",
        "hard_violations": [
         "physically impossible staging (현우의 왼쪽 다리가 트럭의 측면 철판 구조물을 물리적으로 관통하여 렌더링됨)"
        ],
        "physics": "병사들의 물리적 억압으로 상체가 지탱되고 있으나, 하반신이 금속 재질을 뚫고 나오는 형태임."
       },
       {
        "label": "B",
        "direction": "병사들이 현우를 향해 시선을 고정하고 당김.",
        "built_space": "배경의 교회와 우측 트럭 배치됨. 예수상의 하회탈 묘사 누락.",
        "entities": "현우, 하회탈 병사와 함께 프롬프트에 없는 결박된 포로 2명이 추가됨.",
        "hard_violations": [
         "invented people (화면 좌측 하단에 프롬프트가 지시하지 않은 인물 2명 임의 추가)",
         "leaked text (우측 병사의 가슴 패치에 읽을 수 있는 문자 노출)"
        ],
        "physics": "병사들이 양팔을 잡아 현우를 공중에 띄운 상태로 지탱함."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "다리가 트럭 구조물을 관통하는 오류와 예수상 하회탈 누락이 있으나, 인물 구성과 미디엄 샷 프레이밍 지시를 더 잘 따름."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "프롬프트에 명시되지 않은 추가 인물(포로들) 등장 및 병사 의상의 텍스트 누출이라는 치명적 위반으로 실패함."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "두 병사가 현우를 향해 시선을 두고 밖으로 거칠게 당기고 있음.",
        "built_space": "배경 중앙에 교회, 우측에 트럭 적재함 배치. 예수상 얼굴에 하회탈이 씌워지지 않음.",
        "entities": "현우(레퍼런스와 일치)와 하회탈 병사 2명만 정확하게 묘사됨.",
        "hard_violations": [
         "physically impossible staging (현우의 왼쪽 다리가 트럭의 측면 철판 구조물을 물리적으로 관통하여 렌더링됨)"
        ],
        "physics": "병사들의 물리적 억압으로 상체가 지탱되고 있으나, 하반신이 금속 재질을 뚫고 나오는 형태임."
       },
       {
        "label": "B",
        "direction": "병사들이 현우를 향해 시선을 고정하고 당김.",
        "built_space": "배경의 교회와 우측 트럭 배치됨. 예수상의 하회탈 묘사 누락.",
        "entities": "현우, 하회탈 병사와 함께 프롬프트에 없는 결박된 포로 2명이 추가됨.",
        "hard_violations": [
         "invented people (화면 좌측 하단에 프롬프트가 지시하지 않은 인물 2명 임의 추가)",
         "leaked text (우측 병사의 가슴 패치에 읽을 수 있는 문자 노출)"
        ],
        "physics": "병사들이 양팔을 잡아 현우를 공중에 띄운 상태로 지탱함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "명시되지 않은 포로 두 명과 가슴 패치의 판독 가능한 문자가 실격 요소이며, 전신과 전경 포로까지 담아 요구된 상체 중심 미디엄 숏에서도 벗어난다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "병사들의 팔에 붙잡힌 현우가 오른쪽 적재함에서 앞으로 쏠리는 순간과 미디엄 구도를 더 충실히 구현하지만, 의상과 예수상의 하회탈은 맞지 않는다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 머리와 몸은 오른쪽 트럭에서 왼쪽 아래 바깥으로 향한다. 지상 병사는 현우의 위팔을 잡아 바깥쪽으로 당기고, 적재함 안 병사는 반대쪽 팔을 붙잡고 있어 그쪽에서는 끌어내기보다 붙들어 버티는 힘도 읽힌다. 병사들의 얼굴은 현우 쪽을 향하며, 왼쪽 전경 포로들도 현우 쪽을 보고 있다. 조준하는 무기는 없다.",
        "built_space": "오른쪽에 철창 적재함 트럭 한 대, 왼쪽 배경에 종탑과 십자가가 있는 석조 성당 한 채, 그 앞에 양팔을 벌린 조각상 한 기와 계단이 보인다. 주변 목제 난간과 횃불도 장소 참조와 대체로 맞는다. 적재함 가장자리는 오른쪽에서 중앙으로 비스듬히 이어지지만, 성당 전면과 현우의 거의 전신 및 전경 포로까지 보여 요구된 상체 중심 미디엄 숏보다 넓다. 병사 한 명은 지상, 다른 한 명은 적재함 안에 있다.",
        "entities": "현우 외에 가면 병사 두 명과 왼쪽 전경의 포로 두 명이 보인다. 병사들은 숏 텍스트에 있지만 추가 포로들은 없다. 현우는 헝클어진 검은 머리의 젊은 동아시아계 남성으로 읽히나 참조의 남색 티셔츠 대신 밝은 속옷과 열린 청회색 셔츠를 입었다. 얼굴과 옷에 상처 및 핏자국이 있지만 치료 흔적은 뚜렷하지 않다. 팔의 밧줄은 보인다. 병사들의 흰 바탕과 붉은 볼 가면은 B의 목제 하회탈보다 형태 일치가 약하다. 예수상 얼굴에는 요구된 하회탈이 확인되지 않는다. 오른쪽 병사의 가슴 패치에는 판독 가능한 영문·숫자 형태의 문자가 보인다. 신발 속 카드는 확인할 수 없다.",
        "hard_violations": [
         "숏 텍스트에 없는 포로 두 명을 왼쪽 전경에 추가했다.",
         "오른쪽 병사의 가슴 패치에 판독 가능한 문자가 있어 문자 금지 조건을 위반한다."
        ],
        "physics": "현우의 두 발은 공중에 있지만 양쪽 병사가 각각 팔과 어깨 부위를 잡고 있고, 몸이 적재함 가장자리에서 막 내려오는 상황이므로 무지지 부유로 볼 수는 없다. 지상 병사는 벌린 다리와 땅에 닿은 발로 버티며, 다른 병사의 하체는 적재함 안에 가려져 있다. 다리의 비대칭 굽힘은 끌려 떨어지는 순간으로 가능하다. 전경 포로들은 바닥에 앉아 있다."
       },
       {
        "label": "B",
        "direction": "현우의 상체와 고개는 오른쪽 적재함에서 왼쪽 앞·아래로 쏠리고, 시선도 아래를 향한다. 왼쪽 병사는 현우의 위팔을 붙잡은 채 몸을 뒤로 기울여 바깥으로 당긴다. 적재함 안 병사는 현우의 반대쪽 팔을 잡고 내려다본다. 고정된 트럭 가장자리 뒤에 하체가 남아 있어 탑승보다 하차 방향이 명확하다. 병사의 허리 칼은 칼집에 있어 조준 대상이 없다.",
        "built_space": "오른쪽 중경에 철창 트럭 한 대와 비스듬한 적재함 가장자리가 있고, 왼쪽 뒤에는 석조 성당 한 채, 종탑과 십자가 하나, 계단 앞 조각상 한 기가 보인다. 목제 난간과 횃불이 참조 장소의 재료와 배치를 이어간다. 병사 한 명은 트럭 밖, 한 명은 적재함 안에 있으며 현우는 그 경계를 넘어 나온다. 인물의 상체가 화면을 크게 차지하고 하체는 잘려 A보다 요구된 미디엄 숏에 가깝다. 성당은 부분 배경으로 남는다.",
        "entities": "현우 한 명과 숏 텍스트에 명시된 하회탈 병사 두 명만 보인다. 현우는 앳된 동아시아계 남성의 얼굴과 헝클어진 검은 머리로 참조에 비교적 가깝다. 다만 참조의 남색 티셔츠가 아니라 오염된 밝은 회색 티셔츠를 입었다. 볼의 상처는 보이나 치료 흔적은 분명하지 않다. 팔이 뒤로 모이고 밧줄이 위팔 부근에 보여 구속 상태가 읽히지만 손목 매듭은 가려져 있다. 병사들은 목제 질감의 하회탈로 얼굴을 가렸다. 배경 예수상에는 요구된 하회탈이 보이지 않는다. 신발과 숨긴 카드는 프레임 밖이며, 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우의 양팔은 병사들의 손에 붙잡혀 있고 골반과 허벅지 부근은 적재함 끝에 걸쳐 있다. 따라서 앞으로 크게 기운 상체를 지탱하는 접촉이 보인다. 왼쪽 병사는 다리를 벌리고 몸을 뒤로 기울여 당기는 힘을 받으며, 발은 프레임 밖이다. 오른쪽 병사는 적재함 내부에서 몸을 숙여 붙잡는다. 현우의 뒤쪽 다리가 트럭 쪽에 남은 자세는 가장자리를 넘어 끌려 나오는 동작으로 가능하며, 근거 없는 공중 부유는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "명시되지 않은 포로 두 명과 가슴 패치의 판독 가능한 문자가 실격 요소이며, 전신과 전경 포로까지 담아 요구된 상체 중심 미디엄 숏에서도 벗어난다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "병사들의 팔에 붙잡힌 현우가 오른쪽 적재함에서 앞으로 쏠리는 순간과 미디엄 구도를 더 충실히 구현하지만, 의상과 예수상의 하회탈은 맞지 않는다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 머리와 몸은 오른쪽 트럭에서 왼쪽 아래 바깥으로 향한다. 지상 병사는 현우의 위팔을 잡아 바깥쪽으로 당기고, 적재함 안 병사는 반대쪽 팔을 붙잡고 있어 그쪽에서는 끌어내기보다 붙들어 버티는 힘도 읽힌다. 병사들의 얼굴은 현우 쪽을 향하며, 왼쪽 전경 포로들도 현우 쪽을 보고 있다. 조준하는 무기는 없다.",
        "built_space": "오른쪽에 철창 적재함 트럭 한 대, 왼쪽 배경에 종탑과 십자가가 있는 석조 성당 한 채, 그 앞에 양팔을 벌린 조각상 한 기와 계단이 보인다. 주변 목제 난간과 횃불도 장소 참조와 대체로 맞는다. 적재함 가장자리는 오른쪽에서 중앙으로 비스듬히 이어지지만, 성당 전면과 현우의 거의 전신 및 전경 포로까지 보여 요구된 상체 중심 미디엄 숏보다 넓다. 병사 한 명은 지상, 다른 한 명은 적재함 안에 있다.",
        "entities": "현우 외에 가면 병사 두 명과 왼쪽 전경의 포로 두 명이 보인다. 병사들은 숏 텍스트에 있지만 추가 포로들은 없다. 현우는 헝클어진 검은 머리의 젊은 동아시아계 남성으로 읽히나 참조의 남색 티셔츠 대신 밝은 속옷과 열린 청회색 셔츠를 입었다. 얼굴과 옷에 상처 및 핏자국이 있지만 치료 흔적은 뚜렷하지 않다. 팔의 밧줄은 보인다. 병사들의 흰 바탕과 붉은 볼 가면은 B의 목제 하회탈보다 형태 일치가 약하다. 예수상 얼굴에는 요구된 하회탈이 확인되지 않는다. 오른쪽 병사의 가슴 패치에는 판독 가능한 영문·숫자 형태의 문자가 보인다. 신발 속 카드는 확인할 수 없다.",
        "hard_violations": [
         "숏 텍스트에 없는 포로 두 명을 왼쪽 전경에 추가했다.",
         "오른쪽 병사의 가슴 패치에 판독 가능한 문자가 있어 문자 금지 조건을 위반한다."
        ],
        "physics": "현우의 두 발은 공중에 있지만 양쪽 병사가 각각 팔과 어깨 부위를 잡고 있고, 몸이 적재함 가장자리에서 막 내려오는 상황이므로 무지지 부유로 볼 수는 없다. 지상 병사는 벌린 다리와 땅에 닿은 발로 버티며, 다른 병사의 하체는 적재함 안에 가려져 있다. 다리의 비대칭 굽힘은 끌려 떨어지는 순간으로 가능하다. 전경 포로들은 바닥에 앉아 있다."
       },
       {
        "label": "A",
        "direction": "현우의 상체와 고개는 오른쪽 적재함에서 왼쪽 앞·아래로 쏠리고, 시선도 아래를 향한다. 왼쪽 병사는 현우의 위팔을 붙잡은 채 몸을 뒤로 기울여 바깥으로 당긴다. 적재함 안 병사는 현우의 반대쪽 팔을 잡고 내려다본다. 고정된 트럭 가장자리 뒤에 하체가 남아 있어 탑승보다 하차 방향이 명확하다. 병사의 허리 칼은 칼집에 있어 조준 대상이 없다.",
        "built_space": "오른쪽 중경에 철창 트럭 한 대와 비스듬한 적재함 가장자리가 있고, 왼쪽 뒤에는 석조 성당 한 채, 종탑과 십자가 하나, 계단 앞 조각상 한 기가 보인다. 목제 난간과 횃불이 참조 장소의 재료와 배치를 이어간다. 병사 한 명은 트럭 밖, 한 명은 적재함 안에 있으며 현우는 그 경계를 넘어 나온다. 인물의 상체가 화면을 크게 차지하고 하체는 잘려 A보다 요구된 미디엄 숏에 가깝다. 성당은 부분 배경으로 남는다.",
        "entities": "현우 한 명과 숏 텍스트에 명시된 하회탈 병사 두 명만 보인다. 현우는 앳된 동아시아계 남성의 얼굴과 헝클어진 검은 머리로 참조에 비교적 가깝다. 다만 참조의 남색 티셔츠가 아니라 오염된 밝은 회색 티셔츠를 입었다. 볼의 상처는 보이나 치료 흔적은 분명하지 않다. 팔이 뒤로 모이고 밧줄이 위팔 부근에 보여 구속 상태가 읽히지만 손목 매듭은 가려져 있다. 병사들은 목제 질감의 하회탈로 얼굴을 가렸다. 배경 예수상에는 요구된 하회탈이 보이지 않는다. 신발과 숨긴 카드는 프레임 밖이며, 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우의 양팔은 병사들의 손에 붙잡혀 있고 골반과 허벅지 부근은 적재함 끝에 걸쳐 있다. 따라서 앞으로 크게 기운 상체를 지탱하는 접촉이 보인다. 왼쪽 병사는 다리를 벌리고 몸을 뒤로 기울여 당기는 힘을 받으며, 발은 프레임 밖이다. 오른쪽 병사는 적재함 내부에서 몸을 숙여 붙잡는다. 현우의 뒤쪽 다리가 트럭 쪽에 남은 자세는 가장자리를 넘어 끌려 나오는 동작으로 가능하며, 근거 없는 공중 부유는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.0
   },
   "adjusted": {
    "A": 1.75,
    "B": 0.75
   },
   "violations": {
    "A": [
     "[gemini-pro] physically impossible staging (현우의 왼쪽 다리가 트럭의 측면 철판 구조물을 물리적으로 관통하여 렌더링됨)"
    ],
    "B": [
     "[gemini-pro] invented people (화면 좌측 하단에 프롬프트가 지시하지 않은 인물 2명 임의 추가)",
     "[gemini-pro] leaked text (우측 병사의 가슴 패치에 읽을 수 있는 문자 노출)",
     "[gpt-high] 숏 텍스트에 없는 포로 두 명을 왼쪽 전경에 추가했다.",
     "[gpt-high] 오른쪽 병사의 가슴 패치에 판독 가능한 문자가 있어 문자 금지 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 750
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "다리가 트럭 구조물을 관통하는 오류와 예수상 하회탈 누락이 있으나, 인물 구성과 미디엄 샷 프레이밍 지시를 더 잘 따름.  ★위반: [gemini-pro] physically impossible staging (현우의 왼쪽 다리가 트럭의 측면 철판 구조물을 물리적으로 관통하여 렌더링됨)"
   },
   {
    "label": "B",
    "score": 750,
    "verdict_ko": "프롬프트에 명시되지 않은 추가 인물(포로들) 등장 및 병사 의상의 텍스트 누출이라는 치명적 위반으로 실패함.  ★위반: [gemini-pro] invented people (화면 좌측 하단에 프롬프트가 지시하지 않은 인물 2명 임의 추가) / [gemini-pro] leaked text (우측 병사의 가슴 패치에 읽을 수 있는 문자 노출) / [gpt-high] 숏 텍스트에 없는 포로 두 명을 왼쪽 전경에 추가했다. / [gpt-high] 오른쪽 병사의 가슴 패치에 판독 가능한 문자가 있어 문자 금지 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_church_arena_forecourt_18a654.png",
    "asset_id": "1ffdd21b-ada1-4f0a-bec6-264c2a6c4350",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-6812-75cc-8531-888c33f0c224",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S57sh5__bgfirst_bg.png",
   "bg_asset_id": "5771d1f0-905b-4ad6-80b0-1c684d75f479",
   "bg_record_key": "S57sh5::bgfirst_bg",
   "chain_winner": true,
   "authority": "groupbg",
   "group_key": "church_arena_forecourt",
   "groupbg_asset_id": "1ffdd21b-ada1-4f0a-bec6-264c2a6c4350"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S57sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:28:39.496864+00:00",
  "fingerprint": "c8b83ba28d38301b90dca8a18cc5b9508f16016c6197143625048d60bc3acb38",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S57sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S57sh5_sel.png",
  "source_sha256": "8b07788ce5caa1aa255ab63c1108ef6df2b5f7ed6f9234716eb860d8bbdc956f",
  "file": "S57sh5_cine.png",
  "staged_sha256": "0404874c95cd68a25bcb68b2ae3c4aa7c7d6e473f0b4cd350b9b6f6d018e255d",
  "latency_ms": 80591
 },
 "S57sh7::signage": {
  "fp": "d62bf2bea1640978",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S57sh7": {
  "input_fingerprint": "04b0a3b625d366f3",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 양팔을 벌린 거대한 예수 동상의 얼굴에 기괴한 가면이 씌워져 있는 섬뜩한 광경.\n\nLOCATION (lock): Outside the village church, at the large outstretched-arm religious statue fitted with a mask. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Jesus statue (Arms outstretched, with a Hahoe mask covering its face) — Seen from below at a lateral three-quarter angle, with both arms legible; used as Primary architectural subject at the end of the tilt; Hahoe mask on the statue (Covering the statue's face) — Its face is visible obliquely above the camera rather than symmetrically head-on; used as Final point of attention within the wider statue composition; Church-like exterior (Surrounding the statue near the entrance) — Exterior portions remain visible around the upward view; used as Provides architectural scale and negative space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight and restrained tonal contrast let the mask's incongruity carry the unease without introducing an artificial glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The Jesus statue has both arms extended and a Hahoe mask over its face. The transport truck and barred cage remain outside the church-like building.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 양팔을 벌린 거대한 예수 동상의 얼굴에 기괴한 가면이 씌워져 있는 섬뜩한 광경.\n\nLOCATION (lock): Outside the village church, at the large outstretched-arm religious statue fitted with a mask. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Jesus statue (Arms outstretched, with a Hahoe mask covering its face) — Seen from below at a lateral three-quarter angle, with both arms legible; used as Primary architectural subject at the end of the tilt; Hahoe mask on the statue (Covering the statue's face) — Its face is visible obliquely above the camera rather than symmetrically head-on; used as Final point of attention within the wider statue composition; Church-like exterior (Surrounding the statue near the entrance) — Exterior portions remain visible around the upward view; used as Provides architectural scale and negative space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight and restrained tonal contrast let the mask's incongruity carry the unease without introducing an artificial glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The Jesus statue has both arms extended and a Hahoe mask over its face. The transport truck and barred cage remain outside the church-like building.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 양팔을 벌린 거대한 예수 동상의 얼굴에 기괴한 가면이 씌워져 있는 섬뜩한 광경.\n\nLOCATION (lock): Outside the village church, at the large outstretched-arm religious statue fitted with a mask. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Jesus statue (Arms outstretched, with a Hahoe mask covering its face) — Seen from below at a lateral three-quarter angle, with both arms legible; used as Primary architectural subject at the end of the tilt; Hahoe mask on the statue (Covering the statue's face) — Its face is visible obliquely above the camera rather than symmetrically head-on; used as Final point of attention within the wider statue composition; Church-like exterior (Surrounding the statue near the entrance) — Exterior portions remain visible around the upward view; used as Provides architectural scale and negative space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight and restrained tonal contrast let the mask's incongruity carry the unease without introducing an artificial glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The Jesus statue has both arms extended and a Hahoe mask over its face. The transport truck and barred cage remain outside the church-like building.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라는 동상을 아래에서 위로 비스듬히 올려다봄.",
    "built_space": "동상 뒤로 석조 건물이 있으나 레퍼런스의 첨탑 구조와 다름. 왼쪽에 트럭과 계단이 배치됨.",
    "entities": "팔을 벌린 예수 동상과 하회탈이 있음. 사람은 배제됨. 트럭은 있으나 철창 구조가 없음.",
    "hard_violations": [],
    "physics": "동상은 서 있으며 가면은 얼굴에 씌워져 있음."
   },
   {
    "label": "B",
    "direction": "카메라는 동상을 아래에서 위로 올려다봄.",
    "built_space": "왼쪽에 원본 성당이 있고, 오른쪽에 완전히 새로운 거대한 건물이 추가되어 공간이 왜곡됨.",
    "entities": "팔을 벌린 동상과 하회탈이 있음. 사람은 배제되었으나 트럭이 누락됨.",
    "hard_violations": [
     "[gemini-pro] 발명된 구조물: 오른쪽에 레퍼런스에 없는 거대한 건물이 추가됨",
     "[gpt-high] 참조의 교회를 왼쪽에 남겨 둔 채 오른쪽에 별도의 대형 교회 정면과 출입구를 추가하여, 고정된 장소의 건축 구조를 중복·증설했다."
    ],
    "physics": "동상은 서 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "가면 쓴 동상의 앵글과 인물 배제 지시를 잘 따랐으나, 배경 성당 구조가 변형되고 트럭의 철창이 누락되었습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "배경에 존재하지 않는 거대한 건물을 임의로 생성하여 장소의 일관성을 심각하게 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 동상을 아래에서 위로 비스듬히 올려다봄.",
        "built_space": "동상 뒤로 석조 건물이 있으나 레퍼런스의 첨탑 구조와 다름. 왼쪽에 트럭과 계단이 배치됨.",
        "entities": "팔을 벌린 예수 동상과 하회탈이 있음. 사람은 배제됨. 트럭은 있으나 철창 구조가 없음.",
        "hard_violations": [],
        "physics": "동상은 서 있으며 가면은 얼굴에 씌워져 있음."
       },
       {
        "label": "B",
        "direction": "카메라는 동상을 아래에서 위로 올려다봄.",
        "built_space": "왼쪽에 원본 성당이 있고, 오른쪽에 완전히 새로운 거대한 건물이 추가되어 공간이 왜곡됨.",
        "entities": "팔을 벌린 동상과 하회탈이 있음. 사람은 배제되었으나 트럭이 누락됨.",
        "hard_violations": [
         "발명된 구조물: 오른쪽에 레퍼런스에 없는 거대한 건물이 추가됨"
        ],
        "physics": "동상은 서 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "가면 쓴 동상의 앵글과 인물 배제 지시를 잘 따랐으나, 배경 성당 구조가 변형되고 트럭의 철창이 누락되었습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "배경에 존재하지 않는 거대한 건물을 임의로 생성하여 장소의 일관성을 심각하게 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 동상을 아래에서 위로 비스듬히 올려다봄.",
        "built_space": "동상 뒤로 석조 건물이 있으나 레퍼런스의 첨탑 구조와 다름. 왼쪽에 트럭과 계단이 배치됨.",
        "entities": "팔을 벌린 예수 동상과 하회탈이 있음. 사람은 배제됨. 트럭은 있으나 철창 구조가 없음.",
        "hard_violations": [],
        "physics": "동상은 서 있으며 가면은 얼굴에 씌워져 있음."
       },
       {
        "label": "B",
        "direction": "카메라는 동상을 아래에서 위로 올려다봄.",
        "built_space": "왼쪽에 원본 성당이 있고, 오른쪽에 완전히 새로운 거대한 건물이 추가되어 공간이 왜곡됨.",
        "entities": "팔을 벌린 동상과 하회탈이 있음. 사람은 배제되었으나 트럭이 누락됨.",
        "hard_violations": [
         "발명된 구조물: 오른쪽에 레퍼런스에 없는 거대한 건물이 추가됨"
        ],
        "physics": "동상은 서 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "가면과 양팔은 선명하지만, 기존 교회 옆에 별도의 대형 교회 정면을 추가하여 장소의 고정 구조를 바꿨고 얼굴도 요구보다 정면에 가깝다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "하나의 석조 교회를 배경으로 양팔과 비스듬한 하회탈을 함께 보여 장소와 핵심 구도를 더 충실히 구현했지만, 올려다보는 각도와 트럭의 철창 재현은 약하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "동상은 양팔을 화면 좌우로 벌리고 손바닥을 앞쪽으로 내민다. 가면은 화면 왼쪽으로 조금 돌아 있으나 두 눈과 양 볼이 거의 정면으로 보인다. 얼굴이 카메라 위에서 비스듬히 내려다보이는 측면 사선 구도보다는 정면에 가까운 모습이다.",
        "built_space": "왼쪽에는 종탑과 십자가 각 하나, 상부 첨두창 하나와 출입문이 있는 기존 형태의 석조 교회가 보인다. 오른쪽에는 별도의 큰 박공 정면, 상부 직사각형 창 하나, 대형 아치 출입구 하나가 추가되어 있다. 두 건물 앞에 각각 계단이 있으며, 동상은 그 사이 전경에 놓였다. 참조에서 동상 뒤를 이루던 단일 교회 전면과 달리 대형 종교 건물 정면이 두 개로 분리된다.",
        "entities": "긴 머리와 긴 옷을 지닌 석조 예수상 하나, 얼굴을 덮은 갈색 하회탈 하나가 보인다. 탈의 과장된 눈썹과 코, 웃는 입 및 목재 표면은 요구에 맞는다. 살아 있는 사람과 읽을 수 있는 문자는 없다. 운송 트럭과 철창은 화면에 보이지 않아 유지 여부를 확인할 수 없다. 낮의 자연광과 낡은 석재는 참조와 대체로 맞는다.",
        "hard_violations": [
         "참조의 교회를 왼쪽에 남겨 둔 채 오른쪽에 별도의 대형 교회 정면과 출입구를 추가하여, 고정된 장소의 건축 구조를 중복·증설했다."
        ],
        "physics": "가면은 머리 옆으로 이어진 끈으로 고정되어 있으며 공중에 떠 있지 않다. 양팔과 손은 동상의 어깨 및 소매와 연결된다. 동상 하부와 받침은 화면 밖이므로 지지 접점은 확인할 수 없지만, 몸통이 화면 아래로 이어져 떠 있는 형상은 아니다. 움직이는 신체나 물체는 없다."
       },
       {
        "label": "B",
        "direction": "동상은 양팔을 좌우로 벌리고 손바닥을 앞쪽으로 향한다. 가면과 머리는 화면 오른쪽으로 돌아 있어 코와 볼의 돌출이 비스듬히 드러난다. 얼굴은 카메라보다 위에 놓여 요구된 사선 관찰에 더 가깝지만, 아래에서 올려다보는 효과는 강하지 않다.",
        "built_space": "동상 뒤에 하나의 석조 교회가 있고 박공지붕, 상부 첨두창 하나, 하부 오른쪽 첨두창 하나와 목재 출입문 하나가 보인다. 측벽과 여러 부벽이 함께 드러나 측면 사선 시점을 만든다. 왼쪽에는 계단과 낮은 마당, 트럭 한 대가 있다. 참조의 거친 석재와 첨두 개구부를 유지하며 별도의 교회 정면을 추가하지 않는다. 종탑 상부와 동상 받침은 화면 밖이다.",
        "entities": "긴 머리와 긴 옷을 지닌 석조 예수상 하나가 양팔을 벌리고, 갈색 하회탈 하나가 얼굴을 덮는다. 탈의 웃는 입과 굽은 눈썹, 나무의 균열 및 광택이 입체적이다. 살아 있는 사람이나 읽을 수 있는 문자는 없다. 왼쪽의 소형 군용색 트럭은 확인되지만, 참조의 높고 밀폐된 철창 적재함 대신 낮은 난간형 적재함으로 보여 운반 장치의 일치는 부족하다. 낮의 자연광과 석재의 마모는 유지된다.",
        "hard_violations": [],
        "physics": "가면은 머리를 두르는 띠와 옆 끈으로 고정된다. 양팔과 손은 동상 몸체에 정상적으로 연결되어 있다. 동상 하부가 화면 밖으로 이어지므로 받침 접점은 보이지 않지만, 부유를 나타내는 간격은 없다. 트럭은 바퀴로 지면에 서 있고, 주변 건축물과 계단도 지면에 연결된다. 별도의 비행이나 운동 동작은 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "가면과 양팔은 선명하지만, 기존 교회 옆에 별도의 대형 교회 정면을 추가하여 장소의 고정 구조를 바꿨고 얼굴도 요구보다 정면에 가깝다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "하나의 석조 교회를 배경으로 양팔과 비스듬한 하회탈을 함께 보여 장소와 핵심 구도를 더 충실히 구현했지만, 올려다보는 각도와 트럭의 철창 재현은 약하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "동상은 양팔을 화면 좌우로 벌리고 손바닥을 앞쪽으로 내민다. 가면은 화면 왼쪽으로 조금 돌아 있으나 두 눈과 양 볼이 거의 정면으로 보인다. 얼굴이 카메라 위에서 비스듬히 내려다보이는 측면 사선 구도보다는 정면에 가까운 모습이다.",
        "built_space": "왼쪽에는 종탑과 십자가 각 하나, 상부 첨두창 하나와 출입문이 있는 기존 형태의 석조 교회가 보인다. 오른쪽에는 별도의 큰 박공 정면, 상부 직사각형 창 하나, 대형 아치 출입구 하나가 추가되어 있다. 두 건물 앞에 각각 계단이 있으며, 동상은 그 사이 전경에 놓였다. 참조에서 동상 뒤를 이루던 단일 교회 전면과 달리 대형 종교 건물 정면이 두 개로 분리된다.",
        "entities": "긴 머리와 긴 옷을 지닌 석조 예수상 하나, 얼굴을 덮은 갈색 하회탈 하나가 보인다. 탈의 과장된 눈썹과 코, 웃는 입 및 목재 표면은 요구에 맞는다. 살아 있는 사람과 읽을 수 있는 문자는 없다. 운송 트럭과 철창은 화면에 보이지 않아 유지 여부를 확인할 수 없다. 낮의 자연광과 낡은 석재는 참조와 대체로 맞는다.",
        "hard_violations": [
         "참조의 교회를 왼쪽에 남겨 둔 채 오른쪽에 별도의 대형 교회 정면과 출입구를 추가하여, 고정된 장소의 건축 구조를 중복·증설했다."
        ],
        "physics": "가면은 머리 옆으로 이어진 끈으로 고정되어 있으며 공중에 떠 있지 않다. 양팔과 손은 동상의 어깨 및 소매와 연결된다. 동상 하부와 받침은 화면 밖이므로 지지 접점은 확인할 수 없지만, 몸통이 화면 아래로 이어져 떠 있는 형상은 아니다. 움직이는 신체나 물체는 없다."
       },
       {
        "label": "A",
        "direction": "동상은 양팔을 좌우로 벌리고 손바닥을 앞쪽으로 향한다. 가면과 머리는 화면 오른쪽으로 돌아 있어 코와 볼의 돌출이 비스듬히 드러난다. 얼굴은 카메라보다 위에 놓여 요구된 사선 관찰에 더 가깝지만, 아래에서 올려다보는 효과는 강하지 않다.",
        "built_space": "동상 뒤에 하나의 석조 교회가 있고 박공지붕, 상부 첨두창 하나, 하부 오른쪽 첨두창 하나와 목재 출입문 하나가 보인다. 측벽과 여러 부벽이 함께 드러나 측면 사선 시점을 만든다. 왼쪽에는 계단과 낮은 마당, 트럭 한 대가 있다. 참조의 거친 석재와 첨두 개구부를 유지하며 별도의 교회 정면을 추가하지 않는다. 종탑 상부와 동상 받침은 화면 밖이다.",
        "entities": "긴 머리와 긴 옷을 지닌 석조 예수상 하나가 양팔을 벌리고, 갈색 하회탈 하나가 얼굴을 덮는다. 탈의 웃는 입과 굽은 눈썹, 나무의 균열 및 광택이 입체적이다. 살아 있는 사람이나 읽을 수 있는 문자는 없다. 왼쪽의 소형 군용색 트럭은 확인되지만, 참조의 높고 밀폐된 철창 적재함 대신 낮은 난간형 적재함으로 보여 운반 장치의 일치는 부족하다. 낮의 자연광과 석재의 마모는 유지된다.",
        "hard_violations": [],
        "physics": "가면은 머리를 두르는 띠와 옆 끈으로 고정된다. 양팔과 손은 동상 몸체에 정상적으로 연결되어 있다. 동상 하부가 화면 밖으로 이어지므로 받침 접점은 보이지 않지만, 부유를 나타내는 간격은 없다. 트럭은 바퀴로 지면에 서 있고, 주변 건축물과 계단도 지면에 연결된다. 별도의 비행이나 운동 동작은 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.042
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.792
   },
   "violations": {
    "B": [
     "[gemini-pro] 발명된 구조물: 오른쪽에 레퍼런스에 없는 거대한 건물이 추가됨",
     "[gpt-high] 참조의 교회를 왼쪽에 남겨 둔 채 오른쪽에 별도의 대형 교회 정면과 출입구를 추가하여, 고정된 장소의 건축 구조를 중복·증설했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 792
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "가면 쓴 동상의 앵글과 인물 배제 지시를 잘 따랐으나, 배경 성당 구조가 변형되고 트럭의 철창이 누락되었습니다."
   },
   {
    "label": "B",
    "score": 792,
    "verdict_ko": "배경에 존재하지 않는 거대한 건물을 임의로 생성하여 장소의 일관성을 심각하게 위반했습니다.  ★위반: [gemini-pro] 발명된 구조물: 오른쪽에 레퍼런스에 없는 거대한 건물이 추가됨 / [gpt-high] 참조의 교회를 왼쪽에 남겨 둔 채 오른쪽에 별도의 대형 교회 정면과 출입구를 추가하여, 고정된 장소의 건축 구조를 중복·증설했다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S57sh5_sel.png",
    "asset_id": "34b9c624-a74d-45bd-a5e1-396e37ff7018",
    "role": "prev_still"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-6ccd-7349-88da-e3fcc25758bb",
  "ref_mode": "prev만 (배경 전용·공유 계획)",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S57sh5"
  },
  "lane_policy": "ab_select_bypass:bg_only:share_plan_prev_bgonly"
 },
 "S57sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:30:53.550162+00:00",
  "fingerprint": "b92f6760480605bf84fd65f37a8ffeda26286b52cfdf89417cfa9c6a149b006a",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S57sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S57sh7_sel.png",
  "source_sha256": "ab7aaaa02b03088ee13ba1a6c1e19949e472995a6ebab7f82d425024e744a53b",
  "file": "S57sh7_cine.png",
  "staged_sha256": "f21de541ff1574366f2973762e2d6de50a975b2d5cc59fec3f3b0410144066c1",
  "latency_ms": 10400
 },
 "S58sh5::signage": {
  "fp": "7e52833d11ea9ba5",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::ae7a6cc9c201b720": {
  "subjects": [],
  "subject_text": "익산 마을 성당 예배당과 단상\n훼손된 고딕 양식 예배당. 높은 기둥과 2층 발코니, 앞쪽 단상이 있으며 스테인드글라스를 통과한 빛이 내부로 스며든다.",
  "identity": "canonical",
  "scope_id": "L220",
  "scope_role": "location_interior",
  "scope_sha": "73535dd8e86b426d"
 },
 "S58sh5::bgfirst_bg": {
  "input_fingerprint": "cadd1dfc5e244007",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 육중한 황금빛 하회탈을 쓴 거대한 백산이 화려한 의복을 입은 채 단상을 향해 한쪽 발을 공중에 든 mid-action 자세의 위압적인 전신.\n\nLOCATION (lock): On the approach to the raised platform inside the ruined village church, lit by daylight through stained glass.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Platform (Awaiting 백산's arrival) — Its side edge is visible along the rightward route; used as Gives his suspended step a destination without frontal staging; Ruined church interior (Damaged) — Seen laterally behind the approach and bowed crowd; used as Provides scale around 백산's full-body silhouette; Golden Hahoe mask (Worn by 백산) — Visible in three-quarter view, aligned with his route toward the platform; used as Concentrates authority within the full-body composition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Light entering through the stained glass gives the ruined interior a restrained sacred quality, with selective richness in 백산's golden mask and royal clothing.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 육중한 황금빛 하회탈을 쓴 거대한 백산이 화려한 의복을 입은 채 단상을 향해 한쪽 발을 공중에 든 mid-action 자세의 위압적인 전신.\n\nLOCATION (lock): On the approach to the raised platform inside the ruined village church, lit by daylight through stained glass.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Platform (Awaiting 백산's arrival) — Its side edge is visible along the rightward route; used as Gives his suspended step a destination without frontal staging; Ruined church interior (Damaged) — Seen laterally behind the approach and bowed crowd; used as Provides scale around 백산's full-body silhouette; Golden Hahoe mask (Worn by 백산) — Visible in three-quarter view, aligned with his route toward the platform; used as Concentrates authority within the full-body composition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Light entering through the stained glass gives the ruined interior a restrained sacred quality, with selective richness in 백산's golden mask and royal clothing.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S58sh5__bgfirst_bg.png",
  "asset_id": "31941b00-404b-4a16-90fe-5d0f9ded9490",
  "input_asset_ids": [
   "6d80eb92-6d35-4bd5-bfec-92b0a707c02a",
   "b25ec118-d5a5-4461-bcc9-1623494e1fcd"
  ]
 },
 "S58sh5": {
  "input_fingerprint": "5d1767def29db37d",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 육중한 황금빛 하회탈을 쓴 거대한 백산이 화려한 의복을 입은 채 단상을 향해 한쪽 발을 공중에 든 mid-action 자세의 위압적인 전신.\n\nLOCATION (lock): On the approach to the raised platform inside the ruined village church, lit by daylight through stained glass. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Platform (Awaiting 백산's arrival) — Its side edge is visible along the rightward route; used as Gives his suspended step a destination without frontal staging; Ruined church interior (Damaged) — Seen laterally behind the approach and bowed crowd; used as Provides scale around 백산's full-body silhouette; Golden Hahoe mask (Worn by 백산) — Visible in three-quarter view, aligned with his route toward the platform; used as Concentrates authority within the full-body composition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Light entering through the stained glass gives the ruined interior a restrained sacred quality, with selective richness in 백산's golden mask and royal clothing.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Daylight filters through stained glass into the damaged church. Charlie is already confined in a steel cage behind the barred entrance, retaining the oversized straw hat, boots and colorful raincoat from his capture. 백산: He has a massive, obese build and wears regal clothing and an intact golden Hahoe mask. His pre-existing radiation-disfigured face remains concealed beneath it.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 백산 (한국인, 성인, 남성 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 육중한 황금빛 하회탈을 쓴 거대한 백산이 화려한 의복을 입은 채 단상을 향해 한쪽 발을 공중에 든 mid-action 자세의 위압적인 전신.\n\nLOCATION (lock): On the approach to the raised platform inside the ruined village church, lit by daylight through stained glass. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Platform (Awaiting 백산's arrival) — Its side edge is visible along the rightward route; used as Gives his suspended step a destination without frontal staging; Ruined church interior (Damaged) — Seen laterally behind the approach and bowed crowd; used as Provides scale around 백산's full-body silhouette; Golden Hahoe mask (Worn by 백산) — Visible in three-quarter view, aligned with his route toward the platform; used as Concentrates authority within the full-body composition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Light entering through the stained glass gives the ruined interior a restrained sacred quality, with selective richness in 백산's golden mask and royal clothing.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Daylight filters through stained glass into the damaged church. Charlie is already confined in a steel cage behind the barred entrance, retaining the oversized straw hat, boots and colorful raincoat from his capture. 백산: He has a massive, obese build and wears regal clothing and an intact golden Hahoe mask. His pre-existing radiation-disfigured face remains concealed beneath it.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 백산 (한국인, 성인, 남성 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 육중한 황금빛 하회탈을 쓴 거대한 백산이 화려한 의복을 입은 채 단상을 향해 한쪽 발을 공중에 든 mid-action 자세의 위압적인 전신.\n\nLOCATION (lock): On the approach to the raised platform inside the ruined village church, lit by daylight through stained glass. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Platform (Awaiting 백산's arrival) — Its side edge is visible along the rightward route; used as Gives his suspended step a destination without frontal staging; Ruined church interior (Damaged) — Seen laterally behind the approach and bowed crowd; used as Provides scale around 백산's full-body silhouette; Golden Hahoe mask (Worn by 백산) — Visible in three-quarter view, aligned with his route toward the platform; used as Concentrates authority within the full-body composition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Light entering through the stained glass gives the ruined interior a restrained sacred quality, with selective richness in 백산's golden mask and royal clothing.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Daylight filters through stained glass into the damaged church. Charlie is already confined in a steel cage behind the barred entrance, retaining the oversized straw hat, boots and colorful raincoat from his capture. 백산: He has a massive, obese build and wears regal clothing and an intact golden Hahoe mask. His pre-existing radiation-disfigured face remains concealed beneath it.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 백산 (한국인, 성인, 남성 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S58sh5__bgfirst_bg.png",
     "asset_id": "31941b00-404b-4a16-90fe-5d0f9ded9490",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S58sh5.png",
     "asset_id": "6d80eb92-6d35-4bd5-bfec-92b0a707c02a",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 백산: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1241731>",
     "asset_id": "b337b8d8-94d9-4a29-9e49-2e19121379b7",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L220B01.png",
     "asset_id": "b25ec118-d5a5-4461-bcc9-1623494e1fcd",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 백산: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1241731>",
     "asset_id": "b337b8d8-94d9-4a29-9e49-2e19121379b7",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "백산의 시선과 들려진 발은 화면 우측 하단의 석재 단상 모서리를 향하고 있습니다.",
    "built_space": "교회 내부의 좌측 스테인드글라스 아래 벽면이 뚫려 다수의 철창이 임의로 배치되어 있으며, 레퍼런스 우측의 쇠창살 구조는 보이지 않습니다.",
    "entities": "비만 체형의 백산은 황금빛 하회탈과 화려한 의복을 입고 있습니다. 엎드린 군중이 좌측에 있고, 찰리는 좌측 철창 안에서 밀짚모자와 우비를 입고 있습니다.",
    "hard_violations": [
     "[gemini-pro] 지정된 공간 구조 왜곡: 레퍼런스의 우측 철창을 무시하고 좌측 벽면에 없는 철창들을 임의로 생성함",
     "[gpt-high] 백산 외에 다수의 군중과 우리 속 찰리를 등장시켜 인물 제한을 위반했다.",
     "[gpt-high] 기존 후면 단상과 별도로 오른쪽 전경에 석재 단상을 추가하여 목적지 구조물을 중복시켰다."
    ],
    "physics": "백산은 왼발로 지면을 딛고 오른발을 들어 올린 안정적인 보행 자세를 취하고 있습니다."
   },
   {
    "label": "B",
    "direction": "백산의 얼굴과 들려진 오른발은 화면 우측 전경에 배치된 목재 단상을 향하고 있습니다.",
    "built_space": "좌측의 스테인드글라스 벽면, 정면의 제단, 우측의 철창과 쇠창살 문이 레퍼런스와 정확히 일치하며 우측 전경에 단상이 배치되었습니다.",
    "entities": "거구의 백산이 황금 하회탈과 황금빛 의복을 착용했습니다. 좌측에 군중이 엎드려 있으며, 찰리는 우측 본래 철창 위치에서 밀짚모자와 우비를 입고 있습니다.",
    "hard_violations": [
     "[gpt-high] 백산만 등장하도록 제한했는데 고개 숙인 군중과 우리 속 찰리를 추가했다.",
     "[gpt-high] 참조의 후면 석재 단상을 남겨 둔 채 오른쪽 전경에 독립된 목재 단상을 추가하여 목적지 구조물을 중복시켰다."
    ],
    "physics": "오른발로 몸을 지탱하고 왼발을 단상을 향해 들어 올린 자연스러운 mid-action 자세를 보여줍니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "레퍼런스 공간의 구조(우측 철창, 좌측 창문)를 완벽히 유지하면서 찰리의 위치와 백산의 움직임을 정확하게 연출했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "레퍼런스에 없는 다수의 철창을 좌측 창문 쪽에 임의로 생성하여 심각한 공간 왜곡을 일으켰습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "백산의 시선과 들려진 발은 화면 우측 하단의 석재 단상 모서리를 향하고 있습니다.",
        "built_space": "교회 내부의 좌측 스테인드글라스 아래 벽면이 뚫려 다수의 철창이 임의로 배치되어 있으며, 레퍼런스 우측의 쇠창살 구조는 보이지 않습니다.",
        "entities": "비만 체형의 백산은 황금빛 하회탈과 화려한 의복을 입고 있습니다. 엎드린 군중이 좌측에 있고, 찰리는 좌측 철창 안에서 밀짚모자와 우비를 입고 있습니다.",
        "hard_violations": [
         "지정된 공간 구조 왜곡: 레퍼런스의 우측 철창을 무시하고 좌측 벽면에 없는 철창들을 임의로 생성함"
        ],
        "physics": "백산은 왼발로 지면을 딛고 오른발을 들어 올린 안정적인 보행 자세를 취하고 있습니다."
       },
       {
        "label": "B",
        "direction": "백산의 얼굴과 들려진 오른발은 화면 우측 전경에 배치된 목재 단상을 향하고 있습니다.",
        "built_space": "좌측의 스테인드글라스 벽면, 정면의 제단, 우측의 철창과 쇠창살 문이 레퍼런스와 정확히 일치하며 우측 전경에 단상이 배치되었습니다.",
        "entities": "거구의 백산이 황금 하회탈과 황금빛 의복을 착용했습니다. 좌측에 군중이 엎드려 있으며, 찰리는 우측 본래 철창 위치에서 밀짚모자와 우비를 입고 있습니다.",
        "hard_violations": [],
        "physics": "오른발로 몸을 지탱하고 왼발을 단상을 향해 들어 올린 자연스러운 mid-action 자세를 보여줍니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "레퍼런스 공간의 구조(우측 철창, 좌측 창문)를 완벽히 유지하면서 찰리의 위치와 백산의 움직임을 정확하게 연출했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "레퍼런스에 없는 다수의 철창을 좌측 창문 쪽에 임의로 생성하여 심각한 공간 왜곡을 일으켰습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "백산의 시선과 들려진 발은 화면 우측 하단의 석재 단상 모서리를 향하고 있습니다.",
        "built_space": "교회 내부의 좌측 스테인드글라스 아래 벽면이 뚫려 다수의 철창이 임의로 배치되어 있으며, 레퍼런스 우측의 쇠창살 구조는 보이지 않습니다.",
        "entities": "비만 체형의 백산은 황금빛 하회탈과 화려한 의복을 입고 있습니다. 엎드린 군중이 좌측에 있고, 찰리는 좌측 철창 안에서 밀짚모자와 우비를 입고 있습니다.",
        "hard_violations": [
         "지정된 공간 구조 왜곡: 레퍼런스의 우측 철창을 무시하고 좌측 벽면에 없는 철창들을 임의로 생성함"
        ],
        "physics": "백산은 왼발로 지면을 딛고 오른발을 들어 올린 안정적인 보행 자세를 취하고 있습니다."
       },
       {
        "label": "B",
        "direction": "백산의 얼굴과 들려진 오른발은 화면 우측 전경에 배치된 목재 단상을 향하고 있습니다.",
        "built_space": "좌측의 스테인드글라스 벽면, 정면의 제단, 우측의 철창과 쇠창살 문이 레퍼런스와 정확히 일치하며 우측 전경에 단상이 배치되었습니다.",
        "entities": "거구의 백산이 황금 하회탈과 황금빛 의복을 착용했습니다. 좌측에 군중이 엎드려 있으며, 찰리는 우측 본래 철창 위치에서 밀짚모자와 우비를 입고 있습니다.",
        "hard_violations": [],
        "physics": "오른발로 몸을 지탱하고 왼발을 단상을 향해 들어 올린 자연스러운 mid-action 자세를 보여줍니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "오른쪽 단상을 향해 발을 든 전신과 사선의 황금 탈은 구현했지만, 금지된 추가 인물과 별도 단상이 등장하고 백산의 거대한 비만 체형도 부족하다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "추가 인물과 단상 중복 때문에 부적격이지만, A보다 육중한 비만 체형과 화려한 왕실 의복, 오른쪽으로 내딛는 전신 동작을 충실히 구현했다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "백산의 몸과 들린 앞발은 화면 오른쪽 목재 단상 계단을 향한다. 황금 탈도 오른쪽을 향하면서 정면 일부가 보여 요구된 사선 각도에 가깝다. 실제 눈의 시선은 탈 때문에 확인할 수 없다. 왼쪽 군중은 바닥을 향해 고개를 숙이고, 우리 속 인물은 실내 쪽을 바라본다.",
        "built_space": "뒤쪽 중앙에 십자가 하나, 높은 등받이 의자 하나와 넓은 석재 계단식 단상이 있고, 양쪽 상층 난간과 스테인드글라스가 보인다. 오른쪽 철창 출입구 앞에는 철제 우리 하나가 있어 참조의 주요 배치와 가깝다. 그러나 오른쪽 전경에 별도의 목재 계단식 단상을 추가하여, 기존 후면 단상과 합쳐 두 개의 독립된 단상이 된다. 백산은 후면 단상이 아니라 새 목재 단상으로 접근한다.",
        "entities": "백산으로 제시된 성인 남성 한 명은 금빛 자수 예복과 온전한 황금색 하회탈 형태의 가면을 착용한다. 얼굴이 가려져 참조 인물의 얼굴 동일성이나 민족성은 직접 확인할 수 없으며, 노출된 검은 머리는 참조와 양립한다. 키는 크게 표현했지만 몸통은 요구된 거대한 비만 체형보다 가늘다. 왼쪽 전경에는 약 열 명의 추가 성인이 있고, 오른쪽 우리에는 밀짚모자·다채로운 우비·노란 장화를 착용한 찰리로 보이는 인물 한 명이 있다. 백산만 허용한 인물 제한과 충돌한다. 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [
         "백산만 등장하도록 제한했는데 고개 숙인 군중과 우리 속 찰리를 추가했다.",
         "참조의 후면 석재 단상을 남겨 둔 채 오른쪽 전경에 독립된 목재 단상을 추가하여 목적지 구조물을 중복시켰다."
        ],
        "physics": "백산은 화면 왼쪽의 뒤쪽 신발을 바닥에 붙여 체중을 지탱하고, 반대쪽 다리를 앞으로 들어 계단 쪽으로 내민다. 전신이 공중에 떠 있는 것은 아니며 보행 중 한 발을 든 자세로 성립한다. 예복은 어깨와 몸에 걸려 아래로 늘어진다. 군중은 바닥에 무릎이나 발을 대고 있으며, 우리 속 인물의 장화는 우리 바닥에 닿는다. 지지 없는 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "백산의 몸통, 들린 앞발과 황금 탈이 모두 화면 오른쪽 전경의 석재 계단을 향한다. 이동 목표는 분명하지만 탈은 요구된 사선 정면보다 거의 옆모습에 가깝다. 눈의 실제 시선은 확인되지 않는다. 뒤쪽 군중은 바닥을 향해 엎드리고, 우리 속 인물도 머리를 낮추고 있다.",
        "built_space": "뒤쪽에 십자가 하나, 높은 등받이 의자 하나와 넓은 석재 계단식 단상이 있다. 왼쪽에는 상층 난간과 여러 스테인드글라스 창, 철창 출입구와 철제 우리 하나가 보인다. 낡은 회벽과 석재 바닥은 참조 장소의 재질을 따른다. 다만 우리와 출입구가 참조와 달리 왼쪽 창벽에 배치되어 있다. 오른쪽 전경에는 후면 단상과 떨어진 별도의 석재 단상이 있어 독립된 단상이 두 개로 읽히며, 백산은 이 전경 단상을 향한다.",
        "entities": "백산으로 제시된 성인 남성 한 명은 큰 배와 두꺼운 팔다리를 가진 육중한 비만 체형이며, 적색·남색·금색의 자수 예복과 금속 허리 장식을 착용한다. 온전한 황금색 하회탈 형태의 가면이 얼굴을 가리고 있어 참조 얼굴이나 민족성을 직접 검증할 수 없다. 왼쪽과 뒤쪽에는 열 명 이상의 추가 성인이 무릎을 꿇고 있으며, 철제 우리 안에는 큰 밀짚모자와 다채로운 겉옷을 입은 찰리로 보이는 인물이 있다. 이들은 백산만 허용한 제한을 위반한다. 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [
         "백산 외에 다수의 군중과 우리 속 찰리를 등장시켜 인물 제한을 위반했다.",
         "기존 후면 단상과 별도로 오른쪽 전경에 석재 단상을 추가하여 목적지 구조물을 중복시켰다."
        ],
        "physics": "백산의 뒤쪽 장화 앞부분은 바닥에 닿고 뒤꿈치는 들려 있어, 그 발로 지지하며 반대쪽 무릎과 장화를 앞으로 들어 올린 보행 순간으로 읽힌다. 공중에 든 발은 오른쪽 계단 방향으로 착지할 수 있는 위치다. 예복과 망토는 몸에 걸쳐져 주름지며, 손은 자연스럽게 아래로 내려와 있다. 군중은 무릎과 손으로 바닥을 짚고, 우리 속 인물은 우리 내부 바닥 높이에 놓여 있다. 명백하게 지지 없이 떠 있는 인물이나 물체는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "오른쪽 단상을 향해 발을 든 전신과 사선의 황금 탈은 구현했지만, 금지된 추가 인물과 별도 단상이 등장하고 백산의 거대한 비만 체형도 부족하다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "추가 인물과 단상 중복 때문에 부적격이지만, A보다 육중한 비만 체형과 화려한 왕실 의복, 오른쪽으로 내딛는 전신 동작을 충실히 구현했다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "백산의 몸과 들린 앞발은 화면 오른쪽 목재 단상 계단을 향한다. 황금 탈도 오른쪽을 향하면서 정면 일부가 보여 요구된 사선 각도에 가깝다. 실제 눈의 시선은 탈 때문에 확인할 수 없다. 왼쪽 군중은 바닥을 향해 고개를 숙이고, 우리 속 인물은 실내 쪽을 바라본다.",
        "built_space": "뒤쪽 중앙에 십자가 하나, 높은 등받이 의자 하나와 넓은 석재 계단식 단상이 있고, 양쪽 상층 난간과 스테인드글라스가 보인다. 오른쪽 철창 출입구 앞에는 철제 우리 하나가 있어 참조의 주요 배치와 가깝다. 그러나 오른쪽 전경에 별도의 목재 계단식 단상을 추가하여, 기존 후면 단상과 합쳐 두 개의 독립된 단상이 된다. 백산은 후면 단상이 아니라 새 목재 단상으로 접근한다.",
        "entities": "백산으로 제시된 성인 남성 한 명은 금빛 자수 예복과 온전한 황금색 하회탈 형태의 가면을 착용한다. 얼굴이 가려져 참조 인물의 얼굴 동일성이나 민족성은 직접 확인할 수 없으며, 노출된 검은 머리는 참조와 양립한다. 키는 크게 표현했지만 몸통은 요구된 거대한 비만 체형보다 가늘다. 왼쪽 전경에는 약 열 명의 추가 성인이 있고, 오른쪽 우리에는 밀짚모자·다채로운 우비·노란 장화를 착용한 찰리로 보이는 인물 한 명이 있다. 백산만 허용한 인물 제한과 충돌한다. 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [
         "백산만 등장하도록 제한했는데 고개 숙인 군중과 우리 속 찰리를 추가했다.",
         "참조의 후면 석재 단상을 남겨 둔 채 오른쪽 전경에 독립된 목재 단상을 추가하여 목적지 구조물을 중복시켰다."
        ],
        "physics": "백산은 화면 왼쪽의 뒤쪽 신발을 바닥에 붙여 체중을 지탱하고, 반대쪽 다리를 앞으로 들어 계단 쪽으로 내민다. 전신이 공중에 떠 있는 것은 아니며 보행 중 한 발을 든 자세로 성립한다. 예복은 어깨와 몸에 걸려 아래로 늘어진다. 군중은 바닥에 무릎이나 발을 대고 있으며, 우리 속 인물의 장화는 우리 바닥에 닿는다. 지지 없는 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "백산의 몸통, 들린 앞발과 황금 탈이 모두 화면 오른쪽 전경의 석재 계단을 향한다. 이동 목표는 분명하지만 탈은 요구된 사선 정면보다 거의 옆모습에 가깝다. 눈의 실제 시선은 확인되지 않는다. 뒤쪽 군중은 바닥을 향해 엎드리고, 우리 속 인물도 머리를 낮추고 있다.",
        "built_space": "뒤쪽에 십자가 하나, 높은 등받이 의자 하나와 넓은 석재 계단식 단상이 있다. 왼쪽에는 상층 난간과 여러 스테인드글라스 창, 철창 출입구와 철제 우리 하나가 보인다. 낡은 회벽과 석재 바닥은 참조 장소의 재질을 따른다. 다만 우리와 출입구가 참조와 달리 왼쪽 창벽에 배치되어 있다. 오른쪽 전경에는 후면 단상과 떨어진 별도의 석재 단상이 있어 독립된 단상이 두 개로 읽히며, 백산은 이 전경 단상을 향한다.",
        "entities": "백산으로 제시된 성인 남성 한 명은 큰 배와 두꺼운 팔다리를 가진 육중한 비만 체형이며, 적색·남색·금색의 자수 예복과 금속 허리 장식을 착용한다. 온전한 황금색 하회탈 형태의 가면이 얼굴을 가리고 있어 참조 얼굴이나 민족성을 직접 검증할 수 없다. 왼쪽과 뒤쪽에는 열 명 이상의 추가 성인이 무릎을 꿇고 있으며, 철제 우리 안에는 큰 밀짚모자와 다채로운 겉옷을 입은 찰리로 보이는 인물이 있다. 이들은 백산만 허용한 제한을 위반한다. 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [
         "백산 외에 다수의 군중과 우리 속 찰리를 등장시켜 인물 제한을 위반했다.",
         "기존 후면 단상과 별도로 오른쪽 전경에 석재 단상을 추가하여 목적지 구조물을 중복시켰다."
        ],
        "physics": "백산의 뒤쪽 장화 앞부분은 바닥에 닿고 뒤꿈치는 들려 있어, 그 발로 지지하며 반대쪽 무릎과 장화를 앞으로 들어 올린 보행 순간으로 읽힌다. 공중에 든 발은 오른쪽 계단 방향으로 착지할 수 있는 위치다. 예복과 망토는 몸에 걸쳐져 주름지며, 손은 자연스럽게 아래로 내려와 있다. 군중은 무릎과 손으로 바닥을 짚고, 우리 속 인물은 우리 내부 바닥 높이에 놓여 있다. 명백하게 지지 없이 떠 있는 인물이나 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.5,
    "B": 1.667
   },
   "adjusted": {
    "A": 1.25,
    "B": 1.417
   },
   "violations": {
    "A": [
     "[gemini-pro] 지정된 공간 구조 왜곡: 레퍼런스의 우측 철창을 무시하고 좌측 벽면에 없는 철창들을 임의로 생성함",
     "[gpt-high] 백산 외에 다수의 군중과 우리 속 찰리를 등장시켜 인물 제한을 위반했다.",
     "[gpt-high] 기존 후면 단상과 별도로 오른쪽 전경에 석재 단상을 추가하여 목적지 구조물을 중복시켰다."
    ],
    "B": [
     "[gpt-high] 백산만 등장하도록 제한했는데 고개 숙인 군중과 우리 속 찰리를 추가했다.",
     "[gpt-high] 참조의 후면 석재 단상을 남겨 둔 채 오른쪽 전경에 독립된 목재 단상을 추가하여 목적지 구조물을 중복시켰다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1417,
   "A": 1250
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1417,
    "verdict_ko": "레퍼런스 공간의 구조(우측 철창, 좌측 창문)를 완벽히 유지하면서 찰리의 위치와 백산의 움직임을 정확하게 연출했습니다.  ★위반: [gpt-high] 백산만 등장하도록 제한했는데 고개 숙인 군중과 우리 속 찰리를 추가했다. / [gpt-high] 참조의 후면 석재 단상을 남겨 둔 채 오른쪽 전경에 독립된 목재 단상을 추가하여 목적지 구조물을 중복시켰다."
   },
   {
    "label": "A",
    "score": 1250,
    "verdict_ko": "레퍼런스에 없는 다수의 철창을 좌측 창문 쪽에 임의로 생성하여 심각한 공간 왜곡을 일으켰습니다.  ★위반: [gemini-pro] 지정된 공간 구조 왜곡: 레퍼런스의 우측 철창을 무시하고 좌측 벽면에 없는 철창들을 임의로 생성함 / [gpt-high] 백산 외에 다수의 군중과 우리 속 찰리를 등장시켜 인물 제한을 위반했다. / [gpt-high] 기존 후면 단상과 별도로 오른쪽 전경에 석재 단상을 추가하여 목적지 구조물을 중복시켰다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L220B01.png",
    "asset_id": "b25ec118-d5a5-4461-bcc9-1623494e1fcd",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 백산: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1241731>",
    "asset_id": "b337b8d8-94d9-4a29-9e49-2e19121379b7",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-6e7b-7677-96eb-1b3c8e575ae0",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S58sh5__bgfirst_bg.png",
   "bg_asset_id": "31941b00-404b-4a16-90fe-5d0f9ded9490",
   "bg_record_key": "S58sh5::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S58sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:32:25.022629+00:00",
  "fingerprint": "36400d1afbb576744d16fba384d302e32f728f744f8b45f5c7e7be3125145327",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S58sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S58sh5_sel.png",
  "source_sha256": "f751ce36bf414a336ff8e0ce2a260690b80fa784bc298d5a6cadb1454d5ac312",
  "file": "S58sh5_cine.png",
  "staged_sha256": "e0443253a3542fddf3efc5cc5fb9348c4bca076fba0bd9139d6fe54862df148b",
  "latency_ms": 10234
 },
 "S58sh15::signage": {
  "fp": "b1d466408b45ee77",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S58sh15": {
  "input_fingerprint": "1187ec8eefacf43e",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 백산의 손짓에 맞춰 거칠게 열린 철창문 안에서 모습을 드러낸 강철 케이지 속 찰리의 낡은 금속 전신.\n\nLOCATION (lock): At a barred side opening adjoining the village church's platform, where a steel cage appears in stained-glass daylight. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Barred gate (Opening far enough to reveal 찰리's full body) — Seen obliquely from the prisoners' side, with its displaced section beside the opening; used as Revealing edge that no longer cuts through 찰리's silhouette; Steel cage (Containing 찰리 behind the open gate) — Front and side bars are visible in the oblique view; used as Establishes a second enclosure around the revealed figure; Church interior around the gate (Damaged) — The approach floor and architecture around the opening remain visible; used as Separates the cage from the foreground and preserves spatial scale.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the church's established stained-glass daylight with controlled metal highlights and no separate illumination invented for the cage.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The barred entrance is open, revealing Charlie's steel cage inside the damaged, stained-glass-lit church. Charlie retains the oversized straw hat, rubber boots and colorful raincoat worn at capture.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 백산의 손짓에 맞춰 거칠게 열린 철창문 안에서 모습을 드러낸 강철 케이지 속 찰리의 낡은 금속 전신.\n\nLOCATION (lock): At a barred side opening adjoining the village church's platform, where a steel cage appears in stained-glass daylight. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Barred gate (Opening far enough to reveal 찰리's full body) — Seen obliquely from the prisoners' side, with its displaced section beside the opening; used as Revealing edge that no longer cuts through 찰리's silhouette; Steel cage (Containing 찰리 behind the open gate) — Front and side bars are visible in the oblique view; used as Establishes a second enclosure around the revealed figure; Church interior around the gate (Damaged) — The approach floor and architecture around the opening remain visible; used as Separates the cage from the foreground and preserves spatial scale.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the church's established stained-glass daylight with controlled metal highlights and no separate illumination invented for the cage.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The barred entrance is open, revealing Charlie's steel cage inside the damaged, stained-glass-lit church. Charlie retains the oversized straw hat, rubber boots and colorful raincoat worn at capture.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 백산의 손짓에 맞춰 거칠게 열린 철창문 안에서 모습을 드러낸 강철 케이지 속 찰리의 낡은 금속 전신.\n\nLOCATION (lock): At a barred side opening adjoining the village church's platform, where a steel cage appears in stained-glass daylight. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Barred gate (Opening far enough to reveal 찰리's full body) — Seen obliquely from the prisoners' side, with its displaced section beside the opening; used as Revealing edge that no longer cuts through 찰리's silhouette; Steel cage (Containing 찰리 behind the open gate) — Front and side bars are visible in the oblique view; used as Establishes a second enclosure around the revealed figure; Church interior around the gate (Damaged) — The approach floor and architecture around the opening remain visible; used as Separates the cage from the foreground and preserves spatial scale.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the church's established stained-glass daylight with controlled metal highlights and no separate illumination invented for the cage.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The barred entrance is open, revealing Charlie's steel cage inside the damaged, stained-glass-lit church. Charlie retains the oversized straw hat, rubber boots and colorful raincoat worn at capture.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "찰리는 카메라 쪽을 정면으로 향하고 있음.",
    "built_space": "교회 본당 내부. 우측의 철창문이 열려 있고, 그 경계 안쪽에 강철 케이지가 놓여 있음.",
    "entities": "고릴라형 로봇 몸체와 밀짚모자는 일치하나, 컬러풀한 우비는 보이지 않음.",
    "hard_violations": [],
    "physics": "찰리는 두 발로 케이지 바닥을 지지하며 서 있고, 케이지는 바닥에 안정적으로 위치함."
   },
   {
    "label": "B",
    "direction": "인물은 카메라 쪽을 정면으로 향하고 있음.",
    "built_space": "교회 본당 내부. 철창문 일부가 분리되어 있으며, 케이지가 철창문 안쪽이 아닌 본당 쪽으로 나와 있음.",
    "entities": "우비와 장화는 존재하나, 로봇 찰리가 아닌 레퍼런스의 인간형 인물이 그대로 등장함.",
    "hard_violations": [
     "[gemini-pro] 지정되지 않은 인물(이전 샷의 인간) 복사 및 등장",
     "[gemini-pro] 케이지가 철창문 안쪽이 아닌 바깥 공간에 잘못 배치됨",
     "[gpt-high] 기존 출입구의 좌우 철창문 부분을 남겨 둔 채 별도의 분리된 철창 문짝을 전경에 추가하여, 참고 장소의 문짝 구성을 중복시켰습니다."
    ],
    "physics": "인물은 케이지 바닥에 서 있고, 분리된 철창은 바닥에 비스듬히 기대어 지지됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "지시된 로봇 찰리의 기본 외형과 케이지 배치는 충족했으나, 착용해야 할 우비가 누락되었습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "로봇 형태의 찰리 대신 레퍼런스의 인물을 그대로 복사하여 지시를 크게 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 카메라 쪽을 정면으로 향하고 있음.",
        "built_space": "교회 본당 내부. 우측의 철창문이 열려 있고, 그 경계 안쪽에 강철 케이지가 놓여 있음.",
        "entities": "고릴라형 로봇 몸체와 밀짚모자는 일치하나, 컬러풀한 우비는 보이지 않음.",
        "hard_violations": [],
        "physics": "찰리는 두 발로 케이지 바닥을 지지하며 서 있고, 케이지는 바닥에 안정적으로 위치함."
       },
       {
        "label": "B",
        "direction": "인물은 카메라 쪽을 정면으로 향하고 있음.",
        "built_space": "교회 본당 내부. 철창문 일부가 분리되어 있으며, 케이지가 철창문 안쪽이 아닌 본당 쪽으로 나와 있음.",
        "entities": "우비와 장화는 존재하나, 로봇 찰리가 아닌 레퍼런스의 인간형 인물이 그대로 등장함.",
        "hard_violations": [
         "지정되지 않은 인물(이전 샷의 인간) 복사 및 등장",
         "케이지가 철창문 안쪽이 아닌 바깥 공간에 잘못 배치됨"
        ],
        "physics": "인물은 케이지 바닥에 서 있고, 분리된 철창은 바닥에 비스듬히 기대어 지지됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "지시된 로봇 찰리의 기본 외형과 케이지 배치는 충족했으나, 착용해야 할 우비가 누락되었습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "로봇 형태의 찰리 대신 레퍼런스의 인물을 그대로 복사하여 지시를 크게 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 카메라 쪽을 정면으로 향하고 있음.",
        "built_space": "교회 본당 내부. 우측의 철창문이 열려 있고, 그 경계 안쪽에 강철 케이지가 놓여 있음.",
        "entities": "고릴라형 로봇 몸체와 밀짚모자는 일치하나, 컬러풀한 우비는 보이지 않음.",
        "hard_violations": [],
        "physics": "찰리는 두 발로 케이지 바닥을 지지하며 서 있고, 케이지는 바닥에 안정적으로 위치함."
       },
       {
        "label": "B",
        "direction": "인물은 카메라 쪽을 정면으로 향하고 있음.",
        "built_space": "교회 본당 내부. 철창문 일부가 분리되어 있으며, 케이지가 철창문 안쪽이 아닌 본당 쪽으로 나와 있음.",
        "entities": "우비와 장화는 존재하나, 로봇 찰리가 아닌 레퍼런스의 인간형 인물이 그대로 등장함.",
        "hard_violations": [
         "지정되지 않은 인물(이전 샷의 인간) 복사 및 등장",
         "케이지가 철창문 안쪽이 아닌 바깥 공간에 잘못 배치됨"
        ],
        "physics": "인물은 케이지 바닥에 서 있고, 분리된 철창은 바닥에 비스듬히 기대어 지지됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "비스듬한 와이드숏과 의상은 맞지만 찰리가 육중한 금속 로봇이 아닌 사람형 인물로 보이며, 기존 철창문 외에 분리된 문짝까지 추가되어 공간 연속성이 깨집니다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "열린 철창문 옆으로 케이지의 정면·측면과 찰리의 금속 전신을 드러내는 구도와 정체성이 정확하지만, 유지해야 할 알록달록한 우비가 빠졌습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "인물의 얼굴은 화면 왼쪽의 교회 내부를 향하며 특정 상대는 보이지 않습니다. 이동이나 손짓은 없고, 무기나 겨냥하는 물체도 없습니다.",
        "built_space": "바퀴 달린 케이지 한 개의 정면과 왼쪽 측면이 보이고 인물의 전신이 그 안에 들어 있습니다. 뒤쪽 출입구에는 좌우의 수직 철창문 부분이 남아 있는데, 오른쪽 전경에는 별도의 철창 패널 한 개가 비스듬히 놓여 있습니다. 손상된 회벽, 상층 난간, 스테인드글라스와 왼쪽 계단, 넓은 접근 바닥은 이전 장소와 유사합니다. 외부 철창문은 인물의 윤곽을 가리지 않지만 케이지 자체의 창살은 몸 앞을 지납니다.",
        "entities": "케이지 안에는 한 명만 있습니다. 밀짚모자, 다색 우비, 노란 고무장화는 유지되었습니다. 그러나 드러난 얼굴과 손은 사람 또는 사람형 가면처럼 보이고, 캐릭터 참고의 넓은 샌드 베이지 장갑판, 육중하고 긴 기계 팔, 짧은 기계 다리가 없습니다. 성별·연령·민족성을 확정하기 어려우며, 요구된 비인간 로봇 정체성이 확인되지 않습니다. 읽을 수 있는 글자는 보이지 않습니다.",
        "hard_violations": [
         "기존 출입구의 좌우 철창문 부분을 남겨 둔 채 별도의 분리된 철창 문짝을 전경에 추가하여, 참고 장소의 문짝 구성을 중복시켰습니다."
        ],
        "physics": "인물은 두 장화로 케이지 바닥을 딛고 있고 케이지는 바퀴로 교회 바닥에 지지됩니다. 모자는 머리에, 우비는 몸에 걸쳐 있습니다. 분리된 철창 패널은 아래 모서리를 바닥에 대고 위쪽을 출입구 쪽에 기대고 있어 떠 있지는 않습니다. 문이 거칠게 열리는 동작 자체는 보이지 않습니다."
       },
       {
        "label": "B",
        "direction": "찰리의 얼굴은 카메라 정면에서 약간 화면 오른쪽의 교회 내부를 향합니다. 시선의 구체적인 상대는 화면에 없으며, 누군가를 겨냥하거나 이동하는 동작도 없습니다. 오른쪽으로 열린 철창문은 찰리의 전신을 드러내는 방향으로 물러나 있습니다.",
        "built_space": "케이지 한 개의 정면과 왼쪽 측면, 바닥과 바퀴가 보입니다. 외부 출입구는 왼쪽에 물러난 철창 부분과 오른쪽으로 열린 문짝으로 구성되며, 오른쪽 문짝은 찰리의 윤곽을 자르지 않습니다. 찰리는 별도의 내부 케이지 안에 서 있어 이중 수용 구조가 명확합니다. 손상된 회벽과 기둥, 상층 난간, 왼쪽 스테인드글라스 한 개와 그 아래 방열기, 전경의 접근 바닥이 유지됩니다. 반사를 이용한 불가능한 공간 표현은 없습니다.",
        "entities": "찰리 한 개체만 등장합니다. 흰 마스크형 기계 얼굴, 주황색 눈, 각진 샌드 베이지 장갑판, 넓은 상체와 긴 팔, 상대적으로 짧은 다리가 캐릭터 참고와 일치합니다. 밀짚모자와 노란 고무장화는 있지만 알록달록한 우비는 없습니다. 강철 케이지와 열린 외부 철창문도 확인되며, 추가 인물이나 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "찰리의 두 장화는 케이지 바닥에 닿아 체중을 지지하고, 케이지 바퀴는 교회 바닥에 닿습니다. 긴 팔은 어깨와 팔꿈치 관절에 연결되어 자연스럽게 아래로 내려가 있습니다. 모자는 머리에 얹혀 있고, 열린 문짝은 출입구 가장자리의 부착부로 지지됩니다. 지지 없이 떠 있는 몸이나 물체는 없습니다. 문이 열린 직후 드러난 정지 순간으로는 성립하지만 개방 동작의 거친 속도감은 뚜렷하지 않습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "비스듬한 와이드숏과 의상은 맞지만 찰리가 육중한 금속 로봇이 아닌 사람형 인물로 보이며, 기존 철창문 외에 분리된 문짝까지 추가되어 공간 연속성이 깨집니다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "열린 철창문 옆으로 케이지의 정면·측면과 찰리의 금속 전신을 드러내는 구도와 정체성이 정확하지만, 유지해야 할 알록달록한 우비가 빠졌습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "인물의 얼굴은 화면 왼쪽의 교회 내부를 향하며 특정 상대는 보이지 않습니다. 이동이나 손짓은 없고, 무기나 겨냥하는 물체도 없습니다.",
        "built_space": "바퀴 달린 케이지 한 개의 정면과 왼쪽 측면이 보이고 인물의 전신이 그 안에 들어 있습니다. 뒤쪽 출입구에는 좌우의 수직 철창문 부분이 남아 있는데, 오른쪽 전경에는 별도의 철창 패널 한 개가 비스듬히 놓여 있습니다. 손상된 회벽, 상층 난간, 스테인드글라스와 왼쪽 계단, 넓은 접근 바닥은 이전 장소와 유사합니다. 외부 철창문은 인물의 윤곽을 가리지 않지만 케이지 자체의 창살은 몸 앞을 지납니다.",
        "entities": "케이지 안에는 한 명만 있습니다. 밀짚모자, 다색 우비, 노란 고무장화는 유지되었습니다. 그러나 드러난 얼굴과 손은 사람 또는 사람형 가면처럼 보이고, 캐릭터 참고의 넓은 샌드 베이지 장갑판, 육중하고 긴 기계 팔, 짧은 기계 다리가 없습니다. 성별·연령·민족성을 확정하기 어려우며, 요구된 비인간 로봇 정체성이 확인되지 않습니다. 읽을 수 있는 글자는 보이지 않습니다.",
        "hard_violations": [
         "기존 출입구의 좌우 철창문 부분을 남겨 둔 채 별도의 분리된 철창 문짝을 전경에 추가하여, 참고 장소의 문짝 구성을 중복시켰습니다."
        ],
        "physics": "인물은 두 장화로 케이지 바닥을 딛고 있고 케이지는 바퀴로 교회 바닥에 지지됩니다. 모자는 머리에, 우비는 몸에 걸쳐 있습니다. 분리된 철창 패널은 아래 모서리를 바닥에 대고 위쪽을 출입구 쪽에 기대고 있어 떠 있지는 않습니다. 문이 거칠게 열리는 동작 자체는 보이지 않습니다."
       },
       {
        "label": "A",
        "direction": "찰리의 얼굴은 카메라 정면에서 약간 화면 오른쪽의 교회 내부를 향합니다. 시선의 구체적인 상대는 화면에 없으며, 누군가를 겨냥하거나 이동하는 동작도 없습니다. 오른쪽으로 열린 철창문은 찰리의 전신을 드러내는 방향으로 물러나 있습니다.",
        "built_space": "케이지 한 개의 정면과 왼쪽 측면, 바닥과 바퀴가 보입니다. 외부 출입구는 왼쪽에 물러난 철창 부분과 오른쪽으로 열린 문짝으로 구성되며, 오른쪽 문짝은 찰리의 윤곽을 자르지 않습니다. 찰리는 별도의 내부 케이지 안에 서 있어 이중 수용 구조가 명확합니다. 손상된 회벽과 기둥, 상층 난간, 왼쪽 스테인드글라스 한 개와 그 아래 방열기, 전경의 접근 바닥이 유지됩니다. 반사를 이용한 불가능한 공간 표현은 없습니다.",
        "entities": "찰리 한 개체만 등장합니다. 흰 마스크형 기계 얼굴, 주황색 눈, 각진 샌드 베이지 장갑판, 넓은 상체와 긴 팔, 상대적으로 짧은 다리가 캐릭터 참고와 일치합니다. 밀짚모자와 노란 고무장화는 있지만 알록달록한 우비는 없습니다. 강철 케이지와 열린 외부 철창문도 확인되며, 추가 인물이나 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "찰리의 두 장화는 케이지 바닥에 닿아 체중을 지지하고, 케이지 바퀴는 교회 바닥에 닿습니다. 긴 팔은 어깨와 팔꿈치 관절에 연결되어 자연스럽게 아래로 내려가 있습니다. 모자는 머리에 얹혀 있고, 열린 문짝은 출입구 가장자리의 부착부로 지지됩니다. 지지 없이 떠 있는 몸이나 물체는 없습니다. 문이 열린 직후 드러난 정지 순간으로는 성립하지만 개방 동작의 거친 속도감은 뚜렷하지 않습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.875
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.625
   },
   "violations": {
    "B": [
     "[gemini-pro] 지정되지 않은 인물(이전 샷의 인간) 복사 및 등장",
     "[gemini-pro] 케이지가 철창문 안쪽이 아닌 바깥 공간에 잘못 배치됨",
     "[gpt-high] 기존 출입구의 좌우 철창문 부분을 남겨 둔 채 별도의 분리된 철창 문짝을 전경에 추가하여, 참고 장소의 문짝 구성을 중복시켰습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 625
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지시된 로봇 찰리의 기본 외형과 케이지 배치는 충족했으나, 착용해야 할 우비가 누락되었습니다."
   },
   {
    "label": "B",
    "score": 625,
    "verdict_ko": "로봇 형태의 찰리 대신 레퍼런스의 인물을 그대로 복사하여 지시를 크게 위반했습니다.  ★위반: [gemini-pro] 지정되지 않은 인물(이전 샷의 인간) 복사 및 등장 / [gemini-pro] 케이지가 철창문 안쪽이 아닌 바깥 공간에 잘못 배치됨 / [gpt-high] 기존 출입구의 좌우 철창문 부분을 남겨 둔 채 별도의 분리된 철창 문짝을 전경에 추가하여, 참고 장소의 문짝 구성을 중복시켰습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S58sh5_sel.png",
    "asset_id": "5fa6c061-d691-4e36-a1b3-68c340578098",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-71d6-715e-bbf5-431ee38d3eff",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S58sh5"
  }
 },
 "S58sh15::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:33:32.746102+00:00",
  "fingerprint": "4f9d80d5296f5139a64551f623e85a16cbc0d1b820d62c5bf5157e09037a6d3d",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S58sh15_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S58sh15_sel.png",
  "source_sha256": "29caa60c1f0b1b28ed4b28e872ab98068fb4dc0109b99a46c67dda69fa2c1499",
  "file": "S58sh15_cine.png",
  "staged_sha256": "68cf4b493396a7f502d42ee382c3af867b11cf7241c766de47dcb8c477b152e7",
  "latency_ms": 10660
 },
 "S58sh25::signage": {
  "fp": "c011b4c6e8eb8b30",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S58sh25": {
  "input_fingerprint": "602bcb1e02f91b75",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 포박당해 몸이 뒤로 쏠린 mid-action 상태로, 찰리가 갇힌 철창문을 향해 절박하게 한 손을 뻗고 입을 크게 벌려 오열하듯 고정된 앰버의 상체.\n\nLOCATION (lock): On the church floor near the platform and barred cage entrance, under daylight filtering through stained glass. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: closed barred gate fragment in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Barred gate (Closed between 앰버 and 찰리) — Only a narrow oblique section is visible at the right edge; used as Marks the direction of the unreachable destination without obscuring her hand.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the established stained-glass daylight and gentle facial tonal separation, without adding a new emotional spotlight.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the steel cage, barred doorway, damaged church surfaces, and stained-glass daylight. Exclude the earlier open-door state; the barred door has now closed.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie's steel cage door is being closed again, with Charlie restrained inside and still retaining his capture outfit. The damaged church remains lit through its stained-glass windows. 앰버: She is being bound again inside the church.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 포박당해 몸이 뒤로 쏠린 mid-action 상태로, 찰리가 갇힌 철창문을 향해 절박하게 한 손을 뻗고 입을 크게 벌려 오열하듯 고정된 앰버의 상체.\n\nLOCATION (lock): On the church floor near the platform and barred cage entrance, under daylight filtering through stained glass. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: closed barred gate fragment in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Barred gate (Closed between 앰버 and 찰리) — Only a narrow oblique section is visible at the right edge; used as Marks the direction of the unreachable destination without obscuring her hand.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the established stained-glass daylight and gentle facial tonal separation, without adding a new emotional spotlight.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the steel cage, barred doorway, damaged church surfaces, and stained-glass daylight. Exclude the earlier open-door state; the barred door has now closed.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie's steel cage door is being closed again, with Charlie restrained inside and still retaining his capture outfit. The damaged church remains lit through its stained-glass windows. 앰버: She is being bound again inside the church.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 포박당해 몸이 뒤로 쏠린 mid-action 상태로, 찰리가 갇힌 철창문을 향해 절박하게 한 손을 뻗고 입을 크게 벌려 오열하듯 고정된 앰버의 상체.\n\nLOCATION (lock): On the church floor near the platform and barred cage entrance, under daylight filtering through stained glass. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: closed barred gate fragment in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Barred gate (Closed between 앰버 and 찰리) — Only a narrow oblique section is visible at the right edge; used as Marks the direction of the unreachable destination without obscuring her hand.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the established stained-glass daylight and gentle facial tonal separation, without adding a new emotional spotlight.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the steel cage, barred doorway, damaged church surfaces, and stained-glass daylight. Exclude the earlier open-door state; the barred door has now closed.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie's steel cage door is being closed again, with Charlie restrained inside and still retaining his capture outfit. The damaged church remains lit through its stained-glass windows. 앰버: She is being bound again inside the church.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "B",
    "direction": "앰버의 시선과 펼친 손은 화면 오른쪽 전방을 향합니다. 철창문은 어깨와 팔 뒤쪽에 있어, 뻗은 팔이 그 문을 향하는 공간 관계로 보이지 않습니다.",
    "built_space": "왼쪽 스테인드글라스 창 하나와 그 아래 난방기 하나, 위층 난간, 벗겨진 벽면, 중앙 뒤의 철제 우리 하나와 오른쪽 대형 철창문이 보입니다. 장소의 재료와 주요 구조는 참고 이미지에 가깝지만 철창문 전체에 가까운 면적과 바닥까지 드러나며, 요구된 오른쪽 가장자리의 좁고 비스듬한 배경 조각이 아닙니다. 문은 닫힌 상태로 보입니다.",
    "entities": "앰버는 금발의 어린 여자아이이며 큰 눈, 둥근 얼굴과 남색 반팔 상의가 참고와 대체로 맞습니다. 혼혈 정체성 자체는 외모만으로 확정할 수 없습니다. 입을 크게 벌리고 눈물을 흘립니다. 그러나 뒤에 별도 인물의 남색 옷을 입은 몸통과 앰버를 잡는 팔·손이 보입니다. 앰버의 포박용 끈은 보이지 않고 타인이 붙잡는 상태입니다. 찰리와 읽을 수 있는 글자는 보이지 않습니다.",
    "hard_violations": [
     "앰버만 등장하도록 제한했는데, 뒤에서 앰버의 팔과 몸을 붙잡는 추가 인물의 몸통과 팔·손이 등장합니다."
    ],
    "physics": "추가 인물의 손이 앰버의 위팔과 옆구리 부근에 닿아 붙잡고 있습니다. 뻗은 팔은 어깨에 정상적으로 연결되어 있으며 하체의 접지는 화면 밖이므로 공중에 떠 있다고 볼 근거는 없습니다. 다만 몸은 전방으로 기울어 있어 포박당해 뒤로 쏠린 상체라는 지정 동작과 다릅니다."
   },
   {
    "label": "A",
    "direction": "앰버의 시선과 뻗은 손은 화면 오른쪽의 철창문 및 그 너머를 향합니다. 손끝은 철창에 닿기 전의 간격을 남겨, 도달하지 못하는 목적지를 향해 손을 뻗는 관계가 읽힙니다.",
    "built_space": "오른쪽에 닫힌 철창문 한 면과 자물쇠 하나가 크게 보이고, 왼쪽 배경에는 큰 스테인드글라스 창 하나와 더 작은 창 구간, 손상된 벽과 바닥이 보입니다. 교회의 재료와 낮빛은 이어지지만 문이 화면 오른쪽 약 삼분의 일을 차지하는 가까운 구조물로 표현되어 좁은 배경 조각이라는 지시와 다릅니다. 참고보다 색광의 빛줄기도 훨씬 강조되었습니다.",
    "entities": "금발의 어린 여자아이 한 명만 보이며, 큰 눈과 둥근 얼굴, 남색 반팔 상의가 앰버 참고와 대체로 일치합니다. 혼혈 여부는 이미지에서 확정할 수 없습니다. 입을 크게 벌린 울음과 눈물이 보이고, 밧줄이 몸통과 뒤로 묶인 팔을 감싸 포박 상태를 명확히 나타냅니다. 찰리는 노출되지 않으며 읽을 수 있는 글자도 없습니다.",
    "hard_violations": [],
    "physics": "밧줄은 몸통과 뒤쪽 팔에 밀착되어 있고, 앞으로 뻗은 팔은 자유롭게 움직일 수 있는 형태입니다. 하체와 발은 화면 밖이므로 접지 방식은 확인할 수 없지만, 상체가 공중에 떠 있는 증거는 없습니다. 몸을 뒤로 당기는 외부 연결이나 장력은 보이지 않으며, 실제 자세는 뒤로 쏠림보다 문 쪽으로 앞으로 숙인 자세입니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": null,
     "normalized": null,
     "ok": false
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "허용되지 않은 다른 사람의 몸과 손이 등장하며, 손을 뻗는 방향도 뒤쪽 철창문이 아니라 전방이고 상체 클로즈업보다 구도가 넓습니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "앰버 혼자 포박된 채 닫힌 철창으로 손을 뻗는 관계는 맞지만, 뒤로 쏠린 순간 대신 앞으로 숙였고 철창이 오른쪽 배경의 좁은 조각보다 크게 등장합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버의 시선과 펼친 손은 화면 오른쪽 전방을 향합니다. 철창문은 어깨와 팔 뒤쪽에 있어, 뻗은 팔이 그 문을 향하는 공간 관계로 보이지 않습니다.",
        "built_space": "왼쪽 스테인드글라스 창 하나와 그 아래 난방기 하나, 위층 난간, 벗겨진 벽면, 중앙 뒤의 철제 우리 하나와 오른쪽 대형 철창문이 보입니다. 장소의 재료와 주요 구조는 참고 이미지에 가깝지만 철창문 전체에 가까운 면적과 바닥까지 드러나며, 요구된 오른쪽 가장자리의 좁고 비스듬한 배경 조각이 아닙니다. 문은 닫힌 상태로 보입니다.",
        "entities": "앰버는 금발의 어린 여자아이이며 큰 눈, 둥근 얼굴과 남색 반팔 상의가 참고와 대체로 맞습니다. 혼혈 정체성 자체는 외모만으로 확정할 수 없습니다. 입을 크게 벌리고 눈물을 흘립니다. 그러나 뒤에 별도 인물의 남색 옷을 입은 몸통과 앰버를 잡는 팔·손이 보입니다. 앰버의 포박용 끈은 보이지 않고 타인이 붙잡는 상태입니다. 찰리와 읽을 수 있는 글자는 보이지 않습니다.",
        "hard_violations": [
         "앰버만 등장하도록 제한했는데, 뒤에서 앰버의 팔과 몸을 붙잡는 추가 인물의 몸통과 팔·손이 등장합니다."
        ],
        "physics": "추가 인물의 손이 앰버의 위팔과 옆구리 부근에 닿아 붙잡고 있습니다. 뻗은 팔은 어깨에 정상적으로 연결되어 있으며 하체의 접지는 화면 밖이므로 공중에 떠 있다고 볼 근거는 없습니다. 다만 몸은 전방으로 기울어 있어 포박당해 뒤로 쏠린 상체라는 지정 동작과 다릅니다."
       },
       {
        "label": "B",
        "direction": "앰버의 시선과 뻗은 손은 화면 오른쪽의 철창문 및 그 너머를 향합니다. 손끝은 철창에 닿기 전의 간격을 남겨, 도달하지 못하는 목적지를 향해 손을 뻗는 관계가 읽힙니다.",
        "built_space": "오른쪽에 닫힌 철창문 한 면과 자물쇠 하나가 크게 보이고, 왼쪽 배경에는 큰 스테인드글라스 창 하나와 더 작은 창 구간, 손상된 벽과 바닥이 보입니다. 교회의 재료와 낮빛은 이어지지만 문이 화면 오른쪽 약 삼분의 일을 차지하는 가까운 구조물로 표현되어 좁은 배경 조각이라는 지시와 다릅니다. 참고보다 색광의 빛줄기도 훨씬 강조되었습니다.",
        "entities": "금발의 어린 여자아이 한 명만 보이며, 큰 눈과 둥근 얼굴, 남색 반팔 상의가 앰버 참고와 대체로 일치합니다. 혼혈 여부는 이미지에서 확정할 수 없습니다. 입을 크게 벌린 울음과 눈물이 보이고, 밧줄이 몸통과 뒤로 묶인 팔을 감싸 포박 상태를 명확히 나타냅니다. 찰리는 노출되지 않으며 읽을 수 있는 글자도 없습니다.",
        "hard_violations": [],
        "physics": "밧줄은 몸통과 뒤쪽 팔에 밀착되어 있고, 앞으로 뻗은 팔은 자유롭게 움직일 수 있는 형태입니다. 하체와 발은 화면 밖이므로 접지 방식은 확인할 수 없지만, 상체가 공중에 떠 있는 증거는 없습니다. 몸을 뒤로 당기는 외부 연결이나 장력은 보이지 않으며, 실제 자세는 뒤로 쏠림보다 문 쪽으로 앞으로 숙인 자세입니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "허용되지 않은 다른 사람의 몸과 손이 등장하며, 손을 뻗는 방향도 뒤쪽 철창문이 아니라 전방이고 상체 클로즈업보다 구도가 넓습니다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "앰버 혼자 포박된 채 닫힌 철창으로 손을 뻗는 관계는 맞지만, 뒤로 쏠린 순간 대신 앞으로 숙였고 철창이 오른쪽 배경의 좁은 조각보다 크게 등장합니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "앰버의 시선과 펼친 손은 화면 오른쪽 전방을 향합니다. 철창문은 어깨와 팔 뒤쪽에 있어, 뻗은 팔이 그 문을 향하는 공간 관계로 보이지 않습니다.",
        "built_space": "왼쪽 스테인드글라스 창 하나와 그 아래 난방기 하나, 위층 난간, 벗겨진 벽면, 중앙 뒤의 철제 우리 하나와 오른쪽 대형 철창문이 보입니다. 장소의 재료와 주요 구조는 참고 이미지에 가깝지만 철창문 전체에 가까운 면적과 바닥까지 드러나며, 요구된 오른쪽 가장자리의 좁고 비스듬한 배경 조각이 아닙니다. 문은 닫힌 상태로 보입니다.",
        "entities": "앰버는 금발의 어린 여자아이이며 큰 눈, 둥근 얼굴과 남색 반팔 상의가 참고와 대체로 맞습니다. 혼혈 정체성 자체는 외모만으로 확정할 수 없습니다. 입을 크게 벌리고 눈물을 흘립니다. 그러나 뒤에 별도 인물의 남색 옷을 입은 몸통과 앰버를 잡는 팔·손이 보입니다. 앰버의 포박용 끈은 보이지 않고 타인이 붙잡는 상태입니다. 찰리와 읽을 수 있는 글자는 보이지 않습니다.",
        "hard_violations": [
         "앰버만 등장하도록 제한했는데, 뒤에서 앰버의 팔과 몸을 붙잡는 추가 인물의 몸통과 팔·손이 등장합니다."
        ],
        "physics": "추가 인물의 손이 앰버의 위팔과 옆구리 부근에 닿아 붙잡고 있습니다. 뻗은 팔은 어깨에 정상적으로 연결되어 있으며 하체의 접지는 화면 밖이므로 공중에 떠 있다고 볼 근거는 없습니다. 다만 몸은 전방으로 기울어 있어 포박당해 뒤로 쏠린 상체라는 지정 동작과 다릅니다."
       },
       {
        "label": "A",
        "direction": "앰버의 시선과 뻗은 손은 화면 오른쪽의 철창문 및 그 너머를 향합니다. 손끝은 철창에 닿기 전의 간격을 남겨, 도달하지 못하는 목적지를 향해 손을 뻗는 관계가 읽힙니다.",
        "built_space": "오른쪽에 닫힌 철창문 한 면과 자물쇠 하나가 크게 보이고, 왼쪽 배경에는 큰 스테인드글라스 창 하나와 더 작은 창 구간, 손상된 벽과 바닥이 보입니다. 교회의 재료와 낮빛은 이어지지만 문이 화면 오른쪽 약 삼분의 일을 차지하는 가까운 구조물로 표현되어 좁은 배경 조각이라는 지시와 다릅니다. 참고보다 색광의 빛줄기도 훨씬 강조되었습니다.",
        "entities": "금발의 어린 여자아이 한 명만 보이며, 큰 눈과 둥근 얼굴, 남색 반팔 상의가 앰버 참고와 대체로 일치합니다. 혼혈 여부는 이미지에서 확정할 수 없습니다. 입을 크게 벌린 울음과 눈물이 보이고, 밧줄이 몸통과 뒤로 묶인 팔을 감싸 포박 상태를 명확히 나타냅니다. 찰리는 노출되지 않으며 읽을 수 있는 글자도 없습니다.",
        "hard_violations": [],
        "physics": "밧줄은 몸통과 뒤쪽 팔에 밀착되어 있고, 앞으로 뻗은 팔은 자유롭게 움직일 수 있는 형태입니다. 하체와 발은 화면 밖이므로 접지 방식은 확인할 수 없지만, 상체가 공중에 떠 있는 증거는 없습니다. 몸을 뒤로 당기는 외부 연결이나 장력은 보이지 않으며, 실제 자세는 뒤로 쏠림보다 문 쪽으로 앞으로 숙인 자세입니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gemini-pro"
   ],
   "route": "single_reverse"
  },
  "totals": {
   "B": 2,
   "A": 6
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2,
    "verdict_ko": "허용되지 않은 다른 사람의 몸과 손이 등장하며, 손을 뻗는 방향도 뒤쪽 철창문이 아니라 전방이고 상체 클로즈업보다 구도가 넓습니다."
   },
   {
    "label": "A",
    "score": 6,
    "verdict_ko": "앰버 혼자 포박된 채 닫힌 철창으로 손을 뻗는 관계는 맞지만, 뒤로 쏠린 순간 대신 앞으로 숙였고 철창이 오른쪽 배경의 좁은 조각보다 크게 등장합니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S58sh15_sel.png",
    "asset_id": "59fc65b2-08c9-415d-9b90-b66eaa6857d5",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-7386-75f9-add9-2e8e6c05a6b4",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S58sh15"
  }
 },
 "S58sh25::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:34:46.047151+00:00",
  "fingerprint": "78161d36099c09262cdd007434b603ee734ace2de634c7859d26758e47e73049",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S58sh25_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S58sh25_sel.png",
  "source_sha256": "50ef15ada13422a3ce8d562b5fe19cc62b43233ac58af101a3357eaf371e7b10",
  "file": "S58sh25_cine.png",
  "staged_sha256": "91e9b27f23b40b7f37cddff5a10913d3b3edd45d16ea8b1edccbff1e219cb326",
  "latency_ms": 10664
 },
 "S59sh10::signage": {
  "fp": "4ce5d70473e7a9fd",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::59af04c98c5db919": {
  "subjects": [],
  "subject_text": "익산 마을 지하감옥 감방\n굵은 쇠창살로 막힌 어두운 지하 감방. 바닥에는 건초더미와 고인 물이 있으며, 작은 창으로 제한적인 빛이 들어온다.",
  "identity": "canonical",
  "scope_id": "L222",
  "scope_role": "location_interior",
  "scope_sha": "e6618845247b342f"
 },
 "S59sh10::bgfirst_bg": {
  "input_fingerprint": "42c97c67a1138a9c",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 쇠창살 앞의 현우를 향해 둥근 감자를 내민 채 허공에 고정된 수빈의 손.\n\nLOCATION (lock): Inside a bare underground village-prison cell, beside its iron bars in the dim light of the dungeon.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Potato (Held out by 수빈, not yet accepted); used as Small focal object supported by the hand, with both bodies providing scale; Cell bars (Intact after 현우's attempts to force them) — Seen obliquely beside and behind 현우; both people remain on the same interior side; used as Connects the offered food to the immediate fact of confinement.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use ambient illumination appropriate to the cell with controlled contrast and readable skin detail, without specifying an unsupported source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 쇠창살 앞의 현우를 향해 둥근 감자를 내민 채 허공에 고정된 수빈의 손.\n\nLOCATION (lock): Inside a bare underground village-prison cell, beside its iron bars in the dim light of the dungeon.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Potato (Held out by 수빈, not yet accepted); used as Small focal object supported by the hand, with both bodies providing scale; Cell bars (Intact after 현우's attempts to force them) — Seen obliquely beside and behind 현우; both people remain on the same interior side; used as Connects the offered food to the immediate fact of confinement.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use ambient illumination appropriate to the cell with controlled contrast and readable skin detail, without specifying an unsupported source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S59sh10__bgfirst_bg.png",
  "asset_id": "9ab1887b-2da7-40ca-a980-d56e92af21ea",
  "input_asset_ids": [
   "0b9b8bdc-0f3c-40f7-a516-90ff65fd60d8",
   "47d011f7-2e24-447d-9f02-310f84be3871"
  ]
 },
 "S59sh10": {
  "input_fingerprint": "faa2300081c70940",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쇠창살 앞의 현우를 향해 둥근 감자를 내민 채 허공에 고정된 수빈의 손.\n\nLOCATION (lock): Inside a bare underground village-prison cell, beside its iron bars in the dim light of the dungeon. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Potato (Held out by 수빈, not yet accepted); used as Small focal object supported by the hand, with both bodies providing scale; Cell bars (Intact after 현우's attempts to force them) — Seen obliquely beside and behind 현우; both people remain on the same interior side; used as Connects the offered food to the immediate fact of confinement.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use ambient illumination appropriate to the cell with controlled contrast and readable skin detail, without specifying an unsupported source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The cell remains locked, with hay on the floor and pooled water inside; a handful of hay has already been wetted. A potato is being offered inside the cell. 수빈: Her face is scratched and dirty, and pre-existing radiation damage to her torso remains covered by her T-shirt. She holds out a potato.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 수빈 right now, so 수빈's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 수빈: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 수빈 (북한 출신 여성, 20세, 젊은 얼굴, 검은 단발머리, 날렵한 머리 끝선) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쇠창살 앞의 현우를 향해 둥근 감자를 내민 채 허공에 고정된 수빈의 손.\n\nLOCATION (lock): Inside a bare underground village-prison cell, beside its iron bars in the dim light of the dungeon. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Potato (Held out by 수빈, not yet accepted); used as Small focal object supported by the hand, with both bodies providing scale; Cell bars (Intact after 현우's attempts to force them) — Seen obliquely beside and behind 현우; both people remain on the same interior side; used as Connects the offered food to the immediate fact of confinement.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use ambient illumination appropriate to the cell with controlled contrast and readable skin detail, without specifying an unsupported source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The cell remains locked, with hay on the floor and pooled water inside; a handful of hay has already been wetted. A potato is being offered inside the cell. 수빈: Her face is scratched and dirty, and pre-existing radiation damage to her torso remains covered by her T-shirt. She holds out a potato.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 수빈 right now, so 수빈's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 수빈: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 수빈 (북한 출신 여성, 20세, 젊은 얼굴, 검은 단발머리, 날렵한 머리 끝선) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 쇠창살 앞의 현우를 향해 둥근 감자를 내민 채 허공에 고정된 수빈의 손.\n\nLOCATION (lock): Inside a bare underground village-prison cell, beside its iron bars in the dim light of the dungeon. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Potato (Held out by 수빈, not yet accepted); used as Small focal object supported by the hand, with both bodies providing scale; Cell bars (Intact after 현우's attempts to force them) — Seen obliquely beside and behind 현우; both people remain on the same interior side; used as Connects the offered food to the immediate fact of confinement.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use ambient illumination appropriate to the cell with controlled contrast and readable skin detail, without specifying an unsupported source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The cell remains locked, with hay on the floor and pooled water inside; a handful of hay has already been wetted. A potato is being offered inside the cell. 수빈: Her face is scratched and dirty, and pre-existing radiation damage to her torso remains covered by her T-shirt. She holds out a potato.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 수빈 right now, so 수빈's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 수빈: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 수빈 (북한 출신 여성, 20세, 젊은 얼굴, 검은 단발머리, 날렵한 머리 끝선) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S59sh10__bgfirst_bg.png",
     "asset_id": "9ab1887b-2da7-40ca-a980-d56e92af21ea",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S59sh10.png",
     "asset_id": "0b9b8bdc-0f3c-40f7-a516-90ff65fd60d8",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 수빈: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1175078>",
     "asset_id": "62c2787c-8e39-42b7-a857-baae4c7b8abe",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L222B02.png",
     "asset_id": "47d011f7-2e24-447d-9f02-310f84be3871",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 수빈: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1175078>",
     "asset_id": "62c2787c-8e39-42b7-a857-baae4c7b8abe",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "수빈은 쇠창살 안쪽에서 현우를 향해 감자를 쥔 손을 내밀고 있으며, 현우는 감자를 아직 받지 않은 채 시선을 아래로 향하고 있다.",
    "built_space": "물웅덩이와 짚단이 있는 지하던전 감방 내부. 쇠창살이 현우의 오른쪽 뒤편에 비스듬히 위치해 있으며, 수빈과 현우 모두 쇠창살 안쪽의 동일한 내부 공간에 머물고 있다.",
    "entities": "수빈은 참조 이미지의 단발머리 외형과 일치하며 얼굴에 흙과 상처가 있다. 현우는 젊은 남성으로 묘사되었고, 수빈의 손에는 둥근 감자가 들려 있다.",
    "hard_violations": [],
    "physics": "수빈은 바닥에 쪼그려 앉아 오른손으로 감자를 안정적으로 받쳐 들고 있고, 현우는 바닥에 주저앉아 체중을 지탱하고 있다. 공중에 떠 있거나 지지되지 않는 사물/신체는 없다."
   },
   {
    "label": "B",
    "direction": "수빈은 쇠창살 너머의 현우를 향해 팔을 뻗어 감자를 내밀고 있고, 현우는 쇠창살 밖에서 수빈을 응시하고 있다.",
    "built_space": "바닥에 짚단이 깔리고 웅덩이가 있는 감방. 쇠창살이 두 인물 사이를 가로막고 있어 수빈은 쇠창살 안쪽에, 현우는 쇠창살 바깥쪽에 배치되어 완전히 분리되어 있다.",
    "entities": "수빈은 단발머리 여성이며 얼굴에 긁힌 상처가 있다. 현우는 수염 자국이 있는 남성으로 나타나고, 수빈이 둥근 감자를 잡고 있다.",
    "hard_violations": [
     "[gemini-pro] 두 사람이 모두 쇠창살 안쪽의 같은 공간에 머물러야 한다는 프롬프트 지시(both people remain on the same interior side)를 어기고 현우를 쇠창살 바깥에 잘못 배치함.",
     "[gpt-high] 두 사람이 철창을 사이에 두고 반대편에 배치되어, 둘 다 동일한 감방 내부 쪽에 있어야 한다는 명시적 공간 조건을 위반한다."
    ],
    "physics": "수빈이 손가락 끝으로 감자를 쥐고 팔을 뻗고 있으며, 현우는 쪼그려 앉은 자세로 바닥에 지지되어 있다. 해부학적으로 불가능하거나 허공에 떠 있는 요소는 없다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "클로즈업 프레이밍 지시보다 다소 넓게 잡혔으나, 두 인물을 쇠창살 안쪽의 동일한 공간에 배치하라는 핵심적인 공간 지시를 정확히 구현하여 가장 프롬프트에 충실합니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "두 인물이 감방 안쪽의 같은 공간에 있어야 한다는 명시적인 지시를 어기고 쇠창살을 사이에 두고 분리되어 있어 치명적인 공간 배치 위반(Hard Violation)에 해당합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "수빈은 쇠창살 안쪽에서 현우를 향해 감자를 쥔 손을 내밀고 있으며, 현우는 감자를 아직 받지 않은 채 시선을 아래로 향하고 있다.",
        "built_space": "물웅덩이와 짚단이 있는 지하던전 감방 내부. 쇠창살이 현우의 오른쪽 뒤편에 비스듬히 위치해 있으며, 수빈과 현우 모두 쇠창살 안쪽의 동일한 내부 공간에 머물고 있다.",
        "entities": "수빈은 참조 이미지의 단발머리 외형과 일치하며 얼굴에 흙과 상처가 있다. 현우는 젊은 남성으로 묘사되었고, 수빈의 손에는 둥근 감자가 들려 있다.",
        "hard_violations": [],
        "physics": "수빈은 바닥에 쪼그려 앉아 오른손으로 감자를 안정적으로 받쳐 들고 있고, 현우는 바닥에 주저앉아 체중을 지탱하고 있다. 공중에 떠 있거나 지지되지 않는 사물/신체는 없다."
       },
       {
        "label": "B",
        "direction": "수빈은 쇠창살 너머의 현우를 향해 팔을 뻗어 감자를 내밀고 있고, 현우는 쇠창살 밖에서 수빈을 응시하고 있다.",
        "built_space": "바닥에 짚단이 깔리고 웅덩이가 있는 감방. 쇠창살이 두 인물 사이를 가로막고 있어 수빈은 쇠창살 안쪽에, 현우는 쇠창살 바깥쪽에 배치되어 완전히 분리되어 있다.",
        "entities": "수빈은 단발머리 여성이며 얼굴에 긁힌 상처가 있다. 현우는 수염 자국이 있는 남성으로 나타나고, 수빈이 둥근 감자를 잡고 있다.",
        "hard_violations": [
         "두 사람이 모두 쇠창살 안쪽의 같은 공간에 머물러야 한다는 프롬프트 지시(both people remain on the same interior side)를 어기고 현우를 쇠창살 바깥에 잘못 배치함."
        ],
        "physics": "수빈이 손가락 끝으로 감자를 쥐고 팔을 뻗고 있으며, 현우는 쪼그려 앉은 자세로 바닥에 지지되어 있다. 해부학적으로 불가능하거나 허공에 떠 있는 요소는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "클로즈업 프레이밍 지시보다 다소 넓게 잡혔으나, 두 인물을 쇠창살 안쪽의 동일한 공간에 배치하라는 핵심적인 공간 지시를 정확히 구현하여 가장 프롬프트에 충실합니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "두 인물이 감방 안쪽의 같은 공간에 있어야 한다는 명시적인 지시를 어기고 쇠창살을 사이에 두고 분리되어 있어 치명적인 공간 배치 위반(Hard Violation)에 해당합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "수빈은 쇠창살 안쪽에서 현우를 향해 감자를 쥔 손을 내밀고 있으며, 현우는 감자를 아직 받지 않은 채 시선을 아래로 향하고 있다.",
        "built_space": "물웅덩이와 짚단이 있는 지하던전 감방 내부. 쇠창살이 현우의 오른쪽 뒤편에 비스듬히 위치해 있으며, 수빈과 현우 모두 쇠창살 안쪽의 동일한 내부 공간에 머물고 있다.",
        "entities": "수빈은 참조 이미지의 단발머리 외형과 일치하며 얼굴에 흙과 상처가 있다. 현우는 젊은 남성으로 묘사되었고, 수빈의 손에는 둥근 감자가 들려 있다.",
        "hard_violations": [],
        "physics": "수빈은 바닥에 쪼그려 앉아 오른손으로 감자를 안정적으로 받쳐 들고 있고, 현우는 바닥에 주저앉아 체중을 지탱하고 있다. 공중에 떠 있거나 지지되지 않는 사물/신체는 없다."
       },
       {
        "label": "B",
        "direction": "수빈은 쇠창살 너머의 현우를 향해 팔을 뻗어 감자를 내밀고 있고, 현우는 쇠창살 밖에서 수빈을 응시하고 있다.",
        "built_space": "바닥에 짚단이 깔리고 웅덩이가 있는 감방. 쇠창살이 두 인물 사이를 가로막고 있어 수빈은 쇠창살 안쪽에, 현우는 쇠창살 바깥쪽에 배치되어 완전히 분리되어 있다.",
        "entities": "수빈은 단발머리 여성이며 얼굴에 긁힌 상처가 있다. 현우는 수염 자국이 있는 남성으로 나타나고, 수빈이 둥근 감자를 잡고 있다.",
        "hard_violations": [
         "두 사람이 모두 쇠창살 안쪽의 같은 공간에 머물러야 한다는 프롬프트 지시(both people remain on the same interior side)를 어기고 현우를 쇠창살 바깥에 잘못 배치함."
        ],
        "physics": "수빈이 손가락 끝으로 감자를 쥐고 팔을 뻗고 있으며, 현우는 쪼그려 앉은 자세로 바닥에 지지되어 있다. 해부학적으로 불가능하거나 허공에 떠 있는 요소는 없다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "감자를 현우에게 내미는 방향은 맞지만, 두 사람을 쇠창살 양쪽으로 갈라놓아 같은 감방 내부에 있어야 한다는 핵심 배치를 위반하며 손 클로즈업도 아니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "같은 감방 안에서 수빈의 손이 아직 받지 않은 감자를 현우에게 내미는 관계는 정확하지만, 손 중심 클로즈업보다 넓고 감자도 둥글기보다 길쭉하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 수빈의 팔과 감자는 오른쪽 현우의 얼굴 쪽을 향한다. 수빈은 현우를 바라보고 현우도 수빈과 감자가 있는 왼쪽을 본다. 현우는 감자에 손을 대지 않아 아직 받지 않은 순간은 맞다.",
        "built_space": "오른쪽에 온전한 세로 철봉 다수와 가로 보강대가 있는 철창 면 하나, 잠금장치 판 하나가 보인다. 그러나 철창이 두 사람 사이를 지나 수빈은 카메라 쪽, 현우는 철창 너머에 있다. 콘크리트 벽과 천장, 높은 채광 틈, 바닥의 짚과 고인 물은 장소의 재료와 분위기에 부합한다. 수빈의 상반신과 긴 팔 전체를 보여 주는 구도여서 요청한 손 중심 클로즈업보다 넓다.",
        "entities": "두 사람은 수빈과 현우로 읽히며 추가 인물이나 읽을 수 있는 글자는 없다. 수빈은 젊은 동아시아계 여성으로 검은 단발, 더러워지고 긁힌 얼굴, 몸통을 덮는 회색 반팔 티셔츠를 갖췄다. 참고의 짙은 남색 상의와는 다르지만 티셔츠라는 장면 조건은 맞는다. 손은 수빈의 팔에 연결되어 있으며 둥근 작은 감자 한 개를 들고 있다. 현우는 짧은 검은 머리의 동아시아계 남성으로 보인다.",
        "hard_violations": [
         "두 사람이 철창을 사이에 두고 반대편에 배치되어, 둘 다 동일한 감방 내부 쪽에 있어야 한다는 명시적 공간 조건을 위반한다."
        ],
        "physics": "감자는 수빈의 굽힌 손가락과 엄지로 받쳐져 있고 손목과 팔이 몸통까지 자연스럽게 이어진다. 팔을 든 채 멈춘 동작은 실제로 가능하며 감자가 홀로 떠 있지 않다. 하체의 지지점은 프레임 밖이지만 상체에 공중 부유나 불가능한 관절 배치는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "왼쪽 수빈의 손바닥과 감자는 오른쪽 현우의 몸 앞을 향해 뻗어 있다. 두 사람의 시선은 아래쪽 감자와 손 부근을 향한다. 현우의 손은 감자를 받으러 올라오지 않아 제안이 아직 수락되지 않은 순간으로 읽힌다.",
        "built_space": "오른쪽에서 뒤로 이어지는 철창 면 하나와 잠금장치 판 하나가 보이고, 철봉은 휘거나 끊어지지 않았다. 두 사람 모두 철창의 카메라 쪽 감방 내부에 있으며 현우의 옆과 뒤로 철창이 비스듬히 놓인다. 거친 콘크리트 벽·천장·바닥, 뒤편 짚더미 하나, 바닥 물웅덩이와 높은 채광 틈이 보인다. 다만 두 사람의 머리부터 무릎 부근까지 포함해 손 클로즈업보다 넓은 구도다.",
        "entities": "수빈과 현우로 읽히는 두 사람 외에 추가 인물은 없다. 수빈은 젊은 동아시아계 여성으로 검은 단발과 뺨의 긁힌 자국, 피부와 옷의 때를 갖췄다. 짙은 반팔 티셔츠가 몸통을 덮으며 참고의 얼굴·머리·체형 특징과 대체로 부합한다. 수빈의 가느다란 손과 팔이 감자 한 개를 받친다. 감자는 실제 감자 질감이지만 요구한 둥근 형태보다 타원형이다. 현우는 젊은 동아시아계 남성으로 보이며, 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "감자는 수빈의 손바닥 위에 놓이고 굽힌 손가락이 아래와 옆을 받친다. 손목·팔꿈치·어깨가 자연스럽게 연결되어 내민 손을 잠시 멈춘 자세가 가능하다. 두 사람은 무릎을 굽혀 낮게 앉은 자세이며 하체가 화면 아래로 이어진다. 발과 엉덩이의 접촉면은 잘렸지만 공중에 떠 있는 몸이나 지지 없이 떠 있는 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "감자를 현우에게 내미는 방향은 맞지만, 두 사람을 쇠창살 양쪽으로 갈라놓아 같은 감방 내부에 있어야 한다는 핵심 배치를 위반하며 손 클로즈업도 아니다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "같은 감방 안에서 수빈의 손이 아직 받지 않은 감자를 현우에게 내미는 관계는 정확하지만, 손 중심 클로즈업보다 넓고 감자도 둥글기보다 길쭉하다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽 수빈의 팔과 감자는 오른쪽 현우의 얼굴 쪽을 향한다. 수빈은 현우를 바라보고 현우도 수빈과 감자가 있는 왼쪽을 본다. 현우는 감자에 손을 대지 않아 아직 받지 않은 순간은 맞다.",
        "built_space": "오른쪽에 온전한 세로 철봉 다수와 가로 보강대가 있는 철창 면 하나, 잠금장치 판 하나가 보인다. 그러나 철창이 두 사람 사이를 지나 수빈은 카메라 쪽, 현우는 철창 너머에 있다. 콘크리트 벽과 천장, 높은 채광 틈, 바닥의 짚과 고인 물은 장소의 재료와 분위기에 부합한다. 수빈의 상반신과 긴 팔 전체를 보여 주는 구도여서 요청한 손 중심 클로즈업보다 넓다.",
        "entities": "두 사람은 수빈과 현우로 읽히며 추가 인물이나 읽을 수 있는 글자는 없다. 수빈은 젊은 동아시아계 여성으로 검은 단발, 더러워지고 긁힌 얼굴, 몸통을 덮는 회색 반팔 티셔츠를 갖췄다. 참고의 짙은 남색 상의와는 다르지만 티셔츠라는 장면 조건은 맞는다. 손은 수빈의 팔에 연결되어 있으며 둥근 작은 감자 한 개를 들고 있다. 현우는 짧은 검은 머리의 동아시아계 남성으로 보인다.",
        "hard_violations": [
         "두 사람이 철창을 사이에 두고 반대편에 배치되어, 둘 다 동일한 감방 내부 쪽에 있어야 한다는 명시적 공간 조건을 위반한다."
        ],
        "physics": "감자는 수빈의 굽힌 손가락과 엄지로 받쳐져 있고 손목과 팔이 몸통까지 자연스럽게 이어진다. 팔을 든 채 멈춘 동작은 실제로 가능하며 감자가 홀로 떠 있지 않다. 하체의 지지점은 프레임 밖이지만 상체에 공중 부유나 불가능한 관절 배치는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "왼쪽 수빈의 손바닥과 감자는 오른쪽 현우의 몸 앞을 향해 뻗어 있다. 두 사람의 시선은 아래쪽 감자와 손 부근을 향한다. 현우의 손은 감자를 받으러 올라오지 않아 제안이 아직 수락되지 않은 순간으로 읽힌다.",
        "built_space": "오른쪽에서 뒤로 이어지는 철창 면 하나와 잠금장치 판 하나가 보이고, 철봉은 휘거나 끊어지지 않았다. 두 사람 모두 철창의 카메라 쪽 감방 내부에 있으며 현우의 옆과 뒤로 철창이 비스듬히 놓인다. 거친 콘크리트 벽·천장·바닥, 뒤편 짚더미 하나, 바닥 물웅덩이와 높은 채광 틈이 보인다. 다만 두 사람의 머리부터 무릎 부근까지 포함해 손 클로즈업보다 넓은 구도다.",
        "entities": "수빈과 현우로 읽히는 두 사람 외에 추가 인물은 없다. 수빈은 젊은 동아시아계 여성으로 검은 단발과 뺨의 긁힌 자국, 피부와 옷의 때를 갖췄다. 짙은 반팔 티셔츠가 몸통을 덮으며 참고의 얼굴·머리·체형 특징과 대체로 부합한다. 수빈의 가느다란 손과 팔이 감자 한 개를 받친다. 감자는 실제 감자 질감이지만 요구한 둥근 형태보다 타원형이다. 현우는 젊은 동아시아계 남성으로 보이며, 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "감자는 수빈의 손바닥 위에 놓이고 굽힌 손가락이 아래와 옆을 받친다. 손목·팔꿈치·어깨가 자연스럽게 연결되어 내민 손을 잠시 멈춘 자세가 가능하다. 두 사람은 무릎을 굽혀 낮게 앉은 자세이며 하체가 화면 아래로 이어진다. 발과 엉덩이의 접촉면은 잘렸지만 공중에 떠 있는 몸이나 지지 없이 떠 있는 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.714
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.464
   },
   "violations": {
    "B": [
     "[gemini-pro] 두 사람이 모두 쇠창살 안쪽의 같은 공간에 머물러야 한다는 프롬프트 지시(both people remain on the same interior side)를 어기고 현우를 쇠창살 바깥에 잘못 배치함.",
     "[gpt-high] 두 사람이 철창을 사이에 두고 반대편에 배치되어, 둘 다 동일한 감방 내부 쪽에 있어야 한다는 명시적 공간 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 464
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "클로즈업 프레이밍 지시보다 다소 넓게 잡혔으나, 두 인물을 쇠창살 안쪽의 동일한 공간에 배치하라는 핵심적인 공간 지시를 정확히 구현하여 가장 프롬프트에 충실합니다."
   },
   {
    "label": "B",
    "score": 464,
    "verdict_ko": "두 인물이 감방 안쪽의 같은 공간에 있어야 한다는 명시적인 지시를 어기고 쇠창살을 사이에 두고 분리되어 있어 치명적인 공간 배치 위반(Hard Violation)에 해당합니다.  ★위반: [gemini-pro] 두 사람이 모두 쇠창살 안쪽의 같은 공간에 머물러야 한다는 프롬프트 지시(both people remain on the same interior side)를 어기고 현우를 쇠창살 바깥에 잘못 배치함. / [gpt-high] 두 사람이 철창을 사이에 두고 반대편에 배치되어, 둘 다 동일한 감방 내부 쪽에 있어야 한다는 명시적 공간 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L222B02.png",
    "asset_id": "47d011f7-2e24-447d-9f02-310f84be3871",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 수빈: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1175078>",
    "asset_id": "62c2787c-8e39-42b7-a857-baae4c7b8abe",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-7534-771f-b960-b8fa6d0a3358",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S59sh10__bgfirst_bg.png",
   "bg_asset_id": "9ab1887b-2da7-40ca-a980-d56e92af21ea",
   "bg_record_key": "S59sh10::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S59sh10::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:36:47.361785+00:00",
  "fingerprint": "d9ae21422b1cfaa1ed9962b11fe8bfd5acaf7710f2a62db384b0e5a488edcdd3",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S59sh10_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S59sh10_sel.png",
  "source_sha256": "79dbffb5d00473b3a5fcc73ce1f3757441c9c525bd483333cff3f20a0e8ddb8e",
  "file": "S59sh10_cine.png",
  "staged_sha256": "1cc371fc37db43e09b40e690391ac630111dfffc67a030a49a4a58cac2286fde",
  "latency_ms": 11216
 },
 "S59sh18::signage": {
  "fp": "62bd332904ce1413",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S59sh18": {
  "input_fingerprint": "b09565d0b0e0a4c1",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 수빈의 배 위에 선명하게 드러난 검붉고 흉측한 방사능 피폭 흉터 클로즈업.\n\nLOCATION (lock): Inside the sparsely furnished underground prison cell, in subdued dungeon light near the other prisoners. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 수빈의 티셔츠 (Lifted to expose the abdominal radiation injuries) — The raised hem crosses the upper edge of the obliquely viewed torso; used as Retains the revealing action within the close-up rather than isolating the injury as an abstract texture.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained ambient illumination appropriate to the cell, keeping the dark-red scars legible without sensational contrast or added glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 수빈 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The underground cell remains locked, with hay on the floor and pooled water inside. 수빈: Her face is scarred and dirty, and her raised T-shirt exposes radiation-damaged skin across her abdomen.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 수빈 (북한 출신 여성, 20세, 젊은 얼굴, 검은 단발머리, 날렵한 머리 끝선) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 수빈의 배 위에 선명하게 드러난 검붉고 흉측한 방사능 피폭 흉터 클로즈업.\n\nLOCATION (lock): Inside the sparsely furnished underground prison cell, in subdued dungeon light near the other prisoners. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 수빈의 티셔츠 (Lifted to expose the abdominal radiation injuries) — The raised hem crosses the upper edge of the obliquely viewed torso; used as Retains the revealing action within the close-up rather than isolating the injury as an abstract texture.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained ambient illumination appropriate to the cell, keeping the dark-red scars legible without sensational contrast or added glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 수빈 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The underground cell remains locked, with hay on the floor and pooled water inside. 수빈: Her face is scarred and dirty, and her raised T-shirt exposes radiation-damaged skin across her abdomen.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 수빈 (북한 출신 여성, 20세, 젊은 얼굴, 검은 단발머리, 날렵한 머리 끝선) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 수빈의 배 위에 선명하게 드러난 검붉고 흉측한 방사능 피폭 흉터 클로즈업.\n\nLOCATION (lock): Inside the sparsely furnished underground prison cell, in subdued dungeon light near the other prisoners. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 수빈의 티셔츠 (Lifted to expose the abdominal radiation injuries) — The raised hem crosses the upper edge of the obliquely viewed torso; used as Retains the revealing action within the close-up rather than isolating the injury as an abstract texture.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained ambient illumination appropriate to the cell, keeping the dark-red scars legible without sensational contrast or added glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 수빈 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The underground cell remains locked, with hay on the floor and pooled water inside. 수빈: Her face is scarred and dirty, and her raised T-shirt exposes radiation-damaged skin across her abdomen.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 수빈 (북한 출신 여성, 20세, 젊은 얼굴, 검은 단발머리, 날렵한 머리 끝선) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "수빈의 손이 입고 있는 티셔츠의 밑단을 위로 잡아당기고 있으며, 렌즈는 복부의 흉터를 직관적으로 향하고 있음.",
    "built_space": "이전 샷에서 확립된 감옥 내부로, 바닥의 물웅덩이, 뒤쪽의 지푸라기 더미, 오른쪽의 쇠창살이 정확한 위치와 스케일로 나타남.",
    "entities": "수빈의 복부와 짙은 회색 티셔츠. 검붉고 흉측한 방사능 피폭 흉터가 명확히 묘사됨. 프롬프트 지시대로 수빈 외의 다른 인물은 등장하지 않음.",
    "hard_violations": [],
    "physics": "손이 셔츠 자락을 단단히 쥐고 있어 옷이 들려진 상태를 자연스럽게 지탱하고 있으며, 몸의 각도와 무게 중심이 사실적임."
   },
   {
    "label": "B",
    "direction": "수빈의 손이 셔츠를 잡아당겨 복부를 드러내고 있음.",
    "built_space": "감옥 내부의 돌벽과 오른쪽의 쇠창살이 보이나 공간의 깊이감이 다소 평면적임.",
    "entities": "수빈의 복부와 흉터, 회색 티셔츠. 그러나 배경 좌측과 우측에 샷 텍스트에 없는 추가 인물 2명이 뚜렷하게 존재함(특히 좌측 인물은 배제 지시가 있던 이전 샷의 인물과 의상 및 외형이 흡사함).",
    "hard_violations": [
     "[gemini-pro] 샷 텍스트에 명시되지 않은 인물 추가 (PEOPLE 조항의 'never add a person the shot text does not show' 위반)",
     "[gemini-pro] 이전 샷에 등장한 인물(또는 그와 흡사한 인물)을 다시 등장시킴 ('Anyone else visible in that photo is NOT in this shot' 위반)",
     "[gpt-high] 수빈 이외에는 누구도 등장시키지 말라는 지시와 달리, 배경에 다른 인물 두 명의 신체와 의상이 보인다."
    ],
    "physics": "손이 옷을 잡고 있는 형태는 유지되나, 배경 인물들의 포즈가 다소 어색하게 굳어 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 10,
        "verdict_ko": "클로즈업 샷의 프레이밍을 정확히 준수하며, 텍스트에 없는 다른 인물을 배제하고 수빈의 상처와 의상 및 감옥의 배경(물웅덩이, 지푸라기, 쇠창살)을 레퍼런스와 일치하게 완벽히 구현했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "수빈의 흉터는 표현되었으나, 샷 텍스트에 명시되지 않은 다른 인물들을 배경에 추가하여 엄격한 인물 제한 지시를 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "수빈의 손이 입고 있는 티셔츠의 밑단을 위로 잡아당기고 있으며, 렌즈는 복부의 흉터를 직관적으로 향하고 있음.",
        "built_space": "이전 샷에서 확립된 감옥 내부로, 바닥의 물웅덩이, 뒤쪽의 지푸라기 더미, 오른쪽의 쇠창살이 정확한 위치와 스케일로 나타남.",
        "entities": "수빈의 복부와 짙은 회색 티셔츠. 검붉고 흉측한 방사능 피폭 흉터가 명확히 묘사됨. 프롬프트 지시대로 수빈 외의 다른 인물은 등장하지 않음.",
        "hard_violations": [],
        "physics": "손이 셔츠 자락을 단단히 쥐고 있어 옷이 들려진 상태를 자연스럽게 지탱하고 있으며, 몸의 각도와 무게 중심이 사실적임."
       },
       {
        "label": "B",
        "direction": "수빈의 손이 셔츠를 잡아당겨 복부를 드러내고 있음.",
        "built_space": "감옥 내부의 돌벽과 오른쪽의 쇠창살이 보이나 공간의 깊이감이 다소 평면적임.",
        "entities": "수빈의 복부와 흉터, 회색 티셔츠. 그러나 배경 좌측과 우측에 샷 텍스트에 없는 추가 인물 2명이 뚜렷하게 존재함(특히 좌측 인물은 배제 지시가 있던 이전 샷의 인물과 의상 및 외형이 흡사함).",
        "hard_violations": [
         "샷 텍스트에 명시되지 않은 인물 추가 (PEOPLE 조항의 'never add a person the shot text does not show' 위반)",
         "이전 샷에 등장한 인물(또는 그와 흡사한 인물)을 다시 등장시킴 ('Anyone else visible in that photo is NOT in this shot' 위반)"
        ],
        "physics": "손이 옷을 잡고 있는 형태는 유지되나, 배경 인물들의 포즈가 다소 어색하게 굳어 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 10,
        "verdict_ko": "클로즈업 샷의 프레이밍을 정확히 준수하며, 텍스트에 없는 다른 인물을 배제하고 수빈의 상처와 의상 및 감옥의 배경(물웅덩이, 지푸라기, 쇠창살)을 레퍼런스와 일치하게 완벽히 구현했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "수빈의 흉터는 표현되었으나, 샷 텍스트에 명시되지 않은 다른 인물들을 배경에 추가하여 엄격한 인물 제한 지시를 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "수빈의 손이 입고 있는 티셔츠의 밑단을 위로 잡아당기고 있으며, 렌즈는 복부의 흉터를 직관적으로 향하고 있음.",
        "built_space": "이전 샷에서 확립된 감옥 내부로, 바닥의 물웅덩이, 뒤쪽의 지푸라기 더미, 오른쪽의 쇠창살이 정확한 위치와 스케일로 나타남.",
        "entities": "수빈의 복부와 짙은 회색 티셔츠. 검붉고 흉측한 방사능 피폭 흉터가 명확히 묘사됨. 프롬프트 지시대로 수빈 외의 다른 인물은 등장하지 않음.",
        "hard_violations": [],
        "physics": "손이 셔츠 자락을 단단히 쥐고 있어 옷이 들려진 상태를 자연스럽게 지탱하고 있으며, 몸의 각도와 무게 중심이 사실적임."
       },
       {
        "label": "B",
        "direction": "수빈의 손이 셔츠를 잡아당겨 복부를 드러내고 있음.",
        "built_space": "감옥 내부의 돌벽과 오른쪽의 쇠창살이 보이나 공간의 깊이감이 다소 평면적임.",
        "entities": "수빈의 복부와 흉터, 회색 티셔츠. 그러나 배경 좌측과 우측에 샷 텍스트에 없는 추가 인물 2명이 뚜렷하게 존재함(특히 좌측 인물은 배제 지시가 있던 이전 샷의 인물과 의상 및 외형이 흡사함).",
        "hard_violations": [
         "샷 텍스트에 명시되지 않은 인물 추가 (PEOPLE 조항의 'never add a person the shot text does not show' 위반)",
         "이전 샷에 등장한 인물(또는 그와 흡사한 인물)을 다시 등장시킴 ('Anyone else visible in that photo is NOT in this shot' 위반)"
        ],
        "physics": "손이 옷을 잡고 있는 형태는 유지되나, 배경 인물들의 포즈가 다소 어색하게 굳어 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "검붉은 복부 손상과 티셔츠를 들어 올리는 동작은 보이지만, 금지된 다른 인물 두 명이 등장하며 흉터보다 상체와 배경을 넓게 담아 결정적으로 불충실하다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "수빈의 복부 흉터와 들어 올린 티셔츠를 비스듬한 클로즈업으로 담고 감방의 재질과 조명을 유지하며, 흉터 색이 다소 옅은 점만 아쉽다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "수빈의 얼굴은 프레임 밖이라 시선은 보이지 않는다. 화면 위쪽 손은 티셔츠를 위로 당겨 카메라 쪽으로 비스듬히 향한 복부를 드러낸다. 뒤쪽 두 인물의 시선은 가림과 흐림 때문에 확인할 수 없다. 무기나 겨냥하는 물체는 없다.",
        "built_space": "왼쪽의 낡은 벽, 뒤쪽 벽면과 모서리, 오른쪽 철창 구획 하나, 바닥의 짚이 보인다. 고정 설비의 중복은 보이지 않지만, 수빈 뒤 왼쪽에는 앉은 인물 한 명, 중앙에는 몸을 숙인 인물 한 명이 있어 수빈만 보여야 하는 구성을 위반한다. 감방 배경과 상체가 차지하는 면적도 커서 흉터 중심의 클로즈업이 느슨하다.",
        "entities": "전경에는 마른 체형의 인물 한 명의 복부·팔·손과 이전 장면에 부합하는 더러운 짙은 회색 티셔츠가 보인다. 얼굴과 머리가 없어 수빈의 정확한 나이·민족적 외양·얼굴 정체성은 확인할 수 없다. 복부에는 검붉고 울퉁불퉁한 손상이 있으나 넓게 퍼진 흉터보다는 국소적인 열린 상처처럼 보인다. 배경에는 밝은 상의를 입은 추가 인물 두 명과 바닥의 어두운 천 덩어리가 보인다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "수빈 이외에는 누구도 등장시키지 말라는 지시와 달리, 배경에 다른 인물 두 명의 신체와 의상이 보인다."
        ],
        "physics": "손가락이 티셔츠를 실제로 움켜쥐고 있으며 천 주름이 잡아당기는 지점으로 모여 들어 올리는 동작은 성립한다. 전경 몸통은 아래 프레임 밖으로 이어져 부유한 자세가 아니다. 배경 왼쪽 인물은 하체가 가려져 지지 접점을 확인하기 어렵고, 중앙 인물은 다리가 바닥까지 이어져 몸을 숙인 자세로 읽힌다. 바닥의 천은 바닥에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "얼굴과 눈은 프레임 밖이다. 오른쪽 위의 손이 티셔츠를 위쪽으로 당기고, 노출된 복부는 카메라를 향해 약간 비스듬히 놓여 있다. 복부 손상과 그 위를 가로지르는 옷자락이 함께 보여 요구된 드러내기 동작이 명확하다. 무기나 지향성 소품은 없다.",
        "built_space": "왼쪽의 거친 벽면 하나, 뒤쪽 벽면, 오른쪽 철창 구획 하나가 흐린 배경으로 보인다. 바닥에는 짚과 고인 물이 있으며 기존 감방의 재질과 절제된 낮 주변광에 부합한다. 물의 밝은 반사도 바닥 표면에서 자연스럽게 나타난다. 추가 인물이나 중복 설비는 없고, 복부 중심의 가까운 프레임을 유지한다.",
        "entities": "수빈으로 제시된 한 인물의 마른 복부, 팔, 옷을 쥔 손이 보인다. 더럽고 해진 짙은 회색 티셔츠는 이전 장면과 일치한다. 얼굴과 검은 단발머리는 적절히 프레임 밖에 있어 정확한 얼굴 정체성이나 나이·출신은 판별할 수 없다. 복부 전반에 적갈색과 자주색의 불규칙한 융기성 흉터가 퍼져 있으며 피부 곡면과 질감을 따른다. 요구한 검붉은 색보다는 일부가 옅지만 흉터 자체는 선명하다. 다른 사람이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "손이 티셔츠를 단단히 쥐어 들어 올리고 천은 그 손에서 아래로 접혀 늘어진다. 반대쪽 팔은 몸 옆으로 내려가며 몸통과 자연스럽게 연결된다. 하체의 지지 접점은 클로즈업 밖이지만 몸이 공중에 떠 있다는 단서는 없다. 흉터는 피부에 붙어 곡면을 따라가고, 짚과 물은 바닥에 놓여 있어 지지나 접촉의 모순이 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "검붉은 복부 손상과 티셔츠를 들어 올리는 동작은 보이지만, 금지된 다른 인물 두 명이 등장하며 흉터보다 상체와 배경을 넓게 담아 결정적으로 불충실하다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "수빈의 복부 흉터와 들어 올린 티셔츠를 비스듬한 클로즈업으로 담고 감방의 재질과 조명을 유지하며, 흉터 색이 다소 옅은 점만 아쉽다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "수빈의 얼굴은 프레임 밖이라 시선은 보이지 않는다. 화면 위쪽 손은 티셔츠를 위로 당겨 카메라 쪽으로 비스듬히 향한 복부를 드러낸다. 뒤쪽 두 인물의 시선은 가림과 흐림 때문에 확인할 수 없다. 무기나 겨냥하는 물체는 없다.",
        "built_space": "왼쪽의 낡은 벽, 뒤쪽 벽면과 모서리, 오른쪽 철창 구획 하나, 바닥의 짚이 보인다. 고정 설비의 중복은 보이지 않지만, 수빈 뒤 왼쪽에는 앉은 인물 한 명, 중앙에는 몸을 숙인 인물 한 명이 있어 수빈만 보여야 하는 구성을 위반한다. 감방 배경과 상체가 차지하는 면적도 커서 흉터 중심의 클로즈업이 느슨하다.",
        "entities": "전경에는 마른 체형의 인물 한 명의 복부·팔·손과 이전 장면에 부합하는 더러운 짙은 회색 티셔츠가 보인다. 얼굴과 머리가 없어 수빈의 정확한 나이·민족적 외양·얼굴 정체성은 확인할 수 없다. 복부에는 검붉고 울퉁불퉁한 손상이 있으나 넓게 퍼진 흉터보다는 국소적인 열린 상처처럼 보인다. 배경에는 밝은 상의를 입은 추가 인물 두 명과 바닥의 어두운 천 덩어리가 보인다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "수빈 이외에는 누구도 등장시키지 말라는 지시와 달리, 배경에 다른 인물 두 명의 신체와 의상이 보인다."
        ],
        "physics": "손가락이 티셔츠를 실제로 움켜쥐고 있으며 천 주름이 잡아당기는 지점으로 모여 들어 올리는 동작은 성립한다. 전경 몸통은 아래 프레임 밖으로 이어져 부유한 자세가 아니다. 배경 왼쪽 인물은 하체가 가려져 지지 접점을 확인하기 어렵고, 중앙 인물은 다리가 바닥까지 이어져 몸을 숙인 자세로 읽힌다. 바닥의 천은 바닥에 놓여 있다."
       },
       {
        "label": "A",
        "direction": "얼굴과 눈은 프레임 밖이다. 오른쪽 위의 손이 티셔츠를 위쪽으로 당기고, 노출된 복부는 카메라를 향해 약간 비스듬히 놓여 있다. 복부 손상과 그 위를 가로지르는 옷자락이 함께 보여 요구된 드러내기 동작이 명확하다. 무기나 지향성 소품은 없다.",
        "built_space": "왼쪽의 거친 벽면 하나, 뒤쪽 벽면, 오른쪽 철창 구획 하나가 흐린 배경으로 보인다. 바닥에는 짚과 고인 물이 있으며 기존 감방의 재질과 절제된 낮 주변광에 부합한다. 물의 밝은 반사도 바닥 표면에서 자연스럽게 나타난다. 추가 인물이나 중복 설비는 없고, 복부 중심의 가까운 프레임을 유지한다.",
        "entities": "수빈으로 제시된 한 인물의 마른 복부, 팔, 옷을 쥔 손이 보인다. 더럽고 해진 짙은 회색 티셔츠는 이전 장면과 일치한다. 얼굴과 검은 단발머리는 적절히 프레임 밖에 있어 정확한 얼굴 정체성이나 나이·출신은 판별할 수 없다. 복부 전반에 적갈색과 자주색의 불규칙한 융기성 흉터가 퍼져 있으며 피부 곡면과 질감을 따른다. 요구한 검붉은 색보다는 일부가 옅지만 흉터 자체는 선명하다. 다른 사람이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "손이 티셔츠를 단단히 쥐어 들어 올리고 천은 그 손에서 아래로 접혀 늘어진다. 반대쪽 팔은 몸 옆으로 내려가며 몸통과 자연스럽게 연결된다. 하체의 지지 접점은 클로즈업 밖이지만 몸이 공중에 떠 있다는 단서는 없다. 흉터는 피부에 붙어 곡면을 따라가고, 짚과 물은 바닥에 놓여 있어 지지나 접촉의 모순이 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.622
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.372
   },
   "violations": {
    "B": [
     "[gemini-pro] 샷 텍스트에 명시되지 않은 인물 추가 (PEOPLE 조항의 'never add a person the shot text does not show' 위반)",
     "[gemini-pro] 이전 샷에 등장한 인물(또는 그와 흡사한 인물)을 다시 등장시킴 ('Anyone else visible in that photo is NOT in this shot' 위반)",
     "[gpt-high] 수빈 이외에는 누구도 등장시키지 말라는 지시와 달리, 배경에 다른 인물 두 명의 신체와 의상이 보인다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 372
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "클로즈업 샷의 프레이밍을 정확히 준수하며, 텍스트에 없는 다른 인물을 배제하고 수빈의 상처와 의상 및 감옥의 배경(물웅덩이, 지푸라기, 쇠창살)을 레퍼런스와 일치하게 완벽히 구현했습니다."
   },
   {
    "label": "B",
    "score": 372,
    "verdict_ko": "수빈의 흉터는 표현되었으나, 샷 텍스트에 명시되지 않은 다른 인물들을 배경에 추가하여 엄격한 인물 제한 지시를 위반했습니다.  ★위반: [gemini-pro] 샷 텍스트에 명시되지 않은 인물 추가 (PEOPLE 조항의 'never add a person the shot text does not show' 위반) / [gemini-pro] 이전 샷에 등장한 인물(또는 그와 흡사한 인물)을 다시 등장시킴 ('Anyone else visible in that photo is NOT in this shot' 위반) / [gpt-high] 수빈 이외에는 누구도 등장시키지 말라는 지시와 달리, 배경에 다른 인물 두 명의 신체와 의상이 보인다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 수빈 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S59sh10_sel.png",
    "asset_id": "b3efd36d-5586-4be1-9c28-baec4361119e",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 수빈: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1175078>",
    "asset_id": "62c2787c-8e39-42b7-a857-baae4c7b8abe",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-7875-7a85-8064-1cce07e82d23",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S59sh10"
  }
 },
 "S59sh18::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T09:43:12.338727+00:00",
  "fingerprint": "f66e850cb46db3f032dc1841fe3237282676fc356bf3f093bbbb2accfcde6c04",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S59sh18_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S59sh18_sel.png",
  "source_sha256": "aa0ba031deb0ab45afecededa9714c8b0cd865cfc2ef62634d42223584c0f2f3",
  "file": "S59sh18_cine.png",
  "staged_sha256": "3d711031b61b111363f9bef922ac901c89b130726da82ca82c8eaf4a3288ec84",
  "latency_ms": 9526
 },
 "S59sh36::signage": {
  "fp": "d852448c9d03c41c",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S59sh36": {
  "input_fingerprint": "a14dfd6ec0ba32e4",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): (회상/판화) 뚱뚱해진 체형에 번쩍이는 금빛 하회탈을 쓴 채 군중을 굽어보는 백산의 위압적인 목판화 질감 그림.\n\nLOCATION (lock): A stylized woodcut flashback of the village ruler elevated above a crowd; no specific architectural setting is established. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: foreground crowd within the woodcut in the lower-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: 금빛 하회탈 (Intact and worn by the heavyset 백산) — The carved facial front is seen obliquely from below, inclined toward the foreground crowd; used as Provides a compact emblem of authority within the larger figure; 판화 속 전경 군중 (Gathered beneath 백산 within the recollection) — Unevenly overlapping backs and partial profiles face inward toward 백산, with varied head tilts and shoulder levels; used as Preserves the human scale and the upward relationship that gives the print its oppressive force.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Express the remembered scene through controlled woodcut light-and-dark blocks, reserving selective gold brilliance for the mask rather than introducing a realistic light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The woodcut-style flashback depicts Baeksan with an increasingly heavy body and a gold Hahoe mask; the village water has been concentrated in a tower as part of the same illustrated sequence.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 백산 (한국인, 성인, 남성 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): (회상/판화) 뚱뚱해진 체형에 번쩍이는 금빛 하회탈을 쓴 채 군중을 굽어보는 백산의 위압적인 목판화 질감 그림.\n\nLOCATION (lock): A stylized woodcut flashback of the village ruler elevated above a crowd; no specific architectural setting is established. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: foreground crowd within the woodcut in the lower-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: 금빛 하회탈 (Intact and worn by the heavyset 백산) — The carved facial front is seen obliquely from below, inclined toward the foreground crowd; used as Provides a compact emblem of authority within the larger figure; 판화 속 전경 군중 (Gathered beneath 백산 within the recollection) — Unevenly overlapping backs and partial profiles face inward toward 백산, with varied head tilts and shoulder levels; used as Preserves the human scale and the upward relationship that gives the print its oppressive force.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Express the remembered scene through controlled woodcut light-and-dark blocks, reserving selective gold brilliance for the mask rather than introducing a realistic light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The woodcut-style flashback depicts Baeksan with an increasingly heavy body and a gold Hahoe mask; the village water has been concentrated in a tower as part of the same illustrated sequence.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 백산 (한국인, 성인, 남성 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): (회상/판화) 뚱뚱해진 체형에 번쩍이는 금빛 하회탈을 쓴 채 군중을 굽어보는 백산의 위압적인 목판화 질감 그림.\n\nLOCATION (lock): A stylized woodcut flashback of the village ruler elevated above a crowd; no specific architectural setting is established. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: foreground crowd within the woodcut in the lower-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: 금빛 하회탈 (Intact and worn by the heavyset 백산) — The carved facial front is seen obliquely from below, inclined toward the foreground crowd; used as Provides a compact emblem of authority within the larger figure; 판화 속 전경 군중 (Gathered beneath 백산 within the recollection) — Unevenly overlapping backs and partial profiles face inward toward 백산, with varied head tilts and shoulder levels; used as Preserves the human scale and the upward relationship that gives the print its oppressive force.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Express the remembered scene through controlled woodcut light-and-dark blocks, reserving selective gold brilliance for the mask rather than introducing a realistic light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The woodcut-style flashback depicts Baeksan with an increasingly heavy body and a gold Hahoe mask; the village water has been concentrated in a tower as part of the same illustrated sequence.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 백산 (한국인, 성인, 남성 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S59sh36__bgfirst_bg.png",
     "asset_id": "c373b774-4997-4590-9fd6-e9ba287393af",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S59sh36.png",
     "asset_id": "f5ca6eee-4fd8-46c4-909f-e7354ed6208f",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 백산: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1241731>",
     "asset_id": "b337b8d8-94d9-4a29-9e49-2e19121379b7",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 백산의 지배를 묘사한 목판화: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:1581563>",
     "asset_id": "e7bb9df0-db9f-4c6d-9c8a-088105423f0f",
     "role": "prop_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L222B01.png",
     "asset_id": "dfd97579-913b-4622-86a5-0d6cd32f2ea6",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 백산: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1241731>",
     "asset_id": "b337b8d8-94d9-4a29-9e49-2e19121379b7",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 백산의 지배를 묘사한 목판화: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:1581563>",
     "asset_id": "e7bb9df0-db9f-4c6d-9c8a-088105423f0f",
     "role": "prop_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "전경의 군중은 단상 위의 백산을 향해 시선을 던지고, 백산은 아래의 군중을 내려다봄.",
    "built_space": "지정된 나무 벽면 로케이션을 무시하고, 임의의 야외 목조 단상과 2개의 깃대가 세워진 3D 공간을 연출함.",
    "entities": "금빛 하회탈을 쓴 육중한 백산과 다양한 뒷모습을 한 전경의 군중이 묘사됨.",
    "hard_violations": [],
    "physics": "인물들은 바닥과 단상에 정상적으로 서 있으며, 깃발은 깃대에 매달려 있음."
   },
   {
    "label": "B",
    "direction": "그림 내부의 군중들이 중앙에 있는 백산을 향해 시선을 모으고 있음.",
    "built_space": "제시된 로케이션(거친 나무 벽면) 위에 그림이 그려진 사각 종이가 붙어 있는 형태임.",
    "entities": "종이 그림 속에 금빛 하회탈을 쓴 백산과 여러 군중의 모습이 묘사됨.",
    "hard_violations": [
     "[gemini-pro] 절대 금지된 텍스트(프롬프트 지시문)가 이미지 하단에 캡션으로 유출됨",
     "[gpt-high] 판화 하단에 촬영 지시문을 읽을 수 있는 한국어 문장으로 삽입해, 글자·자막·지시문 노출 금지를 위반했다."
    ],
    "physics": "종이 포스터가 나무 벽면에 밀착되어 지탱되고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "지정된 로케이션(나무 벽면)을 무시하고 야외 단상을 임의로 연출했으나, 텍스트 유출과 같은 치명적 위반이 없어 우선됨."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "지시문 텍스트가 이미지 하단에 캡션으로 그대로 노출되는 치명적인 텍스트 유출 위반(Hard Violation)이 발생함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "전경의 군중은 단상 위의 백산을 향해 시선을 던지고, 백산은 아래의 군중을 내려다봄.",
        "built_space": "지정된 나무 벽면 로케이션을 무시하고, 임의의 야외 목조 단상과 2개의 깃대가 세워진 3D 공간을 연출함.",
        "entities": "금빛 하회탈을 쓴 육중한 백산과 다양한 뒷모습을 한 전경의 군중이 묘사됨.",
        "hard_violations": [],
        "physics": "인물들은 바닥과 단상에 정상적으로 서 있으며, 깃발은 깃대에 매달려 있음."
       },
       {
        "label": "B",
        "direction": "그림 내부의 군중들이 중앙에 있는 백산을 향해 시선을 모으고 있음.",
        "built_space": "제시된 로케이션(거친 나무 벽면) 위에 그림이 그려진 사각 종이가 붙어 있는 형태임.",
        "entities": "종이 그림 속에 금빛 하회탈을 쓴 백산과 여러 군중의 모습이 묘사됨.",
        "hard_violations": [
         "절대 금지된 텍스트(프롬프트 지시문)가 이미지 하단에 캡션으로 유출됨"
        ],
        "physics": "종이 포스터가 나무 벽면에 밀착되어 지탱되고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "지정된 로케이션(나무 벽면)을 무시하고 야외 단상을 임의로 연출했으나, 텍스트 유출과 같은 치명적 위반이 없어 우선됨."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "지시문 텍스트가 이미지 하단에 캡션으로 그대로 노출되는 치명적인 텍스트 유출 위반(Hard Violation)이 발생함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "전경의 군중은 단상 위의 백산을 향해 시선을 던지고, 백산은 아래의 군중을 내려다봄.",
        "built_space": "지정된 나무 벽면 로케이션을 무시하고, 임의의 야외 목조 단상과 2개의 깃대가 세워진 3D 공간을 연출함.",
        "entities": "금빛 하회탈을 쓴 육중한 백산과 다양한 뒷모습을 한 전경의 군중이 묘사됨.",
        "hard_violations": [],
        "physics": "인물들은 바닥과 단상에 정상적으로 서 있으며, 깃발은 깃대에 매달려 있음."
       },
       {
        "label": "B",
        "direction": "그림 내부의 군중들이 중앙에 있는 백산을 향해 시선을 모으고 있음.",
        "built_space": "제시된 로케이션(거친 나무 벽면) 위에 그림이 그려진 사각 종이가 붙어 있는 형태임.",
        "entities": "종이 그림 속에 금빛 하회탈을 쓴 백산과 여러 군중의 모습이 묘사됨.",
        "hard_violations": [
         "절대 금지된 텍스트(프롬프트 지시문)가 이미지 하단에 캡션으로 유출됨"
        ],
        "physics": "종이 포스터가 나무 벽면에 밀착되어 지탱되고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "하단에 지시문을 읽을 수 있는 자막으로 그대로 노출한 중대 위반이 있으며, 군중을 내려다보는 비스듬한 가면 대신 정면 가면과 판화 소품 자체를 보여준다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "낮은 시점의 와이드 구도와 높은 백산을 올려다보는 전경 군중의 관계는 더 정확하지만, 인물이 실사 인간이 아닌 평면 판화처럼 보이고 자연광 배경도 지정된 목판 명암 표현과 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "군중의 등과 좌우 옆얼굴은 중앙 백산을 향하며 일부는 고개를 올리고 있다. 반면 백산의 가면은 거의 정면으로 카메라를 향해 있어, 전경 군중을 향해 아래로 기울어진 조각 면이라는 지시는 약하게 구현됐다. 무기나 이동 동작은 없다.",
        "built_space": "거친 검은 목재 표면 위에 테두리가 있는 판화 한 장이 놓여 있다. 판화 내부에는 좌우의 여러 가옥과 왼쪽 뒤 원통형 물탱크 한 개가 보인다. 백산은 중앙 군중 뒤에 있지만 높이를 제공하는 단이나 지면은 보이지 않는다. 목재 바탕은 장소 참조와 가깝지만, 장면 안의 와이드 숏보다 판화 소품을 내려다보는 촬영에 가깝다.",
        "entities": "배가 크게 나온 성인 남성 백산 한 명과 다수의 군중이 판화로 표현됐다. 백산은 얼굴 전체를 덮는 온전한 금빛 웃는 탈과 어두운 겹옷을 착용한다. 탈 때문에 참조 얼굴의 동일성은 확인할 수 없으며, 가려진 얼굴을 노출하지는 않았다. 군중은 등과 옆얼굴, 옷 주름이 구별되지만 실사 인간은 아니다. 물탱크는 보이나 물 자체는 보이지 않는다. 하단에는 한국어 지시문이 명확하게 읽힌다.",
        "hard_violations": [
         "판화 하단에 촬영 지시문을 읽을 수 있는 한국어 문장으로 삽입해, 글자·자막·지시문 노출 금지를 위반했다."
        ],
        "physics": "판화 종이는 목재 표면에 밀착해 지지되는 것으로 보인다. 그림 속 백산의 하체와 군중의 발은 가려져 있어 발 접촉은 확인할 수 없지만, 공중에 떠 있는 자세는 아니다. 탈은 얼굴에 착용돼 있고 옷은 몸을 따라 내려간다. 다만 실제 배우의 체중과 자세를 촬영한 장면이 아니라 인쇄된 그림이다."
       },
       {
        "label": "B",
        "direction": "백산은 머리와 가면을 화면 오른쪽 아래 군중 쪽으로 돌리고 숙이고 있다. 가면의 앞면은 아래에서 비스듬히 보인다. 전경 군중은 중앙 높은 백산을 향해 몸을 모으고 고개를 올리며, 머리 기울기와 어깨 높이도 서로 다르다. 무기나 이동 동작은 없다.",
        "built_space": "중앙에 높은 목제 단 한 개가 있고, 앞쪽에는 판재와 가로 난간, 세 개의 굵은 기둥이 보인다. 단 양쪽에는 검은 천을 단 장대가 하나씩, 총 두 개 있다. 백산은 단 위, 군중은 그 아래 전경에 배치돼 높이 관계가 명확하다. 카메라는 군중 높이보다 낮은 곳에서 올려다본다. 뒤로 하늘과 산, 일부 지붕이 보이며, 참조의 거친 검은 목재 재질은 단에 부분적으로 반영됐지만 같은 목재 표면 중심의 공간은 아니다.",
        "entities": "백산은 배가 나온 성인 남성으로 표현됐고, 온전한 금빛 하회탈과 어두운 긴 겹옷을 착용한다. 얼굴은 가려져 참조 인물의 얼굴 동일성을 확인할 수 없다. 머리는 참조의 짧은 머리 대신 상투 형태다. 아래 군중은 각각 구별되는 머리와 등, 일부 옆얼굴을 갖지만 모두 굵은 판화 선으로 그려져 실사 인물로 읽히지 않는다. 물탱크는 보이지 않으며, 글자나 로고도 보이지 않는다.",
        "hard_violations": [],
        "physics": "백산의 발은 단 앞벽과 옷자락에 가려져 있지만 몸 아래에 지지할 수 있는 단이 있어 부유로 보이지 않는다. 군중의 하체는 프레임 밖이거나 서로 가려져 있으며 서 있는 상체 배치는 가능하다. 두 천은 각각 장대의 가로대에 매달려 지지된다. 다만 백산과 군중의 평면적인 판화 표현이 실제 목재 구조 및 하늘과 분리돼 보여, 신체와 의복의 물리적 실재감은 부족하다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "하단에 지시문을 읽을 수 있는 자막으로 그대로 노출한 중대 위반이 있으며, 군중을 내려다보는 비스듬한 가면 대신 정면 가면과 판화 소품 자체를 보여준다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "낮은 시점의 와이드 구도와 높은 백산을 올려다보는 전경 군중의 관계는 더 정확하지만, 인물이 실사 인간이 아닌 평면 판화처럼 보이고 자연광 배경도 지정된 목판 명암 표현과 다르다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "군중의 등과 좌우 옆얼굴은 중앙 백산을 향하며 일부는 고개를 올리고 있다. 반면 백산의 가면은 거의 정면으로 카메라를 향해 있어, 전경 군중을 향해 아래로 기울어진 조각 면이라는 지시는 약하게 구현됐다. 무기나 이동 동작은 없다.",
        "built_space": "거친 검은 목재 표면 위에 테두리가 있는 판화 한 장이 놓여 있다. 판화 내부에는 좌우의 여러 가옥과 왼쪽 뒤 원통형 물탱크 한 개가 보인다. 백산은 중앙 군중 뒤에 있지만 높이를 제공하는 단이나 지면은 보이지 않는다. 목재 바탕은 장소 참조와 가깝지만, 장면 안의 와이드 숏보다 판화 소품을 내려다보는 촬영에 가깝다.",
        "entities": "배가 크게 나온 성인 남성 백산 한 명과 다수의 군중이 판화로 표현됐다. 백산은 얼굴 전체를 덮는 온전한 금빛 웃는 탈과 어두운 겹옷을 착용한다. 탈 때문에 참조 얼굴의 동일성은 확인할 수 없으며, 가려진 얼굴을 노출하지는 않았다. 군중은 등과 옆얼굴, 옷 주름이 구별되지만 실사 인간은 아니다. 물탱크는 보이나 물 자체는 보이지 않는다. 하단에는 한국어 지시문이 명확하게 읽힌다.",
        "hard_violations": [
         "판화 하단에 촬영 지시문을 읽을 수 있는 한국어 문장으로 삽입해, 글자·자막·지시문 노출 금지를 위반했다."
        ],
        "physics": "판화 종이는 목재 표면에 밀착해 지지되는 것으로 보인다. 그림 속 백산의 하체와 군중의 발은 가려져 있어 발 접촉은 확인할 수 없지만, 공중에 떠 있는 자세는 아니다. 탈은 얼굴에 착용돼 있고 옷은 몸을 따라 내려간다. 다만 실제 배우의 체중과 자세를 촬영한 장면이 아니라 인쇄된 그림이다."
       },
       {
        "label": "A",
        "direction": "백산은 머리와 가면을 화면 오른쪽 아래 군중 쪽으로 돌리고 숙이고 있다. 가면의 앞면은 아래에서 비스듬히 보인다. 전경 군중은 중앙 높은 백산을 향해 몸을 모으고 고개를 올리며, 머리 기울기와 어깨 높이도 서로 다르다. 무기나 이동 동작은 없다.",
        "built_space": "중앙에 높은 목제 단 한 개가 있고, 앞쪽에는 판재와 가로 난간, 세 개의 굵은 기둥이 보인다. 단 양쪽에는 검은 천을 단 장대가 하나씩, 총 두 개 있다. 백산은 단 위, 군중은 그 아래 전경에 배치돼 높이 관계가 명확하다. 카메라는 군중 높이보다 낮은 곳에서 올려다본다. 뒤로 하늘과 산, 일부 지붕이 보이며, 참조의 거친 검은 목재 재질은 단에 부분적으로 반영됐지만 같은 목재 표면 중심의 공간은 아니다.",
        "entities": "백산은 배가 나온 성인 남성으로 표현됐고, 온전한 금빛 하회탈과 어두운 긴 겹옷을 착용한다. 얼굴은 가려져 참조 인물의 얼굴 동일성을 확인할 수 없다. 머리는 참조의 짧은 머리 대신 상투 형태다. 아래 군중은 각각 구별되는 머리와 등, 일부 옆얼굴을 갖지만 모두 굵은 판화 선으로 그려져 실사 인물로 읽히지 않는다. 물탱크는 보이지 않으며, 글자나 로고도 보이지 않는다.",
        "hard_violations": [],
        "physics": "백산의 발은 단 앞벽과 옷자락에 가려져 있지만 몸 아래에 지지할 수 있는 단이 있어 부유로 보이지 않는다. 군중의 하체는 프레임 밖이거나 서로 가려져 있으며 서 있는 상체 배치는 가능하다. 두 천은 각각 장대의 가로대에 매달려 지지된다. 다만 백산과 군중의 평면적인 판화 표현이 실제 목재 구조 및 하늘과 분리돼 보여, 신체와 의복의 물리적 실재감은 부족하다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.0
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.75
   },
   "violations": {
    "B": [
     "[gemini-pro] 절대 금지된 텍스트(프롬프트 지시문)가 이미지 하단에 캡션으로 유출됨",
     "[gpt-high] 판화 하단에 촬영 지시문을 읽을 수 있는 한국어 문장으로 삽입해, 글자·자막·지시문 노출 금지를 위반했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 750
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지정된 로케이션(나무 벽면)을 무시하고 야외 단상을 임의로 연출했으나, 텍스트 유출과 같은 치명적 위반이 없어 우선됨."
   },
   {
    "label": "B",
    "score": 750,
    "verdict_ko": "지시문 텍스트가 이미지 하단에 캡션으로 그대로 노출되는 치명적인 텍스트 유출 위반(Hard Violation)이 발생함.  ★위반: [gemini-pro] 절대 금지된 텍스트(프롬프트 지시문)가 이미지 하단에 캡션으로 유출됨 / [gpt-high] 판화 하단에 촬영 지시문을 읽을 수 있는 한국어 문장으로 삽입해, 글자·자막·지시문 노출 금지를 위반했다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L222B01.png",
    "asset_id": "dfd97579-913b-4622-86a5-0d6cd32f2ea6",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 백산: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1241731>",
    "asset_id": "b337b8d8-94d9-4a29-9e49-2e19121379b7",
    "role": "character_ref"
   },
   {
    "label": "PROP REFERENCE — 백산의 지배를 묘사한 목판화: the exact object appearing in this shot; match its look, material and wear exactly.",
    "path": "<bytes:1581563>",
    "asset_id": "e7bb9df0-db9f-4c6d-9c8a-088105423f0f",
    "role": "prop_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-7a1d-75f5-b59e-40cb63ba0d11",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S59sh36__bgfirst_bg.png",
   "bg_asset_id": "c373b774-4997-4590-9fd6-e9ba287393af",
   "bg_record_key": "S59sh36::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S59sh36::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:45:24.751027+00:00",
  "fingerprint": "23d7fee06eb7062de4f91135ca2e492a9dd59f3961bd63dfe3ef9a44f4a61b0b",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S59sh36_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S59sh36_sel.png",
  "source_sha256": "707e0641b89f8cbfc4182752f6d7a5cd9f9c153105eb5b43f0e35bfec1416ae5",
  "file": "S59sh36_cine.png",
  "staged_sha256": "e9ee06592bb1dab36b60fa954fb028ef73512631c027ee6b44d15856baee84e3",
  "latency_ms": 9771
 },
 "S60sh4::signage": {
  "fp": "d91c2cce8e6ded46",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S60sh4": {
  "input_fingerprint": "3fe250a3a93cae9b",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 찰리의 반대편에 우뚝 선 거대하고 육중한 전투 병기 B-200의 위협적인 실루엣.\n\nLOCATION (lock): On the open-air fighting ground in front of the village church, opposite the arriving robot and surrounded by spectators and torches. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nSTRUCTURE LOOK AUTHORITY: the attached STRUCTURE LOOK photograph is the identity of the fixed structure at this location — wherever that structure appears in the frame, its shape, proportions, openings, materials and colors are LOCKED to it. The LOCATION PHOTOGRAPH remains the authority for this shot's sub-space, surroundings, time of day and lighting. If the two conflict on the structure itself, the STRUCTURE LOOK photo wins; for everything else, the LOCATION PHOTOGRAPH wins.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 열린 케이지 출입구 (Open after 찰리's arrival at ground level) — Only the near side of the opening appears at the far left edge behind 찰리; used as Anchors the camera's departure point without obstructing either robot; 격투장 바닥 (An open interval separates 찰리 and B-200 before their fight); used as Supplies shared ground and a credible distance comparison without foreground scale exaggeration.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the established sunset illumination and arena torchlight to articulate B-200's threatening outline while retaining enough tonal detail to read both bodies.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The arena stands in front of the church, lit by torches and late-day light, with Charlie's raised cage open. Charlie retains the worn Ubik chest logo and his straw hat, colorful raincoat and oversized boots, while the larger, heavily built B-200 has converted gun-hands.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 찰리의 반대편에 우뚝 선 거대하고 육중한 전투 병기 B-200의 위협적인 실루엣.\n\nLOCATION (lock): On the open-air fighting ground in front of the village church, opposite the arriving robot and surrounded by spectators and torches. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 열린 케이지 출입구 (Open after 찰리's arrival at ground level) — Only the near side of the opening appears at the far left edge behind 찰리; used as Anchors the camera's departure point without obstructing either robot; 격투장 바닥 (An open interval separates 찰리 and B-200 before their fight); used as Supplies shared ground and a credible distance comparison without foreground scale exaggeration.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the established sunset illumination and arena torchlight to articulate B-200's threatening outline while retaining enough tonal detail to read both bodies.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The arena stands in front of the church, lit by torches and late-day light, with Charlie's raised cage open. Charlie retains the worn Ubik chest logo and his straw hat, colorful raincoat and oversized boots, while the larger, heavily built B-200 has converted gun-hands.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 찰리의 반대편에 우뚝 선 거대하고 육중한 전투 병기 B-200의 위협적인 실루엣.\n\nLOCATION (lock): On the open-air fighting ground in front of the village church, opposite the arriving robot and surrounded by spectators and torches. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nSTRUCTURE LOOK AUTHORITY: the attached STRUCTURE LOOK photograph is the identity of the fixed structure at this location — wherever that structure appears in the frame, its shape, proportions, openings, materials and colors are LOCKED to it. The LOCATION PHOTOGRAPH remains the authority for this shot's sub-space, surroundings, time of day and lighting. If the two conflict on the structure itself, the STRUCTURE LOOK photo wins; for everything else, the LOCATION PHOTOGRAPH wins.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 열린 케이지 출입구 (Open after 찰리's arrival at ground level) — Only the near side of the opening appears at the far left edge behind 찰리; used as Anchors the camera's departure point without obstructing either robot; 격투장 바닥 (An open interval separates 찰리 and B-200 before their fight); used as Supplies shared ground and a credible distance comparison without foreground scale exaggeration.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the established sunset illumination and arena torchlight to articulate B-200's threatening outline while retaining enough tonal detail to read both bodies.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The arena stands in front of the church, lit by torches and late-day light, with Charlie's raised cage open. Charlie retains the worn Ubik chest logo and his straw hat, colorful raincoat and oversized boots, while the larger, heavily built B-200 has converted gun-hands.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S60sh4__bgfirst_bg.png",
     "asset_id": "ab433a49-41f6-42f3-9c2c-d68dd6bf3b32",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S60sh4.png",
     "asset_id": "6a6601d1-910a-4eba-a636-eb4b443cdfb9",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — B-200: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1339855>",
     "asset_id": "8091d94b-e8e7-407e-97d6-c030f55a73f9",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its spatial layout, surroundings, fixed features, time of day and lighting mood are spatial truth; stage the moment inside this place. If a STRUCTURE LOOK photograph is also attached, that photo wins for the fixed structure itself — this photograph wins for everything around it. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L229B01.png",
     "asset_id": "239110df-f90f-4ca2-b3ef-4a8f15c13fe9",
     "role": "location_plate"
    },
    {
     "label": "STRUCTURE LOOK — the confirmed photograph of the fixed structure at this location: wherever the structure appears in the frame, its shape, proportions, materials, colors and openings are LOCKED to this photo. Never copy its camera framing, time of day or lighting — the shot text and the LOCATION PHOTOGRAPH are the authorities for those.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_village_church_arena_sel.png",
     "asset_id": "7cc6d446-071f-4801-886a-a504b66fe88c",
     "role": "structure_seed_look"
    },
    {
     "label": "CHARACTER REFERENCE — B-200: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1339855>",
     "asset_id": "8091d94b-e8e7-407e-97d6-c030f55a73f9",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "카메라는 찰리의 등 뒤에서 B-200을 바라봅니다. 찰리는 B-200을 향해 시선을 두고 걸어가며, B-200 역시 찰리를 마주보고 있습니다.",
    "built_space": "격투장 바닥과 관람석은 위치 레퍼런스와 일치하며, 왼쪽 가장자리에 케이지 출입구가 찰리 뒤에 정확히 배치되어 프레임을 잡아줍니다. 하지만 뒷배경의 성당이 STRUCTURE LOOK의 폐허가 아닌 LOCATION 사진의 온전한 건물로 렌더링되어 지시를 위반했습니다.",
    "entities": "찰리는 밀짚모자, 화려한 우비, 큰 부츠를 착용한 인간으로 묘사되었으며(로고는 가슴이 아닌 등에 위치함), B-200은 무기 팔과 육중한 장갑을 갖춘 레퍼런스 디자인과 완벽하게 일치합니다.",
    "hard_violations": [
     "[gpt-high] 찰리의 우비 등에 Ubik이라는 읽을 수 있는 로고가 노출되어, 판독 가능한 글자와 로고를 금지한 조건을 위반했습니다."
    ],
    "physics": "찰리는 왼발로 지면을 딛고 오른발을 들어 올리며 걷는 자연스러운 체중 이동을 보여주며, B-200은 양발로 격투장 바닥을 단단히 딛고 서 있습니다."
   },
   {
    "label": "B",
    "direction": "카메라는 B-200의 등 뒤에서 찰리와 성당을 바라봅니다. B-200은 찰리를 향해 있고, 찰리는 B-200과 카메라 쪽을 정면으로 마주보고 있습니다.",
    "built_space": "성당 건물은 원형 탑과 부서진 창문 등 STRUCTURE LOOK의 폐허 상태를 훌륭하게 반영했습니다. 하지만 케이지가 찰리의 뒤가 아닌 화면 좌측 전면에 배치되어 프레이밍 요건을 어겼고, 관람석에는 많은 군중이 앉아 있습니다.",
    "entities": "B-200은 육중한 뒷모습을 잘 보여줍니다. 하지만 찰리가 인간이 아닌 빛나는 눈과 금속 얼굴을 가진 로봇으로 렌더링되었으며, 우비와 부츠 등의 의상만 착용하고 있습니다.",
    "hard_violations": [
     "[gemini-pro] invented people (지문 텍스트에 명시되지 않은 관중들을 관람석에 임의로 추가함)",
     "[gemini-pro] physically impossible anatomy (모든 캐릭터는 인간이어야 한다는 지시를 어기고 찰리를 금속 피부와 빛나는 눈을 가진 로봇으로 묘사함)",
     "[gpt-high] 등장인물 제한에서 허용하지 않은 다수의 관중을 관람석과 교회 앞 계단에 추가했습니다."
    ],
    "physics": "B-200과 찰리 모두 지면에 두 발을 붙이고 서 있으나, 찰리는 자연스러운 동작 없이 지나치게 경직된 차렷 자세(stiff attention stance)를 취하고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "제시된 와이드 샷 프레이밍과 캐릭터 요건을 훌륭히 따랐으나, STRUCTURE LOOK 레퍼런스를 무시하고 온전한 상태의 성당을 배치해 공간 일치도(Priority 3)에서 감점되었습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "성당의 폐허 모습은 정확히 반영했으나, 인간이어야 할 찰리를 로봇으로 묘사하고 샷 텍스트에 없는 관중을 추가하여 심각한 하드 위반을 범했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 찰리의 등 뒤에서 B-200을 바라봅니다. 찰리는 B-200을 향해 시선을 두고 걸어가며, B-200 역시 찰리를 마주보고 있습니다.",
        "built_space": "격투장 바닥과 관람석은 위치 레퍼런스와 일치하며, 왼쪽 가장자리에 케이지 출입구가 찰리 뒤에 정확히 배치되어 프레임을 잡아줍니다. 하지만 뒷배경의 성당이 STRUCTURE LOOK의 폐허가 아닌 LOCATION 사진의 온전한 건물로 렌더링되어 지시를 위반했습니다.",
        "entities": "찰리는 밀짚모자, 화려한 우비, 큰 부츠를 착용한 인간으로 묘사되었으며(로고는 가슴이 아닌 등에 위치함), B-200은 무기 팔과 육중한 장갑을 갖춘 레퍼런스 디자인과 완벽하게 일치합니다.",
        "hard_violations": [],
        "physics": "찰리는 왼발로 지면을 딛고 오른발을 들어 올리며 걷는 자연스러운 체중 이동을 보여주며, B-200은 양발로 격투장 바닥을 단단히 딛고 서 있습니다."
       },
       {
        "label": "B",
        "direction": "카메라는 B-200의 등 뒤에서 찰리와 성당을 바라봅니다. B-200은 찰리를 향해 있고, 찰리는 B-200과 카메라 쪽을 정면으로 마주보고 있습니다.",
        "built_space": "성당 건물은 원형 탑과 부서진 창문 등 STRUCTURE LOOK의 폐허 상태를 훌륭하게 반영했습니다. 하지만 케이지가 찰리의 뒤가 아닌 화면 좌측 전면에 배치되어 프레이밍 요건을 어겼고, 관람석에는 많은 군중이 앉아 있습니다.",
        "entities": "B-200은 육중한 뒷모습을 잘 보여줍니다. 하지만 찰리가 인간이 아닌 빛나는 눈과 금속 얼굴을 가진 로봇으로 렌더링되었으며, 우비와 부츠 등의 의상만 착용하고 있습니다.",
        "hard_violations": [
         "invented people (지문 텍스트에 명시되지 않은 관중들을 관람석에 임의로 추가함)",
         "physically impossible anatomy (모든 캐릭터는 인간이어야 한다는 지시를 어기고 찰리를 금속 피부와 빛나는 눈을 가진 로봇으로 묘사함)"
        ],
        "physics": "B-200과 찰리 모두 지면에 두 발을 붙이고 서 있으나, 찰리는 자연스러운 동작 없이 지나치게 경직된 차렷 자세(stiff attention stance)를 취하고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "제시된 와이드 샷 프레이밍과 캐릭터 요건을 훌륭히 따랐으나, STRUCTURE LOOK 레퍼런스를 무시하고 온전한 상태의 성당을 배치해 공간 일치도(Priority 3)에서 감점되었습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "성당의 폐허 모습은 정확히 반영했으나, 인간이어야 할 찰리를 로봇으로 묘사하고 샷 텍스트에 없는 관중을 추가하여 심각한 하드 위반을 범했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 찰리의 등 뒤에서 B-200을 바라봅니다. 찰리는 B-200을 향해 시선을 두고 걸어가며, B-200 역시 찰리를 마주보고 있습니다.",
        "built_space": "격투장 바닥과 관람석은 위치 레퍼런스와 일치하며, 왼쪽 가장자리에 케이지 출입구가 찰리 뒤에 정확히 배치되어 프레임을 잡아줍니다. 하지만 뒷배경의 성당이 STRUCTURE LOOK의 폐허가 아닌 LOCATION 사진의 온전한 건물로 렌더링되어 지시를 위반했습니다.",
        "entities": "찰리는 밀짚모자, 화려한 우비, 큰 부츠를 착용한 인간으로 묘사되었으며(로고는 가슴이 아닌 등에 위치함), B-200은 무기 팔과 육중한 장갑을 갖춘 레퍼런스 디자인과 완벽하게 일치합니다.",
        "hard_violations": [],
        "physics": "찰리는 왼발로 지면을 딛고 오른발을 들어 올리며 걷는 자연스러운 체중 이동을 보여주며, B-200은 양발로 격투장 바닥을 단단히 딛고 서 있습니다."
       },
       {
        "label": "B",
        "direction": "카메라는 B-200의 등 뒤에서 찰리와 성당을 바라봅니다. B-200은 찰리를 향해 있고, 찰리는 B-200과 카메라 쪽을 정면으로 마주보고 있습니다.",
        "built_space": "성당 건물은 원형 탑과 부서진 창문 등 STRUCTURE LOOK의 폐허 상태를 훌륭하게 반영했습니다. 하지만 케이지가 찰리의 뒤가 아닌 화면 좌측 전면에 배치되어 프레이밍 요건을 어겼고, 관람석에는 많은 군중이 앉아 있습니다.",
        "entities": "B-200은 육중한 뒷모습을 잘 보여줍니다. 하지만 찰리가 인간이 아닌 빛나는 눈과 금속 얼굴을 가진 로봇으로 렌더링되었으며, 우비와 부츠 등의 의상만 착용하고 있습니다.",
        "hard_violations": [
         "invented people (지문 텍스트에 명시되지 않은 관중들을 관람석에 임의로 추가함)",
         "physically impossible anatomy (모든 캐릭터는 인간이어야 한다는 지시를 어기고 찰리를 금속 피부와 빛나는 눈을 가진 로봇으로 묘사함)"
        ],
        "physics": "B-200과 찰리 모두 지면에 두 발을 붙이고 서 있으나, 찰리는 자연스러운 동작 없이 지나치게 경직된 차렷 자세(stiff attention stance)를 취하고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "폐허 교회와 로봇 찰리는 더 충실하지만, 제한된 등장인물 외 군중을 추가하고 B-200을 거대한 전경 후면으로 배치해 지정된 출입구 기준 구도와 크기 비교를 훼손했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "왼쪽 출입구의 찰리와 맞은편 B-200이라는 배치는 더 가깝지만, 읽히는 Ubik 표기와 인간으로 표현된 찰리, 달라진 교회 및 과장된 전경 때문에 재촬영이 필요합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "B-200은 카메라에 등을 보이고 화면 왼쪽 중경의 찰리 쪽으로 몸을 돌렸습니다. 찰리도 B-200 쪽을 보고 있습니다. 보이는 왼쪽 포신 묶음은 찰리 오른편의 바닥 방향으로 내려가며 찰리 몸을 직접 겨누지는 않습니다. 다른 팔의 포구 방향은 몸에 가려 명확하지 않습니다. 이 장면은 대치를 요구하며 직접 조준을 명시하지는 않습니다.",
        "built_space": "왼쪽에 열린 철망 케이지 한 개, 중앙 뒤에 교회 한 채와 진입 계단, 좌우에 관람석과 난간, 최소 다섯 곳의 횃불이 보입니다. 교회의 부서진 박공, 원형 창, 오른쪽 종탑은 구조 참조에 비교적 가깝습니다. 그러나 출입구의 가까운 가장자리만 보이는 대신 케이지의 상당 부분이 노출되고, 찰리는 그 바로 앞이 아니라 멀리 중앙 왼쪽에 있습니다. B-200이 오른쪽 전경을 거의 가득 채워 공통 지면에서의 크기 비교보다 원근 확대가 지배합니다.",
        "entities": "B-200 한 대는 짙은 회색 장갑, 굵은 관절, 넓은 어깨와 팔에 결합된 포신을 갖춰 참조의 중장비형 로봇에 가깝습니다. 찰리 한 대는 기계 얼굴, 밀짚모자, 다색 우비, 큰 부츠와 가슴의 낡은 표식을 지닙니다. 가슴 글자는 확실하게 판독되지 않습니다. 관람석과 계단에는 다양한 연령의 남녀로 보이는 군중이 다수 있습니다. 군중은 장소 설명에는 부합하지만 별도의 등장인물 제한에는 어긋납니다.",
        "hard_violations": [
         "등장인물 제한에서 허용하지 않은 다수의 관중을 관람석과 교회 앞 계단에 추가했습니다."
        ],
        "physics": "두 로봇 모두 양발이 흙바닥에 닿아 몸을 지탱합니다. B-200의 무거운 몸통은 벌어진 다리와 기계 관절 위에 놓이며 포신은 팔에 연결되어 있습니다. 관중도 계단이나 좌석에 서거나 앉아 있고, 횃불은 고정된 받침대가 지지합니다. 지지 없이 떠 있는 몸이나 물체는 보이지 않습니다."
       },
       {
        "label": "B",
        "direction": "왼쪽 찰리는 오른쪽 뒤편의 B-200을 향해 고개와 몸을 돌렸고, B-200도 찰리 쪽을 향해 서 있습니다. 양팔의 포신은 화면 왼쪽으로 향하지만 수평에 가까워, 낮은 전경의 찰리 몸에 정확히 조준선이 닿는다고 단정하기는 어렵습니다. 찰리의 이동 방향은 열린 케이지에서 격투장 안쪽입니다.",
        "built_space": "프레임 맨 왼쪽에 철망 출입구의 가까운 부분 한 곳이 있고 바로 앞에 찰리가 있습니다. 오른쪽 중경의 B-200과 찰리 사이에는 넓은 흙바닥이 열려 있어 지정된 배치에 더 가깝습니다. 중앙 교회 한 채, 정면 계단 한 벌, 좌우 관람석, 다섯 곳의 횃불이 보입니다. 다만 교회는 구조 참조의 무너진 비대칭 외벽과 오른쪽 종탑 대신 정돈된 높은 고딕 입면으로 바뀌었습니다. 찰리를 매우 가까운 전경에 두어 화면상으로 B-200보다 크게 만든 점도 원근 과장 금지에 맞지 않습니다.",
        "entities": "B-200 한 대는 회색 장갑과 굵은 관절, 양손의 다연장 포신을 갖췄으며 정면의 장갑 구성도 참조에 대체로 가깝습니다. 찰리는 밀짚모자, 다색 우비와 큰 부츠를 착용했지만 노출된 손과 종아리가 인간의 피부로 보여 로봇이라는 설정과 다릅니다. 우비 등에는 Ubik이라는 글자가 명확히 읽힙니다. 별도의 관중은 없으며 관람석은 비어 있습니다.",
        "hard_violations": [
         "찰리의 우비 등에 Ubik이라는 읽을 수 있는 로고가 노출되어, 판독 가능한 글자와 로고를 금지한 조건을 위반했습니다."
        ],
        "physics": "찰리는 앞쪽 부츠에 체중을 옮기며 뒤쪽 부츠의 앞부분을 바닥에 둔 보행 자세로 보입니다. B-200은 벌어진 두 발을 지면에 붙이고 서 있으며 포신은 양팔에 기계적으로 연결됩니다. 철망과 횃불도 각각 틀과 받침대에 고정되어 있습니다. 지지 없이 떠 있는 몸이나 물체는 보이지 않습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "폐허 교회와 로봇 찰리는 더 충실하지만, 제한된 등장인물 외 군중을 추가하고 B-200을 거대한 전경 후면으로 배치해 지정된 출입구 기준 구도와 크기 비교를 훼손했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "왼쪽 출입구의 찰리와 맞은편 B-200이라는 배치는 더 가깝지만, 읽히는 Ubik 표기와 인간으로 표현된 찰리, 달라진 교회 및 과장된 전경 때문에 재촬영이 필요합니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "B-200은 카메라에 등을 보이고 화면 왼쪽 중경의 찰리 쪽으로 몸을 돌렸습니다. 찰리도 B-200 쪽을 보고 있습니다. 보이는 왼쪽 포신 묶음은 찰리 오른편의 바닥 방향으로 내려가며 찰리 몸을 직접 겨누지는 않습니다. 다른 팔의 포구 방향은 몸에 가려 명확하지 않습니다. 이 장면은 대치를 요구하며 직접 조준을 명시하지는 않습니다.",
        "built_space": "왼쪽에 열린 철망 케이지 한 개, 중앙 뒤에 교회 한 채와 진입 계단, 좌우에 관람석과 난간, 최소 다섯 곳의 횃불이 보입니다. 교회의 부서진 박공, 원형 창, 오른쪽 종탑은 구조 참조에 비교적 가깝습니다. 그러나 출입구의 가까운 가장자리만 보이는 대신 케이지의 상당 부분이 노출되고, 찰리는 그 바로 앞이 아니라 멀리 중앙 왼쪽에 있습니다. B-200이 오른쪽 전경을 거의 가득 채워 공통 지면에서의 크기 비교보다 원근 확대가 지배합니다.",
        "entities": "B-200 한 대는 짙은 회색 장갑, 굵은 관절, 넓은 어깨와 팔에 결합된 포신을 갖춰 참조의 중장비형 로봇에 가깝습니다. 찰리 한 대는 기계 얼굴, 밀짚모자, 다색 우비, 큰 부츠와 가슴의 낡은 표식을 지닙니다. 가슴 글자는 확실하게 판독되지 않습니다. 관람석과 계단에는 다양한 연령의 남녀로 보이는 군중이 다수 있습니다. 군중은 장소 설명에는 부합하지만 별도의 등장인물 제한에는 어긋납니다.",
        "hard_violations": [
         "등장인물 제한에서 허용하지 않은 다수의 관중을 관람석과 교회 앞 계단에 추가했습니다."
        ],
        "physics": "두 로봇 모두 양발이 흙바닥에 닿아 몸을 지탱합니다. B-200의 무거운 몸통은 벌어진 다리와 기계 관절 위에 놓이며 포신은 팔에 연결되어 있습니다. 관중도 계단이나 좌석에 서거나 앉아 있고, 횃불은 고정된 받침대가 지지합니다. 지지 없이 떠 있는 몸이나 물체는 보이지 않습니다."
       },
       {
        "label": "A",
        "direction": "왼쪽 찰리는 오른쪽 뒤편의 B-200을 향해 고개와 몸을 돌렸고, B-200도 찰리 쪽을 향해 서 있습니다. 양팔의 포신은 화면 왼쪽으로 향하지만 수평에 가까워, 낮은 전경의 찰리 몸에 정확히 조준선이 닿는다고 단정하기는 어렵습니다. 찰리의 이동 방향은 열린 케이지에서 격투장 안쪽입니다.",
        "built_space": "프레임 맨 왼쪽에 철망 출입구의 가까운 부분 한 곳이 있고 바로 앞에 찰리가 있습니다. 오른쪽 중경의 B-200과 찰리 사이에는 넓은 흙바닥이 열려 있어 지정된 배치에 더 가깝습니다. 중앙 교회 한 채, 정면 계단 한 벌, 좌우 관람석, 다섯 곳의 횃불이 보입니다. 다만 교회는 구조 참조의 무너진 비대칭 외벽과 오른쪽 종탑 대신 정돈된 높은 고딕 입면으로 바뀌었습니다. 찰리를 매우 가까운 전경에 두어 화면상으로 B-200보다 크게 만든 점도 원근 과장 금지에 맞지 않습니다.",
        "entities": "B-200 한 대는 회색 장갑과 굵은 관절, 양손의 다연장 포신을 갖췄으며 정면의 장갑 구성도 참조에 대체로 가깝습니다. 찰리는 밀짚모자, 다색 우비와 큰 부츠를 착용했지만 노출된 손과 종아리가 인간의 피부로 보여 로봇이라는 설정과 다릅니다. 우비 등에는 Ubik이라는 글자가 명확히 읽힙니다. 별도의 관중은 없으며 관람석은 비어 있습니다.",
        "hard_violations": [
         "찰리의 우비 등에 Ubik이라는 읽을 수 있는 로고가 노출되어, 판독 가능한 글자와 로고를 금지한 조건을 위반했습니다."
        ],
        "physics": "찰리는 앞쪽 부츠에 체중을 옮기며 뒤쪽 부츠의 앞부분을 바닥에 둔 보행 자세로 보입니다. B-200은 벌어진 두 발을 지면에 붙이고 서 있으며 포신은 양팔에 기계적으로 연결됩니다. 철망과 횃불도 각각 틀과 받침대에 고정되어 있습니다. 지지 없이 떠 있는 몸이나 물체는 보이지 않습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.952
   },
   "adjusted": {
    "A": 1.75,
    "B": 0.702
   },
   "violations": {
    "B": [
     "[gemini-pro] invented people (지문 텍스트에 명시되지 않은 관중들을 관람석에 임의로 추가함)",
     "[gemini-pro] physically impossible anatomy (모든 캐릭터는 인간이어야 한다는 지시를 어기고 찰리를 금속 피부와 빛나는 눈을 가진 로봇으로 묘사함)",
     "[gpt-high] 등장인물 제한에서 허용하지 않은 다수의 관중을 관람석과 교회 앞 계단에 추가했습니다."
    ],
    "A": [
     "[gpt-high] 찰리의 우비 등에 Ubik이라는 읽을 수 있는 로고가 노출되어, 판독 가능한 글자와 로고를 금지한 조건을 위반했습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 702
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "제시된 와이드 샷 프레이밍과 캐릭터 요건을 훌륭히 따랐으나, STRUCTURE LOOK 레퍼런스를 무시하고 온전한 상태의 성당을 배치해 공간 일치도(Priority 3)에서 감점되었습니다.  ★위반: [gpt-high] 찰리의 우비 등에 Ubik이라는 읽을 수 있는 로고가 노출되어, 판독 가능한 글자와 로고를 금지한 조건을 위반했습니다."
   },
   {
    "label": "B",
    "score": 702,
    "verdict_ko": "성당의 폐허 모습은 정확히 반영했으나, 인간이어야 할 찰리를 로봇으로 묘사하고 샷 텍스트에 없는 관중을 추가하여 심각한 하드 위반을 범했습니다.  ★위반: [gemini-pro] invented people (지문 텍스트에 명시되지 않은 관중들을 관람석에 임의로 추가함) / [gemini-pro] physically impossible anatomy (모든 캐릭터는 인간이어야 한다는 지시를 어기고 찰리를 금속 피부와 빛나는 눈을 가진 로봇으로 묘사함) / [gpt-high] 등장인물 제한에서 허용하지 않은 다수의 관중을 관람석과 교회 앞 계단에 추가했습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its spatial layout, surroundings, fixed features, time of day and lighting mood are spatial truth; stage the moment inside this place. If a STRUCTURE LOOK photograph is also attached, that photo wins for the fixed structure itself — this photograph wins for everything around it. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L229B01.png",
    "asset_id": "239110df-f90f-4ca2-b3ef-4a8f15c13fe9",
    "role": "location_plate"
   },
   {
    "label": "STRUCTURE LOOK — the confirmed photograph of the fixed structure at this location: wherever the structure appears in the frame, its shape, proportions, materials, colors and openings are LOCKED to this photo. Never copy its camera framing, time of day or lighting — the shot text and the LOCATION PHOTOGRAPH are the authorities for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_village_church_arena_sel.png",
    "asset_id": "7cc6d446-071f-4801-886a-a504b66fe88c",
    "role": "structure_seed_look"
   },
   {
    "label": "CHARACTER REFERENCE — B-200: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1339855>",
    "asset_id": "8091d94b-e8e7-407e-97d6-c030f55a73f9",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-7d72-74a6-901e-102c8054a873",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S60sh4__bgfirst_bg.png",
   "bg_asset_id": "ab433a49-41f6-42f3-9c2c-d68dd6bf3b32",
   "bg_record_key": "S60sh4::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate",
   "seed_attached": true
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  },
  "lane_policy": "ab_select_ready"
 },
 "S60sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:48:04.731749+00:00",
  "fingerprint": "9eb9dbe99ab7b9cfe5a9a347dfe60380ae41d1f9cc7c476d5160621ac5ee2f30",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S60sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S60sh4_sel.png",
  "source_sha256": "0990e4df1fb7c6615402cfbed7bc24f2e0380ad49b1fa5ce9b2abf4a8aaa6b3e",
  "file": "S60sh4_cine.png",
  "staged_sha256": "460cc257b648260fca38f5c7a0980603149807f1067a14eadf060eed61315779",
  "latency_ms": 10821
 },
 "S60sh52::signage": {
  "fp": "2a94887d8e4d47b7",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::6c9ae0e716405f00": {
  "subjects": [],
  "subject_text": "익산 마을 성당 앞 격투장과 관중석\n성당 앞 마을 중심에 마련된 야외 격투장. 중앙 경기 구역 둘레에 관중석과 횃불이 배치되고 바닥에는 케이지 승강구가 있다.",
  "identity": "canonical",
  "scope_id": "L229",
  "scope_role": "location_exterior",
  "scope_sha": "cd835905025bc38b"
 },
 "S60sh52::bgfirst_bg": {
  "input_fingerprint": "a6abcd17ac433680",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 바닥에 주저앉은 백산의 절반쯤 깨진 황금 가면 아래로 드러난, 흉측하게 녹아내린 화상 얼굴 클로즈업.\n\nLOCATION (lock): On the wet ground beside the collapsed village water-tank tower, amid the wreckage in the evening light.\n\nTIME OF DAY (lock): sunset.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 반쯤 깨진 황금 가면 (Half broken, exposing the previously concealed facial damage) — The remaining outer facial surface and broken edge are visible obliquely beside the exposed face; used as Keeps the failed concealment and the evidence beneath it in the same focal plane.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the established evening sunset give restrained warmth to the wet face and broken gold mask, maintaining readable disfigurement without exaggerated horror lighting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 바닥에 주저앉은 백산의 절반쯤 깨진 황금 가면 아래로 드러난, 흉측하게 녹아내린 화상 얼굴 클로즈업.\n\nLOCATION (lock): On the wet ground beside the collapsed village water-tank tower, amid the wreckage in the evening light.\n\nTIME OF DAY (lock): sunset.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 반쯤 깨진 황금 가면 (Half broken, exposing the previously concealed facial damage) — The remaining outer facial surface and broken edge are visible obliquely beside the exposed face; used as Keeps the failed concealment and the evidence beneath it in the same focal plane.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the established evening sunset give restrained warmth to the wet face and broken gold mask, maintaining readable disfigurement without exaggerated horror lighting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S60sh52__bgfirst_bg.png",
  "asset_id": "aeb46e86-9443-4079-ae58-e593ffa40546",
  "input_asset_ids": [
   "4d52529a-4346-46e6-91fd-c388e0037459",
   "1f832f0c-4b90-4195-8434-227eb2c715f6"
  ]
 },
 "S60sh52": {
  "input_fingerprint": "dbf679453d8687c5",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 바닥에 주저앉은 백산의 절반쯤 깨진 황금 가면 아래로 드러난, 흉측하게 녹아내린 화상 얼굴 클로즈업.\n\nLOCATION (lock): On the wet ground beside the collapsed village water-tank tower, amid the wreckage in the evening light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 반쯤 깨진 황금 가면 (Half broken, exposing the previously concealed facial damage) — The remaining outer facial surface and broken edge are visible obliquely beside the exposed face; used as Keeps the failed concealment and the evidence beneath it in the same focal plane.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the established evening sunset give restrained warmth to the wet face and broken gold mask, maintaining readable disfigurement without exaggerated horror lighting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The water tower has been shot through and toppled, releasing a large flood of stored water in the sunset light. Charlie bears accumulated dents and holes from the fighting, while B-200's deployed gun-hands remain intact. 백산: He is soaking wet, with his gold Hahoe mask half broken and his radiation-disfigured face exposed. His royal-style clothing has not yet been stripped away.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 백산 (한국인, 성인, 남성 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 바닥에 주저앉은 백산의 절반쯤 깨진 황금 가면 아래로 드러난, 흉측하게 녹아내린 화상 얼굴 클로즈업.\n\nLOCATION (lock): On the wet ground beside the collapsed village water-tank tower, amid the wreckage in the evening light. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 반쯤 깨진 황금 가면 (Half broken, exposing the previously concealed facial damage) — The remaining outer facial surface and broken edge are visible obliquely beside the exposed face; used as Keeps the failed concealment and the evidence beneath it in the same focal plane.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the established evening sunset give restrained warmth to the wet face and broken gold mask, maintaining readable disfigurement without exaggerated horror lighting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The water tower has been shot through and toppled, releasing a large flood of stored water in the sunset light. Charlie bears accumulated dents and holes from the fighting, while B-200's deployed gun-hands remain intact. 백산: He is soaking wet, with his gold Hahoe mask half broken and his radiation-disfigured face exposed. His royal-style clothing has not yet been stripped away.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 백산 (한국인, 성인, 남성 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 바닥에 주저앉은 백산의 절반쯤 깨진 황금 가면 아래로 드러난, 흉측하게 녹아내린 화상 얼굴 클로즈업.\n\nLOCATION (lock): On the wet ground beside the collapsed village water-tank tower, amid the wreckage in the evening light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 반쯤 깨진 황금 가면 (Half broken, exposing the previously concealed facial damage) — The remaining outer facial surface and broken edge are visible obliquely beside the exposed face; used as Keeps the failed concealment and the evidence beneath it in the same focal plane.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Let the established evening sunset give restrained warmth to the wet face and broken gold mask, maintaining readable disfigurement without exaggerated horror lighting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The water tower has been shot through and toppled, releasing a large flood of stored water in the sunset light. Charlie bears accumulated dents and holes from the fighting, while B-200's deployed gun-hands remain intact. 백산: He is soaking wet, with his gold Hahoe mask half broken and his radiation-disfigured face exposed. His royal-style clothing has not yet been stripped away.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 백산 (한국인, 성인, 남성 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S60sh52__bgfirst_bg.png",
     "asset_id": "aeb46e86-9443-4079-ae58-e593ffa40546",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S60sh52.png",
     "asset_id": "4d52529a-4346-46e6-91fd-c388e0037459",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 백산: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1241731>",
     "asset_id": "b337b8d8-94d9-4a29-9e49-2e19121379b7",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L229B03.png",
     "asset_id": "1f832f0c-4b90-4195-8434-227eb2c715f6",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 백산: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1241731>",
     "asset_id": "b337b8d8-94d9-4a29-9e49-2e19121379b7",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "피사체는 화면 오른쪽 밖을 멍하게 응시하고 있음.",
    "built_space": "레퍼런스 이미지의 교회 건물, 물탱크 파이프, 젖은 바닥이 배경에 올바르게 배치되어 있음.",
    "entities": "백산이 반쯤 깨진 가면을 쓰고 있으나 왼쪽 눈동자가 하얗게 변색됨. 오른쪽 배경에 지시되지 않은 다른 인물의 다리가 등장함.",
    "hard_violations": [
     "[gemini-pro] 프롬프트에 명시되지 않은 인물(오른쪽 부츠 신은 다리)이 추가됨",
     "[gemini-pro] 왼쪽 눈동자가 하얗게 변색되어 인체 해부학 및 표정 지시사항을 위반함",
     "[gpt-high] 화면 오른쪽에 검은 바지와 부츠를 착용한 별도 인물의 하체를 추가했다. 백산 이외의 사람이나 신체 일부를 금지한 장면 조건을 위반한다."
    ],
    "physics": "가면은 피사체의 얼굴에 밀착되어 지지되고 있으며, 오른쪽의 불청객은 바닥을 딛고 서 있음."
   },
   {
    "label": "B",
    "direction": "피사체는 화면 왼쪽 아래 밖을 응시하고 있음.",
    "built_space": "젖은 바닥과 무너진 구조물 잔해들이 보이나 레퍼런스의 교회 형태는 명확하지 않음.",
    "entities": "백산의 얼굴에 심각한 화상 흉터가 사실적으로 묘사되었고 눈동자는 정상임. 왕실 의복과 반쯤 깨진 황금 가면이 존재함.",
    "hard_violations": [
     "[gemini-pro] 가면이 끈이나 손의 지지 없이 피사체의 얼굴 앞에 물리적으로 불가능하게 떠 있음"
    ],
    "physics": "반쯤 깨진 가면이 피사체의 턱이나 코 근처에 닿아 있는 듯 보이나, 실제로는 중력을 무시한 채 허공에 떠 있음. 피사체는 바닥에 주저앉아 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "화상 입은 얼굴과 클로즈업 프레이밍은 프롬프트에 부합하나, 반쯤 깨진 가면이 아무런 지지 없이 허공에 떠 있어 물리적 오류로 인해 최종 탈락함."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지시되지 않은 인물의 다리가 등장하고 눈동자가 하얗게 변색되는 등 치명적인 프롬프트 위반이 다수 발생하여 사용이 불가함."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "피사체는 화면 왼쪽 아래 밖을 응시하고 있음.",
        "built_space": "젖은 바닥과 무너진 구조물 잔해들이 보이나 레퍼런스의 교회 형태는 명확하지 않음.",
        "entities": "백산의 얼굴에 심각한 화상 흉터가 사실적으로 묘사되었고 눈동자는 정상임. 왕실 의복과 반쯤 깨진 황금 가면이 존재함.",
        "hard_violations": [
         "가면이 끈이나 손의 지지 없이 피사체의 얼굴 앞에 물리적으로 불가능하게 떠 있음"
        ],
        "physics": "반쯤 깨진 가면이 피사체의 턱이나 코 근처에 닿아 있는 듯 보이나, 실제로는 중력을 무시한 채 허공에 떠 있음. 피사체는 바닥에 주저앉아 있음."
       },
       {
        "label": "A",
        "direction": "피사체는 화면 오른쪽 밖을 멍하게 응시하고 있음.",
        "built_space": "레퍼런스 이미지의 교회 건물, 물탱크 파이프, 젖은 바닥이 배경에 올바르게 배치되어 있음.",
        "entities": "백산이 반쯤 깨진 가면을 쓰고 있으나 왼쪽 눈동자가 하얗게 변색됨. 오른쪽 배경에 지시되지 않은 다른 인물의 다리가 등장함.",
        "hard_violations": [
         "프롬프트에 명시되지 않은 인물(오른쪽 부츠 신은 다리)이 추가됨",
         "왼쪽 눈동자가 하얗게 변색되어 인체 해부학 및 표정 지시사항을 위반함"
        ],
        "physics": "가면은 피사체의 얼굴에 밀착되어 지지되고 있으며, 오른쪽의 불청객은 바닥을 딛고 서 있음."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "화상 입은 얼굴과 클로즈업 프레이밍은 프롬프트에 부합하나, 반쯤 깨진 가면이 아무런 지지 없이 허공에 떠 있어 물리적 오류로 인해 최종 탈락함."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지시되지 않은 인물의 다리가 등장하고 눈동자가 하얗게 변색되는 등 치명적인 프롬프트 위반이 다수 발생하여 사용이 불가함."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "피사체는 화면 왼쪽 아래 밖을 응시하고 있음.",
        "built_space": "젖은 바닥과 무너진 구조물 잔해들이 보이나 레퍼런스의 교회 형태는 명확하지 않음.",
        "entities": "백산의 얼굴에 심각한 화상 흉터가 사실적으로 묘사되었고 눈동자는 정상임. 왕실 의복과 반쯤 깨진 황금 가면이 존재함.",
        "hard_violations": [
         "가면이 끈이나 손의 지지 없이 피사체의 얼굴 앞에 물리적으로 불가능하게 떠 있음"
        ],
        "physics": "반쯤 깨진 가면이 피사체의 턱이나 코 근처에 닿아 있는 듯 보이나, 실제로는 중력을 무시한 채 허공에 떠 있음. 피사체는 바닥에 주저앉아 있음."
       },
       {
        "label": "A",
        "direction": "피사체는 화면 오른쪽 밖을 멍하게 응시하고 있음.",
        "built_space": "레퍼런스 이미지의 교회 건물, 물탱크 파이프, 젖은 바닥이 배경에 올바르게 배치되어 있음.",
        "entities": "백산이 반쯤 깨진 가면을 쓰고 있으나 왼쪽 눈동자가 하얗게 변색됨. 오른쪽 배경에 지시되지 않은 다른 인물의 다리가 등장함.",
        "hard_violations": [
         "프롬프트에 명시되지 않은 인물(오른쪽 부츠 신은 다리)이 추가됨",
         "왼쪽 눈동자가 하얗게 변색되어 인체 해부학 및 표정 지시사항을 위반함"
        ],
        "physics": "가면은 피사체의 얼굴에 밀착되어 지지되고 있으며, 오른쪽의 불청객은 바닥을 딛고 서 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "젖은 화상 얼굴과 깨진 금빛 가면을 밀착 클로즈업으로 담고 백산만 등장시키지만, 가면 파편이 과하게 벌어져 반쯤 깨진 하회탈의 형태가 불명확하다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "장소와 반파된 가면은 잘 재현했으나, 오른쪽에 허용되지 않은 다른 사람의 하체를 추가한 것이 결정적인 위반이다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "백산은 고개를 숙이고 화면 오른쪽 아래의 가까운 지면 쪽으로 눈을 내린다. 특정 인물이나 물체를 바라본다고 단정할 수 없으며, 정면 렌즈 응시는 아니다. 가면의 남은 바깥 면과 파단면이 노출된 얼굴 옆에서 비스듬히 보인다.",
        "built_space": "뒤에는 쓰러진 탱크로 보이는 금속 구조물 일부와 여러 지지대, 콘크리트 잔해, 물이 고인 바닥이 보인다. 클로즈업 때문에 탱크 전체나 고정 시설의 총수는 확인할 수 없다. 참고 장소의 젖은 철골·콘크리트 폐허와는 부합하지만 교회와 관람석은 식별되지 않아 정확한 장소 일치의 증거는 제한적이다. 낮게 웅크린 상체 위치는 지면에 주저앉은 상황과 양립한다.",
        "entities": "검은 머리의 성인 동아시아계 남성 한 명만 보이며 백산의 기본 인상과 대체로 맞는다. 젖어 붙은 머리카락, 물기 있는 심한 화상 변형, 금색 문양의 왕실풍 옷이 보인다. 금빛 가면은 깨져 있지만 바깥으로 벌어진 얼굴 모양 파편 때문에 하나의 반파된 하회탈이라는 형태가 다소 혼란스럽다. 눈은 자연스러운 홍채와 동공을 유지한다. 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "머리는 숙인 목과 상체로 지지된다. 엉덩이와 다리는 화면 밖이므로 바닥 접촉 자체는 확인되지 않지만 공중에 떠 있다는 증거는 없다. 가면의 이마 부분은 얼굴에 붙어 있고 돌출 파편은 아래쪽 가면 부분과 이어져 보인다. 부착 방식은 불명확하나 독립적으로 떠 있는 조각으로 단정할 수는 없다. 물은 피부를 따라 흐르고 바닥의 낮은 곳에 고여 있다."
       },
       {
        "label": "B",
        "direction": "백산의 얼굴과 노출된 눈은 화면 오른쪽 위를 향한다. 오른쪽에 서 있는 다른 사람 쪽을 올려다보는 관계로 읽히지만, 그 상대는 이 장면에 허용되지 않았다. 남은 가면은 얼굴에 정상 방향으로 씌워져 있고 바깥 표면과 깨진 가장자리가 카메라에 보인다.",
        "built_space": "왼쪽에 기울어진 물탱크 구조물 한 기와 물을 쏟는 굵은 배관 하나가 보인다. 뒤에는 중앙 첨두형 출입구 하나, 그 앞 계단과 의자 하나가 있어 참고 장소의 배치가 명확하다. 젖은 바닥과 잔해도 일치한다. 백산은 전경의 낮은 위치에 있고, 다른 사람의 검은 바지와 부츠가 오른쪽 뒤에 서 있어 요구된 단독 인물 배치를 위반한다.",
        "entities": "중앙에는 검은 머리의 성인 동아시아계 남성 백산이 보이며, 드러난 코와 입의 인상은 참고 인물과 대체로 맞는다. 반쪽 가까이 깨진 금빛 가면, 얼굴 한쪽의 심한 화상 흉터, 젖은 왕실풍 자수 의상이 있다. 노출된 눈에는 홍채와 동공이 보인다. 오른쪽의 다른 사람 하체는 백산의 신체일 수 없는 별도 인물이다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "화면 오른쪽에 검은 바지와 부츠를 착용한 별도 인물의 하체를 추가했다. 백산 이외의 사람이나 신체 일부를 금지한 장면 조건을 위반한다."
        ],
        "physics": "백산의 머리는 목과 기울어진 상체가 지지하며, 하체는 프레임 밖이라 착석 접촉점은 보이지 않는다. 가면은 얼굴 곡면을 따라 밀착되어 있어 떠 있는 물체로 보이지 않는다. 추가 인물의 부츠는 젖은 지면에 닿아 있다. 탱크와 배관에서 나온 물은 아래로 떨어져 지면에 모이며 물리적으로 자연스럽다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "젖은 화상 얼굴과 깨진 금빛 가면을 밀착 클로즈업으로 담고 백산만 등장시키지만, 가면 파편이 과하게 벌어져 반쯤 깨진 하회탈의 형태가 불명확하다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "장소와 반파된 가면은 잘 재현했으나, 오른쪽에 허용되지 않은 다른 사람의 하체를 추가한 것이 결정적인 위반이다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "백산은 고개를 숙이고 화면 오른쪽 아래의 가까운 지면 쪽으로 눈을 내린다. 특정 인물이나 물체를 바라본다고 단정할 수 없으며, 정면 렌즈 응시는 아니다. 가면의 남은 바깥 면과 파단면이 노출된 얼굴 옆에서 비스듬히 보인다.",
        "built_space": "뒤에는 쓰러진 탱크로 보이는 금속 구조물 일부와 여러 지지대, 콘크리트 잔해, 물이 고인 바닥이 보인다. 클로즈업 때문에 탱크 전체나 고정 시설의 총수는 확인할 수 없다. 참고 장소의 젖은 철골·콘크리트 폐허와는 부합하지만 교회와 관람석은 식별되지 않아 정확한 장소 일치의 증거는 제한적이다. 낮게 웅크린 상체 위치는 지면에 주저앉은 상황과 양립한다.",
        "entities": "검은 머리의 성인 동아시아계 남성 한 명만 보이며 백산의 기본 인상과 대체로 맞는다. 젖어 붙은 머리카락, 물기 있는 심한 화상 변형, 금색 문양의 왕실풍 옷이 보인다. 금빛 가면은 깨져 있지만 바깥으로 벌어진 얼굴 모양 파편 때문에 하나의 반파된 하회탈이라는 형태가 다소 혼란스럽다. 눈은 자연스러운 홍채와 동공을 유지한다. 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "머리는 숙인 목과 상체로 지지된다. 엉덩이와 다리는 화면 밖이므로 바닥 접촉 자체는 확인되지 않지만 공중에 떠 있다는 증거는 없다. 가면의 이마 부분은 얼굴에 붙어 있고 돌출 파편은 아래쪽 가면 부분과 이어져 보인다. 부착 방식은 불명확하나 독립적으로 떠 있는 조각으로 단정할 수는 없다. 물은 피부를 따라 흐르고 바닥의 낮은 곳에 고여 있다."
       },
       {
        "label": "A",
        "direction": "백산의 얼굴과 노출된 눈은 화면 오른쪽 위를 향한다. 오른쪽에 서 있는 다른 사람 쪽을 올려다보는 관계로 읽히지만, 그 상대는 이 장면에 허용되지 않았다. 남은 가면은 얼굴에 정상 방향으로 씌워져 있고 바깥 표면과 깨진 가장자리가 카메라에 보인다.",
        "built_space": "왼쪽에 기울어진 물탱크 구조물 한 기와 물을 쏟는 굵은 배관 하나가 보인다. 뒤에는 중앙 첨두형 출입구 하나, 그 앞 계단과 의자 하나가 있어 참고 장소의 배치가 명확하다. 젖은 바닥과 잔해도 일치한다. 백산은 전경의 낮은 위치에 있고, 다른 사람의 검은 바지와 부츠가 오른쪽 뒤에 서 있어 요구된 단독 인물 배치를 위반한다.",
        "entities": "중앙에는 검은 머리의 성인 동아시아계 남성 백산이 보이며, 드러난 코와 입의 인상은 참고 인물과 대체로 맞는다. 반쪽 가까이 깨진 금빛 가면, 얼굴 한쪽의 심한 화상 흉터, 젖은 왕실풍 자수 의상이 있다. 노출된 눈에는 홍채와 동공이 보인다. 오른쪽의 다른 사람 하체는 백산의 신체일 수 없는 별도 인물이다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "화면 오른쪽에 검은 바지와 부츠를 착용한 별도 인물의 하체를 추가했다. 백산 이외의 사람이나 신체 일부를 금지한 장면 조건을 위반한다."
        ],
        "physics": "백산의 머리는 목과 기울어진 상체가 지지하며, 하체는 프레임 밖이라 착석 접촉점은 보이지 않는다. 가면은 얼굴 곡면을 따라 밀착되어 있어 떠 있는 물체로 보이지 않는다. 추가 인물의 부츠는 젖은 지면에 닿아 있다. 탱크와 배관에서 나온 물은 아래로 떨어져 지면에 모이며 물리적으로 자연스럽다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.179,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.929,
    "B": 1.75
   },
   "violations": {
    "B": [
     "[gemini-pro] 가면이 끈이나 손의 지지 없이 피사체의 얼굴 앞에 물리적으로 불가능하게 떠 있음"
    ],
    "A": [
     "[gemini-pro] 프롬프트에 명시되지 않은 인물(오른쪽 부츠 신은 다리)이 추가됨",
     "[gemini-pro] 왼쪽 눈동자가 하얗게 변색되어 인체 해부학 및 표정 지시사항을 위반함",
     "[gpt-high] 화면 오른쪽에 검은 바지와 부츠를 착용한 별도 인물의 하체를 추가했다. 백산 이외의 사람이나 신체 일부를 금지한 장면 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 1750,
   "A": 929
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "화상 입은 얼굴과 클로즈업 프레이밍은 프롬프트에 부합하나, 반쯤 깨진 가면이 아무런 지지 없이 허공에 떠 있어 물리적 오류로 인해 최종 탈락함.  ★위반: [gemini-pro] 가면이 끈이나 손의 지지 없이 피사체의 얼굴 앞에 물리적으로 불가능하게 떠 있음"
   },
   {
    "label": "A",
    "score": 929,
    "verdict_ko": "지시되지 않은 인물의 다리가 등장하고 눈동자가 하얗게 변색되는 등 치명적인 프롬프트 위반이 다수 발생하여 사용이 불가함.  ★위반: [gemini-pro] 프롬프트에 명시되지 않은 인물(오른쪽 부츠 신은 다리)이 추가됨 / [gemini-pro] 왼쪽 눈동자가 하얗게 변색되어 인체 해부학 및 표정 지시사항을 위반함 / [gpt-high] 화면 오른쪽에 검은 바지와 부츠를 착용한 별도 인물의 하체를 추가했다. 백산 이외의 사람이나 신체 일부를 금지한 장면 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L229B03.png",
    "asset_id": "1f832f0c-4b90-4195-8434-227eb2c715f6",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 백산: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1241731>",
    "asset_id": "b337b8d8-94d9-4a29-9e49-2e19121379b7",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-80d9-7c8c-bb39-a42ceacc0a59",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S60sh52__bgfirst_bg.png",
   "bg_asset_id": "aeb46e86-9443-4079-ae58-e593ffa40546",
   "bg_record_key": "S60sh52::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S60sh52::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:41:37.483985+00:00",
  "fingerprint": "8221d7edbe79ed13b09bf7f135deae5d286b9966251af354b6062f0d75870506",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S60sh52_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S60sh52_sel.png",
  "source_sha256": "982e0c9207d847abd46d972bc6001f31a855c763c19b170eb7f80bcdbcbd24b1",
  "file": "S60sh52_cine.png",
  "staged_sha256": "01cd49e274f715263a82fb793987977dc48b3a8e6944dc75249954be6c787175",
  "latency_ms": 12329
 },
 "S60sh63::signage": {
  "fp": "5b70c7e5af8ef902",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S60sh63::bgfirst_bg": {
  "input_fingerprint": "1da2083e740ebd82",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 노을 빛 아래, 서로를 향해 뻗은 현우의 거친 손과 찰리의 육중한 금속 손이 단단히 맞잡힌 찰나의 클로즈업.\n\nLOCATION (lock): On a ridge overlooking the damaged village at sunset, where the youth and robot sit together.\n\nTIME OF DAY (lock): sunset.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 능선의 지면 (Visible only in the small interval beneath the seated pair's forearms); used as Retains the physical setting behind the handclasp without competing with it.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Soft sunset warmth joins the tactile detail of 현우's skin and 찰리's metal hand with gentle contrast and an intimate, unforced tenderness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 노을 빛 아래, 서로를 향해 뻗은 현우의 거친 손과 찰리의 육중한 금속 손이 단단히 맞잡힌 찰나의 클로즈업.\n\nLOCATION (lock): On a ridge overlooking the damaged village at sunset, where the youth and robot sit together.\n\nTIME OF DAY (lock): sunset.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 능선의 지면 (Visible only in the small interval beneath the seated pair's forearms); used as Retains the physical setting behind the handclasp without competing with it.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Soft sunset warmth joins the tactile detail of 현우's skin and 찰리's metal hand with gentle contrast and an intimate, unforced tenderness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S60sh63__bgfirst_bg.png",
  "asset_id": "1a5d48e7-71af-4630-8d5b-960f4796ee43",
  "input_asset_ids": [
   "3f4ba729-6aad-496e-a709-774fcddbcb83",
   "5382b8c1-3827-43d2-bb92-334028ead33d"
  ]
 },
 "S60sh63": {
  "input_fingerprint": "69f0beab193dd836",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 노을 빛 아래, 서로를 향해 뻗은 현우의 거친 손과 찰리의 육중한 금속 손이 단단히 맞잡힌 찰나의 클로즈업.\n\nLOCATION (lock): On a ridge overlooking the damaged village at sunset, where the youth and robot sit together. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 능선의 지면 (Visible only in the small interval beneath the seated pair's forearms); used as Retains the physical setting behind the handclasp without competing with it.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Soft sunset warmth joins the tactile detail of 현우's skin and 찰리's metal hand with gentle contrast and an intimate, unforced tenderness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Sunset lights the ridge above the devastated village. Charlie remains badly dented and punctured, with malfunctioning sensors, and extends a metal hand. 현우: He sits on the ridge, visibly battered from the fighting, with a brighter expression and one hand extended.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 노을 빛 아래, 서로를 향해 뻗은 현우의 거친 손과 찰리의 육중한 금속 손이 단단히 맞잡힌 찰나의 클로즈업.\n\nLOCATION (lock): On a ridge overlooking the damaged village at sunset, where the youth and robot sit together. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 능선의 지면 (Visible only in the small interval beneath the seated pair's forearms); used as Retains the physical setting behind the handclasp without competing with it.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Soft sunset warmth joins the tactile detail of 현우's skin and 찰리's metal hand with gentle contrast and an intimate, unforced tenderness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Sunset lights the ridge above the devastated village. Charlie remains badly dented and punctured, with malfunctioning sensors, and extends a metal hand. 현우: He sits on the ridge, visibly battered from the fighting, with a brighter expression and one hand extended.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 노을 빛 아래, 서로를 향해 뻗은 현우의 거친 손과 찰리의 육중한 금속 손이 단단히 맞잡힌 찰나의 클로즈업.\n\nLOCATION (lock): On a ridge overlooking the damaged village at sunset, where the youth and robot sit together. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 능선의 지면 (Visible only in the small interval beneath the seated pair's forearms); used as Retains the physical setting behind the handclasp without competing with it.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Soft sunset warmth joins the tactile detail of 현우's skin and 찰리's metal hand with gentle contrast and an intimate, unforced tenderness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Sunset lights the ridge above the devastated village. Charlie remains badly dented and punctured, with malfunctioning sensors, and extends a metal hand. 현우: He sits on the ridge, visibly battered from the fighting, with a brighter expression and one hand extended.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S60sh63__bgfirst_bg.png",
     "asset_id": "1a5d48e7-71af-4630-8d5b-960f4796ee43",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S60sh63.png",
     "asset_id": "3f4ba729-6aad-496e-a709-774fcddbcb83",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L229B02.png",
     "asset_id": "5382b8c1-3827-43d2-bb92-334028ead33d",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "두 사람의 팔이 화면 중앙을 향해 뻗어 있으며, 양손이 단단히 맞잡힌 상태를 향함.",
    "built_space": "산등성이 위. 뒤로 노을 진 산맥이 보이며, 두 인물이 지면에 안정적으로 앉아 있음.",
    "entities": "현우(상처 난 거친 팔, 낡은 옷), 찰리(샌드 베이지 장갑판, 점과 선이 있는 흰색 마스크형 얼굴 일부).",
    "hard_violations": [],
    "physics": "지면에 앉아 체중을 지탱하고 있으며, 맞잡은 두 손은 양쪽 팔과 어깨의 근육 및 관절로 자연스럽게 지탱됨."
   },
   {
    "label": "B",
    "direction": "화면 양끝에서 뻗어 나온 두 팔이 중앙에서 맞닿아 있음.",
    "built_space": "산등성이 위. 아래로 파괴된 마을과 산맥이 보이나, 원근감이 다소 왜곡되어 마을이 미니어처처럼 보임.",
    "entities": "현우(두꺼운 팔과 거친 피부), 찰리(샌드 베이지 금속 손과 팔).",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 해부학: 사람의 손에 손가락이 묘사되지 않고 절단된 듯 뭉툭하게 로봇 손과 융합됨."
    ],
    "physics": "공중에 뜬 두 손이 맞닿아 있으나, 화면 밖으로 이어지는 양팔에 의해 지탱됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "클로즈업 지시보다 넓게 샷이 잡혀 인물들의 상반신과 다리가 노출되었으나, 해부학적으로 정확하고 완벽한 악수 장면을 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "손에 집중한 클로즈업 프레임은 잘 따랐으나, 사람의 손가락이 묘사되지 않은 치명적인 해부학적 오류가 있음."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 사람의 팔이 화면 중앙을 향해 뻗어 있으며, 양손이 단단히 맞잡힌 상태를 향함.",
        "built_space": "산등성이 위. 뒤로 노을 진 산맥이 보이며, 두 인물이 지면에 안정적으로 앉아 있음.",
        "entities": "현우(상처 난 거친 팔, 낡은 옷), 찰리(샌드 베이지 장갑판, 점과 선이 있는 흰색 마스크형 얼굴 일부).",
        "hard_violations": [],
        "physics": "지면에 앉아 체중을 지탱하고 있으며, 맞잡은 두 손은 양쪽 팔과 어깨의 근육 및 관절로 자연스럽게 지탱됨."
       },
       {
        "label": "B",
        "direction": "화면 양끝에서 뻗어 나온 두 팔이 중앙에서 맞닿아 있음.",
        "built_space": "산등성이 위. 아래로 파괴된 마을과 산맥이 보이나, 원근감이 다소 왜곡되어 마을이 미니어처처럼 보임.",
        "entities": "현우(두꺼운 팔과 거친 피부), 찰리(샌드 베이지 금속 손과 팔).",
        "hard_violations": [
         "물리적으로 불가능한 해부학: 사람의 손에 손가락이 묘사되지 않고 절단된 듯 뭉툭하게 로봇 손과 융합됨."
        ],
        "physics": "공중에 뜬 두 손이 맞닿아 있으나, 화면 밖으로 이어지는 양팔에 의해 지탱됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "클로즈업 지시보다 넓게 샷이 잡혀 인물들의 상반신과 다리가 노출되었으나, 해부학적으로 정확하고 완벽한 악수 장면을 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "손에 집중한 클로즈업 프레임은 잘 따랐으나, 사람의 손가락이 묘사되지 않은 치명적인 해부학적 오류가 있음."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "두 사람의 팔이 화면 중앙을 향해 뻗어 있으며, 양손이 단단히 맞잡힌 상태를 향함.",
        "built_space": "산등성이 위. 뒤로 노을 진 산맥이 보이며, 두 인물이 지면에 안정적으로 앉아 있음.",
        "entities": "현우(상처 난 거친 팔, 낡은 옷), 찰리(샌드 베이지 장갑판, 점과 선이 있는 흰색 마스크형 얼굴 일부).",
        "hard_violations": [],
        "physics": "지면에 앉아 체중을 지탱하고 있으며, 맞잡은 두 손은 양쪽 팔과 어깨의 근육 및 관절로 자연스럽게 지탱됨."
       },
       {
        "label": "B",
        "direction": "화면 양끝에서 뻗어 나온 두 팔이 중앙에서 맞닿아 있음.",
        "built_space": "산등성이 위. 아래로 파괴된 마을과 산맥이 보이나, 원근감이 다소 왜곡되어 마을이 미니어처처럼 보임.",
        "entities": "현우(두꺼운 팔과 거친 피부), 찰리(샌드 베이지 금속 손과 팔).",
        "hard_violations": [
         "물리적으로 불가능한 해부학: 사람의 손에 손가락이 묘사되지 않고 절단된 듯 뭉툭하게 로봇 손과 융합됨."
        ],
        "physics": "공중에 뜬 두 손이 맞닿아 있으나, 화면 밖으로 이어지는 양팔에 의해 지탱됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "손의 접촉과 손상된 금속 질감은 맞지만, 마을·하늘·넓은 지면이 드러나 손 맞잡기의 밀착된 구도와 제한된 배경 지시에서 더 멀다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "단단히 맞잡은 손과 양옆의 앉은 자세, 팔 아래 좁은 지면이 더 충실하지만, 팔 위로 넓게 보이는 노을 풍경은 배경 제한을 벗어난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 팔은 왼쪽에서 오른쪽으로, 찰리의 팔은 오른쪽에서 왼쪽으로 뻗어 중앙에서 만난다. 금속 손가락이 사람 손을 아래와 위에서 감싸지만, 현우의 손가락은 비교적 길게 펴져 있어 상호 악력은 덜 뚜렷하다. 얼굴과 시선은 보이지 않는다.",
        "built_space": "인공 좌석이나 고정 시설은 보이지 않는다. 돌과 마른 풀이 있는 능선 지면은 참고 장소와 어울리지만 화면 하단에 넓게 펼쳐지고, 손 뒤로 여러 파손 건물과 산, 하늘까지 노출된다. 지면을 두 팔 아래 작은 틈에만 보이게 하라는 구도와 다르다. 앉은 자세의 지지점은 프레임 밖이다.",
        "entities": "사람 손 하나와 육중한 로봇 손 하나가 보인다. 현우의 손과 팔에는 흙, 긁힘과 상처가 있으며, 손만으로 정확한 나이와 한국계 미국인 정체성을 확인할 수는 없다. 어두운 긴 소매는 참고의 남색 반소매와 다르다. 찰리의 샌드 베이지 장갑판과 검은 관절, 찍힘과 관통 흔적은 요구에 부합한다. 얼굴과 센서는 제외되어 평가할 수 없다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "사람 손은 손목과 팔뚝에, 금속 손은 손목 관절과 장갑 전완에 이어져 각각 지지된다. 두 손의 접촉과 금속 손가락의 굽힘은 가능한 자세이며, 독립적으로 떠 있는 물체는 없다. 좌면이나 골반이 보이지 않는 것은 이 클로즈업만으로 물리적 오류라고 판단할 수 없다."
       },
       {
        "label": "B",
        "direction": "왼쪽 현우의 손과 오른쪽 찰리의 손이 서로를 향해 뻗어 화면 중앙 아래에서 맞잡힌다. 사람 손가락과 금속 손가락이 상대 손을 감싸며 단단한 악수를 더 명확하게 만든다. 찰리 얼굴 일부는 왼쪽을 향하지만 눈은 잘려 있어 시선의 도착점은 확인되지 않는다.",
        "built_space": "왼쪽에는 현우의 굽힌 허벅지, 오른쪽에는 찰리의 무릎 장갑이 보여 능선에 함께 앉은 배치가 읽힌다. 두 인물 사이 손 아래에 돌과 마른 풀이 있는 지면이 좁게 보이며 참고 장소의 재질과 맞는다. 인공 좌석이나 고정 시설은 없다. 다만 팔 위에도 산과 하늘이 넓게 보여 배경을 작은 틈으로 제한한 지시는 완전히 지키지 못했다.",
        "entities": "상처 난 사람 손 하나와 큰 금속 손 하나가 중심이다. 사람 팔은 청년 남성 설정과 모순되지 않지만 얼굴이 없어 현우의 정확한 신원은 확인할 수 없다. 해진 회갈색 긴 소매는 참고의 남색 반소매와 다르다. 찰리는 베이지 장갑, 검은 기계 관절, 흰 각진 얼굴 일부가 참고와 일치하며 장갑에 긁힘과 패인 손상이 보인다. 센서 고장은 이 크롭에서 확인할 수 없다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 손은 각자의 전완과 손목 관절에 연결되어 있고, 구부러진 손가락이 상대 손에 접촉한다. 팔을 앞으로 내밀어 악수하는 동작으로 가능한 구조다. 양옆의 굽힌 다리와 바로 아래 능선 지면은 앉은 배치를 뒷받침한다. 골반 접촉점은 잘렸지만 근거 없이 공중에 떠 있는 신체나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "손의 접촉과 손상된 금속 질감은 맞지만, 마을·하늘·넓은 지면이 드러나 손 맞잡기의 밀착된 구도와 제한된 배경 지시에서 더 멀다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "단단히 맞잡은 손과 양옆의 앉은 자세, 팔 아래 좁은 지면이 더 충실하지만, 팔 위로 넓게 보이는 노을 풍경은 배경 제한을 벗어난다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 팔은 왼쪽에서 오른쪽으로, 찰리의 팔은 오른쪽에서 왼쪽으로 뻗어 중앙에서 만난다. 금속 손가락이 사람 손을 아래와 위에서 감싸지만, 현우의 손가락은 비교적 길게 펴져 있어 상호 악력은 덜 뚜렷하다. 얼굴과 시선은 보이지 않는다.",
        "built_space": "인공 좌석이나 고정 시설은 보이지 않는다. 돌과 마른 풀이 있는 능선 지면은 참고 장소와 어울리지만 화면 하단에 넓게 펼쳐지고, 손 뒤로 여러 파손 건물과 산, 하늘까지 노출된다. 지면을 두 팔 아래 작은 틈에만 보이게 하라는 구도와 다르다. 앉은 자세의 지지점은 프레임 밖이다.",
        "entities": "사람 손 하나와 육중한 로봇 손 하나가 보인다. 현우의 손과 팔에는 흙, 긁힘과 상처가 있으며, 손만으로 정확한 나이와 한국계 미국인 정체성을 확인할 수는 없다. 어두운 긴 소매는 참고의 남색 반소매와 다르다. 찰리의 샌드 베이지 장갑판과 검은 관절, 찍힘과 관통 흔적은 요구에 부합한다. 얼굴과 센서는 제외되어 평가할 수 없다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "사람 손은 손목과 팔뚝에, 금속 손은 손목 관절과 장갑 전완에 이어져 각각 지지된다. 두 손의 접촉과 금속 손가락의 굽힘은 가능한 자세이며, 독립적으로 떠 있는 물체는 없다. 좌면이나 골반이 보이지 않는 것은 이 클로즈업만으로 물리적 오류라고 판단할 수 없다."
       },
       {
        "label": "A",
        "direction": "왼쪽 현우의 손과 오른쪽 찰리의 손이 서로를 향해 뻗어 화면 중앙 아래에서 맞잡힌다. 사람 손가락과 금속 손가락이 상대 손을 감싸며 단단한 악수를 더 명확하게 만든다. 찰리 얼굴 일부는 왼쪽을 향하지만 눈은 잘려 있어 시선의 도착점은 확인되지 않는다.",
        "built_space": "왼쪽에는 현우의 굽힌 허벅지, 오른쪽에는 찰리의 무릎 장갑이 보여 능선에 함께 앉은 배치가 읽힌다. 두 인물 사이 손 아래에 돌과 마른 풀이 있는 지면이 좁게 보이며 참고 장소의 재질과 맞는다. 인공 좌석이나 고정 시설은 없다. 다만 팔 위에도 산과 하늘이 넓게 보여 배경을 작은 틈으로 제한한 지시는 완전히 지키지 못했다.",
        "entities": "상처 난 사람 손 하나와 큰 금속 손 하나가 중심이다. 사람 팔은 청년 남성 설정과 모순되지 않지만 얼굴이 없어 현우의 정확한 신원은 확인할 수 없다. 해진 회갈색 긴 소매는 참고의 남색 반소매와 다르다. 찰리는 베이지 장갑, 검은 기계 관절, 흰 각진 얼굴 일부가 참고와 일치하며 장갑에 긁힘과 패인 손상이 보인다. 센서 고장은 이 크롭에서 확인할 수 없다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 손은 각자의 전완과 손목 관절에 연결되어 있고, 구부러진 손가락이 상대 손에 접촉한다. 팔을 앞으로 내밀어 악수하는 동작으로 가능한 구조다. 양옆의 굽힌 다리와 바로 아래 능선 지면은 앉은 배치를 뒷받침한다. 골반 접촉점은 잘렸지만 근거 없이 공중에 떠 있는 신체나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.179
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.929
   },
   "violations": {
    "B": [
     "[gemini-pro] 물리적으로 불가능한 해부학: 사람의 손에 손가락이 묘사되지 않고 절단된 듯 뭉툭하게 로봇 손과 융합됨."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 929
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "클로즈업 지시보다 넓게 샷이 잡혀 인물들의 상반신과 다리가 노출되었으나, 해부학적으로 정확하고 완벽한 악수 장면을 구현함."
   },
   {
    "label": "B",
    "score": 929,
    "verdict_ko": "손에 집중한 클로즈업 프레임은 잘 따랐으나, 사람의 손가락이 묘사되지 않은 치명적인 해부학적 오류가 있음.  ★위반: [gemini-pro] 물리적으로 불가능한 해부학: 사람의 손에 손가락이 묘사되지 않고 절단된 듯 뭉툭하게 로봇 손과 융합됨."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L229B02.png",
    "asset_id": "5382b8c1-3827-43d2-bb92-334028ead33d",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-8433-76bd-8e88-3eafce215d47",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S60sh63__bgfirst_bg.png",
   "bg_asset_id": "1a5d48e7-71af-4630-8d5b-960f4796ee43",
   "bg_record_key": "S60sh63::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S60sh63::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:44:01.675220+00:00",
  "fingerprint": "6f556f1302b234ba72c1c404ee56c001a4ffb8146d963a6485f348461a8017fc",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S60sh63_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S60sh63_sel.png",
  "source_sha256": "282fd52dbe6386b47ed932327ab02221aa3fc73f1a06b736919d1b6ac30f999e",
  "file": "S60sh63_cine.png",
  "staged_sha256": "a63729ccb139461f6016d1dcd5d206f7ae5163094eb364e8cbfb3a913338b021",
  "latency_ms": 13087
 },
 "S61sh1::signage": {
  "fp": "3d752f0da3adcbcb",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S61sh1": {
  "input_fingerprint": "ef931fbedb2e1656",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, moonlit.\n\nSHOT TEXT (authoritative, Korean): 어두운 밤, 늪지대 진흙 속에 박힌 낡은 캠핑카의 외관 전경.\n\nLOCATION (lock): At the marsh's muddy vehicle-stranding point at night, where the abandoned camper remains sunk in the ground. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 낡은 캠핑카 (Old and lodged in the swamp mud) — The near end, adjacent side, and part of the upper body are visible from the elevated diagonal viewpoint; used as Establishes the discovered vehicle at a readable environmental scale; 늪지대 진흙 (Surrounding and holding the vehicle's wheels); used as Occupies the foreground and side margins, explaining why the vehicle remains here.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the scene's dark nighttime ambience with restrained tonal separation sufficient to distinguish the vehicle from the swamp, without adding an unestablished light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old camper remains stuck in the swamp at night, with no successful extraction established. A rusty radiation-zone sign stands nearby, with crows gathered around it.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, moonlit.\n\nSHOT TEXT (authoritative, Korean): 어두운 밤, 늪지대 진흙 속에 박힌 낡은 캠핑카의 외관 전경.\n\nLOCATION (lock): At the marsh's muddy vehicle-stranding point at night, where the abandoned camper remains sunk in the ground. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 낡은 캠핑카 (Old and lodged in the swamp mud) — The near end, adjacent side, and part of the upper body are visible from the elevated diagonal viewpoint; used as Establishes the discovered vehicle at a readable environmental scale; 늪지대 진흙 (Surrounding and holding the vehicle's wheels); used as Occupies the foreground and side margins, explaining why the vehicle remains here.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the scene's dark nighttime ambience with restrained tonal separation sufficient to distinguish the vehicle from the swamp, without adding an unestablished light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old camper remains stuck in the swamp at night, with no successful extraction established. A rusty radiation-zone sign stands nearby, with crows gathered around it.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, moonlit.\n\nSHOT TEXT (authoritative, Korean): 어두운 밤, 늪지대 진흙 속에 박힌 낡은 캠핑카의 외관 전경.\n\nLOCATION (lock): At the marsh's muddy vehicle-stranding point at night, where the abandoned camper remains sunk in the ground. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 낡은 캠핑카 (Old and lodged in the swamp mud) — The near end, adjacent side, and part of the upper body are visible from the elevated diagonal viewpoint; used as Establishes the discovered vehicle at a readable environmental scale; 늪지대 진흙 (Surrounding and holding the vehicle's wheels); used as Occupies the foreground and side margins, explaining why the vehicle remains here.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the scene's dark nighttime ambience with restrained tonal separation sufficient to distinguish the vehicle from the swamp, without adding an unestablished light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old camper remains stuck in the swamp at night, with no successful extraction established. A rusty radiation-zone sign stands nearby, with crows gathered around it.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "특정 방향을 가리키는 요구나 시선 없음.",
    "built_space": "늪지대 진흙 속에 캠핑카 1대가 박혀 있고, 우측에 표지판 2개가 세워져 있음. 높은 대각선 시점이 잘 구현됨.",
    "entities": "낡은 캠핑카, 방사능 표지판, 까마귀, 진흙 등 프롬프트가 요구한 객체들이 모두 존재하며 인물은 등장하지 않음.",
    "hard_violations": [],
    "physics": "캠핑카는 진흙 지표면에 단단히 박혀 있으며, 까마귀들은 표지판에 자연스럽게 앉아 지지받고 있음."
   },
   {
    "label": "B",
    "direction": "특정 방향을 가리키는 요구나 시선 없음.",
    "built_space": "늪지대 환경에 캠핑카 1대와 표지판 2개가 있으나, 낮은 시점과 전경의 진흙 바닥 형태가 참조 이미지와 완전히 동일함.",
    "entities": "낡은 캠핑카, 방사능 표지판, 까마귀가 존재하며 인물은 없음.",
    "hard_violations": [],
    "physics": "캠핑카와 까마귀 모두 중력에 맞게 안정적으로 지지되어 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "요구된 높은 대각선 시점을 정확히 반영하였으며, 참조 이미지의 구도를 모방하지 않고 늪지대 환경과 야간 분위기를 성공적으로 연출함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "참조 이미지의 카메라 구도와 전경(진흙 바닥)을 그대로 복사하여 '높은 대각선 시점' 및 구도 복사 금지 지시를 명백히 위반함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "특정 방향을 가리키는 요구나 시선 없음.",
        "built_space": "늪지대 진흙 속에 캠핑카 1대가 박혀 있고, 우측에 표지판 2개가 세워져 있음. 높은 대각선 시점이 잘 구현됨.",
        "entities": "낡은 캠핑카, 방사능 표지판, 까마귀, 진흙 등 프롬프트가 요구한 객체들이 모두 존재하며 인물은 등장하지 않음.",
        "hard_violations": [],
        "physics": "캠핑카는 진흙 지표면에 단단히 박혀 있으며, 까마귀들은 표지판에 자연스럽게 앉아 지지받고 있음."
       },
       {
        "label": "B",
        "direction": "특정 방향을 가리키는 요구나 시선 없음.",
        "built_space": "늪지대 환경에 캠핑카 1대와 표지판 2개가 있으나, 낮은 시점과 전경의 진흙 바닥 형태가 참조 이미지와 완전히 동일함.",
        "entities": "낡은 캠핑카, 방사능 표지판, 까마귀가 존재하며 인물은 없음.",
        "hard_violations": [],
        "physics": "캠핑카와 까마귀 모두 중력에 맞게 안정적으로 지지되어 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "요구된 높은 대각선 시점을 정확히 반영하였으며, 참조 이미지의 구도를 모방하지 않고 늪지대 환경과 야간 분위기를 성공적으로 연출함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "참조 이미지의 카메라 구도와 전경(진흙 바닥)을 그대로 복사하여 '높은 대각선 시점' 및 구도 복사 금지 지시를 명백히 위반함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "특정 방향을 가리키는 요구나 시선 없음.",
        "built_space": "늪지대 진흙 속에 캠핑카 1대가 박혀 있고, 우측에 표지판 2개가 세워져 있음. 높은 대각선 시점이 잘 구현됨.",
        "entities": "낡은 캠핑카, 방사능 표지판, 까마귀, 진흙 등 프롬프트가 요구한 객체들이 모두 존재하며 인물은 등장하지 않음.",
        "hard_violations": [],
        "physics": "캠핑카는 진흙 지표면에 단단히 박혀 있으며, 까마귀들은 표지판에 자연스럽게 앉아 지지받고 있음."
       },
       {
        "label": "B",
        "direction": "특정 방향을 가리키는 요구나 시선 없음.",
        "built_space": "늪지대 환경에 캠핑카 1대와 표지판 2개가 있으나, 낮은 시점과 전경의 진흙 바닥 형태가 참조 이미지와 완전히 동일함.",
        "entities": "낡은 캠핑카, 방사능 표지판, 까마귀가 존재하며 인물은 없음.",
        "hard_violations": [],
        "physics": "캠핑카와 까마귀 모두 중력에 맞게 안정적으로 지지되어 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "밤의 늪과 고착된 캠핑카는 구현했지만, 높은 사선 시점이 약하고 표지판이 차량 뒤쪽에 놓여 참조의 장소 배치 연속성이 떨어집니다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "높은 사선의 와이드숏에서 차량 앞면·측면·지붕과 매몰된 바퀴를 보여주며, 앞쪽의 두 표지판과 까마귀까지 참조 장소에 가깝게 유지합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "캠핑카 앞쪽은 화면 오른쪽을 향하고 카메라에는 뒤쪽 끝과 출입문이 있는 측면이 보입니다. 까마귀들은 표지판과 지붕 위에서 서로 다른 방향을 보고 있습니다. 요구된 조준이나 특정 시선 대상은 없습니다.",
        "built_space": "차량 한 대의 뒤창 하나, 측면 큰 창 두 개, 창이 달린 문 하나, 뒤쪽 사다리 하나와 지붕 설비가 보입니다. 원형 표지판 하나와 방사능 표지판 하나는 모두 차량 왼쪽 뒤쪽에 있습니다. 참조의 차량 앞쪽 오른편 표지판 배치와 차이가 있습니다. 진흙·고인 물·갈대와 전경 바퀴 자국은 유지되지만, 시점은 비교적 낮아 지붕이 조금만 드러납니다.",
        "entities": "낡고 녹슨 밝은색 캠핑카, 붉은 가로 띠, 진흙 늪, 녹슨 방사능 표지판, 원형 진입금지 표지판과 여러 까마귀가 있습니다. 사람이나 얼굴은 없습니다. 달과 차가운 야간 조명이 보이며 읽을 수 있는 글자는 없습니다. 차량의 뒤쪽 형상과 사다리는 참조에서 확인되지 않는 부분입니다.",
        "hard_violations": [],
        "physics": "차량 바퀴와 하부가 진흙에 잠겨 지면에 지지되며 견인되거나 떠 있지 않습니다. 표지판 기둥은 땅에 박혀 있습니다. 까마귀는 표지판 윗부분과 차량 지붕에 발을 대고 앉아 있고, 지면의 새도 진흙 위에 서 있습니다. 고인 물의 반사도 수면 위치와 부합합니다."
       },
       {
        "label": "B",
        "direction": "캠핑카 앞면은 화면 오른쪽 아래를 향하며, 카메라는 높은 사선에서 앞면과 출입문 쪽 측면을 내려다봅니다. 표지판 앞면도 카메라 쪽으로 비스듬히 보입니다. 까마귀들은 표지판 주변에서 좌우를 바라보거나 짧게 날고 있으며, 특정 대상을 향해야 하는 행동은 없습니다.",
        "built_space": "캠핑카 한 대에 앞 유리, 측면 큰 창 두 개, 창이 달린 출입문 하나와 지붕의 큰 설비 하나가 보입니다. 원형 표지판 하나와 기울어진 사각 방사능 표지판 하나가 차량 앞쪽 오른편에 있어 참조의 상대 배치를 잘 유지합니다. 차량 주변과 전경에 진흙과 얕은 물이 넓게 확보되고, 높은 사선 시점에서 지붕까지 분명히 보입니다. 표지판의 크기도 가까운 위치로 설명됩니다.",
        "entities": "참조와 유사한 전면 형상, 밝은 차체, 붉은 측면 띠와 부식 흔적을 지닌 낡은 캠핑카가 있습니다. 늪 진흙, 갈대, 녹슨 방사능 표지판과 원형 표지판, 그 주변의 까마귀가 모두 보입니다. 사람은 없고 읽을 수 있는 글자도 없습니다. 구름 사이 달빛에 의한 어두운 밤으로 표현되어 있습니다.",
        "hard_violations": [],
        "physics": "차량의 바퀴와 하부가 진흙에 깊이 묻혀 있고 차체는 지면에 지지됩니다. 두 표지판은 흙에 박힌 기둥으로 지탱됩니다. 앉은 까마귀는 표지판 가장자리를 발로 딛고 있으며 지면의 새는 땅 위에 있습니다. 공중의 새는 날개를 펼친 비행 자세여서 지지 없이 정지한 물체로 보이지 않습니다. 수면의 광택과 반사는 자연스럽습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "밤의 늪과 고착된 캠핑카는 구현했지만, 높은 사선 시점이 약하고 표지판이 차량 뒤쪽에 놓여 참조의 장소 배치 연속성이 떨어집니다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "높은 사선의 와이드숏에서 차량 앞면·측면·지붕과 매몰된 바퀴를 보여주며, 앞쪽의 두 표지판과 까마귀까지 참조 장소에 가깝게 유지합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "캠핑카 앞쪽은 화면 오른쪽을 향하고 카메라에는 뒤쪽 끝과 출입문이 있는 측면이 보입니다. 까마귀들은 표지판과 지붕 위에서 서로 다른 방향을 보고 있습니다. 요구된 조준이나 특정 시선 대상은 없습니다.",
        "built_space": "차량 한 대의 뒤창 하나, 측면 큰 창 두 개, 창이 달린 문 하나, 뒤쪽 사다리 하나와 지붕 설비가 보입니다. 원형 표지판 하나와 방사능 표지판 하나는 모두 차량 왼쪽 뒤쪽에 있습니다. 참조의 차량 앞쪽 오른편 표지판 배치와 차이가 있습니다. 진흙·고인 물·갈대와 전경 바퀴 자국은 유지되지만, 시점은 비교적 낮아 지붕이 조금만 드러납니다.",
        "entities": "낡고 녹슨 밝은색 캠핑카, 붉은 가로 띠, 진흙 늪, 녹슨 방사능 표지판, 원형 진입금지 표지판과 여러 까마귀가 있습니다. 사람이나 얼굴은 없습니다. 달과 차가운 야간 조명이 보이며 읽을 수 있는 글자는 없습니다. 차량의 뒤쪽 형상과 사다리는 참조에서 확인되지 않는 부분입니다.",
        "hard_violations": [],
        "physics": "차량 바퀴와 하부가 진흙에 잠겨 지면에 지지되며 견인되거나 떠 있지 않습니다. 표지판 기둥은 땅에 박혀 있습니다. 까마귀는 표지판 윗부분과 차량 지붕에 발을 대고 앉아 있고, 지면의 새도 진흙 위에 서 있습니다. 고인 물의 반사도 수면 위치와 부합합니다."
       },
       {
        "label": "A",
        "direction": "캠핑카 앞면은 화면 오른쪽 아래를 향하며, 카메라는 높은 사선에서 앞면과 출입문 쪽 측면을 내려다봅니다. 표지판 앞면도 카메라 쪽으로 비스듬히 보입니다. 까마귀들은 표지판 주변에서 좌우를 바라보거나 짧게 날고 있으며, 특정 대상을 향해야 하는 행동은 없습니다.",
        "built_space": "캠핑카 한 대에 앞 유리, 측면 큰 창 두 개, 창이 달린 출입문 하나와 지붕의 큰 설비 하나가 보입니다. 원형 표지판 하나와 기울어진 사각 방사능 표지판 하나가 차량 앞쪽 오른편에 있어 참조의 상대 배치를 잘 유지합니다. 차량 주변과 전경에 진흙과 얕은 물이 넓게 확보되고, 높은 사선 시점에서 지붕까지 분명히 보입니다. 표지판의 크기도 가까운 위치로 설명됩니다.",
        "entities": "참조와 유사한 전면 형상, 밝은 차체, 붉은 측면 띠와 부식 흔적을 지닌 낡은 캠핑카가 있습니다. 늪 진흙, 갈대, 녹슨 방사능 표지판과 원형 표지판, 그 주변의 까마귀가 모두 보입니다. 사람은 없고 읽을 수 있는 글자도 없습니다. 구름 사이 달빛에 의한 어두운 밤으로 표현되어 있습니다.",
        "hard_violations": [],
        "physics": "차량의 바퀴와 하부가 진흙에 깊이 묻혀 있고 차체는 지면에 지지됩니다. 두 표지판은 흙에 박힌 기둥으로 지탱됩니다. 앉은 까마귀는 표지판 가장자리를 발로 딛고 있으며 지면의 새는 땅 위에 있습니다. 공중의 새는 날개를 펼친 비행 자세여서 지지 없이 정지한 물체로 보이지 않습니다. 수면의 광택과 반사는 자연스럽습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.095
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.095
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1095
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "요구된 높은 대각선 시점을 정확히 반영하였으며, 참조 이미지의 구도를 모방하지 않고 늪지대 환경과 야간 분위기를 성공적으로 연출함."
   },
   {
    "label": "B",
    "score": 1095,
    "verdict_ko": "참조 이미지의 카메라 구도와 전경(진흙 바닥)을 그대로 복사하여 '높은 대각선 시점' 및 구도 복사 금지 지시를 명백히 위반함."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S56sh5_sel.png",
    "asset_id": "d4c32470-0402-4a2a-9107-1b88173ecf42",
    "role": "prev_still"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-878b-74bf-93b3-2ffc1b8b400f",
  "ref_mode": "prev만 (배경 전용·공유 계획)",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S56sh5"
  },
  "lane_policy": "ab_select_bypass:bg_only:share_plan_prev_bgonly"
 },
 "S61sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:45:10.941078+00:00",
  "fingerprint": "4b2d1096002c281b17190a369a7bb2174984bfb271e2b7599d42fc4ac1db5133",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S61sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S61sh1_sel.png",
  "source_sha256": "6b199b9f8662adb91cf9ffde624676962eac8391c2e1ec051a88913b1fbdb3bf",
  "file": "S61sh1_cine.png",
  "staged_sha256": "5fa1b519a45dd2068aac71660bc9928031272c4828488f0efaa40ab63aae5836",
  "latency_ms": 11573
 },
 "S61sh5::signage": {
  "fp": "f2414ebd8b88048f",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S61sh5": {
  "input_fingerprint": "fa371b19545641b2",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, moonlit.\n\nSHOT TEXT (authoritative, Korean): 한쪽 입꼬리를 비스듬히 올린 채 거만하게 웃는 박철진의 얼굴 클로즈업.\n\nLOCATION (lock): Beside the stranded camper in the dark marsh, near a corroded radiation warning sign. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 자동차 옆면 일부 (Stationary beside the inspection) — A small, softly focused section of the near side remains behind 박철진's shoulder; used as Connects the close-up to the preceding approach without revealing additional people.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the established nighttime ambience with controlled facial contrast that makes the crooked smile readable without introducing a new source or color cue.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the camper embedded in swamp mud and the same dark nighttime environment. Exclude the children and any daytime sunlight from the earlier swamp sequence.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old camper remains stranded in swamp mud beside the rusty radiation-zone sign. It is night, and crows remain gathered near the sign before taking flight. 박철진: He remains at the swamp inspection site.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, moonlit.\n\nSHOT TEXT (authoritative, Korean): 한쪽 입꼬리를 비스듬히 올린 채 거만하게 웃는 박철진의 얼굴 클로즈업.\n\nLOCATION (lock): Beside the stranded camper in the dark marsh, near a corroded radiation warning sign. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 자동차 옆면 일부 (Stationary beside the inspection) — A small, softly focused section of the near side remains behind 박철진's shoulder; used as Connects the close-up to the preceding approach without revealing additional people.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the established nighttime ambience with controlled facial contrast that makes the crooked smile readable without introducing a new source or color cue.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the camper embedded in swamp mud and the same dark nighttime environment. Exclude the children and any daytime sunlight from the earlier swamp sequence.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old camper remains stranded in swamp mud beside the rusty radiation-zone sign. It is night, and crows remain gathered near the sign before taking flight. 박철진: He remains at the swamp inspection site.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, moonlit.\n\nSHOT TEXT (authoritative, Korean): 한쪽 입꼬리를 비스듬히 올린 채 거만하게 웃는 박철진의 얼굴 클로즈업.\n\nLOCATION (lock): Beside the stranded camper in the dark marsh, near a corroded radiation warning sign. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 자동차 옆면 일부 (Stationary beside the inspection) — A small, softly focused section of the near side remains behind 박철진's shoulder; used as Connects the close-up to the preceding approach without revealing additional people.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the established nighttime ambience with controlled facial contrast that makes the crooked smile readable without introducing a new source or color cue.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the camper embedded in swamp mud and the same dark nighttime environment. Exclude the children and any daytime sunlight from the earlier swamp sequence.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old camper remains stranded in swamp mud beside the rusty radiation-zone sign. It is night, and crows remain gathered near the sign before taking flight. 박철진: He remains at the swamp inspection site.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "시선은 카메라 우측 밖의 대상을 향하고 있음.",
    "built_space": "어두운 늪지대 환경이며, 인물 뒤로 지시된 캠핑카의 측면 일부가 부드럽게 아웃포커싱되어 배치됨.",
    "entities": "레퍼런스와 이목구비 및 헤어스타일이 일치하는 박철진. 프롬프트가 요구한 대로 한쪽 입꼬리가 비스듬히 올라간 거만한 미소를 명확하게 짓고 있음. 의상은 검은색 재킷으로 변경됨.",
    "hard_violations": [],
    "physics": "인물이 땅에 안정적으로 서 있는 자세임."
   },
   {
    "label": "B",
    "direction": "시선은 카메라 우측을 향하고 있음.",
    "built_space": "인물 뒤쪽으로 캠핑카의 측면과 창문이 아웃포커싱된 상태로 자리 잡고 있음.",
    "entities": "레퍼런스와 일치하는 박철진의 얼굴과 헤어스타일, 그리고 동일한 정장과 넥타이를 착용함. 미소를 띠고 있으나 한쪽 입꼬리가 비스듬히 올라간 거만한 느낌이 다소 부족함.",
    "hard_violations": [],
    "physics": "자연스럽게 서 있는 상태로 물리적 오류 없음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "샷 텍스트가 명시한 '한쪽 입꼬리를 비스듬히 올린 거만한 웃음'을 완벽하게 표현하였으며, 기존 늪지대의 야간 조명 분위기를 지시대로 자연스럽게 유지하여 가장 훌륭한 결과물입니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "레퍼런스의 정장 의상은 정확히 반영했으나, 요구된 거만한 표정 연출이 밋밋하고 얼굴에 새로운 강한 조명이 더해져 지정된 야간 무드를 해쳤습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 카메라 우측 밖의 대상을 향하고 있음.",
        "built_space": "어두운 늪지대 환경이며, 인물 뒤로 지시된 캠핑카의 측면 일부가 부드럽게 아웃포커싱되어 배치됨.",
        "entities": "레퍼런스와 이목구비 및 헤어스타일이 일치하는 박철진. 프롬프트가 요구한 대로 한쪽 입꼬리가 비스듬히 올라간 거만한 미소를 명확하게 짓고 있음. 의상은 검은색 재킷으로 변경됨.",
        "hard_violations": [],
        "physics": "인물이 땅에 안정적으로 서 있는 자세임."
       },
       {
        "label": "B",
        "direction": "시선은 카메라 우측을 향하고 있음.",
        "built_space": "인물 뒤쪽으로 캠핑카의 측면과 창문이 아웃포커싱된 상태로 자리 잡고 있음.",
        "entities": "레퍼런스와 일치하는 박철진의 얼굴과 헤어스타일, 그리고 동일한 정장과 넥타이를 착용함. 미소를 띠고 있으나 한쪽 입꼬리가 비스듬히 올라간 거만한 느낌이 다소 부족함.",
        "hard_violations": [],
        "physics": "자연스럽게 서 있는 상태로 물리적 오류 없음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "샷 텍스트가 명시한 '한쪽 입꼬리를 비스듬히 올린 거만한 웃음'을 완벽하게 표현하였으며, 기존 늪지대의 야간 조명 분위기를 지시대로 자연스럽게 유지하여 가장 훌륭한 결과물입니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "레퍼런스의 정장 의상은 정확히 반영했으나, 요구된 거만한 표정 연출이 밋밋하고 얼굴에 새로운 강한 조명이 더해져 지정된 야간 무드를 해쳤습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 카메라 우측 밖의 대상을 향하고 있음.",
        "built_space": "어두운 늪지대 환경이며, 인물 뒤로 지시된 캠핑카의 측면 일부가 부드럽게 아웃포커싱되어 배치됨.",
        "entities": "레퍼런스와 이목구비 및 헤어스타일이 일치하는 박철진. 프롬프트가 요구한 대로 한쪽 입꼬리가 비스듬히 올라간 거만한 미소를 명확하게 짓고 있음. 의상은 검은색 재킷으로 변경됨.",
        "hard_violations": [],
        "physics": "인물이 땅에 안정적으로 서 있는 자세임."
       },
       {
        "label": "B",
        "direction": "시선은 카메라 우측을 향하고 있음.",
        "built_space": "인물 뒤쪽으로 캠핑카의 측면과 창문이 아웃포커싱된 상태로 자리 잡고 있음.",
        "entities": "레퍼런스와 일치하는 박철진의 얼굴과 헤어스타일, 그리고 동일한 정장과 넥타이를 착용함. 미소를 띠고 있으나 한쪽 입꼬리가 비스듬히 올라간 거만한 느낌이 다소 부족함.",
        "hard_violations": [],
        "physics": "자연스럽게 서 있는 상태로 물리적 오류 없음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "얼굴 중심의 클로즈업과 한쪽 입꼬리를 올린 거만한 미소가 명확하고, 참조의 정장·셔츠·넥타이까지 유지해 더 충실하다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "비대칭 미소와 달빛 아래 클로즈업은 잘 구현했지만, 참조의 정장 차림을 검은 지퍼 재킷으로 바꾼 점이 뚜렷한 불일치다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴은 화면 오른쪽으로 약간 돌아가 있고 눈도 카메라 오른쪽의 화면 밖을 향한다. 시선 대상은 보이지 않으며, 지문은 특정 대상을 지정하지 않는다. 화면 오른쪽 입꼬리가 더 올라가 비스듬한 미소가 읽힌다. 무기나 방향을 확인할 휴대 소품은 없다.",
        "built_space": "어깨 뒤 화면 오른쪽에 캠핑카 측면 일부가 흐릿하게 보인다. 창 하나가 있는 출입문 한 개, 손잡이와 측면의 적갈색 띠가 확인된다. 인물은 차 밖 옆에 있으며 차체와 겹치는 방식이 자연스럽다. 배경 차체가 다소 넓게 보이지만 차량 전체를 드러내지는 않는다. 표지판·까마귀·바퀴의 진흙 접촉부는 클로즈업 밖이므로 확인할 수 없다. 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "중년의 한국인 남성으로 설정된 참조와 부합하는 외형의 인물 한 명만 보인다. 짧게 넘긴 검은 머리, 얼굴 윤곽, 작은 귀걸이와 수염 자국이 참조에 가깝고, 남색 정장·흰 셔츠·줄무늬 넥타이를 유지한다. 낡고 얼룩진 밝은 캠핑카 측면과 어두운 청색 계열 야간 환경이 보인다. 다른 사람이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리는 목과 어깨에 정상적으로 연결되어 있고 옷은 몸 위에 자연스럽게 걸쳐져 있다. 하체와 발은 프레임 밖이라 지면 접촉은 확인되지 않지만, 공중에 뜨거나 비정상적으로 지지된 모습은 없다. 손에 든 물체나 운동 중인 물체도 없다."
       },
       {
        "label": "B",
        "direction": "얼굴과 눈은 화면 오른쪽의 보이지 않는 대상을 향한다. 특정 시선 대상이 지문에 정해져 있지 않아 방향상 충돌은 없다. 화면 오른쪽 입꼬리를 치켜올린 미소와 살짝 좁힌 눈이 거만한 표정을 만든다. 무기나 방향성 있는 휴대 소품은 없다.",
        "built_space": "인물 뒤 오른쪽에는 얼룩진 캠핑카 측면과 화면 가장자리에 잘린 창 한 개가 보이고, 왼쪽에는 흐릿한 습지 수면과 풀이 보인다. 인물은 차체 바깥 옆에 서 있는 배치로 읽힌다. 측면 패널이 배경에서 비교적 넓은 면적을 차지한다. 표지판·까마귀·차량 하부는 프레임 밖이며, 보이는 범위에 중복 설비나 불가능한 반사는 없다.",
        "entities": "인물은 한 명이며 중년 남성의 얼굴, 짧은 검은 머리와 작은 귀걸이는 참조에 가깝다. 그러나 보이는 의상은 참조의 남색 정장·흰 셔츠·넥타이가 아니라 목을 드러낸 검은 지퍼 재킷이다. 낡은 캠핑카와 어두운 습지, 차가운 야간 색감은 장소 설정에 부합한다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리와 목, 어깨의 연결 및 자세는 해부학적으로 자연스럽다. 재킷은 어깨와 몸통에 지지되어 접히며, 떠 있는 물체는 없다. 발과 지면은 클로즈업 밖이라 직접 확인할 수 없지만 부유나 불가능한 동작의 징후는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "얼굴 중심의 클로즈업과 한쪽 입꼬리를 올린 거만한 미소가 명확하고, 참조의 정장·셔츠·넥타이까지 유지해 더 충실하다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "비대칭 미소와 달빛 아래 클로즈업은 잘 구현했지만, 참조의 정장 차림을 검은 지퍼 재킷으로 바꾼 점이 뚜렷한 불일치다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴은 화면 오른쪽으로 약간 돌아가 있고 눈도 카메라 오른쪽의 화면 밖을 향한다. 시선 대상은 보이지 않으며, 지문은 특정 대상을 지정하지 않는다. 화면 오른쪽 입꼬리가 더 올라가 비스듬한 미소가 읽힌다. 무기나 방향을 확인할 휴대 소품은 없다.",
        "built_space": "어깨 뒤 화면 오른쪽에 캠핑카 측면 일부가 흐릿하게 보인다. 창 하나가 있는 출입문 한 개, 손잡이와 측면의 적갈색 띠가 확인된다. 인물은 차 밖 옆에 있으며 차체와 겹치는 방식이 자연스럽다. 배경 차체가 다소 넓게 보이지만 차량 전체를 드러내지는 않는다. 표지판·까마귀·바퀴의 진흙 접촉부는 클로즈업 밖이므로 확인할 수 없다. 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "중년의 한국인 남성으로 설정된 참조와 부합하는 외형의 인물 한 명만 보인다. 짧게 넘긴 검은 머리, 얼굴 윤곽, 작은 귀걸이와 수염 자국이 참조에 가깝고, 남색 정장·흰 셔츠·줄무늬 넥타이를 유지한다. 낡고 얼룩진 밝은 캠핑카 측면과 어두운 청색 계열 야간 환경이 보인다. 다른 사람이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리는 목과 어깨에 정상적으로 연결되어 있고 옷은 몸 위에 자연스럽게 걸쳐져 있다. 하체와 발은 프레임 밖이라 지면 접촉은 확인되지 않지만, 공중에 뜨거나 비정상적으로 지지된 모습은 없다. 손에 든 물체나 운동 중인 물체도 없다."
       },
       {
        "label": "A",
        "direction": "얼굴과 눈은 화면 오른쪽의 보이지 않는 대상을 향한다. 특정 시선 대상이 지문에 정해져 있지 않아 방향상 충돌은 없다. 화면 오른쪽 입꼬리를 치켜올린 미소와 살짝 좁힌 눈이 거만한 표정을 만든다. 무기나 방향성 있는 휴대 소품은 없다.",
        "built_space": "인물 뒤 오른쪽에는 얼룩진 캠핑카 측면과 화면 가장자리에 잘린 창 한 개가 보이고, 왼쪽에는 흐릿한 습지 수면과 풀이 보인다. 인물은 차체 바깥 옆에 서 있는 배치로 읽힌다. 측면 패널이 배경에서 비교적 넓은 면적을 차지한다. 표지판·까마귀·차량 하부는 프레임 밖이며, 보이는 범위에 중복 설비나 불가능한 반사는 없다.",
        "entities": "인물은 한 명이며 중년 남성의 얼굴, 짧은 검은 머리와 작은 귀걸이는 참조에 가깝다. 그러나 보이는 의상은 참조의 남색 정장·흰 셔츠·넥타이가 아니라 목을 드러낸 검은 지퍼 재킷이다. 낡은 캠핑카와 어두운 습지, 차가운 야간 색감은 장소 설정에 부합한다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리와 목, 어깨의 연결 및 자세는 해부학적으로 자연스럽다. 재킷은 어깨와 몸통에 지지되어 접히며, 떠 있는 물체는 없다. 발과 지면은 클로즈업 밖이라 직접 확인할 수 없지만 부유나 불가능한 동작의 징후는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.778,
    "B": 1.667
   },
   "adjusted": {
    "A": 1.778,
    "B": 1.667
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1778,
   "B": 1667
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1778,
    "verdict_ko": "샷 텍스트가 명시한 '한쪽 입꼬리를 비스듬히 올린 거만한 웃음'을 완벽하게 표현하였으며, 기존 늪지대의 야간 조명 분위기를 지시대로 자연스럽게 유지하여 가장 훌륭한 결과물입니다."
   },
   {
    "label": "B",
    "score": 1667,
    "verdict_ko": "레퍼런스의 정장 의상은 정확히 반영했으나, 요구된 거만한 표정 연출이 밋밋하고 얼굴에 새로운 강한 조명이 더해져 지정된 야간 무드를 해쳤습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S61sh1_sel.png",
    "asset_id": "c99d73f2-c3bf-49d9-9c70-2169957b5561",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1401722>",
    "asset_id": "fee7383c-fb61-4b3a-ba7c-79f2555de00b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-8937-7e2c-91a6-39c9970f0622",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S61sh1"
  }
 },
 "S61sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:46:09.163994+00:00",
  "fingerprint": "445be08771a4deb2937a11c1ddfbfc6d6e30355ee74cbe3318290df2dfff10af",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S61sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S61sh5_sel.png",
  "source_sha256": "5177bfeb452bdc0a769c9a1050b87e4b26ddfb071994cd3f43fa628043ac6c0b",
  "file": "S61sh5_cine.png",
  "staged_sha256": "c295212eea211c912667526664eccdd6a9548247a5420cfd971d6a82423bd193",
  "latency_ms": 11993
 },
 "S61sh7::signage": {
  "fp": "f4d270f2c24fcfab",
  "inscriptions": [],
  "cues": [
   {
    "text_native": "",
    "source": "scene_text_implied",
    "source_quote": "방사능 구역 팻말"
   }
  ],
  "dropped": []
 },
 "S61sh7": {
  "input_fingerprint": "6bc43ce8b28622b8",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, moonlit.\n\nSHOT TEXT (authoritative, Korean): 어둠 속 진흙 바닥에 꽂힌 낡고 녹슨 방사능 구역 팻말 클로즈업.\n\nLOCATION (lock): At the rusted radiation warning sign planted in the muddy marsh ground at night. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 방사능 구역 팻말 (Old and rusted, planted in the muddy ground) — The camera sees the front bearing the radiation-area warning at a shallow oblique angle; the rear is not presented; used as Provides the focused narrative evidence while remaining surrounded by visible environmental context; 팻말 아래 진흙 바닥 (The sign support is embedded in the mud); used as Anchors the warning physically in the swamp and prevents an isolated floating-sign composition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the dark nighttime ambience while separating the rusty warning face from its surroundings just enough for the warning to read.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rusty radiation-zone sign remains planted beside the swamp, where the old camper is still stranded. Night surrounds the site as the crows take flight and the moon becomes visible.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, moonlit.\n\nSHOT TEXT (authoritative, Korean): 어둠 속 진흙 바닥에 꽂힌 낡고 녹슨 방사능 구역 팻말 클로즈업.\n\nLOCATION (lock): At the rusted radiation warning sign planted in the muddy marsh ground at night. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 방사능 구역 팻말 (Old and rusted, planted in the muddy ground) — The camera sees the front bearing the radiation-area warning at a shallow oblique angle; the rear is not presented; used as Provides the focused narrative evidence while remaining surrounded by visible environmental context; 팻말 아래 진흙 바닥 (The sign support is embedded in the mud); used as Anchors the warning physically in the swamp and prevents an isolated floating-sign composition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the dark nighttime ambience while separating the rusty warning face from its surroundings just enough for the warning to read.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rusty radiation-zone sign remains planted beside the swamp, where the old camper is still stranded. Night surrounds the site as the crows take flight and the moon becomes visible.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, moonlit.\n\nSHOT TEXT (authoritative, Korean): 어둠 속 진흙 바닥에 꽂힌 낡고 녹슨 방사능 구역 팻말 클로즈업.\n\nLOCATION (lock): At the rusted radiation warning sign planted in the muddy marsh ground at night. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 방사능 구역 팻말 (Old and rusted, planted in the muddy ground) — The camera sees the front bearing the radiation-area warning at a shallow oblique angle; the rear is not presented; used as Provides the focused narrative evidence while remaining surrounded by visible environmental context; 팻말 아래 진흙 바닥 (The sign support is embedded in the mud); used as Anchors the warning physically in the swamp and prevents an isolated floating-sign composition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the dark nighttime ambience while separating the rusty warning face from its surroundings just enough for the warning to read.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rusty radiation-zone sign remains planted beside the swamp, where the old camper is still stranded. Night surrounds the site as the crows take flight and the moon becomes visible.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라가 팻말의 앞면 방사능 마크를 향함.",
    "built_space": "진흙 바닥의 습지, 좌측에 캠핑카 일부 위치.",
    "entities": "녹슨 방사능 팻말, 진흙, 달, 팻말을 잡은 인물.",
    "hard_violations": [
     "[gemini-pro] 프롬프트에 없는 인물 생성",
     "[gpt-high] 사람이 없어야 하는 장면에 사람의 손·팔·몸을 추가했습니다. 이 장면에는 물체를 잡는 행동이 지정되지 않았으므로 손을 허용하는 예외도 적용되지 않습니다."
    ],
    "physics": "팻말은 바닥 기둥에 지지되며 우측을 사람 손이 잡고 있음."
   },
   {
    "label": "B",
    "direction": "카메라는 팻말 앞면을 비스듬히 향함.",
    "built_space": "진흙과 물웅덩이가 있는 습지, 뒤쪽에 진입 금지 표지판 위치.",
    "entities": "칠이 벗겨진 방사능 팻말, 진흙, 쓰러진 둥근 표지판.",
    "hard_violations": [],
    "physics": "팻말은 기둥으로 진흙에 박혀 지지되며 뒤쪽 표지판은 바닥에 쓰러져 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "인물 등장 금지 지침을 어기고 팻말을 잡고 있는 사람을 생성하여 심각한 위반을 범했습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "인물 없이 지시된 클로즈업 구도를 잘 구현했으나 배경의 표지판이 쓰러져 있어 레퍼런스와 다소 차이가 있습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라가 팻말의 앞면 방사능 마크를 향함.",
        "built_space": "진흙 바닥의 습지, 좌측에 캠핑카 일부 위치.",
        "entities": "녹슨 방사능 팻말, 진흙, 달, 팻말을 잡은 인물.",
        "hard_violations": [
         "프롬프트에 없는 인물 생성"
        ],
        "physics": "팻말은 바닥 기둥에 지지되며 우측을 사람 손이 잡고 있음."
       },
       {
        "label": "B",
        "direction": "카메라는 팻말 앞면을 비스듬히 향함.",
        "built_space": "진흙과 물웅덩이가 있는 습지, 뒤쪽에 진입 금지 표지판 위치.",
        "entities": "칠이 벗겨진 방사능 팻말, 진흙, 쓰러진 둥근 표지판.",
        "hard_violations": [],
        "physics": "팻말은 기둥으로 진흙에 박혀 지지되며 뒤쪽 표지판은 바닥에 쓰러져 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "인물 등장 금지 지침을 어기고 팻말을 잡고 있는 사람을 생성하여 심각한 위반을 범했습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "인물 없이 지시된 클로즈업 구도를 잘 구현했으나 배경의 표지판이 쓰러져 있어 레퍼런스와 다소 차이가 있습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라가 팻말의 앞면 방사능 마크를 향함.",
        "built_space": "진흙 바닥의 습지, 좌측에 캠핑카 일부 위치.",
        "entities": "녹슨 방사능 팻말, 진흙, 달, 팻말을 잡은 인물.",
        "hard_violations": [
         "프롬프트에 없는 인물 생성"
        ],
        "physics": "팻말은 바닥 기둥에 지지되며 우측을 사람 손이 잡고 있음."
       },
       {
        "label": "B",
        "direction": "카메라는 팻말 앞면을 비스듬히 향함.",
        "built_space": "진흙과 물웅덩이가 있는 습지, 뒤쪽에 진입 금지 표지판 위치.",
        "entities": "칠이 벗겨진 방사능 팻말, 진흙, 쓰러진 둥근 표지판.",
        "hard_violations": [],
        "physics": "팻말은 기둥으로 진흙에 박혀 지지되며 뒤쪽 표지판은 바닥에 쓰러져 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "사람 없이 녹슨 팻말 앞면을 얕은 사선 클로즈업으로 보여 주고 밤의 진흙 습지를 유지하지만, 지지대가 진흙에 박힌 접점은 프레임 아래로 잘려 명확하지 않습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "팻말의 방향과 야간 습지 질감은 맞지만, 금지된 사람의 몸과 팻말을 잡는 손을 추가하여 장소만 보여 주라는 핵심 조건을 위반합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "방사능 기호가 있는 앞면이 카메라를 향해 얕게 비스듬히 놓였고 뒷면은 보이지 않습니다. 왼쪽 아래 원형 진입금지 표지도 앞면이 보입니다. 사람의 시선이나 이동하는 몸은 없습니다.",
        "built_space": "전경에 사각 방사능 팻말 하나와 지지대 하나, 왼쪽 아래에 원형 표지 하나와 그 지지대가 보입니다. 주변은 물이 고인 진흙과 마른 습지 식생으로 참조 장소와 연결됩니다. 원형 표지는 참조보다 낮고 납작하게 보입니다. 캠핑카는 이 좁은 프레임에 없으며, 방사능 팻말 지지대의 지면 접점은 하단 밖입니다.",
        "entities": "낡은 황색 금속판, 검은 방사능 기호, 심한 녹과 벗겨진 도장이 확인됩니다. 원형 진입금지 표지는 참조에도 있는 물체입니다. 사람과 까마귀는 보이지 않고, 읽을 수 있는 글자나 워터마크도 없습니다. 어두운 청회색 습지와 물의 반사가 야간 분위기에 부합합니다.",
        "hard_violations": [],
        "physics": "금속판은 뒤쪽 지지대에 부착되어 있고 지지대가 아래로 이어져 떠 있는 물체로 보이지 않습니다. 다만 진흙에 실제로 박힌 끝부분은 잘려 있어 접촉 상태를 직접 확인할 수 없습니다. 원형 표지 역시 별도 지지대가 받치고 있으며 공중에 뜬 몸이나 물체는 없습니다."
       },
       {
        "label": "B",
        "direction": "방사능 팻말의 앞면이 카메라 쪽으로 약간 비스듬히 향합니다. 오른쪽에서 들어온 손은 팻말의 위쪽 가장자리를 잡고 있습니다. 사람의 얼굴과 눈은 보이지 않아 시선은 확인할 수 없습니다.",
        "built_space": "중앙에 사각 방사능 팻말 하나와 아래로 이어지는 지지대 하나가 있습니다. 왼쪽 가장자리에는 캠핑카 앞부분이, 뒤에는 갈대와 물웅덩이가 보입니다. 오른쪽에는 사람의 몸 일부가 공간을 차지합니다. 지지대의 진흙 접점은 하단 밖이며, 참조의 원형 표지는 이 프레임에서 보이지 않습니다.",
        "entities": "녹슨 금속 방사능 팻말과 검은 방사능 기호, 오래된 캠핑카 일부, 달, 습지가 보입니다. 검은 겉옷을 입은 사람의 손·팔·몸 일부가 추가되었으며 얼굴이 없어 나이·성별·민족성은 판별할 수 없습니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "사람이 없어야 하는 장면에 사람의 손·팔·몸을 추가했습니다. 이 장면에는 물체를 잡는 행동이 지정되지 않았으므로 손을 허용하는 예외도 적용되지 않습니다."
        ],
        "physics": "팻말은 아래 지지대와 위 가장자리를 잡는 손에 의해 지지되어 떠 있지 않습니다. 손과 팔은 오른쪽 몸에 연결되어 있으며 물리적으로 불가능한 잡기 동작은 보이지 않습니다. 그러나 사람이 팻말을 잡는 행동 자체가 요청되지 않은 연출입니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "사람 없이 녹슨 팻말 앞면을 얕은 사선 클로즈업으로 보여 주고 밤의 진흙 습지를 유지하지만, 지지대가 진흙에 박힌 접점은 프레임 아래로 잘려 명확하지 않습니다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "팻말의 방향과 야간 습지 질감은 맞지만, 금지된 사람의 몸과 팻말을 잡는 손을 추가하여 장소만 보여 주라는 핵심 조건을 위반합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "방사능 기호가 있는 앞면이 카메라를 향해 얕게 비스듬히 놓였고 뒷면은 보이지 않습니다. 왼쪽 아래 원형 진입금지 표지도 앞면이 보입니다. 사람의 시선이나 이동하는 몸은 없습니다.",
        "built_space": "전경에 사각 방사능 팻말 하나와 지지대 하나, 왼쪽 아래에 원형 표지 하나와 그 지지대가 보입니다. 주변은 물이 고인 진흙과 마른 습지 식생으로 참조 장소와 연결됩니다. 원형 표지는 참조보다 낮고 납작하게 보입니다. 캠핑카는 이 좁은 프레임에 없으며, 방사능 팻말 지지대의 지면 접점은 하단 밖입니다.",
        "entities": "낡은 황색 금속판, 검은 방사능 기호, 심한 녹과 벗겨진 도장이 확인됩니다. 원형 진입금지 표지는 참조에도 있는 물체입니다. 사람과 까마귀는 보이지 않고, 읽을 수 있는 글자나 워터마크도 없습니다. 어두운 청회색 습지와 물의 반사가 야간 분위기에 부합합니다.",
        "hard_violations": [],
        "physics": "금속판은 뒤쪽 지지대에 부착되어 있고 지지대가 아래로 이어져 떠 있는 물체로 보이지 않습니다. 다만 진흙에 실제로 박힌 끝부분은 잘려 있어 접촉 상태를 직접 확인할 수 없습니다. 원형 표지 역시 별도 지지대가 받치고 있으며 공중에 뜬 몸이나 물체는 없습니다."
       },
       {
        "label": "A",
        "direction": "방사능 팻말의 앞면이 카메라 쪽으로 약간 비스듬히 향합니다. 오른쪽에서 들어온 손은 팻말의 위쪽 가장자리를 잡고 있습니다. 사람의 얼굴과 눈은 보이지 않아 시선은 확인할 수 없습니다.",
        "built_space": "중앙에 사각 방사능 팻말 하나와 아래로 이어지는 지지대 하나가 있습니다. 왼쪽 가장자리에는 캠핑카 앞부분이, 뒤에는 갈대와 물웅덩이가 보입니다. 오른쪽에는 사람의 몸 일부가 공간을 차지합니다. 지지대의 진흙 접점은 하단 밖이며, 참조의 원형 표지는 이 프레임에서 보이지 않습니다.",
        "entities": "녹슨 금속 방사능 팻말과 검은 방사능 기호, 오래된 캠핑카 일부, 달, 습지가 보입니다. 검은 겉옷을 입은 사람의 손·팔·몸 일부가 추가되었으며 얼굴이 없어 나이·성별·민족성은 판별할 수 없습니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "사람이 없어야 하는 장면에 사람의 손·팔·몸을 추가했습니다. 이 장면에는 물체를 잡는 행동이 지정되지 않았으므로 손을 허용하는 예외도 적용되지 않습니다."
        ],
        "physics": "팻말은 아래 지지대와 위 가장자리를 잡는 손에 의해 지지되어 떠 있지 않습니다. 손과 팔은 오른쪽 몸에 연결되어 있으며 물리적으로 불가능한 잡기 동작은 보이지 않습니다. 그러나 사람이 팻말을 잡는 행동 자체가 요청되지 않은 연출입니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.679,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.429,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 프롬프트에 없는 인물 생성",
     "[gpt-high] 사람이 없어야 하는 장면에 사람의 손·팔·몸을 추가했습니다. 이 장면에는 물체를 잡는 행동이 지정되지 않았으므로 손을 허용하는 예외도 적용되지 않습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "A": 429,
   "B": 2000
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 429,
    "verdict_ko": "인물 등장 금지 지침을 어기고 팻말을 잡고 있는 사람을 생성하여 심각한 위반을 범했습니다.  ★위반: [gemini-pro] 프롬프트에 없는 인물 생성 / [gpt-high] 사람이 없어야 하는 장면에 사람의 손·팔·몸을 추가했습니다. 이 장면에는 물체를 잡는 행동이 지정되지 않았으므로 손을 허용하는 예외도 적용되지 않습니다."
   },
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "인물 없이 지시된 클로즈업 구도를 잘 구현했으나 배경의 표지판이 쓰러져 있어 레퍼런스와 다소 차이가 있습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S61sh1_sel.png",
    "asset_id": "c99d73f2-c3bf-49d9-9c70-2169957b5561",
    "role": "prev_still"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-8ae1-7975-84b0-a259d7c12894",
  "ref_mode": "prev만 (배경 전용·공유 계획)",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S61sh1"
  },
  "lane_policy": "share_plan_prev_bgonly"
 },
 "S61sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:47:19.965274+00:00",
  "fingerprint": "bd34c0a0ffe0c1d5f6d0193995d10e2aefbd4d9849381204390c2f2d2cfce1b0",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S61sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S61sh7_sel.png",
  "source_sha256": "a796798a7c409bb67df5230648647f883a08fe0ae61ef65ced2e850cc1e49049",
  "file": "S61sh7_cine.png",
  "staged_sha256": "ffda4ddbaf8d81cbe4d7be1d7b76ea854386c246c942a185f8d667101453275a",
  "latency_ms": 11324
 },
 "S62sh3::signage": {
  "fp": "a03021a1549dfde5",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S62sh3": {
  "input_fingerprint": "47b4842d7fa356b7",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 활짝 열린 철문 너머, 쌀과 감자 포대 몇 개만 덩그러니 놓인 텅 빈 식량창고 내부.\n\nLOCATION (lock): Inside the village food storehouse, with a few rice and potato sacks illuminated by daylight through the open metal door. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 식량창고 철문 (Fully open) — The open door leaf is seen obliquely along the left edge, leaving the inward view unobstructed; used as Creates a threshold frame and establishes that the empty space is the revealed storage interior; 쌀 포대 몇 개와 감자 한 포대 (The few remaining provisions in the otherwise empty warehouse); used as Forms a small, readable supply cluster whose modest scale is measured against the surrounding empty floor; 비어 있는 창고 내부 (Largely empty) — The diagonal view reveals the interior floor extending beyond the remaining sacks; used as Makes absence, rather than the sacks themselves, the main spatial statement.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral ambient daylight appropriate to the open storage entrance, preserving enough interior detail for the sparse supplies and surrounding emptiness to register.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The food store's old iron door is open, revealing only a few sacks of rice and one sack of potatoes. Elsewhere in the village square, an old truck is undergoing starting attempts; Charlie's dents, holes and sensor damage remain unrepaired.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 활짝 열린 철문 너머, 쌀과 감자 포대 몇 개만 덩그러니 놓인 텅 빈 식량창고 내부.\n\nLOCATION (lock): Inside the village food storehouse, with a few rice and potato sacks illuminated by daylight through the open metal door. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 식량창고 철문 (Fully open) — The open door leaf is seen obliquely along the left edge, leaving the inward view unobstructed; used as Creates a threshold frame and establishes that the empty space is the revealed storage interior; 쌀 포대 몇 개와 감자 한 포대 (The few remaining provisions in the otherwise empty warehouse); used as Forms a small, readable supply cluster whose modest scale is measured against the surrounding empty floor; 비어 있는 창고 내부 (Largely empty) — The diagonal view reveals the interior floor extending beyond the remaining sacks; used as Makes absence, rather than the sacks themselves, the main spatial statement.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral ambient daylight appropriate to the open storage entrance, preserving enough interior detail for the sparse supplies and surrounding emptiness to register.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The food store's old iron door is open, revealing only a few sacks of rice and one sack of potatoes. Elsewhere in the village square, an old truck is undergoing starting attempts; Charlie's dents, holes and sensor damage remain unrepaired.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 활짝 열린 철문 너머, 쌀과 감자 포대 몇 개만 덩그러니 놓인 텅 빈 식량창고 내부.\n\nLOCATION (lock): Inside the village food storehouse, with a few rice and potato sacks illuminated by daylight through the open metal door. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 식량창고 철문 (Fully open) — The open door leaf is seen obliquely along the left edge, leaving the inward view unobstructed; used as Creates a threshold frame and establishes that the empty space is the revealed storage interior; 쌀 포대 몇 개와 감자 한 포대 (The few remaining provisions in the otherwise empty warehouse); used as Forms a small, readable supply cluster whose modest scale is measured against the surrounding empty floor; 비어 있는 창고 내부 (Largely empty) — The diagonal view reveals the interior floor extending beyond the remaining sacks; used as Makes absence, rather than the sacks themselves, the main spatial statement.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral ambient daylight appropriate to the open storage entrance, preserving enough interior detail for the sparse supplies and surrounding emptiness to register.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The food store's old iron door is open, revealing only a few sacks of rice and one sack of potatoes. Elsewhere in the village square, an old truck is undergoing starting attempts; Charlie's dents, holes and sensor damage remain unrepaired.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라는 창고 내부를 비추며 중앙 안쪽의 또 다른 출입구를 향함.",
    "built_space": "창고 내부. 왼쪽에 열린 문이 있고, 정면 안쪽에 또 다른 열린 문이 보이나 바깥 풍경이 레퍼런스와 일치하지 않음.",
    "entities": "바닥에 여러 개의 포대가 모여 있음. 사람 없음. 포대 겉면에 'POTATO' 등의 영어 단어가 명확하게 적혀 있음.",
    "hard_violations": [
     "[gemini-pro] 읽을 수 있는 텍스트 포함 (포대의 'POTATO' 등 글귀)",
     "[gpt-high] 감자 포대에 ‘POTATO’라는 영문이 선명하게 읽혀, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 조건을 위반한다."
    ],
    "physics": "포대들이 바닥에 자연스럽게 놓여 있음."
   },
   {
    "label": "B",
    "direction": "카메라는 창고 내부를 향해 사선으로 넓게 비추고 있으며, 시선은 바닥에 놓인 포대들과 텅 빈 공간으로 향함.",
    "built_space": "창고 내부. 화면 왼쪽에 활짝 열린 철문이 비스듬히 자리 잡고 있으며, 문 너머로 레퍼런스의 녹슨 급수탑과 주변 건물이 올바른 위치 관계로 보임. 내부는 넓고 비어 있음.",
    "entities": "바닥에 쌀 포대와 감자가 노출된 포대가 놓여 있음. 사람 없음. 포대에 적힌 글씨는 뭉개져 있어 읽을 수 없음.",
    "hard_violations": [
     "[gpt-high] 오른쪽 뒤 벽에 기대 놓은 판재 두 개는 포대 몇 개만 남아 있어야 하는 창고에 추가된 별도 물건이다."
    ],
    "physics": "포대들이 바닥에 안정적으로 놓여 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "금지된 식별 가능한 텍스트가 없으며, 열린 문 너머로 레퍼런스의 급수탑을 배치하여 지정된 로케이션을 충실히 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "프롬프트에서 명시적으로 금지한 읽을 수 있는 텍스트가 포함되어 심각한 위반에 해당합니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "카메라는 창고 내부를 향해 사선으로 넓게 비추고 있으며, 시선은 바닥에 놓인 포대들과 텅 빈 공간으로 향함.",
        "built_space": "창고 내부. 화면 왼쪽에 활짝 열린 철문이 비스듬히 자리 잡고 있으며, 문 너머로 레퍼런스의 녹슨 급수탑과 주변 건물이 올바른 위치 관계로 보임. 내부는 넓고 비어 있음.",
        "entities": "바닥에 쌀 포대와 감자가 노출된 포대가 놓여 있음. 사람 없음. 포대에 적힌 글씨는 뭉개져 있어 읽을 수 없음.",
        "hard_violations": [],
        "physics": "포대들이 바닥에 안정적으로 놓여 있음."
       },
       {
        "label": "A",
        "direction": "카메라는 창고 내부를 비추며 중앙 안쪽의 또 다른 출입구를 향함.",
        "built_space": "창고 내부. 왼쪽에 열린 문이 있고, 정면 안쪽에 또 다른 열린 문이 보이나 바깥 풍경이 레퍼런스와 일치하지 않음.",
        "entities": "바닥에 여러 개의 포대가 모여 있음. 사람 없음. 포대 겉면에 'POTATO' 등의 영어 단어가 명확하게 적혀 있음.",
        "hard_violations": [
         "읽을 수 있는 텍스트 포함 (포대의 'POTATO' 등 글귀)"
        ],
        "physics": "포대들이 바닥에 자연스럽게 놓여 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "금지된 식별 가능한 텍스트가 없으며, 열린 문 너머로 레퍼런스의 급수탑을 배치하여 지정된 로케이션을 충실히 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "프롬프트에서 명시적으로 금지한 읽을 수 있는 텍스트가 포함되어 심각한 위반에 해당합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "카메라는 창고 내부를 향해 사선으로 넓게 비추고 있으며, 시선은 바닥에 놓인 포대들과 텅 빈 공간으로 향함.",
        "built_space": "창고 내부. 화면 왼쪽에 활짝 열린 철문이 비스듬히 자리 잡고 있으며, 문 너머로 레퍼런스의 녹슨 급수탑과 주변 건물이 올바른 위치 관계로 보임. 내부는 넓고 비어 있음.",
        "entities": "바닥에 쌀 포대와 감자가 노출된 포대가 놓여 있음. 사람 없음. 포대에 적힌 글씨는 뭉개져 있어 읽을 수 없음.",
        "hard_violations": [],
        "physics": "포대들이 바닥에 안정적으로 놓여 있음."
       },
       {
        "label": "A",
        "direction": "카메라는 창고 내부를 비추며 중앙 안쪽의 또 다른 출입구를 향함.",
        "built_space": "창고 내부. 왼쪽에 열린 문이 있고, 정면 안쪽에 또 다른 열린 문이 보이나 바깥 풍경이 레퍼런스와 일치하지 않음.",
        "entities": "바닥에 여러 개의 포대가 모여 있음. 사람 없음. 포대 겉면에 'POTATO' 등의 영어 단어가 명확하게 적혀 있음.",
        "hard_violations": [
         "읽을 수 있는 텍스트 포함 (포대의 'POTATO' 등 글귀)"
        ],
        "physics": "포대들이 바닥에 자연스럽게 놓여 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "작은 식량 더미 너머로 넓게 이어지는 빈 바닥과 사선 구도는 더 충실하지만, 뒤쪽에 추가한 판재가 ‘포대 몇 개만 놓인 창고’ 조건을 위반한다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "왼쪽 철문과 낮의 빈 창고는 표현했지만, 감자 포대의 선명한 영문 글자가 읽을 수 있는 글자 금지 조건을 명백히 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "인물이나 조준·이동하는 물체는 없다. 카메라는 왼쪽 출입구를 옆에 두고 창고 깊숙한 오른쪽 뒤편으로 비스듬히 향한다. 포대 뒤로 빈 바닥이 길게 이어지지만, 문턱 밖에서 내부를 바라보기보다는 이미 실내에 들어온 시점에 가깝다.",
        "built_space": "왼쪽 출입구 하나에 열린 철문짝 두 개가 보인다. 벗겨진 미장 벽, 콘크리트 바닥, 노출 지붕 트러스, 오른쪽 높은 창 일부, 뒤쪽 왼편의 좁은 문 형태 개구부가 보인다. 외부 급수탑은 참조 장소와 연결된다. 참조 사진은 내부를 보여주지 않아 내부 트러스나 뒤쪽 개구부의 정확한 일치 여부는 확인할 수 없다. 반사상은 없다.",
        "entities": "사람은 없다. 중앙 아래에 곡물 포대로 보이는 네 포대와 감자가 직접 드러난 열린 감자 포대 하나가 있다. 닫힌 곡물 포대의 내용물이 쌀인지는 시각적으로 확정할 수 없다. 포대 인쇄는 보이지만 명확히 읽히는 단어는 확인하기 어렵다. 오른쪽 뒤 벽에는 식량과 무관한 좁은 판재 두 개가 추가되어 있다.",
        "hard_violations": [
         "오른쪽 뒤 벽에 기대 놓은 판재 두 개는 포대 몇 개만 남아 있어야 하는 창고에 추가된 별도 물건이다."
        ],
        "physics": "맨 아래 포대와 감자 포대는 바닥에 놓였고, 쌓인 포대는 아래 포대가 받친다. 감자는 포대 안에 담겨 있다. 철문은 경첩으로 지지되며 판재는 바닥과 벽에 기대어 있다. 지지 없이 떠 있는 물체는 없다."
       },
       {
        "label": "B",
        "direction": "인물이나 조준·이동하는 물체는 없다. 카메라는 왼쪽 열린 철문 옆에서 창고 안쪽 벽과 중앙 오른쪽 식량 더미를 향한다. 내부를 가리는 물체는 없지만, 뒤 벽을 비교적 정면으로 보아 A보다 바닥의 사선 깊이가 덜 강조된다.",
        "built_space": "왼쪽 출입구에 철문짝 두 개가 열려 있고 가까운 문짝이 화면 왼쪽을 크게 차지한다. 콘크리트 바닥, 낡은 미장 벽, 박공지붕 아래 목재 트러스, 오른쪽 상단 창 하나가 보인다. 문과 벽의 재질·색은 참조 창고와 대체로 부합하지만, 참조에서 보이지 않는 내부 구조까지 동일하다고 확인할 수는 없다. 반사상은 없다.",
        "entities": "사람은 없다. 중앙 오른쪽에 곡물 포대로 보이는 세 포대와 감자용으로 표시된 포대 하나가 모여 있다. 모두 닫혀 있어 쌀과 감자 자체는 보이지 않는다. 앞쪽 오른쪽 포대의 ‘POTATO’ 글자는 분명히 읽힌다. 주변 바닥에는 별도 저장 물품이 없다.",
        "hard_violations": [
         "감자 포대에 ‘POTATO’라는 영문이 선명하게 읽혀, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 조건을 위반한다."
        ],
        "physics": "포대 네 개는 바닥에 닿아 있고 서로 기대어 서 있다. 철문은 문틀의 경첩으로 지지되며 지붕 부재는 벽과 트러스에 연결된다. 떠 있거나 지지 관계가 불가능한 물체는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "작은 식량 더미 너머로 넓게 이어지는 빈 바닥과 사선 구도는 더 충실하지만, 뒤쪽에 추가한 판재가 ‘포대 몇 개만 놓인 창고’ 조건을 위반한다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "왼쪽 철문과 낮의 빈 창고는 표현했지만, 감자 포대의 선명한 영문 글자가 읽을 수 있는 글자 금지 조건을 명백히 위반한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "인물이나 조준·이동하는 물체는 없다. 카메라는 왼쪽 출입구를 옆에 두고 창고 깊숙한 오른쪽 뒤편으로 비스듬히 향한다. 포대 뒤로 빈 바닥이 길게 이어지지만, 문턱 밖에서 내부를 바라보기보다는 이미 실내에 들어온 시점에 가깝다.",
        "built_space": "왼쪽 출입구 하나에 열린 철문짝 두 개가 보인다. 벗겨진 미장 벽, 콘크리트 바닥, 노출 지붕 트러스, 오른쪽 높은 창 일부, 뒤쪽 왼편의 좁은 문 형태 개구부가 보인다. 외부 급수탑은 참조 장소와 연결된다. 참조 사진은 내부를 보여주지 않아 내부 트러스나 뒤쪽 개구부의 정확한 일치 여부는 확인할 수 없다. 반사상은 없다.",
        "entities": "사람은 없다. 중앙 아래에 곡물 포대로 보이는 네 포대와 감자가 직접 드러난 열린 감자 포대 하나가 있다. 닫힌 곡물 포대의 내용물이 쌀인지는 시각적으로 확정할 수 없다. 포대 인쇄는 보이지만 명확히 읽히는 단어는 확인하기 어렵다. 오른쪽 뒤 벽에는 식량과 무관한 좁은 판재 두 개가 추가되어 있다.",
        "hard_violations": [
         "오른쪽 뒤 벽에 기대 놓은 판재 두 개는 포대 몇 개만 남아 있어야 하는 창고에 추가된 별도 물건이다."
        ],
        "physics": "맨 아래 포대와 감자 포대는 바닥에 놓였고, 쌓인 포대는 아래 포대가 받친다. 감자는 포대 안에 담겨 있다. 철문은 경첩으로 지지되며 판재는 바닥과 벽에 기대어 있다. 지지 없이 떠 있는 물체는 없다."
       },
       {
        "label": "A",
        "direction": "인물이나 조준·이동하는 물체는 없다. 카메라는 왼쪽 열린 철문 옆에서 창고 안쪽 벽과 중앙 오른쪽 식량 더미를 향한다. 내부를 가리는 물체는 없지만, 뒤 벽을 비교적 정면으로 보아 A보다 바닥의 사선 깊이가 덜 강조된다.",
        "built_space": "왼쪽 출입구에 철문짝 두 개가 열려 있고 가까운 문짝이 화면 왼쪽을 크게 차지한다. 콘크리트 바닥, 낡은 미장 벽, 박공지붕 아래 목재 트러스, 오른쪽 상단 창 하나가 보인다. 문과 벽의 재질·색은 참조 창고와 대체로 부합하지만, 참조에서 보이지 않는 내부 구조까지 동일하다고 확인할 수는 없다. 반사상은 없다.",
        "entities": "사람은 없다. 중앙 오른쪽에 곡물 포대로 보이는 세 포대와 감자용으로 표시된 포대 하나가 모여 있다. 모두 닫혀 있어 쌀과 감자 자체는 보이지 않는다. 앞쪽 오른쪽 포대의 ‘POTATO’ 글자는 분명히 읽힌다. 주변 바닥에는 별도 저장 물품이 없다.",
        "hard_violations": [
         "감자 포대에 ‘POTATO’라는 영문이 선명하게 읽혀, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 조건을 위반한다."
        ],
        "physics": "포대 네 개는 바닥에 닿아 있고 서로 기대어 서 있다. 철문은 문틀의 경첩으로 지지되며 지붕 부재는 벽과 트러스에 연결된다. 떠 있거나 지지 관계가 불가능한 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.029,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.779,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 읽을 수 있는 텍스트 포함 (포대의 'POTATO' 등 글귀)",
     "[gpt-high] 감자 포대에 ‘POTATO’라는 영문이 선명하게 읽혀, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 조건을 위반한다."
    ],
    "B": [
     "[gpt-high] 오른쪽 뒤 벽에 기대 놓은 판재 두 개는 포대 몇 개만 남아 있어야 하는 창고에 추가된 별도 물건이다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 1750,
   "A": 779
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "금지된 식별 가능한 텍스트가 없으며, 열린 문 너머로 레퍼런스의 급수탑을 배치하여 지정된 로케이션을 충실히 구현했습니다.  ★위반: [gpt-high] 오른쪽 뒤 벽에 기대 놓은 판재 두 개는 포대 몇 개만 남아 있어야 하는 창고에 추가된 별도 물건이다."
   },
   {
    "label": "A",
    "score": 779,
    "verdict_ko": "프롬프트에서 명시적으로 금지한 읽을 수 있는 텍스트가 포함되어 심각한 위반에 해당합니다.  ★위반: [gemini-pro] 읽을 수 있는 텍스트 포함 (포대의 'POTATO' 등 글귀) / [gpt-high] 감자 포대에 ‘POTATO’라는 영문이 선명하게 읽혀, 이미지 어디에도 읽을 수 있는 글자를 두지 말라는 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_ruined_village_square_sel.png",
    "asset_id": "e716b949-d46e-4479-90bb-be052601519b",
    "role": "location_seed_bg"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-8c89-769a-9fdb-bf1c7d1d36b8",
  "ref_mode": "seed-bg만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  },
  "lane_policy": "ab_select_bypass:bg_only"
 },
 "S62sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:48:18.117363+00:00",
  "fingerprint": "d3ce69862ad9cd5c243e8ef886d8b4d3412f4a6754c427b0bc67c438f31ec097",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S62sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S62sh3_sel.png",
  "source_sha256": "34406e5983bcdb55e4cbfc76daf6f35a73c8476c960789fdf2a1edf78b8a7d18",
  "file": "S62sh3_cine.png",
  "staged_sha256": "3774df59c680a458c61476915a20248bc75c674602f456bf2f5e87176f106f3d",
  "latency_ms": 11921
 },
 "S62sh15::signage": {
  "fp": "c230a55faf0cfd98",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S62sh15": {
  "input_fingerprint": "54f104ac62589c2a",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바닥에 투사된 빛나는 홀로그램 지도 위로 서남쪽 위치의 붉은 점이 밝게 켜져 있는 찰리의 시점 쇼트.\n\nLOCATION (lock): On the ground of the village's open square, beside the old truck being repaired. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: navigation map projected onto the ground in the middle-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: 바닥에 투사된 홀로그램 지도 (Active, with the southwest destination marked by a bright red point) — The map-bearing ground surface faces upward and is viewed steeply from above, with the southwest marker in the map's lower-left sector; used as Carries 찰리's navigation information while retaining visible ground around the projection; 공터 바닥 (Visible beneath and around the projected map); used as Establishes the projection's physical placement without adding interface borders or device hardware.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain daytime ambient visibility while allowing the projected map and its bright red southwest marker to read as localized emitted light without darkening the entire setting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old truck remains in the village square during repair attempts, and Charlie retains his battle damage and malfunctioning sensors. Charlie's POV contains a navigation map identifying a department store and supermarket 25.6 km to the southwest, not a floor-projected hologram.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바닥에 투사된 빛나는 홀로그램 지도 위로 서남쪽 위치의 붉은 점이 밝게 켜져 있는 찰리의 시점 쇼트.\n\nLOCATION (lock): On the ground of the village's open square, beside the old truck being repaired. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: navigation map projected onto the ground in the middle-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: 바닥에 투사된 홀로그램 지도 (Active, with the southwest destination marked by a bright red point) — The map-bearing ground surface faces upward and is viewed steeply from above, with the southwest marker in the map's lower-left sector; used as Carries 찰리's navigation information while retaining visible ground around the projection; 공터 바닥 (Visible beneath and around the projected map); used as Establishes the projection's physical placement without adding interface borders or device hardware.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain daytime ambient visibility while allowing the projected map and its bright red southwest marker to read as localized emitted light without darkening the entire setting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old truck remains in the village square during repair attempts, and Charlie retains his battle damage and malfunctioning sensors. Charlie's POV contains a navigation map identifying a department store and supermarket 25.6 km to the southwest, not a floor-projected hologram.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바닥에 투사된 빛나는 홀로그램 지도 위로 서남쪽 위치의 붉은 점이 밝게 켜져 있는 찰리의 시점 쇼트.\n\nLOCATION (lock): On the ground of the village's open square, beside the old truck being repaired. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: navigation map projected onto the ground in the middle-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: 바닥에 투사된 홀로그램 지도 (Active, with the southwest destination marked by a bright red point) — The map-bearing ground surface faces upward and is viewed steeply from above, with the southwest marker in the map's lower-left sector; used as Carries 찰리's navigation information while retaining visible ground around the projection; 공터 바닥 (Visible beneath and around the projected map); used as Establishes the projection's physical placement without adding interface borders or device hardware.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain daytime ambient visibility while allowing the projected map and its bright red southwest marker to read as localized emitted light without darkening the entire setting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old truck remains in the village square during repair attempts, and Charlie retains his battle damage and malfunctioning sensors. Charlie's POV contains a navigation map identifying a department store and supermarket 25.6 km to the southwest, not a floor-projected hologram.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라는 바닥을 비스듬히 내려다보고 있으며, 시선은 중앙의 홀로그램 지도를 향함.",
    "built_space": "흙바닥의 공터이며, 배경에 수리 중인 낡은 트럭과 공구들이 놓여 있음.",
    "entities": "바닥의 질감이 비치는 홀로그램 지도, 지도 좌측 하단(서남쪽)의 밝은 붉은 점, 낡은 트럭 모두 지시사항에 부합함.",
    "hard_violations": [
     "[gpt-high] 장치 하드웨어를 넣지 말라는 구도 지시와 달리, 전경에 금속 프레임과 노출 배선이 추가되어 있다."
    ],
    "physics": "홀로그램이 바닥의 굴곡을 따라 자연스럽게 투사되어 있음."
   },
   {
    "label": "B",
    "direction": "1인칭 시점으로 카메라는 바닥의 홀로그램 지도를 내려다보고 있음.",
    "built_space": "바닥에 갈라진 틈이 있는 공터이며, 우측 상단에 트럭의 바퀴가 보임.",
    "entities": "홀로그램 지도와 붉은 점이 있으나, 지도 위에 프롬프트 텍스트가 그대로 노출됨. 화자의 팔과 장갑이 보임.",
    "hard_violations": [
     "[gemini-pro] 읽을 수 있는 글씨 금지(No readable writing) 지시를 어기고 프롬프트의 텍스트('바닥에 투사된 홀로그램 지도', '서남쪽 위치의 붉은 점' 등)가 그대로 유출됨.",
     "[gpt-high] 지도 위에 '바닥에 투사된 홀로그램 지도', '서남쪽 위치의 붉은 점', '2069', '25.6 km' 등 읽을 수 있는 문구와 숫자가 노출되어 있다.",
     "[gpt-high] 명시적으로 배제한 인터페이스 테두리와 조작 아이콘이 지도에 추가되어 있다."
    ],
    "physics": "홀로그램이 평평하게 투사되었고, 화자의 팔은 자연스러운 자세를 취하고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "지시된 시점과 홀로그램 지도의 붉은 점 위치를 정확히 묘사했으며, 금지된 텍스트 노출 없이 프롬프트를 잘 따랐습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "프롬프트의 지시문이 그대로 이미지 안에 읽을 수 있는 텍스트로 유출되는 치명적인 오류가 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 바닥을 비스듬히 내려다보고 있으며, 시선은 중앙의 홀로그램 지도를 향함.",
        "built_space": "흙바닥의 공터이며, 배경에 수리 중인 낡은 트럭과 공구들이 놓여 있음.",
        "entities": "바닥의 질감이 비치는 홀로그램 지도, 지도 좌측 하단(서남쪽)의 밝은 붉은 점, 낡은 트럭 모두 지시사항에 부합함.",
        "hard_violations": [],
        "physics": "홀로그램이 바닥의 굴곡을 따라 자연스럽게 투사되어 있음."
       },
       {
        "label": "B",
        "direction": "1인칭 시점으로 카메라는 바닥의 홀로그램 지도를 내려다보고 있음.",
        "built_space": "바닥에 갈라진 틈이 있는 공터이며, 우측 상단에 트럭의 바퀴가 보임.",
        "entities": "홀로그램 지도와 붉은 점이 있으나, 지도 위에 프롬프트 텍스트가 그대로 노출됨. 화자의 팔과 장갑이 보임.",
        "hard_violations": [
         "읽을 수 있는 글씨 금지(No readable writing) 지시를 어기고 프롬프트의 텍스트('바닥에 투사된 홀로그램 지도', '서남쪽 위치의 붉은 점' 등)가 그대로 유출됨."
        ],
        "physics": "홀로그램이 평평하게 투사되었고, 화자의 팔은 자연스러운 자세를 취하고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "지시된 시점과 홀로그램 지도의 붉은 점 위치를 정확히 묘사했으며, 금지된 텍스트 노출 없이 프롬프트를 잘 따랐습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "프롬프트의 지시문이 그대로 이미지 안에 읽을 수 있는 텍스트로 유출되는 치명적인 오류가 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 바닥을 비스듬히 내려다보고 있으며, 시선은 중앙의 홀로그램 지도를 향함.",
        "built_space": "흙바닥의 공터이며, 배경에 수리 중인 낡은 트럭과 공구들이 놓여 있음.",
        "entities": "바닥의 질감이 비치는 홀로그램 지도, 지도 좌측 하단(서남쪽)의 밝은 붉은 점, 낡은 트럭 모두 지시사항에 부합함.",
        "hard_violations": [],
        "physics": "홀로그램이 바닥의 굴곡을 따라 자연스럽게 투사되어 있음."
       },
       {
        "label": "B",
        "direction": "1인칭 시점으로 카메라는 바닥의 홀로그램 지도를 내려다보고 있음.",
        "built_space": "바닥에 갈라진 틈이 있는 공터이며, 우측 상단에 트럭의 바퀴가 보임.",
        "entities": "홀로그램 지도와 붉은 점이 있으나, 지도 위에 프롬프트 텍스트가 그대로 노출됨. 화자의 팔과 장갑이 보임.",
        "hard_violations": [
         "읽을 수 있는 글씨 금지(No readable writing) 지시를 어기고 프롬프트의 텍스트('바닥에 투사된 홀로그램 지도', '서남쪽 위치의 붉은 점' 등)가 그대로 유출됨."
        ],
        "physics": "홀로그램이 평평하게 투사되었고, 화자의 팔은 자연스러운 자세를 취하고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 1,
        "verdict_ko": "서남쪽 붉은 점과 내려다보는 구도는 맞지만, 읽을 수 있는 설명문과 지도 테두리·조작 아이콘이 명시적 금지사항을 위반한다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "지면이 비치는 발광 지도와 좌하단 붉은 점은 더 충실하지만, 불필요한 전경 기계장치와 넓고 낮아진 시점 때문에 요구된 클로즈업을 완전히 실현하지 못한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 바닥의 지도를 비교적 가파르게 내려다본다. 밝은 붉은 점은 지도 좌하단에 있어 지정된 서남쪽 위치와 맞는다. 지도 면은 위를 향하며 관찰자에게 보인다. 인물의 눈이나 무기는 보이지 않는다.",
        "built_space": "갈라진 포장 바닥 중앙에 직사각형 지도 하나가 있고, 오른쪽 위에는 트럭 바퀴 하나와 그 뒤의 받침이 보인다. 바퀴 주변에는 공구와 부품이 여러 개 놓여 있다. 지도에 청록색 외곽 테두리와 왼쪽 조작 아이콘 열이 추가되어, 경계 없는 바닥 투사라는 요구와 다르다. 전경의 팔이 바닥 일부를 가린다.",
        "entities": "발광 지도, 좌하단 붉은 점, 공터 바닥, 낡은 트럭 일부가 보인다. 지도에는 상점 표식과 거리 숫자가 있지만, 설명 문구와 연도까지 선명하게 읽혀 글자 금지를 어긴다. 찢어진 소매와 장갑 낀 팔은 보이나 찰리의 정체나 센서 고장을 확인할 수는 없다. 인물 얼굴은 없고 별도의 참조 이미지도 제공되지 않았다.",
        "hard_violations": [
         "지도 위에 '바닥에 투사된 홀로그램 지도', '서남쪽 위치의 붉은 점', '2069', '25.6 km' 등 읽을 수 있는 문구와 숫자가 노출되어 있다.",
         "명시적으로 배제한 인터페이스 테두리와 조작 아이콘이 지도에 추가되어 있다."
        ],
        "physics": "지도는 포장 바닥에 밀착된 빛으로 표현되며, 붉은 점 주변에는 국소적인 발광이 있다. 바퀴와 공구는 지면에 놓여 있다. 전경 팔은 화면 밖 신체로 이어지는 구도이며, 떠 있는 독립 신체로 보이지 않는다. 다만 불투명한 위성사진 면과 선명한 설명문은 실제 지면 투사보다 화면 합성처럼 보인다."
       },
       {
        "label": "B",
        "direction": "카메라는 흙바닥 지도를 비스듬히 내려다보고, 지도 좌하단의 밝은 붉은 점으로 밝은 경로가 이어진다. 목적지 위치는 서남쪽 요구에 맞는다. 다만 먼 트럭과 배경까지 보이는 각도여서 요구된 가파른 하향 시점보다 낮다. 인물의 시선이나 무기는 없다.",
        "built_space": "흙과 자갈, 타이어 자국이 있는 공터 중앙에 지도 하나가 투사되어 있다. 위쪽 배경에는 트럭 한 대와 보이는 바퀴 두 개, 그 앞의 수리 도구들이 있다. 지도 주변 지면이 충분히 보이고 뚜렷한 인터페이스 외곽 틀은 없다. 그러나 아래쪽과 오른쪽 전경에 금속 프레임과 배선이 크게 들어와 장치 하드웨어 없는 구도와 충돌한다.",
        "entities": "투명한 도로 지도, 좌하단 붉은 목적지 점, 공터 바닥, 수리 중인 것으로 보이는 낡은 트럭이 있다. 지도에는 작은 문자형 표식들이 있지만 백화점·슈퍼마켓과 25.6km라는 정보는 확실히 판독되지 않는다. 사람은 보이지 않는다. 전경의 손상된 듯한 기계 부품을 찰리의 센서라고 확정할 근거는 없다.",
        "hard_violations": [
         "장치 하드웨어를 넣지 말라는 구도 지시와 달리, 전경에 금속 프레임과 노출 배선이 추가되어 있다."
        ],
        "physics": "지도 선 아래로 흙의 질감이 비치고 붉은 빛이 주변 지면에 번져 바닥 투사로 읽힌다. 트럭 바퀴와 수리 도구는 지면에 놓여 있다. 전경 배선은 화면 가장자리의 기계 구조에 연결되어 있으며, 지지 없이 떠 있는 물체나 인체는 보이지 않는다. 낮의 주변 밝기도 유지된다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "서남쪽 붉은 점과 내려다보는 구도는 맞지만, 읽을 수 있는 설명문과 지도 테두리·조작 아이콘이 명시적 금지사항을 위반한다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "지면이 비치는 발광 지도와 좌하단 붉은 점은 더 충실하지만, 불필요한 전경 기계장치와 넓고 낮아진 시점 때문에 요구된 클로즈업을 완전히 실현하지 못한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "카메라는 바닥의 지도를 비교적 가파르게 내려다본다. 밝은 붉은 점은 지도 좌하단에 있어 지정된 서남쪽 위치와 맞는다. 지도 면은 위를 향하며 관찰자에게 보인다. 인물의 눈이나 무기는 보이지 않는다.",
        "built_space": "갈라진 포장 바닥 중앙에 직사각형 지도 하나가 있고, 오른쪽 위에는 트럭 바퀴 하나와 그 뒤의 받침이 보인다. 바퀴 주변에는 공구와 부품이 여러 개 놓여 있다. 지도에 청록색 외곽 테두리와 왼쪽 조작 아이콘 열이 추가되어, 경계 없는 바닥 투사라는 요구와 다르다. 전경의 팔이 바닥 일부를 가린다.",
        "entities": "발광 지도, 좌하단 붉은 점, 공터 바닥, 낡은 트럭 일부가 보인다. 지도에는 상점 표식과 거리 숫자가 있지만, 설명 문구와 연도까지 선명하게 읽혀 글자 금지를 어긴다. 찢어진 소매와 장갑 낀 팔은 보이나 찰리의 정체나 센서 고장을 확인할 수는 없다. 인물 얼굴은 없고 별도의 참조 이미지도 제공되지 않았다.",
        "hard_violations": [
         "지도 위에 '바닥에 투사된 홀로그램 지도', '서남쪽 위치의 붉은 점', '2069', '25.6 km' 등 읽을 수 있는 문구와 숫자가 노출되어 있다.",
         "명시적으로 배제한 인터페이스 테두리와 조작 아이콘이 지도에 추가되어 있다."
        ],
        "physics": "지도는 포장 바닥에 밀착된 빛으로 표현되며, 붉은 점 주변에는 국소적인 발광이 있다. 바퀴와 공구는 지면에 놓여 있다. 전경 팔은 화면 밖 신체로 이어지는 구도이며, 떠 있는 독립 신체로 보이지 않는다. 다만 불투명한 위성사진 면과 선명한 설명문은 실제 지면 투사보다 화면 합성처럼 보인다."
       },
       {
        "label": "A",
        "direction": "카메라는 흙바닥 지도를 비스듬히 내려다보고, 지도 좌하단의 밝은 붉은 점으로 밝은 경로가 이어진다. 목적지 위치는 서남쪽 요구에 맞는다. 다만 먼 트럭과 배경까지 보이는 각도여서 요구된 가파른 하향 시점보다 낮다. 인물의 시선이나 무기는 없다.",
        "built_space": "흙과 자갈, 타이어 자국이 있는 공터 중앙에 지도 하나가 투사되어 있다. 위쪽 배경에는 트럭 한 대와 보이는 바퀴 두 개, 그 앞의 수리 도구들이 있다. 지도 주변 지면이 충분히 보이고 뚜렷한 인터페이스 외곽 틀은 없다. 그러나 아래쪽과 오른쪽 전경에 금속 프레임과 배선이 크게 들어와 장치 하드웨어 없는 구도와 충돌한다.",
        "entities": "투명한 도로 지도, 좌하단 붉은 목적지 점, 공터 바닥, 수리 중인 것으로 보이는 낡은 트럭이 있다. 지도에는 작은 문자형 표식들이 있지만 백화점·슈퍼마켓과 25.6km라는 정보는 확실히 판독되지 않는다. 사람은 보이지 않는다. 전경의 손상된 듯한 기계 부품을 찰리의 센서라고 확정할 근거는 없다.",
        "hard_violations": [
         "장치 하드웨어를 넣지 말라는 구도 지시와 달리, 전경에 금속 프레임과 노출 배선이 추가되어 있다."
        ],
        "physics": "지도 선 아래로 흙의 질감이 비치고 붉은 빛이 주변 지면에 번져 바닥 투사로 읽힌다. 트럭 바퀴와 수리 도구는 지면에 놓여 있다. 전경 배선은 화면 가장자리의 기계 구조에 연결되어 있으며, 지지 없이 떠 있는 물체나 인체는 보이지 않는다. 낮의 주변 밝기도 유지된다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.5
   },
   "adjusted": {
    "A": 1.75,
    "B": 0.25
   },
   "violations": {
    "B": [
     "[gemini-pro] 읽을 수 있는 글씨 금지(No readable writing) 지시를 어기고 프롬프트의 텍스트('바닥에 투사된 홀로그램 지도', '서남쪽 위치의 붉은 점' 등)가 그대로 유출됨.",
     "[gpt-high] 지도 위에 '바닥에 투사된 홀로그램 지도', '서남쪽 위치의 붉은 점', '2069', '25.6 km' 등 읽을 수 있는 문구와 숫자가 노출되어 있다.",
     "[gpt-high] 명시적으로 배제한 인터페이스 테두리와 조작 아이콘이 지도에 추가되어 있다."
    ],
    "A": [
     "[gpt-high] 장치 하드웨어를 넣지 말라는 구도 지시와 달리, 전경에 금속 프레임과 노출 배선이 추가되어 있다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 250
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "지시된 시점과 홀로그램 지도의 붉은 점 위치를 정확히 묘사했으며, 금지된 텍스트 노출 없이 프롬프트를 잘 따랐습니다.  ★위반: [gpt-high] 장치 하드웨어를 넣지 말라는 구도 지시와 달리, 전경에 금속 프레임과 노출 배선이 추가되어 있다."
   },
   {
    "label": "B",
    "score": 250,
    "verdict_ko": "프롬프트의 지시문이 그대로 이미지 안에 읽을 수 있는 텍스트로 유출되는 치명적인 오류가 발생했습니다.  ★위반: [gemini-pro] 읽을 수 있는 글씨 금지(No readable writing) 지시를 어기고 프롬프트의 텍스트('바닥에 투사된 홀로그램 지도', '서남쪽 위치의 붉은 점' 등)가 그대로 유출됨. / [gpt-high] 지도 위에 '바닥에 투사된 홀로그램 지도', '서남쪽 위치의 붉은 점', '2069', '25.6 km' 등 읽을 수 있는 문구와 숫자가 노출되어 있다. / [gpt-high] 명시적으로 배제한 인터페이스 테두리와 조작 아이콘이 지도에 추가되어 있다."
   }
  ],
  "refs": [],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-8e33-7dd1-9052-8535d95ff2a0",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S62sh15::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:49:11.399577+00:00",
  "fingerprint": "894de6c2ff9503a055b819b25203d8b0eed713492aae5a377de71af2dd7e1859",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S62sh15_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S62sh15_sel.png",
  "source_sha256": "c3d2fe42bc3bdb3e666265c1dde589f700b5edadb32f227e59b3d942581ba402",
  "file": "S62sh15_cine.png",
  "staged_sha256": "f5f5e48cc691b37f8c29c63576332017de2600b607b1eced6dfd6e29e31c59e1",
  "latency_ms": 13116
 },
 "S62sh18::signage": {
  "fp": "f3f240bb7116c54f",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S62sh18": {
  "input_fingerprint": "88ec2acc8a369177",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 자신감에 찬 표정으로 한쪽 입꼬리를 올린 채 미소 짓고 있는 현우의 얼굴 클로즈업.\n\nLOCATION (lock): In the village's open square near the stalled old truck, where the food-search plan is discussed. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Village clearing (The conversation remains in the clearing beside the truck); used as A narrow, softly resolved background margin preserves the location without competing with the half-smile.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight appropriate to the open clearing maintains restrained contrast and natural facial detail without turning the confident smile into a heroic lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old truck remains the vehicle being prepared in the village square, with an old football in use nearby. Charlie's punctured, dented body and sensor damage have not been repaired. 현우: He stands in the village square, retaining his battle injuries.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 자신감에 찬 표정으로 한쪽 입꼬리를 올린 채 미소 짓고 있는 현우의 얼굴 클로즈업.\n\nLOCATION (lock): In the village's open square near the stalled old truck, where the food-search plan is discussed. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Village clearing (The conversation remains in the clearing beside the truck); used as A narrow, softly resolved background margin preserves the location without competing with the half-smile.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight appropriate to the open clearing maintains restrained contrast and natural facial detail without turning the confident smile into a heroic lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old truck remains the vehicle being prepared in the village square, with an old football in use nearby. Charlie's punctured, dented body and sensor damage have not been repaired. 현우: He stands in the village square, retaining his battle injuries.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 자신감에 찬 표정으로 한쪽 입꼬리를 올린 채 미소 짓고 있는 현우의 얼굴 클로즈업.\n\nLOCATION (lock): In the village's open square near the stalled old truck, where the food-search plan is discussed. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Village clearing (The conversation remains in the clearing beside the truck); used as A narrow, softly resolved background margin preserves the location without competing with the half-smile.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight appropriate to the open clearing maintains restrained contrast and natural facial detail without turning the confident smile into a heroic lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old truck remains the vehicle being prepared in the village square, with an old football in use nearby. Charlie's punctured, dented body and sensor damage have not been repaired. 현우: He stands in the village square, retaining his battle injuries.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선은 화면 우측 밖을 향하고 있음.",
    "built_space": "배경에 낡은 트럭이 위치하며 얕은 심도로 마을 공터가 표현됨.",
    "entities": "현우의 얼굴, 파란 티셔츠, 뺨의 상처가 확인되나 우측 전경에 프롬프트에 없는 인물의 뒷모습이 걸쳐 있음.",
    "hard_violations": [
     "[gemini-pro] 프롬프트에 없는 인물(우측 전경) 추가",
     "[gpt-high] 현우만 허용된 장면의 오른쪽 전경에 다른 인물의 머리와 얼굴 일부를 추가했습니다."
    ],
    "physics": "현우의 자세와 머리는 자연스럽게 신체에 의해 지탱됨."
   },
   {
    "label": "B",
    "direction": "현우의 시선은 화면 좌측 밖을 향하고 있음.",
    "built_space": "우측 배경에 낡은 트럭과 마을 공터가 보임.",
    "entities": "현우의 의상 색상이 다르고 전투 상처가 없으며, 좌측 배경에 프롬프트에 없는 인물 2명이 서 있음.",
    "hard_violations": [
     "[gemini-pro] 프롬프트에 없는 인물들(좌측 배경) 추가",
     "[gpt-high] 현우만 허용된 장면에 최소 두 명의 배경 인물을 추가했습니다."
    ],
    "physics": "신체는 자연스럽게 서 있는 상태로 지탱됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "요청된 표정(한쪽 입꼬리 올림), 상처, 의상은 잘 구현되었으나, 화면 우측 전경에 지시되지 않은 인물이 추가되어 치명적인 위반입니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "한쪽 입꼬리만 올린 미소가 아니며 상처가 누락되었고, 배경에 지시되지 않은 인물들이 추가되어 역시 치명적인 위반입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선은 화면 우측 밖을 향하고 있음.",
        "built_space": "배경에 낡은 트럭이 위치하며 얕은 심도로 마을 공터가 표현됨.",
        "entities": "현우의 얼굴, 파란 티셔츠, 뺨의 상처가 확인되나 우측 전경에 프롬프트에 없는 인물의 뒷모습이 걸쳐 있음.",
        "hard_violations": [
         "프롬프트에 없는 인물(우측 전경) 추가"
        ],
        "physics": "현우의 자세와 머리는 자연스럽게 신체에 의해 지탱됨."
       },
       {
        "label": "B",
        "direction": "현우의 시선은 화면 좌측 밖을 향하고 있음.",
        "built_space": "우측 배경에 낡은 트럭과 마을 공터가 보임.",
        "entities": "현우의 의상 색상이 다르고 전투 상처가 없으며, 좌측 배경에 프롬프트에 없는 인물 2명이 서 있음.",
        "hard_violations": [
         "프롬프트에 없는 인물들(좌측 배경) 추가"
        ],
        "physics": "신체는 자연스럽게 서 있는 상태로 지탱됨."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "요청된 표정(한쪽 입꼬리 올림), 상처, 의상은 잘 구현되었으나, 화면 우측 전경에 지시되지 않은 인물이 추가되어 치명적인 위반입니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "한쪽 입꼬리만 올린 미소가 아니며 상처가 누락되었고, 배경에 지시되지 않은 인물들이 추가되어 역시 치명적인 위반입니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선은 화면 우측 밖을 향하고 있음.",
        "built_space": "배경에 낡은 트럭이 위치하며 얕은 심도로 마을 공터가 표현됨.",
        "entities": "현우의 얼굴, 파란 티셔츠, 뺨의 상처가 확인되나 우측 전경에 프롬프트에 없는 인물의 뒷모습이 걸쳐 있음.",
        "hard_violations": [
         "프롬프트에 없는 인물(우측 전경) 추가"
        ],
        "physics": "현우의 자세와 머리는 자연스럽게 신체에 의해 지탱됨."
       },
       {
        "label": "B",
        "direction": "현우의 시선은 화면 좌측 밖을 향하고 있음.",
        "built_space": "우측 배경에 낡은 트럭과 마을 공터가 보임.",
        "entities": "현우의 의상 색상이 다르고 전투 상처가 없으며, 좌측 배경에 프롬프트에 없는 인물 2명이 서 있음.",
        "hard_violations": [
         "프롬프트에 없는 인물들(좌측 배경) 추가"
        ],
        "physics": "신체는 자연스럽게 서 있는 상태로 지탱됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "배경 인물 추가로 부적격이지만, 현우 단독 얼굴 중심의 클로즈업과 한쪽 입꼬리 미소는 전경 인물을 끼운 B보다 지시된 구도에 가깝습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "상처와 남색 상의는 잘 반영했지만, 허용되지 않은 상대의 머리를 전경에 넣어 현우 얼굴 클로즈업을 대화 상대 너머 구도로 바꿨습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 얼굴은 거의 정면이고 눈은 화면 왼쪽 바깥을 향합니다. 시선 대상은 보이지 않으며, 샷 지문은 특정 대상을 지정하지 않습니다. 화면 오른쪽 입꼬리가 조금 더 올라간 미소가 보입니다. 무기나 방향성 있는 휴대 물체는 없습니다.",
        "built_space": "흙바닥 광장 뒤로 낮은 마을 건물들이 있고, 화면 오른쪽에 낡은 트럭 한 대의 운전실과 앞바퀴가 보입니다. 현우는 트럭 앞쪽 공간에 서 있는 배치로 읽힙니다. 얼굴과 어깨를 담았으나 배경이 양옆으로 넓게 드러나, 요구한 좁고 부드러운 배경 여백보다 트럭과 마을의 비중이 큽니다. 중복 차량이나 불가능한 반사는 보이지 않습니다.",
        "entities": "중앙 인물은 앳된 동아시아계 남성으로, 헝클어진 검은 머리와 얼굴 특징이 현우 참조에 대체로 부합합니다. 한국계 미국인이라는 국적·배경은 외관만으로 확인할 수 없습니다. 상의는 참조의 남색보다 회흑색에 가깝고, 뚜렷한 전투 상처는 식별되지 않습니다. 왼쪽 배경에 최소 두 명의 추가 인물이 보입니다. 낡은 트럭은 있으나 축구공과 찰리는 이 얼굴 중심 구도에서 식별되지 않으며, 이를 누락으로 보지는 않습니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "현우만 허용된 장면에 최소 두 명의 배경 인물을 추가했습니다."
        ],
        "physics": "현우의 머리는 목과 어깨에 자연스럽게 이어지고, 서 있는 상체 자세에 해부학적 모순은 없습니다. 발과 지면 접촉은 프레임 밖입니다. 트럭의 보이는 앞바퀴는 흙바닥에 놓여 있으며, 배경 인물들도 지면 위에 서 있는 것으로 보입니다. 떠 있는 물체나 지지 없는 신체는 보이지 않습니다."
       },
       {
        "label": "B",
        "direction": "현우는 화면 오른쪽 전경에 있는 상대의 얼굴을 향해 고개와 눈을 돌리고 있습니다. 시선과 상대 위치는 일치하지만, 그 상대는 지문이 허용하지 않은 인물입니다. 한쪽 입꼬리를 살짝 올린 자신감 있는 미소가 보이며, 무기나 휴대 물체는 없습니다.",
        "built_space": "흙바닥 광장과 낮은 건물, 뒤쪽의 낡은 트럭 한 대가 흐리게 보입니다. 현우는 트럭보다 카메라 가까이에 있고, 추가 인물의 머리가 오른쪽 전경을 크게 가립니다. 얼굴은 크게 담겼지만 상대 너머로 현우를 보는 구도가 되어, 현우 얼굴과 좁은 배경 여백만을 중심으로 한 지시에서 벗어납니다. 중복 차량이나 불가능한 반사는 없습니다.",
        "entities": "현우는 앳된 동아시아계 남성으로 보이고 검은 머리, 얼굴 특징, 남색 둥근 목 상의가 참조와 대체로 맞습니다. 화면 왼쪽 뺨의 긁힌 상처가 전투 부상 유지 조건을 반영합니다. 오른쪽 전경에는 별도 인물의 검은 머리와 귀, 턱 일부가 보이며, 이 인물의 신원과 나이는 확인할 수 없습니다. 트럭은 보이지만 축구공과 찰리는 식별되지 않으며, 클로즈업 밖의 항목을 누락으로 평가하지 않습니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "현우만 허용된 장면의 오른쪽 전경에 다른 인물의 머리와 얼굴 일부를 추가했습니다."
        ],
        "physics": "현우의 머리와 목, 어깨 연결 및 가벼운 고개 회전은 자연스럽습니다. 전경 인물은 몸통이 화면 밖에 있어 머리 일부만 보이는 것으로 읽히며, 부유하는 신체로 볼 근거는 없습니다. 트럭은 바퀴로 지면에 지지됩니다. 점프나 투척, 지지 없는 물체는 없습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "배경 인물 추가로 부적격이지만, 현우 단독 얼굴 중심의 클로즈업과 한쪽 입꼬리 미소는 전경 인물을 끼운 B보다 지시된 구도에 가깝습니다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "상처와 남색 상의는 잘 반영했지만, 허용되지 않은 상대의 머리를 전경에 넣어 현우 얼굴 클로즈업을 대화 상대 너머 구도로 바꿨습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 얼굴은 거의 정면이고 눈은 화면 왼쪽 바깥을 향합니다. 시선 대상은 보이지 않으며, 샷 지문은 특정 대상을 지정하지 않습니다. 화면 오른쪽 입꼬리가 조금 더 올라간 미소가 보입니다. 무기나 방향성 있는 휴대 물체는 없습니다.",
        "built_space": "흙바닥 광장 뒤로 낮은 마을 건물들이 있고, 화면 오른쪽에 낡은 트럭 한 대의 운전실과 앞바퀴가 보입니다. 현우는 트럭 앞쪽 공간에 서 있는 배치로 읽힙니다. 얼굴과 어깨를 담았으나 배경이 양옆으로 넓게 드러나, 요구한 좁고 부드러운 배경 여백보다 트럭과 마을의 비중이 큽니다. 중복 차량이나 불가능한 반사는 보이지 않습니다.",
        "entities": "중앙 인물은 앳된 동아시아계 남성으로, 헝클어진 검은 머리와 얼굴 특징이 현우 참조에 대체로 부합합니다. 한국계 미국인이라는 국적·배경은 외관만으로 확인할 수 없습니다. 상의는 참조의 남색보다 회흑색에 가깝고, 뚜렷한 전투 상처는 식별되지 않습니다. 왼쪽 배경에 최소 두 명의 추가 인물이 보입니다. 낡은 트럭은 있으나 축구공과 찰리는 이 얼굴 중심 구도에서 식별되지 않으며, 이를 누락으로 보지는 않습니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "현우만 허용된 장면에 최소 두 명의 배경 인물을 추가했습니다."
        ],
        "physics": "현우의 머리는 목과 어깨에 자연스럽게 이어지고, 서 있는 상체 자세에 해부학적 모순은 없습니다. 발과 지면 접촉은 프레임 밖입니다. 트럭의 보이는 앞바퀴는 흙바닥에 놓여 있으며, 배경 인물들도 지면 위에 서 있는 것으로 보입니다. 떠 있는 물체나 지지 없는 신체는 보이지 않습니다."
       },
       {
        "label": "A",
        "direction": "현우는 화면 오른쪽 전경에 있는 상대의 얼굴을 향해 고개와 눈을 돌리고 있습니다. 시선과 상대 위치는 일치하지만, 그 상대는 지문이 허용하지 않은 인물입니다. 한쪽 입꼬리를 살짝 올린 자신감 있는 미소가 보이며, 무기나 휴대 물체는 없습니다.",
        "built_space": "흙바닥 광장과 낮은 건물, 뒤쪽의 낡은 트럭 한 대가 흐리게 보입니다. 현우는 트럭보다 카메라 가까이에 있고, 추가 인물의 머리가 오른쪽 전경을 크게 가립니다. 얼굴은 크게 담겼지만 상대 너머로 현우를 보는 구도가 되어, 현우 얼굴과 좁은 배경 여백만을 중심으로 한 지시에서 벗어납니다. 중복 차량이나 불가능한 반사는 없습니다.",
        "entities": "현우는 앳된 동아시아계 남성으로 보이고 검은 머리, 얼굴 특징, 남색 둥근 목 상의가 참조와 대체로 맞습니다. 화면 왼쪽 뺨의 긁힌 상처가 전투 부상 유지 조건을 반영합니다. 오른쪽 전경에는 별도 인물의 검은 머리와 귀, 턱 일부가 보이며, 이 인물의 신원과 나이는 확인할 수 없습니다. 트럭은 보이지만 축구공과 찰리는 식별되지 않으며, 클로즈업 밖의 항목을 누락으로 평가하지 않습니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "현우만 허용된 장면의 오른쪽 전경에 다른 인물의 머리와 얼굴 일부를 추가했습니다."
        ],
        "physics": "현우의 머리와 목, 어깨 연결 및 가벼운 고개 회전은 자연스럽습니다. 전경 인물은 몸통이 화면 밖에 있어 머리 일부만 보이는 것으로 읽히며, 부유하는 신체로 볼 근거는 없습니다. 트럭은 바퀴로 지면에 지지됩니다. 점프나 투척, 지지 없는 물체는 없습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.667,
    "B": 1.667
   },
   "adjusted": {
    "A": 1.417,
    "B": 1.417
   },
   "violations": {
    "A": [
     "[gemini-pro] 프롬프트에 없는 인물(우측 전경) 추가",
     "[gpt-high] 현우만 허용된 장면의 오른쪽 전경에 다른 인물의 머리와 얼굴 일부를 추가했습니다."
    ],
    "B": [
     "[gemini-pro] 프롬프트에 없는 인물들(좌측 배경) 추가",
     "[gpt-high] 현우만 허용된 장면에 최소 두 명의 배경 인물을 추가했습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1417,
   "B": 1417
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1417,
    "verdict_ko": "요청된 표정(한쪽 입꼬리 올림), 상처, 의상은 잘 구현되었으나, 화면 우측 전경에 지시되지 않은 인물이 추가되어 치명적인 위반입니다.  ★위반: [gemini-pro] 프롬프트에 없는 인물(우측 전경) 추가 / [gpt-high] 현우만 허용된 장면의 오른쪽 전경에 다른 인물의 머리와 얼굴 일부를 추가했습니다."
   },
   {
    "label": "B",
    "score": 1417,
    "verdict_ko": "한쪽 입꼬리만 올린 미소가 아니며 상처가 누락되었고, 배경에 지시되지 않은 인물들이 추가되어 역시 치명적인 위반입니다.  ★위반: [gemini-pro] 프롬프트에 없는 인물들(좌측 배경) 추가 / [gpt-high] 현우만 허용된 장면에 최소 두 명의 배경 인물을 추가했습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features, lighting mood and each person's clothing are LOCKED to this photo; never copy its camera framing. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S62sh15_sel.png",
    "asset_id": "7350440c-3f64-4d57-b9de-c162a9138d92",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-8fd6-7909-a24d-747650458a1a",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S62sh15"
  }
 },
 "S62sh18::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:50:14.723117+00:00",
  "fingerprint": "2e3d5ccec51858b62307097cccea9b83939c41452f076937685eb6762553ceb5",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S62sh18_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S62sh18_sel.png",
  "source_sha256": "91eddddad9145c25bfc94d5fefb1dec0797cf872fc3aaa21b5326c0cbd279640",
  "file": "S62sh18_cine.png",
  "staged_sha256": "0f51c05691e13d1b9d4054c81cce78792e7177c3b9a6e3d9f60153dbc1f0d679",
  "latency_ms": 12500
 },
 "S63sh1::signage": {
  "fp": "f28921543ee49df5",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S63sh1": {
  "input_fingerprint": "3722c1480235373d",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 밝은 낮, 황량한 흙먼지 도로 위를 달리며 바퀴 뒤로 흙먼지를 일으키고 있는 낡은 트럭의 외관 전경.\n\nLOCATION (lock): On a dusty rural road outside the village, where the old supply truck travels through the exposed landscape. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Old truck (Traveling along the road and raising dust behind its wheels) — The rear and passenger-side exterior are visible from an elevated rear-quarter angle; used as The moving spatial anchor, kept modest in size within the surrounding road; Barren dirt road (Extends ahead of the traveling truck) — Runs from the lower foreground past the truck toward the upper-center distance; used as Establishes forward travel and provides the visual direction to preserve in the subsequent cab cut; Wheel-raised dust (Trailing behind the moving truck); used as Makes the vehicle's movement readable without obscuring its exterior.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Bright daytime light makes the barren road and wheel-raised dust legible with controlled contrast and an understated tonal range.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old village truck is traveling along the road in daylight.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 밝은 낮, 황량한 흙먼지 도로 위를 달리며 바퀴 뒤로 흙먼지를 일으키고 있는 낡은 트럭의 외관 전경.\n\nLOCATION (lock): On a dusty rural road outside the village, where the old supply truck travels through the exposed landscape. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Old truck (Traveling along the road and raising dust behind its wheels) — The rear and passenger-side exterior are visible from an elevated rear-quarter angle; used as The moving spatial anchor, kept modest in size within the surrounding road; Barren dirt road (Extends ahead of the traveling truck) — Runs from the lower foreground past the truck toward the upper-center distance; used as Establishes forward travel and provides the visual direction to preserve in the subsequent cab cut; Wheel-raised dust (Trailing behind the moving truck); used as Makes the vehicle's movement readable without obscuring its exterior.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Bright daytime light makes the barren road and wheel-raised dust legible with controlled contrast and an understated tonal range.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old village truck is traveling along the road in daylight.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 밝은 낮, 황량한 흙먼지 도로 위를 달리며 바퀴 뒤로 흙먼지를 일으키고 있는 낡은 트럭의 외관 전경.\n\nLOCATION (lock): On a dusty rural road outside the village, where the old supply truck travels through the exposed landscape. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Old truck (Traveling along the road and raising dust behind its wheels) — The rear and passenger-side exterior are visible from an elevated rear-quarter angle; used as The moving spatial anchor, kept modest in size within the surrounding road; Barren dirt road (Extends ahead of the traveling truck) — Runs from the lower foreground past the truck toward the upper-center distance; used as Establishes forward travel and provides the visual direction to preserve in the subsequent cab cut; Wheel-raised dust (Trailing behind the moving truck); used as Makes the vehicle's movement readable without obscuring its exterior.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Bright daytime light makes the barren road and wheel-raised dust legible with controlled contrast and an understated tonal range.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old village truck is traveling along the road in daylight.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "트럭이 화면 중앙 먼 곳을 향해 주행하며 바퀴 뒤로 흙먼지를 일으키고 있습니다.",
    "built_space": "참고 이미지의 좁은 바위 지형이 배경에 배치되었으나, 바위의 하단부가 도로 지면과 이어지지 않고 일직선으로 잘려 있습니다.",
    "entities": "낡은 트럭, 흙먼지가 날리는 황량한 비포장 도로는 프롬프트의 설명과 일치합니다. 인물은 등장하지 않습니다.",
    "hard_violations": [
     "[gemini-pro] 콜라주/합성 위반 (참고 이미지의 바위 지형을 2D 스티커처럼 배경에 오려 붙임)"
    ],
    "physics": "트럭과 흙먼지는 지면에 닿아 있으나, 배경의 바위는 지지되는 바닥 형태 없이 지평선 위에 부자연스럽게 떠 있습니다."
   },
   {
    "label": "B",
    "direction": "트럭이 화면 전방의 흙길을 따라 주행 중입니다.",
    "built_space": "도로 좌측에 참고 이미지의 바위 지형이 얹혀 있으며 주변 풍경과 어색하게 결합되어 있습니다.",
    "entities": "트럭과 흙길은 존재하나, 트럭의 후면에 '낡은 트럭'이라는 명확한 한글 텍스트가 노출되어 있습니다.",
    "hard_violations": [
     "[gemini-pro] 금지된 텍스트 노출 (트럭 후면의 '낡은 트럭' 글자)",
     "[gemini-pro] 콜라주/합성 위반 (트럭 아래의 흙먼지가 뚜렷한 직사각형 경계를 가진 그래픽 오버레이로 덮여 있음)",
     "[gpt-high] 적재함 후면에 읽을 수 있는 한글 문구가 있어, 어디에도 읽을 수 있는 글자를 넣지 말라는 조건을 위반한다."
    ],
    "physics": "트럭 아래 발생하는 흙먼지가 자연스럽게 흩날리지 않고 직사각형 형태의 평면적인 덩어리로 바닥 위에 떠 있어 물리적으로 불가능합니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "프롬프트의 차량과 도로 묘사는 따랐으나, 참고 이미지를 평면적인 스티커처럼 배경에 그대로 오려 붙여 실사 요건을 심각하게 위반했습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "프롬프트가 명시적으로 금지한 텍스트('낡은 트럭')가 포함되었으며, 흙먼지가 직사각형 그래픽으로 조잡하게 합성되어 기각 대상입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "트럭이 화면 중앙 먼 곳을 향해 주행하며 바퀴 뒤로 흙먼지를 일으키고 있습니다.",
        "built_space": "참고 이미지의 좁은 바위 지형이 배경에 배치되었으나, 바위의 하단부가 도로 지면과 이어지지 않고 일직선으로 잘려 있습니다.",
        "entities": "낡은 트럭, 흙먼지가 날리는 황량한 비포장 도로는 프롬프트의 설명과 일치합니다. 인물은 등장하지 않습니다.",
        "hard_violations": [
         "콜라주/합성 위반 (참고 이미지의 바위 지형을 2D 스티커처럼 배경에 오려 붙임)"
        ],
        "physics": "트럭과 흙먼지는 지면에 닿아 있으나, 배경의 바위는 지지되는 바닥 형태 없이 지평선 위에 부자연스럽게 떠 있습니다."
       },
       {
        "label": "B",
        "direction": "트럭이 화면 전방의 흙길을 따라 주행 중입니다.",
        "built_space": "도로 좌측에 참고 이미지의 바위 지형이 얹혀 있으며 주변 풍경과 어색하게 결합되어 있습니다.",
        "entities": "트럭과 흙길은 존재하나, 트럭의 후면에 '낡은 트럭'이라는 명확한 한글 텍스트가 노출되어 있습니다.",
        "hard_violations": [
         "금지된 텍스트 노출 (트럭 후면의 '낡은 트럭' 글자)",
         "콜라주/합성 위반 (트럭 아래의 흙먼지가 뚜렷한 직사각형 경계를 가진 그래픽 오버레이로 덮여 있음)"
        ],
        "physics": "트럭 아래 발생하는 흙먼지가 자연스럽게 흩날리지 않고 직사각형 형태의 평면적인 덩어리로 바닥 위에 떠 있어 물리적으로 불가능합니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "프롬프트의 차량과 도로 묘사는 따랐으나, 참고 이미지를 평면적인 스티커처럼 배경에 그대로 오려 붙여 실사 요건을 심각하게 위반했습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "프롬프트가 명시적으로 금지한 텍스트('낡은 트럭')가 포함되었으며, 흙먼지가 직사각형 그래픽으로 조잡하게 합성되어 기각 대상입니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "트럭이 화면 중앙 먼 곳을 향해 주행하며 바퀴 뒤로 흙먼지를 일으키고 있습니다.",
        "built_space": "참고 이미지의 좁은 바위 지형이 배경에 배치되었으나, 바위의 하단부가 도로 지면과 이어지지 않고 일직선으로 잘려 있습니다.",
        "entities": "낡은 트럭, 흙먼지가 날리는 황량한 비포장 도로는 프롬프트의 설명과 일치합니다. 인물은 등장하지 않습니다.",
        "hard_violations": [
         "콜라주/합성 위반 (참고 이미지의 바위 지형을 2D 스티커처럼 배경에 오려 붙임)"
        ],
        "physics": "트럭과 흙먼지는 지면에 닿아 있으나, 배경의 바위는 지지되는 바닥 형태 없이 지평선 위에 부자연스럽게 떠 있습니다."
       },
       {
        "label": "B",
        "direction": "트럭이 화면 전방의 흙길을 따라 주행 중입니다.",
        "built_space": "도로 좌측에 참고 이미지의 바위 지형이 얹혀 있으며 주변 풍경과 어색하게 결합되어 있습니다.",
        "entities": "트럭과 흙길은 존재하나, 트럭의 후면에 '낡은 트럭'이라는 명확한 한글 텍스트가 노출되어 있습니다.",
        "hard_violations": [
         "금지된 텍스트 노출 (트럭 후면의 '낡은 트럭' 글자)",
         "콜라주/합성 위반 (트럭 아래의 흙먼지가 뚜렷한 직사각형 경계를 가진 그래픽 오버레이로 덮여 있음)"
        ],
        "physics": "트럭 아래 발생하는 흙먼지가 자연스럽게 흩날리지 않고 직사각형 형태의 평면적인 덩어리로 바닥 위에 떠 있어 물리적으로 불가능합니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "후면·조수석 측면과 주행 먼지는 맞지만, 적재함의 읽을 수 있는 한글이 명시적 금지 사항을 위반하고 참조 바위가 과도하게 큰 전경을 차지한다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "문자 없이 밝은 낮의 주행 와이드숏을 구현했지만, 조수석이 아닌 운전석 측면이 보이고 도로가 상단 중앙 대신 왼쪽으로 향하며 참조 바위의 규모와 연결이 부자연스럽다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "트럭 앞부분은 화면 오른쪽 위의 마을과 도로 소실점을 향한다. 먼지는 뒷바퀴에서 왼쪽 아래로 이어져 전진 방향과 맞는다. 후면과 오른쪽 조수석 측면이 보인다. 다만 도로의 목적 방향은 요구된 상단 중앙보다 오른쪽에 치우친다. 식별 가능한 인물이나 시선은 없다.",
        "built_space": "비포장도로 하나가 하단 전경에서 트럭을 지나 우측 상단 마을로 이어진다. 왼쪽에는 참조의 큰 좌측 바위, 우측 바위 덩어리, 위를 가로지르는 바위와 안쪽 돌출석으로 이루어진 틈이 재현되어 있다. 그러나 이 바위군이 화면 왼쪽 절반을 압도하며 도로변 지면과의 연결도 사진 조각을 붙인 듯 급격하다. 원경에는 여러 주택과 전신주가 있다. 참조만으로 이 추가 배치나 바위와 도로의 정확한 공간 관계는 확인되지 않는다.",
        "entities": "낡고 도장이 벗겨진 파란 소형 화물트럭 한 대, 흙길, 마른 주변 지형, 바퀴 뒤 흙먼지가 보인다. 적재함에는 자루와 묶인 짐이 있다. 인물은 식별되지 않아 장면 수준 인물 목록을 불필요하게 추가하지 않았다. 적재함 후면에는 읽을 수 있는 한글 표기가 있어 문자 금지 조건에 어긋난다. 참조 바위의 무늬와 틈 형태는 닮았지만 전체 장면 속 규모와 재질 연결은 자연스럽지 않다.",
        "hard_violations": [
         "적재함 후면에 읽을 수 있는 한글 문구가 있어, 어디에도 읽을 수 있는 글자를 넣지 말라는 조건을 위반한다."
        ],
        "physics": "보이는 앞뒤 타이어가 도로에 닿아 차체를 지지하며, 적재물은 적재함 바닥 위에 놓여 있다. 먼지는 바퀴 주변에서 일어나 차량 뒤로 퍼져 주행으로 발생한 것으로 읽힌다. 먼지가 하부 일부를 가리지만 트럭 외관은 대체로 보인다. 명백히 지지 없이 떠 있는 차량이나 짐은 없다."
       },
       {
        "label": "B",
        "direction": "트럭 앞부분은 화면 왼쪽 위로 이어지는 도로를 향하고 먼지는 오른쪽 아래 후방으로 흐른다. 전진과 먼지 방향은 일치한다. 그러나 보이는 것은 후면과 왼쪽 운전석 측면으로, 지정된 조수석 측면과 반대다. 도로도 상단 중앙이 아니라 왼쪽 원경으로 빠진다. 식별 가능한 사람의 시선은 없다.",
        "built_space": "도로 하나가 하단 오른쪽에서 트럭 아래를 지나 상단 왼쪽으로 굽는다. 주변은 메마른 흙과 돌, 드문 나무로 이루어져 있다. 상단 중앙에는 참조의 좁은 바위 틈과 이를 둘러싼 좌우 바위, 상부 가로 바위 및 안쪽 돌출석이 보인다. 다만 바위군이 먼 배경에 비해 지나치게 크고, 바닥 경계와 주변 지형의 연결이 부자연스럽다. 인공 건축물이나 고정 설비는 보이지 않는다.",
        "entities": "녹슨 파란 화물트럭 한 대와 비포장도로, 바퀴 뒤의 흙먼지가 확인된다. 적재함에는 낮게 놓인 밝은색 물체 일부가 보이지만 종류는 확정할 수 없다. 인물과 읽을 수 있는 문자는 보이지 않는다. 낡은 트럭과 황량한 낮 풍경은 요청에 맞으며, 바위 표면과 틈 형태는 참조를 따르지만 주변 풍경에 자연스럽게 통합되지는 않는다.",
        "hard_violations": [],
        "physics": "왼쪽 앞뒤 타이어가 노면에 닿아 트럭을 지지한다. 적재함 안 물체도 바닥 위에 놓인 것으로 보인다. 흙먼지가 뒷바퀴 부근에서 솟아 뒤로 길게 퍼져 주행 동작을 설명한다. 먼지가 후면 하부를 일부 가리지만 차체 대부분은 드러난다. 지지 없이 공중에 떠 있는 물체나 불가능한 차체 자세는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "후면·조수석 측면과 주행 먼지는 맞지만, 적재함의 읽을 수 있는 한글이 명시적 금지 사항을 위반하고 참조 바위가 과도하게 큰 전경을 차지한다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "문자 없이 밝은 낮의 주행 와이드숏을 구현했지만, 조수석이 아닌 운전석 측면이 보이고 도로가 상단 중앙 대신 왼쪽으로 향하며 참조 바위의 규모와 연결이 부자연스럽다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "트럭 앞부분은 화면 오른쪽 위의 마을과 도로 소실점을 향한다. 먼지는 뒷바퀴에서 왼쪽 아래로 이어져 전진 방향과 맞는다. 후면과 오른쪽 조수석 측면이 보인다. 다만 도로의 목적 방향은 요구된 상단 중앙보다 오른쪽에 치우친다. 식별 가능한 인물이나 시선은 없다.",
        "built_space": "비포장도로 하나가 하단 전경에서 트럭을 지나 우측 상단 마을로 이어진다. 왼쪽에는 참조의 큰 좌측 바위, 우측 바위 덩어리, 위를 가로지르는 바위와 안쪽 돌출석으로 이루어진 틈이 재현되어 있다. 그러나 이 바위군이 화면 왼쪽 절반을 압도하며 도로변 지면과의 연결도 사진 조각을 붙인 듯 급격하다. 원경에는 여러 주택과 전신주가 있다. 참조만으로 이 추가 배치나 바위와 도로의 정확한 공간 관계는 확인되지 않는다.",
        "entities": "낡고 도장이 벗겨진 파란 소형 화물트럭 한 대, 흙길, 마른 주변 지형, 바퀴 뒤 흙먼지가 보인다. 적재함에는 자루와 묶인 짐이 있다. 인물은 식별되지 않아 장면 수준 인물 목록을 불필요하게 추가하지 않았다. 적재함 후면에는 읽을 수 있는 한글 표기가 있어 문자 금지 조건에 어긋난다. 참조 바위의 무늬와 틈 형태는 닮았지만 전체 장면 속 규모와 재질 연결은 자연스럽지 않다.",
        "hard_violations": [
         "적재함 후면에 읽을 수 있는 한글 문구가 있어, 어디에도 읽을 수 있는 글자를 넣지 말라는 조건을 위반한다."
        ],
        "physics": "보이는 앞뒤 타이어가 도로에 닿아 차체를 지지하며, 적재물은 적재함 바닥 위에 놓여 있다. 먼지는 바퀴 주변에서 일어나 차량 뒤로 퍼져 주행으로 발생한 것으로 읽힌다. 먼지가 하부 일부를 가리지만 트럭 외관은 대체로 보인다. 명백히 지지 없이 떠 있는 차량이나 짐은 없다."
       },
       {
        "label": "A",
        "direction": "트럭 앞부분은 화면 왼쪽 위로 이어지는 도로를 향하고 먼지는 오른쪽 아래 후방으로 흐른다. 전진과 먼지 방향은 일치한다. 그러나 보이는 것은 후면과 왼쪽 운전석 측면으로, 지정된 조수석 측면과 반대다. 도로도 상단 중앙이 아니라 왼쪽 원경으로 빠진다. 식별 가능한 사람의 시선은 없다.",
        "built_space": "도로 하나가 하단 오른쪽에서 트럭 아래를 지나 상단 왼쪽으로 굽는다. 주변은 메마른 흙과 돌, 드문 나무로 이루어져 있다. 상단 중앙에는 참조의 좁은 바위 틈과 이를 둘러싼 좌우 바위, 상부 가로 바위 및 안쪽 돌출석이 보인다. 다만 바위군이 먼 배경에 비해 지나치게 크고, 바닥 경계와 주변 지형의 연결이 부자연스럽다. 인공 건축물이나 고정 설비는 보이지 않는다.",
        "entities": "녹슨 파란 화물트럭 한 대와 비포장도로, 바퀴 뒤의 흙먼지가 확인된다. 적재함에는 낮게 놓인 밝은색 물체 일부가 보이지만 종류는 확정할 수 없다. 인물과 읽을 수 있는 문자는 보이지 않는다. 낡은 트럭과 황량한 낮 풍경은 요청에 맞으며, 바위 표면과 틈 형태는 참조를 따르지만 주변 풍경에 자연스럽게 통합되지는 않는다.",
        "hard_violations": [],
        "physics": "왼쪽 앞뒤 타이어가 노면에 닿아 트럭을 지지한다. 적재함 안 물체도 바닥 위에 놓인 것으로 보인다. 흙먼지가 뒷바퀴 부근에서 솟아 뒤로 길게 퍼져 주행 동작을 설명한다. 먼지가 후면 하부를 일부 가리지만 차체 대부분은 드러난다. 지지 없이 공중에 떠 있는 물체나 불가능한 차체 자세는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.067
   },
   "adjusted": {
    "A": 1.75,
    "B": 0.817
   },
   "violations": {
    "A": [
     "[gemini-pro] 콜라주/합성 위반 (참고 이미지의 바위 지형을 2D 스티커처럼 배경에 오려 붙임)"
    ],
    "B": [
     "[gemini-pro] 금지된 텍스트 노출 (트럭 후면의 '낡은 트럭' 글자)",
     "[gemini-pro] 콜라주/합성 위반 (트럭 아래의 흙먼지가 뚜렷한 직사각형 경계를 가진 그래픽 오버레이로 덮여 있음)",
     "[gpt-high] 적재함 후면에 읽을 수 있는 한글 문구가 있어, 어디에도 읽을 수 있는 글자를 넣지 말라는 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 817
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "프롬프트의 차량과 도로 묘사는 따랐으나, 참고 이미지를 평면적인 스티커처럼 배경에 그대로 오려 붙여 실사 요건을 심각하게 위반했습니다.  ★위반: [gemini-pro] 콜라주/합성 위반 (참고 이미지의 바위 지형을 2D 스티커처럼 배경에 오려 붙임)"
   },
   {
    "label": "B",
    "score": 817,
    "verdict_ko": "프롬프트가 명시적으로 금지한 텍스트('낡은 트럭')가 포함되었으며, 흙먼지가 직사각형 그래픽으로 조잡하게 합성되어 기각 대상입니다.  ★위반: [gemini-pro] 금지된 텍스트 노출 (트럭 후면의 '낡은 트럭' 글자) / [gemini-pro] 콜라주/합성 위반 (트럭 아래의 흙먼지가 뚜렷한 직사각형 경계를 가진 그래픽 오버레이로 덮여 있음) / [gpt-high] 적재함 후면에 읽을 수 있는 한글 문구가 있어, 어디에도 읽을 수 있는 글자를 넣지 말라는 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_open_roof_truck_sel.png",
    "asset_id": "c94c53e3-76a2-4f78-b4ae-95204318d3c2",
    "role": "location_seed_bg"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-917a-76ac-a41d-3ca7bc8e46ff",
  "ref_mode": "seed-bg만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  },
  "lane_policy": "ab_select_bypass:bg_only"
 },
 "S63sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:51:42.117417+00:00",
  "fingerprint": "6f4e62cb2af24c1e47195b254cefa47e712465881adb3b1bc4a8464264a8e2a2",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S63sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S63sh1_sel.png",
  "source_sha256": "cb6a30f70a1a2b3aa42444ad36e33c3981d8badcbd4a2bf473bae330dec20701",
  "file": "S63sh1_cine.png",
  "staged_sha256": "3f791cb1c06bee978a66f76876c3dc5dcd43b1bb9c68b73d41ab0eeb80b9e1f6",
  "latency_ms": 13048
 },
 "S63sh16::signage": {
  "fp": "609d812f5f8a669e",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S63sh16": {
  "input_fingerprint": "8c6cb9ed5b9d8ad4",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): (회상) 어두운 방 안, 핏기 없는 창백한 얼굴로 눈을 반쯤 감은 채 누워있는 미연의 얼굴 클로즈업.\n\nLOCATION (lock): On a floating container panel over the inundated refugee settlement at dawn, at the mother's final resting place in the flashback. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Room interior (Only an indistinct portion surrounding the reclining figure is visible); used as Provides minimal spatial context without introducing furnishings absent from the scene.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The remembered room remains dark, with restrained facial exposure preserving 미연's pallor and half-closed eyes without adding a dream glow or a separate color treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Miyeon is stretched out on a floating container panel, her body supported by the panel and her abdomen bleeding from a puncture wound. The source does not specify whether she rests on her back or side, her head's direction, or the arrangement of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): In the recalled scene, a container panel floats on the open sea at dawn; this is not an indoor deathbed. 미연: She is stretched out on a floating container panel with a bleeding puncture wound in her abdomen and a face swollen from the earlier beating. She is barely conscious, immediately before her death.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): (회상) 어두운 방 안, 핏기 없는 창백한 얼굴로 눈을 반쯤 감은 채 누워있는 미연의 얼굴 클로즈업.\n\nLOCATION (lock): On a floating container panel over the inundated refugee settlement at dawn, at the mother's final resting place in the flashback. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Room interior (Only an indistinct portion surrounding the reclining figure is visible); used as Provides minimal spatial context without introducing furnishings absent from the scene.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The remembered room remains dark, with restrained facial exposure preserving 미연's pallor and half-closed eyes without adding a dream glow or a separate color treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Miyeon is stretched out on a floating container panel, her body supported by the panel and her abdomen bleeding from a puncture wound. The source does not specify whether she rests on her back or side, her head's direction, or the arrangement of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): In the recalled scene, a container panel floats on the open sea at dawn; this is not an indoor deathbed. 미연: She is stretched out on a floating container panel with a bleeding puncture wound in her abdomen and a face swollen from the earlier beating. She is barely conscious, immediately before her death.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): (회상) 어두운 방 안, 핏기 없는 창백한 얼굴로 눈을 반쯤 감은 채 누워있는 미연의 얼굴 클로즈업.\n\nLOCATION (lock): On a floating container panel over the inundated refugee settlement at dawn, at the mother's final resting place in the flashback. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Room interior (Only an indistinct portion surrounding the reclining figure is visible); used as Provides minimal spatial context without introducing furnishings absent from the scene.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The remembered room remains dark, with restrained facial exposure preserving 미연's pallor and half-closed eyes without adding a dream glow or a separate color treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 미연 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Miyeon is stretched out on a floating container panel, her body supported by the panel and her abdomen bleeding from a puncture wound. The source does not specify whether she rests on her back or side, her head's direction, or the arrangement of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): In the recalled scene, a container panel floats on the open sea at dawn; this is not an indoor deathbed. 미연: She is stretched out on a floating container panel with a bleeding puncture wound in her abdomen and a face swollen from the earlier beating. She is barely conscious, immediately before her death.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 미연 (한국인 여성, 40대 후반, 중년의 얼굴 윤곽, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "반쯤 감긴 눈으로 초점 없이 위쪽 허공을 응시함.",
    "built_space": "녹슨 철제 컨테이너 패널 위에 누워 있으며, 배경은 지시된 대로 어둡고 불분명하게 처리됨.",
    "entities": "미연(40대 후반 한국인 여성)이 이전 샷과 동일한 회색 티셔츠를 입고 창백한 안색을 띰.",
    "hard_violations": [],
    "physics": "머리와 몸통이 철제 패널 표면에 완전히 밀착되어 중력에 맞게 지탱됨."
   },
   {
    "label": "B",
    "direction": "반쯤 감긴 눈으로 초점 없이 비스듬히 허공을 바라봄.",
    "built_space": "골판형 철판 벽면이 보이는 컨테이너 내부에 위치하며, 바닥에 캔버스 자루가 깔려 있음.",
    "entities": "미연(40대 후반 한국인 여성)이 이전 샷과 다른 파란색 옷을 입고 입가에 피가 묻어 있음.",
    "hard_violations": [
     "[gpt-high] 참조에 없는 커튼, 목재 받침 구조, 머리 받침 천과 경첩 달린 문틀을 추가하여, 고정된 패널 주변 장소를 별도의 실내 공간으로 변경했습니다."
    ],
    "physics": "머리가 바닥의 캔버스 자루 위에 놓여 안정적으로 지탱됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "이전 샷의 의상(회색 티셔츠)과 녹슨 컨테이너 바닥을 정확히 유지했고, 어두운 배경 속 창백한 얼굴 클로즈업 지시를 충실히 따름."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "얼굴 묘사는 적절하나 이전 샷으로 고정(Lock)된 의상 지시를 어기고 파란 옷을 입혔으며, 바닥에 임의의 천을 추가함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "반쯤 감긴 눈으로 초점 없이 위쪽 허공을 응시함.",
        "built_space": "녹슨 철제 컨테이너 패널 위에 누워 있으며, 배경은 지시된 대로 어둡고 불분명하게 처리됨.",
        "entities": "미연(40대 후반 한국인 여성)이 이전 샷과 동일한 회색 티셔츠를 입고 창백한 안색을 띰.",
        "hard_violations": [],
        "physics": "머리와 몸통이 철제 패널 표면에 완전히 밀착되어 중력에 맞게 지탱됨."
       },
       {
        "label": "B",
        "direction": "반쯤 감긴 눈으로 초점 없이 비스듬히 허공을 바라봄.",
        "built_space": "골판형 철판 벽면이 보이는 컨테이너 내부에 위치하며, 바닥에 캔버스 자루가 깔려 있음.",
        "entities": "미연(40대 후반 한국인 여성)이 이전 샷과 다른 파란색 옷을 입고 입가에 피가 묻어 있음.",
        "hard_violations": [],
        "physics": "머리가 바닥의 캔버스 자루 위에 놓여 안정적으로 지탱됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "이전 샷의 의상(회색 티셔츠)과 녹슨 컨테이너 바닥을 정확히 유지했고, 어두운 배경 속 창백한 얼굴 클로즈업 지시를 충실히 따름."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "얼굴 묘사는 적절하나 이전 샷으로 고정(Lock)된 의상 지시를 어기고 파란 옷을 입혔으며, 바닥에 임의의 천을 추가함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "반쯤 감긴 눈으로 초점 없이 위쪽 허공을 응시함.",
        "built_space": "녹슨 철제 컨테이너 패널 위에 누워 있으며, 배경은 지시된 대로 어둡고 불분명하게 처리됨.",
        "entities": "미연(40대 후반 한국인 여성)이 이전 샷과 동일한 회색 티셔츠를 입고 창백한 안색을 띰.",
        "hard_violations": [],
        "physics": "머리와 몸통이 철제 패널 표면에 완전히 밀착되어 중력에 맞게 지탱됨."
       },
       {
        "label": "B",
        "direction": "반쯤 감긴 눈으로 초점 없이 비스듬히 허공을 바라봄.",
        "built_space": "골판형 철판 벽면이 보이는 컨테이너 내부에 위치하며, 바닥에 캔버스 자루가 깔려 있음.",
        "entities": "미연(40대 후반 한국인 여성)이 이전 샷과 다른 파란색 옷을 입고 입가에 피가 묻어 있음.",
        "hard_violations": [],
        "physics": "머리가 바닥의 캔버스 자루 위에 놓여 안정적으로 지탱됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "창백한 얼굴과 반쯤 감긴 눈의 클로즈업은 맞지만, 커튼·목재·문틀·받침 천을 추가하여 고정된 컨테이너 패널 주변 공간을 다른 장소로 바꿨습니다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "패널에 지지된 얼굴 클로즈업과 어두운 최소 배경, 창백함과 반쯤 감긴 눈을 잘 구현했으나, 회색 상의와 약해진 얼굴 상처는 참조와 다릅니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "머리는 화면 왼쪽, 몸통은 오른쪽으로 이어지며 얼굴이 카메라 쪽으로 돌아가 있습니다. 반쯤 열린 눈은 카메라 근처를 향하지만 특정 대상을 응시하지 않는 흐릿한 시선입니다. 지시된 무기나 방향성 소품은 없습니다.",
        "built_space": "머리 아래 녹슨 금속 면과 접힌 천 하나, 왼쪽의 늘어진 천 하나, 머리 뒤 목재 가로대 두 줄, 뒤쪽 골판 금속 벽, 오른쪽 전경의 문틀과 경첩 하나가 보입니다. 참조의 노출된 컨테이너 패널 대신 문이 있는 작은 실내 공간을 구체적으로 만들었으며, 배경을 불분명하게 최소화하라는 지시에도 어긋납니다.",
        "entities": "검은 머리의 중년 동아시아계 여성 한 명만 있으며 미연의 얼굴 윤곽과 대체로 부합합니다. 창백한 피부, 반쯤 감긴 정상적인 눈, 입가의 피가 보입니다. 목 주변의 해진 푸른 옷은 참조의 짙은 남색 티셔츠와 다릅니다. 복부 상처는 올바른 얼굴 중심 구도 밖이므로 평가할 수 없습니다. 추가 인물이나 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "참조에 없는 커튼, 목재 받침 구조, 머리 받침 천과 경첩 달린 문틀을 추가하여, 고정된 패널 주변 장소를 별도의 실내 공간으로 변경했습니다."
        ],
        "physics": "머리와 머리카락은 금속 면 위에 놓인 접힌 천에 눌려 지지됩니다. 목은 오른쪽의 몸통으로 자연스럽게 이어지며 공중에 들린 신체 부위는 보이지 않습니다. 추가된 받침의 연속성 문제와 별개로, 보이는 자세 자체는 중력에 부합합니다."
       },
       {
        "label": "B",
        "direction": "머리는 화면 왼쪽, 몸통은 오른쪽 아래로 이어집니다. 얼굴은 위쪽과 카메라 쪽을 향하고 반쯤 열린 눈은 카메라 부근으로 초점 없이 향합니다. 특정 응시 대상이 없는 쇠약한 상태로, 요청된 순간에 부합합니다.",
        "built_space": "인물 아래에 녹슨 컨테이너 패널 하나와 보강대가 보이고, 머리 뒤에는 구멍이 난 고정 철물 하나가 있습니다. 패널의 재질과 마모, 철물은 이전 장면과 잘 이어집니다. 나머지 배경은 어두워 방의 구조나 별도 가구를 확정할 수 없으며, 요청한 최소 공간 정보에 가깝습니다. 바다와 수평선은 클로즈업 밖입니다.",
        "entities": "검은 머리의 중년 동아시아계 여성 한 명으로, 미연의 연령대와 기본 외형에 대체로 부합합니다. 창백한 얼굴과 반쯤 감긴 눈은 명확하고 눈의 해부학적 변형은 없습니다. 다만 참조보다 얼굴의 부기와 입가 출혈이 약하며, 보이는 상의는 참조의 남색보다 회색으로 읽힙니다. 복부와 그 상처는 구도 밖입니다. 다른 사람이나 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "뒤통수와 늘어진 머리카락이 패널에 닿고, 어깨와 몸통도 같은 면 위에 누워 있습니다. 목과 머리를 능동적으로 들어 올린 모습이나 지지 없이 떠 있는 부위는 보이지 않습니다. 패널 전체와 수면 접촉은 화면 밖이지만, 보이는 신체의 지지는 자연스럽습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "창백한 얼굴과 반쯤 감긴 눈의 클로즈업은 맞지만, 커튼·목재·문틀·받침 천을 추가하여 고정된 컨테이너 패널 주변 공간을 다른 장소로 바꿨습니다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "패널에 지지된 얼굴 클로즈업과 어두운 최소 배경, 창백함과 반쯤 감긴 눈을 잘 구현했으나, 회색 상의와 약해진 얼굴 상처는 참조와 다릅니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "머리는 화면 왼쪽, 몸통은 오른쪽으로 이어지며 얼굴이 카메라 쪽으로 돌아가 있습니다. 반쯤 열린 눈은 카메라 근처를 향하지만 특정 대상을 응시하지 않는 흐릿한 시선입니다. 지시된 무기나 방향성 소품은 없습니다.",
        "built_space": "머리 아래 녹슨 금속 면과 접힌 천 하나, 왼쪽의 늘어진 천 하나, 머리 뒤 목재 가로대 두 줄, 뒤쪽 골판 금속 벽, 오른쪽 전경의 문틀과 경첩 하나가 보입니다. 참조의 노출된 컨테이너 패널 대신 문이 있는 작은 실내 공간을 구체적으로 만들었으며, 배경을 불분명하게 최소화하라는 지시에도 어긋납니다.",
        "entities": "검은 머리의 중년 동아시아계 여성 한 명만 있으며 미연의 얼굴 윤곽과 대체로 부합합니다. 창백한 피부, 반쯤 감긴 정상적인 눈, 입가의 피가 보입니다. 목 주변의 해진 푸른 옷은 참조의 짙은 남색 티셔츠와 다릅니다. 복부 상처는 올바른 얼굴 중심 구도 밖이므로 평가할 수 없습니다. 추가 인물이나 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "참조에 없는 커튼, 목재 받침 구조, 머리 받침 천과 경첩 달린 문틀을 추가하여, 고정된 패널 주변 장소를 별도의 실내 공간으로 변경했습니다."
        ],
        "physics": "머리와 머리카락은 금속 면 위에 놓인 접힌 천에 눌려 지지됩니다. 목은 오른쪽의 몸통으로 자연스럽게 이어지며 공중에 들린 신체 부위는 보이지 않습니다. 추가된 받침의 연속성 문제와 별개로, 보이는 자세 자체는 중력에 부합합니다."
       },
       {
        "label": "A",
        "direction": "머리는 화면 왼쪽, 몸통은 오른쪽 아래로 이어집니다. 얼굴은 위쪽과 카메라 쪽을 향하고 반쯤 열린 눈은 카메라 부근으로 초점 없이 향합니다. 특정 응시 대상이 없는 쇠약한 상태로, 요청된 순간에 부합합니다.",
        "built_space": "인물 아래에 녹슨 컨테이너 패널 하나와 보강대가 보이고, 머리 뒤에는 구멍이 난 고정 철물 하나가 있습니다. 패널의 재질과 마모, 철물은 이전 장면과 잘 이어집니다. 나머지 배경은 어두워 방의 구조나 별도 가구를 확정할 수 없으며, 요청한 최소 공간 정보에 가깝습니다. 바다와 수평선은 클로즈업 밖입니다.",
        "entities": "검은 머리의 중년 동아시아계 여성 한 명으로, 미연의 연령대와 기본 외형에 대체로 부합합니다. 창백한 얼굴과 반쯤 감긴 눈은 명확하고 눈의 해부학적 변형은 없습니다. 다만 참조보다 얼굴의 부기와 입가 출혈이 약하며, 보이는 상의는 참조의 남색보다 회색으로 읽힙니다. 복부와 그 상처는 구도 밖입니다. 다른 사람이나 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "뒤통수와 늘어진 머리카락이 패널에 닿고, 어깨와 몸통도 같은 면 위에 누워 있습니다. 목과 머리를 능동적으로 들어 올린 모습이나 지지 없이 떠 있는 부위는 보이지 않습니다. 패널 전체와 수면 접촉은 화면 밖이지만, 보이는 신체의 지지는 자연스럽습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.946
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.696
   },
   "violations": {
    "B": [
     "[gpt-high] 참조에 없는 커튼, 목재 받침 구조, 머리 받침 천과 경첩 달린 문틀을 추가하여, 고정된 패널 주변 장소를 별도의 실내 공간으로 변경했습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 696
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "이전 샷의 의상(회색 티셔츠)과 녹슨 컨테이너 바닥을 정확히 유지했고, 어두운 배경 속 창백한 얼굴 클로즈업 지시를 충실히 따름."
   },
   {
    "label": "B",
    "score": 696,
    "verdict_ko": "얼굴 묘사는 적절하나 이전 샷으로 고정(Lock)된 의상 지시를 어기고 파란 옷을 입혔으며, 바닥에 임의의 천을 추가함.  ★위반: [gpt-high] 참조에 없는 커튼, 목재 받침 구조, 머리 받침 천과 경첩 달린 문틀을 추가하여, 고정된 패널 주변 장소를 별도의 실내 공간으로 변경했습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 미연 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S38sh7_sel.png",
    "asset_id": "f4087254-1c5a-4a73-b72d-40eea4ee6c67",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 미연: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1113064>",
    "asset_id": "f3646b1e-66ff-47c9-87d4-cdad1d05d96b",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-9330-7c99-b85a-0f1fc5d24bd6",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S38sh7"
  }
 },
 "S63sh16::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:52:44.915963+00:00",
  "fingerprint": "f49036aa443a4789539a1dcb6e0d4959ad76c38e28147331ca43326cb217ec11",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S63sh16_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S63sh16_sel.png",
  "source_sha256": "1629cc4cae54a84831adf4dbc394d27fa1c667254fbf315f19ffac1f13f41f13",
  "file": "S63sh16_cine.png",
  "staged_sha256": "b1094d0b6effbcc6484304e052fe316c2434922b42ced5e61ac930d42bb36710",
  "latency_ms": 11600
 },
 "S63sh17::confined_fp_apt": {
  "applies": true,
  "reason_ko": "이 샷은 트럭 운전석이라는 밀폐된 차량 내부 공간에서 진행됩니다. 현우가 운전석에 앉아 있고 빛이 조수석 창문에서 들어오는 구체적인 방향성이 제시되어 있으므로 인물의 착석 위치와 창문의 상대적인 위치를 정확하게 배치하지 않으면 화면의 연속성과 개연성이 깨지는 심각한 오류가 발생할 수 있습니다.",
  "input_fingerprint": "93c5004a51d6e787"
 },
 "S63sh17::signage": {
  "fp": "fb25ace90e92090b",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "confinedfp::529f9dcd3f66": {
  "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/confinedfp_base_529f9dcd3f66.png",
  "place_text": "At the wheel inside the old truck's compact cab, with warm daylight entering through the passenger-side window.",
  "input_fingerprint": "ceeaf8d945ef443b"
 },
 "S63sh17::confined_fp": {
  "reads": {
   "controls": "A steering wheel is located at the driver's station on the left.",
   "mirrors": "No mirrors are depicted in the diagram.",
   "camera": "Positioned near the center windshield area on the passenger side, angled leftward and pointing directly at the driver.",
   "occupants": "Hyun-woo occupies the driver seat."
  },
  "mismatches": [],
  "scene_description_en": "The camera, positioned on the passenger side near the windshield, points diagonally left toward the driver's seat in a close-up framing. Hyun-woo occupies the driver's seat on the left side of the screen, facing forward toward the left edge of the frame, presenting his right profile to the lens. The steering wheel sits directly in front of him on the left. Sunlight illuminates the right side of his face, cast from the passenger-side window located off-screen to the right. The blurred interior of the driver's side cabin wall serves as the background behind him on the far left.",
  "fixed": false,
  "input_fingerprint": "ee553b8a16535dd4"
 },
 "era_assess::0d4b3c45090d2aca": {
  "subjects": [],
  "subject_text": "현우와 수빈이 사용하는 낡은 트럭 운전석\n운전석과 조수석이 나란히 놓인 낡은 운전 공간. 앞유리 아래로 운전대와 계기판이 있고 양옆 창문으로 외부가 보인다.",
  "identity": "canonical",
  "scope_id": "L232",
  "scope_role": "location_exterior",
  "scope_sha": "9aec4edb9cd1cee5"
 },
 "S63sh17": {
  "input_fingerprint": "9bfa1e985fbcf82a",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 조수석 창문으로 들어온 따스한 햇빛을 받으며 입술을 꽉 다문 현우의 결의에 찬 얼굴 클로즈업.\n\nLOCATION (lock): At the wheel inside the old truck's compact cab, with warm daylight entering through the passenger-side window. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Truck cab interior (현우 remains seated in the driver's position) — A narrow interior portion behind the driver is seen from the passenger side; used as Maintains the restored present-tense location without bringing 수빈 into this close-up.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Warm sunlight entering through the passenger-side window reaches 현우's face, with controlled highlights preserving the tension around his closed lips.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old village truck continues along the road in daylight. 현우: He remains in the driver's seat, with his fighting injuries still present and a resolute expression.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera, positioned on the passenger side near the windshield, points diagonally left toward the driver's seat in a close-up framing. Hyun-woo occupies the driver's seat on the left side of the screen, facing forward toward the left edge of the frame, presenting his right profile to the lens. The steering wheel sits directly in front of him on the left. Sunlight illuminates the right side of his face, cast from the passenger-side window located off-screen to the right. The blurred interior of the driver's side cabin wall serves as the background behind him on the far left.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 조수석 창문으로 들어온 따스한 햇빛을 받으며 입술을 꽉 다문 현우의 결의에 찬 얼굴 클로즈업.\n\nLOCATION (lock): At the wheel inside the old truck's compact cab, with warm daylight entering through the passenger-side window. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Warm sunlight entering through the passenger-side window reaches 현우's face, with controlled highlights preserving the tension around his closed lips.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old village truck continues along the road in daylight. 현우: He remains in the driver's seat, with his fighting injuries still present and a resolute expression.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera, positioned on the passenger side near the windshield, points diagonally left toward the driver's seat in a close-up framing. Hyun-woo occupies the driver's seat on the left side of the screen, facing forward toward the left edge of the frame, presenting his right profile to the lens. The steering wheel sits directly in front of him on the left. Sunlight illuminates the right side of his face, cast from the passenger-side window located off-screen to the right. The blurred interior of the driver's side cabin wall serves as the background behind him on the far left.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 조수석 창문으로 들어온 따스한 햇빛을 받으며 입술을 꽉 다문 현우의 결의에 찬 얼굴 클로즈업.\n\nLOCATION (lock): At the wheel inside the old truck's compact cab, with warm daylight entering through the passenger-side window. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Warm sunlight entering through the passenger-side window reaches 현우's face, with controlled highlights preserving the tension around his closed lips.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old village truck continues along the road in daylight. 현우: He remains in the driver's seat, with his fighting injuries still present and a resolute expression.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S63sh17_confinedfp.png",
     "asset_id": null,
     "role": null
    },
    {
     "label": "현우",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S63sh17_confinedfp.png",
     "asset_id": null,
     "role": null
    },
    {
     "label": "현우",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "프레임 우측 앞 유리창 쪽을 주시하고 있다.",
    "built_space": "카메라가 조수석에 위치해 운전자의 우측면을 촬영하며, 우측에 운전대가 있어 평면도와 일치한다.",
    "entities": "현우의 얼굴과 꽉 다문 입술 표정은 일치하나, 요구된 얼굴의 싸움 상처는 생략되었다.",
    "hard_violations": [],
    "physics": "좌석에 기대어 앉은 자세가 중력에 맞게 정상적으로 지지된다."
   },
   {
    "label": "B",
    "direction": "프레임 좌측 앞 유리창 쪽을 주시하고 있다.",
    "built_space": "운전자의 좌측면이 보이고 좌측에 운전대가 있어, 조수석 시점이라는 지시와 평면도의 공간 구조에 위배된다.",
    "entities": "현우의 외모, 표정, 그리고 뺨의 상처가 지시사항과 잘 일치한다.",
    "hard_violations": [
     "[gemini-pro] 명시된 조수석 카메라 시점 및 좌핸들 차량 공간 구조를 반대 방향으로 왜곡하여 렌더링함.",
     "[gpt-high] 운전대·전방·운전자 사이의 좌우 관계가 평면도의 좌측 운전석 및 조수석 카메라 배치와 반대로 구현되어, 지정된 촬영 위치와 양립하지 않습니다."
    ],
    "physics": "좌석에 앉은 자세가 정상적으로 지지된다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "평면도의 카메라 시점과 조명 방향을 정확히 구현했으나, 요구된 상처 묘사가 누락되었습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "상처와 표정 묘사는 우수하나, 카메라 위치와 트럭 내부 구조를 반대로 렌더링한 치명적 공간 오류가 있습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "프레임 우측 앞 유리창 쪽을 주시하고 있다.",
        "built_space": "카메라가 조수석에 위치해 운전자의 우측면을 촬영하며, 우측에 운전대가 있어 평면도와 일치한다.",
        "entities": "현우의 얼굴과 꽉 다문 입술 표정은 일치하나, 요구된 얼굴의 싸움 상처는 생략되었다.",
        "hard_violations": [],
        "physics": "좌석에 기대어 앉은 자세가 중력에 맞게 정상적으로 지지된다."
       },
       {
        "label": "B",
        "direction": "프레임 좌측 앞 유리창 쪽을 주시하고 있다.",
        "built_space": "운전자의 좌측면이 보이고 좌측에 운전대가 있어, 조수석 시점이라는 지시와 평면도의 공간 구조에 위배된다.",
        "entities": "현우의 외모, 표정, 그리고 뺨의 상처가 지시사항과 잘 일치한다.",
        "hard_violations": [
         "명시된 조수석 카메라 시점 및 좌핸들 차량 공간 구조를 반대 방향으로 왜곡하여 렌더링함."
        ],
        "physics": "좌석에 앉은 자세가 정상적으로 지지된다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "평면도의 카메라 시점과 조명 방향을 정확히 구현했으나, 요구된 상처 묘사가 누락되었습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "상처와 표정 묘사는 우수하나, 카메라 위치와 트럭 내부 구조를 반대로 렌더링한 치명적 공간 오류가 있습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "프레임 우측 앞 유리창 쪽을 주시하고 있다.",
        "built_space": "카메라가 조수석에 위치해 운전자의 우측면을 촬영하며, 우측에 운전대가 있어 평면도와 일치한다.",
        "entities": "현우의 얼굴과 꽉 다문 입술 표정은 일치하나, 요구된 얼굴의 싸움 상처는 생략되었다.",
        "hard_violations": [],
        "physics": "좌석에 기대어 앉은 자세가 중력에 맞게 정상적으로 지지된다."
       },
       {
        "label": "B",
        "direction": "프레임 좌측 앞 유리창 쪽을 주시하고 있다.",
        "built_space": "운전자의 좌측면이 보이고 좌측에 운전대가 있어, 조수석 시점이라는 지시와 평면도의 공간 구조에 위배된다.",
        "entities": "현우의 외모, 표정, 그리고 뺨의 상처가 지시사항과 잘 일치한다.",
        "hard_violations": [
         "명시된 조수석 카메라 시점 및 좌핸들 차량 공간 구조를 반대 방향으로 왜곡하여 렌더링함."
        ],
        "physics": "좌석에 앉은 자세가 정상적으로 지지된다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "얼굴 클로즈업과 싸움의 상처는 잘 드러나지만, 운전대와 전방의 좌우 배치가 지정된 조수석 촬영 시점 및 좌측 운전석 배치와 맞지 않습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "조수석에서 운전자를 바라보는 공간 관계와 전방 주시는 맞지만, 얼굴보다 상체와 실내를 넓게 담았고 입술을 꽉 다문 결의는 약합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 얼굴과 눈은 화면 왼쪽, 운전대 너머 차량 전방을 향합니다. 카메라를 응시하지는 않습니다. 따뜻한 빛이 볼과 귀, 목에 닿지만 조수석 창에서 들어오는 방향인지는 이 배치로 확인하기 어렵습니다.",
        "built_space": "화면 왼쪽 아래에 운전대 하나와 계기판 일부, 뒤쪽에 큰 측면 창 하나와 왼쪽 가장자리의 다른 유리 일부, 그 사이 기둥에 안전띠 고정부 하나가 보입니다. 얼굴은 오른쪽에 크게 배치됩니다. 전방과 운전대가 왼쪽으로 놓인 이 구도는 제시된 좌측 운전석을 조수석에서 보는 배치와 반대이며, 운전석 바깥쪽에서 본 구도 또는 좌우가 뒤집힌 운전실처럼 읽힙니다.",
        "entities": "인물은 한 명이며, 앳된 동아시아계 남성의 얼굴과 헝클어진 검은 머리, 남색 둥근목 티셔츠가 참조와 대체로 일치합니다. 국적은 외형으로 확인할 수 없습니다. 관자놀이와 볼의 긁힌 상처가 명확합니다. 입술은 닫혀 있지만 꽉 눌러 다문 긴장감은 약합니다. 낡은 트럭의 금속과 창틀은 실물 재질로 보이며 다른 사람이나 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "운전대·전방·운전자 사이의 좌우 관계가 평면도의 좌측 운전석 및 조수석 카메라 배치와 반대로 구현되어, 지정된 촬영 위치와 양립하지 않습니다."
        ],
        "physics": "머리와 목은 어깨 위에 자연스럽게 연결되어 있으며 부유하거나 비정상적으로 꺾인 부분은 없습니다. 좌석과 골반, 손은 클로즈업 밖이므로 착석 접촉이나 운전대 파지는 확인할 수 없습니다. 운전대는 조향축에 연결된 것으로 보이고, 안전띠는 기둥의 고정부에 매달려 있습니다."
       },
       {
        "label": "B",
        "direction": "현우는 화면 오른쪽의 운전대와 앞유리 너머 전방을 바라봅니다. 시선이 렌즈로 향하지 않아 운전 중인 상황에 맞습니다. 카메라 쪽에서 얼굴에 닿는 따뜻한 낮빛은 조수석 쪽 채광과 양립합니다.",
        "built_space": "운전대 하나가 현우의 앞쪽인 화면 오른쪽에 있고, 그 너머 계기판 일부가 보입니다. 뒤에는 낡은 등받이 하나와 인접 좌석 일부가 있으며, 측면 창 하나와 작은 보조 유리 구획, 오른쪽 앞유리 일부, 왼쪽 뒤창 일부가 보입니다. 안전띠 고정부 하나와 선바이저 하나도 있습니다. 조수석에서 좌측 운전자를 보는 관계는 평면도에 맞지만, 어깨와 가슴 및 실내가 많이 포함되어 요구된 얼굴 클로즈업보다 넓습니다.",
        "entities": "한 명의 앳된 동아시아계 남성이 보이며 검은 헝클어진 머리, 얼굴 윤곽, 남색 티셔츠가 참조와 대체로 맞습니다. 한국계 미국인이라는 국적·배경 자체는 영상만으로 판별할 수 없습니다. 볼에는 붉은 자국과 작은 상처가 있으나 싸움의 부상이라는 인상은 A보다 약합니다. 입술은 거의 닫혀 있지만 힘주어 다문 표정은 아닙니다. 수빈이나 다른 사람은 없고, 창과 선바이저의 작은 표지는 글자를 읽을 수 없는 수준입니다.",
        "hard_violations": [],
        "physics": "상체는 좌석 등받이 앞에 자연스럽게 세워져 있어 운전석에 앉은 자세와 양립합니다. 골반과 손은 화면 밖이므로 좌판 접촉과 운전대 파지는 직접 확인할 수 없습니다. 운전대는 조향축에 지지되고 좌석·창·선바이저도 차체에 정상적으로 고정되어 있으며, 지지 없이 떠 있는 물체나 불가능한 자세는 보이지 않습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "얼굴 클로즈업과 싸움의 상처는 잘 드러나지만, 운전대와 전방의 좌우 배치가 지정된 조수석 촬영 시점 및 좌측 운전석 배치와 맞지 않습니다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "조수석에서 운전자를 바라보는 공간 관계와 전방 주시는 맞지만, 얼굴보다 상체와 실내를 넓게 담았고 입술을 꽉 다문 결의는 약합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 얼굴과 눈은 화면 왼쪽, 운전대 너머 차량 전방을 향합니다. 카메라를 응시하지는 않습니다. 따뜻한 빛이 볼과 귀, 목에 닿지만 조수석 창에서 들어오는 방향인지는 이 배치로 확인하기 어렵습니다.",
        "built_space": "화면 왼쪽 아래에 운전대 하나와 계기판 일부, 뒤쪽에 큰 측면 창 하나와 왼쪽 가장자리의 다른 유리 일부, 그 사이 기둥에 안전띠 고정부 하나가 보입니다. 얼굴은 오른쪽에 크게 배치됩니다. 전방과 운전대가 왼쪽으로 놓인 이 구도는 제시된 좌측 운전석을 조수석에서 보는 배치와 반대이며, 운전석 바깥쪽에서 본 구도 또는 좌우가 뒤집힌 운전실처럼 읽힙니다.",
        "entities": "인물은 한 명이며, 앳된 동아시아계 남성의 얼굴과 헝클어진 검은 머리, 남색 둥근목 티셔츠가 참조와 대체로 일치합니다. 국적은 외형으로 확인할 수 없습니다. 관자놀이와 볼의 긁힌 상처가 명확합니다. 입술은 닫혀 있지만 꽉 눌러 다문 긴장감은 약합니다. 낡은 트럭의 금속과 창틀은 실물 재질로 보이며 다른 사람이나 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "운전대·전방·운전자 사이의 좌우 관계가 평면도의 좌측 운전석 및 조수석 카메라 배치와 반대로 구현되어, 지정된 촬영 위치와 양립하지 않습니다."
        ],
        "physics": "머리와 목은 어깨 위에 자연스럽게 연결되어 있으며 부유하거나 비정상적으로 꺾인 부분은 없습니다. 좌석과 골반, 손은 클로즈업 밖이므로 착석 접촉이나 운전대 파지는 확인할 수 없습니다. 운전대는 조향축에 연결된 것으로 보이고, 안전띠는 기둥의 고정부에 매달려 있습니다."
       },
       {
        "label": "A",
        "direction": "현우는 화면 오른쪽의 운전대와 앞유리 너머 전방을 바라봅니다. 시선이 렌즈로 향하지 않아 운전 중인 상황에 맞습니다. 카메라 쪽에서 얼굴에 닿는 따뜻한 낮빛은 조수석 쪽 채광과 양립합니다.",
        "built_space": "운전대 하나가 현우의 앞쪽인 화면 오른쪽에 있고, 그 너머 계기판 일부가 보입니다. 뒤에는 낡은 등받이 하나와 인접 좌석 일부가 있으며, 측면 창 하나와 작은 보조 유리 구획, 오른쪽 앞유리 일부, 왼쪽 뒤창 일부가 보입니다. 안전띠 고정부 하나와 선바이저 하나도 있습니다. 조수석에서 좌측 운전자를 보는 관계는 평면도에 맞지만, 어깨와 가슴 및 실내가 많이 포함되어 요구된 얼굴 클로즈업보다 넓습니다.",
        "entities": "한 명의 앳된 동아시아계 남성이 보이며 검은 헝클어진 머리, 얼굴 윤곽, 남색 티셔츠가 참조와 대체로 맞습니다. 한국계 미국인이라는 국적·배경 자체는 영상만으로 판별할 수 없습니다. 볼에는 붉은 자국과 작은 상처가 있으나 싸움의 부상이라는 인상은 A보다 약합니다. 입술은 거의 닫혀 있지만 힘주어 다문 표정은 아닙니다. 수빈이나 다른 사람은 없고, 창과 선바이저의 작은 표지는 글자를 읽을 수 없는 수준입니다.",
        "hard_violations": [],
        "physics": "상체는 좌석 등받이 앞에 자연스럽게 세워져 있어 운전석에 앉은 자세와 양립합니다. 골반과 손은 화면 밖이므로 좌판 접촉과 운전대 파지는 직접 확인할 수 없습니다. 운전대는 조향축에 지지되고 좌석·창·선바이저도 차체에 정상적으로 고정되어 있으며, 지지 없이 떠 있는 물체나 불가능한 자세는 보이지 않습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.857
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.607
   },
   "violations": {
    "B": [
     "[gemini-pro] 명시된 조수석 카메라 시점 및 좌핸들 차량 공간 구조를 반대 방향으로 왜곡하여 렌더링함.",
     "[gpt-high] 운전대·전방·운전자 사이의 좌우 관계가 평면도의 좌측 운전석 및 조수석 카메라 배치와 반대로 구현되어, 지정된 촬영 위치와 양립하지 않습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 607
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "평면도의 카메라 시점과 조명 방향을 정확히 구현했으나, 요구된 상처 묘사가 누락되었습니다."
   },
   {
    "label": "B",
    "score": 607,
    "verdict_ko": "상처와 표정 묘사는 우수하나, 카메라 위치와 트럭 내부 구조를 반대로 렌더링한 치명적 공간 오류가 있습니다.  ★위반: [gemini-pro] 명시된 조수석 카메라 시점 및 좌핸들 차량 공간 구조를 반대 방향으로 왜곡하여 렌더링함. / [gpt-high] 운전대·전방·운전자 사이의 좌우 관계가 평면도의 좌측 운전석 및 조수석 카메라 배치와 반대로 구현되어, 지정된 촬영 위치와 양립하지 않습니다."
   }
  ],
  "refs": [
   {
    "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S63sh17_confinedfp.png",
    "asset_id": null,
    "role": null
   },
   {
    "label": "현우",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-94d9-789d-838a-82b72f122c9c",
  "confined_fp": {
   "base_key": "confinedfp::529f9dcd3f66",
   "apt_reason": "이 샷은 트럭 운전석이라는 밀폐된 차량 내부 공간에서 진행됩니다. 현우가 운전석에 앉아 있고 빛이 조수석 창문에서 들어오는 구체적인 방향성이 제시되어 있으므로 인물의 착석 위치와 창문의 상대적인 위치를 정확하게 배치하지 않으면 화면의 연속성과 개연성이 깨지는 심각한 오류가 발생할 수 있습니다.",
   "fixed": false,
   "mismatches": []
  },
  "ref_mode": "confined_fp: 도면+장면설명+엔티티",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S63sh17::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T07:55:21.071948+00:00",
  "fingerprint": "e7079f6c6f38b08fb0d730600f9e6efce779bba291379f49a011d21898f07578",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S63sh17_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S63sh17_sel.png",
  "source_sha256": "a21365ef61bb2f9c688996a50e7ad1f39a687c7604927e88ddcd11342c8eb9d8",
  "file": "S63sh17_cine.png",
  "staged_sha256": "b9bd95eb26ecc55429cb6bcf9aa4c463024293f5239aae2e7de41bebb6320bb1",
  "latency_ms": 8621
 },
 "S64sh7::signage": {
  "fp": "516d630c928dd2fb",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S64sh7": {
  "input_fingerprint": "a51caf4d67bf796e",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리를 향해 두꺼운 팔뚝의 고사포 총구를 매섭게 정조준한 B-200의 위협적인 자세.\n\nLOCATION (lock): In a shadowed corner inside the village's large machine-parts warehouse, with daylight entering through its half-open door. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: B-200's forearm gun (Raised and aimed at 찰리) — The barrel is viewed obliquely from the side, with its muzzle directed toward the left foreground figure, not the camera; used as Creates a readable threat line between the characters while remaining proportionate to B-200's body; Pile of broken robots (Heaped in a corner of the warehouse) — Irregular portions of the piled bodies remain visible beyond the confrontation; used as Provides subdued warehouse context and a background against which B-200's living movement registers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the warehouse corner's established darkness while retaining enough restrained tonal separation to read the gun, crouched body, and intervening space.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The machinery warehouse has a half-open door and a heap of broken robots; baby birds are already concealed behind a box in B-200's corner. B-200 crouches in the darkness with converted gun-hands raised, while Charlie has brought in a football and retains his damaged body, chest ring and impaired systems.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by B-200 right now, so B-200's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to B-200: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리를 향해 두꺼운 팔뚝의 고사포 총구를 매섭게 정조준한 B-200의 위협적인 자세.\n\nLOCATION (lock): In a shadowed corner inside the village's large machine-parts warehouse, with daylight entering through its half-open door. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: B-200's forearm gun (Raised and aimed at 찰리) — The barrel is viewed obliquely from the side, with its muzzle directed toward the left foreground figure, not the camera; used as Creates a readable threat line between the characters while remaining proportionate to B-200's body; Pile of broken robots (Heaped in a corner of the warehouse) — Irregular portions of the piled bodies remain visible beyond the confrontation; used as Provides subdued warehouse context and a background against which B-200's living movement registers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the warehouse corner's established darkness while retaining enough restrained tonal separation to read the gun, crouched body, and intervening space.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The machinery warehouse has a half-open door and a heap of broken robots; baby birds are already concealed behind a box in B-200's corner. B-200 crouches in the darkness with converted gun-hands raised, while Charlie has brought in a football and retains his damaged body, chest ring and impaired systems.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by B-200 right now, so B-200's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to B-200: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리를 향해 두꺼운 팔뚝의 고사포 총구를 매섭게 정조준한 B-200의 위협적인 자세.\n\nLOCATION (lock): In a shadowed corner inside the village's large machine-parts warehouse, with daylight entering through its half-open door. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: B-200's forearm gun (Raised and aimed at 찰리) — The barrel is viewed obliquely from the side, with its muzzle directed toward the left foreground figure, not the camera; used as Creates a readable threat line between the characters while remaining proportionate to B-200's body; Pile of broken robots (Heaped in a corner of the warehouse) — Irregular portions of the piled bodies remain visible beyond the confrontation; used as Provides subdued warehouse context and a background against which B-200's living movement registers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the warehouse corner's established darkness while retaining enough restrained tonal separation to read the gun, crouched body, and intervening space.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The machinery warehouse has a half-open door and a heap of broken robots; baby birds are already concealed behind a box in B-200's corner. B-200 crouches in the darkness with converted gun-hands raised, while Charlie has brought in a football and retains his damaged body, chest ring and impaired systems.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by B-200 right now, so B-200's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to B-200: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S64sh7__bgfirst_bg.png",
     "asset_id": "2c542339-d519-40a3-aee9-40b352e75792",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S64sh7.png",
     "asset_id": "cbafc64b-e6e6-4acf-8c80-6fb23ad1a5d9",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — B-200: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1339855>",
     "asset_id": "8091d94b-e8e7-407e-97d6-c030f55a73f9",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L238B01.png",
     "asset_id": "e1a9f48b-d9e5-4c90-99a8-fadcfb66b2d8",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — B-200: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1339855>",
     "asset_id": "8091d94b-e8e7-407e-97d6-c030f55a73f9",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "B-200이 오른쪽 팔뚝의 총구를 왼쪽 전경의 찰리를 향해 정조준하고 있음.",
    "built_space": "참조 이미지와 동일한 창고 내부. 왼쪽에 반쯤 열린 문과 선반, 상자들이 정확한 위치에 배치됨.",
    "entities": "B-200은 참조와 일치하는 무기 형태를 지니고 웅크림. 찰리는 등에 빛나는 링과 기계 팔을 가진 채 축구공을 들고 있으며, 우측 상자 뒤에 아기 새들이 있음. 전체적으로 2D 일러스트 화풍임.",
    "hard_violations": [],
    "physics": "B-200과 찰리 모두 바닥에 안정적으로 서 있으며, 찰리의 오른손이 축구공을 정상적으로 쥐고 있음."
   },
   {
    "label": "B",
    "direction": "B-200이 오른쪽의 거대한 총구를 왼쪽 전경의 찰리 머리를 향해 겨누고 있음.",
    "built_space": "참조 이미지의 창고 내부를 실사로 구현함. 왼쪽 문 앞에 부서진 로봇 잔해가 쌓여 있음.",
    "entities": "B-200의 우측 총구가 단일형 거대 포신으로 변형되어 참조와 다르고 서 있는 자세임. 찰리는 형체를 알아보기 힘든 초점 나간 실루엣으로 뭉개져 묘사됨. 축구공과 아기 새는 보이지 않음.",
    "hard_violations": [
     "[gpt-high] 어깨 장갑에 '01'로 판독 가능한 숫자가 노출되어 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
    ],
    "physics": "B-200은 바닥을 딛고 서 있으며, 주변의 로봇 잔해들도 바닥에 정상적으로 놓여 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "사진 질감(포토리얼리즘)을 무시한 2D 화풍이나, B-200의 웅크린 자세, 찰리의 손상된 신체와 축구공, 아기 새 등 요구된 디테일을 정확히 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "실사풍이나 찰리를 초점이 나간 실루엣으로 뭉개어 표현한 금지 사항을 어겼으며, B-200의 우측 총구가 크게 변형되었고 소품이 누락됨."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "B-200이 오른쪽 팔뚝의 총구를 왼쪽 전경의 찰리를 향해 정조준하고 있음.",
        "built_space": "참조 이미지와 동일한 창고 내부. 왼쪽에 반쯤 열린 문과 선반, 상자들이 정확한 위치에 배치됨.",
        "entities": "B-200은 참조와 일치하는 무기 형태를 지니고 웅크림. 찰리는 등에 빛나는 링과 기계 팔을 가진 채 축구공을 들고 있으며, 우측 상자 뒤에 아기 새들이 있음. 전체적으로 2D 일러스트 화풍임.",
        "hard_violations": [],
        "physics": "B-200과 찰리 모두 바닥에 안정적으로 서 있으며, 찰리의 오른손이 축구공을 정상적으로 쥐고 있음."
       },
       {
        "label": "B",
        "direction": "B-200이 오른쪽의 거대한 총구를 왼쪽 전경의 찰리 머리를 향해 겨누고 있음.",
        "built_space": "참조 이미지의 창고 내부를 실사로 구현함. 왼쪽 문 앞에 부서진 로봇 잔해가 쌓여 있음.",
        "entities": "B-200의 우측 총구가 단일형 거대 포신으로 변형되어 참조와 다르고 서 있는 자세임. 찰리는 형체를 알아보기 힘든 초점 나간 실루엣으로 뭉개져 묘사됨. 축구공과 아기 새는 보이지 않음.",
        "hard_violations": [],
        "physics": "B-200은 바닥을 딛고 서 있으며, 주변의 로봇 잔해들도 바닥에 정상적으로 놓여 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "사진 질감(포토리얼리즘)을 무시한 2D 화풍이나, B-200의 웅크린 자세, 찰리의 손상된 신체와 축구공, 아기 새 등 요구된 디테일을 정확히 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "실사풍이나 찰리를 초점이 나간 실루엣으로 뭉개어 표현한 금지 사항을 어겼으며, B-200의 우측 총구가 크게 변형되었고 소품이 누락됨."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "B-200이 오른쪽 팔뚝의 총구를 왼쪽 전경의 찰리를 향해 정조준하고 있음.",
        "built_space": "참조 이미지와 동일한 창고 내부. 왼쪽에 반쯤 열린 문과 선반, 상자들이 정확한 위치에 배치됨.",
        "entities": "B-200은 참조와 일치하는 무기 형태를 지니고 웅크림. 찰리는 등에 빛나는 링과 기계 팔을 가진 채 축구공을 들고 있으며, 우측 상자 뒤에 아기 새들이 있음. 전체적으로 2D 일러스트 화풍임.",
        "hard_violations": [],
        "physics": "B-200과 찰리 모두 바닥에 안정적으로 서 있으며, 찰리의 오른손이 축구공을 정상적으로 쥐고 있음."
       },
       {
        "label": "B",
        "direction": "B-200이 오른쪽의 거대한 총구를 왼쪽 전경의 찰리 머리를 향해 겨누고 있음.",
        "built_space": "참조 이미지의 창고 내부를 실사로 구현함. 왼쪽 문 앞에 부서진 로봇 잔해가 쌓여 있음.",
        "entities": "B-200의 우측 총구가 단일형 거대 포신으로 변형되어 참조와 다르고 서 있는 자세임. 찰리는 형체를 알아보기 힘든 초점 나간 실루엣으로 뭉개져 묘사됨. 축구공과 아기 새는 보이지 않음.",
        "hard_violations": [],
        "physics": "B-200은 바닥을 딛고 서 있으며, 주변의 로봇 잔해들도 바닥에 정상적으로 놓여 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "찰리를 향한 조준선은 성립하지만 어깨의 판독 가능한 숫자가 금지 조건을 위반하며, 과대해진 단일 포신과 거의 전신을 담은 구도가 지정된 미디엄 숏에서 벗어난다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "비례가 유지된 팔 포신으로 찰리를 겨누는 웅크린 자세는 더 충실하지만, 전신 위주의 넓은 구도와 등에 붙은 가슴 링, 노출된 아기 새는 지시와 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "B-200의 머리는 왼쪽 전경의 찰리를 향한다. 크게 돌출된 포신은 비스듬한 측면으로 보이며 왼쪽 아래로 향해, 연장하면 찰리의 머리 부근에 닿는다. 카메라를 정면 조준하지는 않는다. 화면 오른쪽의 다른 다연장 포신은 전방 아래쪽을 향하며 찰리를 직접 겨누지는 않는다. 찰리는 등을 보이고 B-200 쪽을 바라본다.",
        "built_space": "왼쪽에 열린 대형 출입문 하나와 높은 창 하나, 천장에 매달린 등 세 개, 뒤쪽의 연속된 부품 선반과 오른쪽 선반이 보인다. 콘크리트 바닥과 낡은 철골·벽체는 장소 참조와 유사하다. B-200은 오른쪽 선반 앞에 있고, 부서진 로봇 더미는 왼쪽 출입문 가까운 바닥과 중앙의 콘크리트 구조물 앞에 놓여 있어 어두운 구석의 배경 더미라는 배치는 약하다. B-200의 발까지 대부분 보여 미디엄 숏보다 넓다.",
        "entities": "B-200은 짙은 회색 장갑판, 붉은 센서, 굵은 관절을 가진 중장비형 로봇으로 참조의 정체성이 대체로 유지된다. 다만 조준하는 팔은 참조의 다연장 구성 대신 유난히 굵고 긴 단일 포신으로 바뀌었다. 왼쪽 전경에는 짧은 검은 머리의 어린 남성형 찰리의 뒷머리와 상체 일부가 보이며, 얼굴과 손상 상태는 흐림과 가림 때문에 확인하기 어렵다. 축구공과 가슴 링은 해당 크롭에서 보이지 않아 누락으로 단정할 수 없다. 고장 난 로봇 더미는 보이고 아기 새는 노출되지 않는다. 화면 오른쪽 어깨 장갑에는 '01'로 읽히는 흰 숫자가 있다.",
        "hard_violations": [
         "어깨 장갑에 '01'로 판독 가능한 숫자가 노출되어 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
        ],
        "physics": "B-200은 벌린 다리와 바닥에 닿은 발로 체중을 받는다. 무릎은 굽었지만 몸통이 높아 깊이 웅크린 자세보다는 낮춘 기립 자세에 가깝다. 포는 기계 팔과 관절에 연결되어 지지되며 떠 있지 않다. 로봇 잔해는 바닥과 서로의 몸체에 기대어 쌓여 있다. 찰리의 하체는 프레임 밖이라 발의 접지는 확인할 수 없지만 부유를 시사하는 모습은 없다."
       },
       {
        "label": "B",
        "direction": "B-200의 머리와 높이 든 포신은 왼쪽 전경의 찰리 쪽으로 향한다. 높은 포신의 축을 왼쪽으로 연장하면 찰리의 상체 부근에 닿으며, 총열 측면이 보여 카메라 정면 조준과 구별된다. 낮은 쪽 포신도 왼쪽 전경 방향을 향한다. 찰리는 몸과 고개를 B-200 쪽으로 돌리고 있다.",
        "built_space": "왼쪽에 열린 대형 출입문 하나와 높은 창 하나, 뒤쪽 부품 선반, 오른쪽 벽의 창 하나와 가장자리 선반이 보인다. 천장에는 매달린 등 두 개가 보이며 나머지 영역은 크롭과 인물에 가려진다. 콘크리트 바닥과 벽, 철골 천장은 장소 참조에 부합한다. B-200은 오른쪽 상자들 옆에서 웅크리고 있고, 로봇 잔해는 그 뒤 오른쪽 구석에 쌓여 있다. 다만 B-200의 양발과 넓은 바닥까지 담아 지정된 미디엄 숏보다 넓으며, 구석의 어둠도 다소 약하다.",
        "entities": "B-200의 짙은 회색 장갑, 굵은 관절, 붉은 센서와 양팔의 다연장 포신은 참조와 대체로 일치한다. 찰리는 검은 단발의 어린 남성형 인물이며 팔의 손상과 기계 부품이 드러나지만, 얼굴이 돌아가 있어 구체적인 외모는 판별하기 어렵다. 손에 축구공을 들고 있으나 지시된 가슴 링은 가슴이 아니라 등 위쪽에 붙어 있다. 오른쪽에는 고장 난 로봇 더미와 나무 상자들이 있으며, 상자 틈의 아기 새 세 마리가 보여 이미 숨겨져 있어야 한다는 상태와 어긋난다. 판독 가능한 문자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "B-200은 양발을 바닥에 넓게 딛고 무릎과 고관절을 굽혀 몸통을 지지한다. 양팔 포신은 팔 관절에 연결되어 있고, 올린 팔의 조준 자세도 기계 구조상 성립한다. 찰리의 손은 축구공 표면을 감싸 잡고 있어 공의 지지점이 보인다. 상자와 잔해는 바닥 또는 아래 상자에 놓여 있으며, 아기 새도 상자 안쪽 턱에 올라 있다. 근거 없이 떠 있는 몸이나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "찰리를 향한 조준선은 성립하지만 어깨의 판독 가능한 숫자가 금지 조건을 위반하며, 과대해진 단일 포신과 거의 전신을 담은 구도가 지정된 미디엄 숏에서 벗어난다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "비례가 유지된 팔 포신으로 찰리를 겨누는 웅크린 자세는 더 충실하지만, 전신 위주의 넓은 구도와 등에 붙은 가슴 링, 노출된 아기 새는 지시와 다르다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "B-200의 머리는 왼쪽 전경의 찰리를 향한다. 크게 돌출된 포신은 비스듬한 측면으로 보이며 왼쪽 아래로 향해, 연장하면 찰리의 머리 부근에 닿는다. 카메라를 정면 조준하지는 않는다. 화면 오른쪽의 다른 다연장 포신은 전방 아래쪽을 향하며 찰리를 직접 겨누지는 않는다. 찰리는 등을 보이고 B-200 쪽을 바라본다.",
        "built_space": "왼쪽에 열린 대형 출입문 하나와 높은 창 하나, 천장에 매달린 등 세 개, 뒤쪽의 연속된 부품 선반과 오른쪽 선반이 보인다. 콘크리트 바닥과 낡은 철골·벽체는 장소 참조와 유사하다. B-200은 오른쪽 선반 앞에 있고, 부서진 로봇 더미는 왼쪽 출입문 가까운 바닥과 중앙의 콘크리트 구조물 앞에 놓여 있어 어두운 구석의 배경 더미라는 배치는 약하다. B-200의 발까지 대부분 보여 미디엄 숏보다 넓다.",
        "entities": "B-200은 짙은 회색 장갑판, 붉은 센서, 굵은 관절을 가진 중장비형 로봇으로 참조의 정체성이 대체로 유지된다. 다만 조준하는 팔은 참조의 다연장 구성 대신 유난히 굵고 긴 단일 포신으로 바뀌었다. 왼쪽 전경에는 짧은 검은 머리의 어린 남성형 찰리의 뒷머리와 상체 일부가 보이며, 얼굴과 손상 상태는 흐림과 가림 때문에 확인하기 어렵다. 축구공과 가슴 링은 해당 크롭에서 보이지 않아 누락으로 단정할 수 없다. 고장 난 로봇 더미는 보이고 아기 새는 노출되지 않는다. 화면 오른쪽 어깨 장갑에는 '01'로 읽히는 흰 숫자가 있다.",
        "hard_violations": [
         "어깨 장갑에 '01'로 판독 가능한 숫자가 노출되어 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
        ],
        "physics": "B-200은 벌린 다리와 바닥에 닿은 발로 체중을 받는다. 무릎은 굽었지만 몸통이 높아 깊이 웅크린 자세보다는 낮춘 기립 자세에 가깝다. 포는 기계 팔과 관절에 연결되어 지지되며 떠 있지 않다. 로봇 잔해는 바닥과 서로의 몸체에 기대어 쌓여 있다. 찰리의 하체는 프레임 밖이라 발의 접지는 확인할 수 없지만 부유를 시사하는 모습은 없다."
       },
       {
        "label": "A",
        "direction": "B-200의 머리와 높이 든 포신은 왼쪽 전경의 찰리 쪽으로 향한다. 높은 포신의 축을 왼쪽으로 연장하면 찰리의 상체 부근에 닿으며, 총열 측면이 보여 카메라 정면 조준과 구별된다. 낮은 쪽 포신도 왼쪽 전경 방향을 향한다. 찰리는 몸과 고개를 B-200 쪽으로 돌리고 있다.",
        "built_space": "왼쪽에 열린 대형 출입문 하나와 높은 창 하나, 뒤쪽 부품 선반, 오른쪽 벽의 창 하나와 가장자리 선반이 보인다. 천장에는 매달린 등 두 개가 보이며 나머지 영역은 크롭과 인물에 가려진다. 콘크리트 바닥과 벽, 철골 천장은 장소 참조에 부합한다. B-200은 오른쪽 상자들 옆에서 웅크리고 있고, 로봇 잔해는 그 뒤 오른쪽 구석에 쌓여 있다. 다만 B-200의 양발과 넓은 바닥까지 담아 지정된 미디엄 숏보다 넓으며, 구석의 어둠도 다소 약하다.",
        "entities": "B-200의 짙은 회색 장갑, 굵은 관절, 붉은 센서와 양팔의 다연장 포신은 참조와 대체로 일치한다. 찰리는 검은 단발의 어린 남성형 인물이며 팔의 손상과 기계 부품이 드러나지만, 얼굴이 돌아가 있어 구체적인 외모는 판별하기 어렵다. 손에 축구공을 들고 있으나 지시된 가슴 링은 가슴이 아니라 등 위쪽에 붙어 있다. 오른쪽에는 고장 난 로봇 더미와 나무 상자들이 있으며, 상자 틈의 아기 새 세 마리가 보여 이미 숨겨져 있어야 한다는 상태와 어긋난다. 판독 가능한 문자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "B-200은 양발을 바닥에 넓게 딛고 무릎과 고관절을 굽혀 몸통을 지지한다. 양팔 포신은 팔 관절에 연결되어 있고, 올린 팔의 조준 자세도 기계 구조상 성립한다. 찰리의 손은 축구공 표면을 감싸 잡고 있어 공의 지지점이 보인다. 상자와 잔해는 바닥 또는 아래 상자에 놓여 있으며, 아기 새도 상자 안쪽 턱에 올라 있다. 근거 없이 떠 있는 몸이나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.929
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.679
   },
   "violations": {
    "B": [
     "[gpt-high] 어깨 장갑에 '01'로 판독 가능한 숫자가 노출되어 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 679
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "사진 질감(포토리얼리즘)을 무시한 2D 화풍이나, B-200의 웅크린 자세, 찰리의 손상된 신체와 축구공, 아기 새 등 요구된 디테일을 정확히 구현함."
   },
   {
    "label": "B",
    "score": 679,
    "verdict_ko": "실사풍이나 찰리를 초점이 나간 실루엣으로 뭉개어 표현한 금지 사항을 어겼으며, B-200의 우측 총구가 크게 변형되었고 소품이 누락됨.  ★위반: [gpt-high] 어깨 장갑에 '01'로 판독 가능한 숫자가 노출되어 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L238B01.png",
    "asset_id": "e1a9f48b-d9e5-4c90-99a8-fadcfb66b2d8",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — B-200: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1339855>",
    "asset_id": "8091d94b-e8e7-407e-97d6-c030f55a73f9",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-981f-7ada-b958-6f6ad639aadf",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S64sh7__bgfirst_bg.png",
   "bg_asset_id": "2c542339-d519-40a3-aee9-40b352e75792",
   "bg_record_key": "S64sh7::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S64sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:50:07.125955+00:00",
  "fingerprint": "7c4b39fa8f03cd0c9624ae3f203556a3ff2f729aff5b09aa473aa0ea9f268375",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S64sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S64sh7_sel.png",
  "source_sha256": "b498933a0ab6a1c3a4d820a6c70e32319163b0a3d30eeba006b16e681ae6d357",
  "file": "S64sh7_cine.png",
  "staged_sha256": "6fed6a58c966f5d5ed0594a1a2242bfb453e7398997beca2d4658dec9845173c",
  "latency_ms": 10595
 },
 "S64sh15::signage": {
  "fp": "1b6306e467d60a7b",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S64sh15": {
  "input_fingerprint": "ce43bac8f6d92eb5",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 금속 가슴 장갑에 달린 둥근 링을 향해 두꺼운 손가락을 뻗은 B-200의 손 클로즈업.\n\nLOCATION (lock): In the same dim corner of the large robot-parts warehouse, near discarded machines and the half-open daylight doorway. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Ring attached to 찰리's chest (Attached to the chest armor and being indicated by B-200) — Its outward-facing shape is readable at a slight oblique angle on 찰리's chest; used as Provides the small shared focal point, kept in focus with the fingertip and visibly attached to the surrounding torso; 찰리's metal chest armor (Visible around the attached ring) — Seen from the open side of the seated pair, with the chest surface receding obliquely; used as Supplies bodily context and scale so the ring does not become an isolated enlarged object.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the warehouse's subdued ambient exposure, allowing restrained surface highlights to separate the pointing hand from 찰리's metal chest armor.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of B-200 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The warehouse door remains half open, the broken robots remain piled in a corner, and the baby birds are still concealed behind the box. B-200 retains his converted gun-hands, while Charlie remains seated nearby with a damaged chest ring, dents, holes and an impaired language system.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 금속 가슴 장갑에 달린 둥근 링을 향해 두꺼운 손가락을 뻗은 B-200의 손 클로즈업.\n\nLOCATION (lock): In the same dim corner of the large robot-parts warehouse, near discarded machines and the half-open daylight doorway. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Ring attached to 찰리's chest (Attached to the chest armor and being indicated by B-200) — Its outward-facing shape is readable at a slight oblique angle on 찰리's chest; used as Provides the small shared focal point, kept in focus with the fingertip and visibly attached to the surrounding torso; 찰리's metal chest armor (Visible around the attached ring) — Seen from the open side of the seated pair, with the chest surface receding obliquely; used as Supplies bodily context and scale so the ring does not become an isolated enlarged object.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the warehouse's subdued ambient exposure, allowing restrained surface highlights to separate the pointing hand from 찰리's metal chest armor.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of B-200 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The warehouse door remains half open, the broken robots remain piled in a corner, and the baby birds are still concealed behind the box. B-200 retains his converted gun-hands, while Charlie remains seated nearby with a damaged chest ring, dents, holes and an impaired language system.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 금속 가슴 장갑에 달린 둥근 링을 향해 두꺼운 손가락을 뻗은 B-200의 손 클로즈업.\n\nLOCATION (lock): In the same dim corner of the large robot-parts warehouse, near discarded machines and the half-open daylight doorway. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Ring attached to 찰리's chest (Attached to the chest armor and being indicated by B-200) — Its outward-facing shape is readable at a slight oblique angle on 찰리's chest; used as Provides the small shared focal point, kept in focus with the fingertip and visibly attached to the surrounding torso; 찰리's metal chest armor (Visible around the attached ring) — Seen from the open side of the seated pair, with the chest surface receding obliquely; used as Supplies bodily context and scale so the ring does not become an isolated enlarged object.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the warehouse's subdued ambient exposure, allowing restrained surface highlights to separate the pointing hand from 찰리's metal chest armor.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of B-200 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The warehouse door remains half open, the broken robots remain piled in a corner, and the baby birds are still concealed behind the box. B-200 retains his converted gun-hands, while Charlie remains seated nearby with a damaged chest ring, dents, holes and an impaired language system.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "짙은 회색 기계 손의 검지가 베이지색 로봇(찰리)의 가슴 중앙에 있는 둥근 원형 구조물을 정확히 가리키고 있습니다.",
    "built_space": "창고 내부. 왼쪽에 열린 문이 있고, 오른쪽에는 나무 상자와 파란색 로봇 부품들이 있으며 상자 위에는 아기 새들이 놓여 있어 이전 샷의 배치와 일치합니다.",
    "entities": "화면으로 들어오는 손은 짙은 회색 장갑판을 지녀 B-200의 외형과 일치하며, 가리킴을 받는 로봇은 샌드 베이지색 몸체와 흰색 마스크를 지닌 찰리입니다. 가슴의 링은 뚫린 고리보다는 내장된 원형 패널에 가깝게 표현되었습니다.",
    "hard_violations": [],
    "physics": "화면 왼쪽 밖에서 뻗어 나온 팔이 기계 손의 무게를 지탱하고 있습니다."
   },
   {
    "label": "B",
    "direction": "베이지색 기계 손의 검지가 베이지색 로봇(찰리)의 가슴에 부착된 둥근 금속 링을 가리키고 있습니다.",
    "built_space": "창고 내부. 왼쪽에 문이 있으나, 이전 샷에서 오른쪽에 있던 파란색 로봇 부품들이 왼쪽 문의 바로 옆으로 이동하여 공간의 일관성이 깨졌습니다.",
    "entities": "찰리의 가슴에 부착된 금속 링의 형태는 정확하나, 지시하는 손이 짙은 회색이 아닌 베이지색으로 렌더링되어 B-200의 정체성과 불일치합니다.",
    "hard_violations": [],
    "physics": "화면 왼쪽 밖에서 뻗어 나온 팔이 기계 손을 안정적으로 지탱하고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "B-200의 짙은 회색 손과 찰리의 베이지색 몸체를 정확히 구분했으며, 이전 샷의 배경 요소(오른쪽의 상자와 새 등)를 훌륭하게 유지하여 지시문을 잘 충족했습니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "가슴의 링 형태는 명확하게 표현되었으나, B-200의 손을 찰리와 같은 베이지색으로 잘못 렌더링하고 배경 부품의 위치를 임의로 변경하여 아쉽습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "짙은 회색 기계 손의 검지가 베이지색 로봇(찰리)의 가슴 중앙에 있는 둥근 원형 구조물을 정확히 가리키고 있습니다.",
        "built_space": "창고 내부. 왼쪽에 열린 문이 있고, 오른쪽에는 나무 상자와 파란색 로봇 부품들이 있으며 상자 위에는 아기 새들이 놓여 있어 이전 샷의 배치와 일치합니다.",
        "entities": "화면으로 들어오는 손은 짙은 회색 장갑판을 지녀 B-200의 외형과 일치하며, 가리킴을 받는 로봇은 샌드 베이지색 몸체와 흰색 마스크를 지닌 찰리입니다. 가슴의 링은 뚫린 고리보다는 내장된 원형 패널에 가깝게 표현되었습니다.",
        "hard_violations": [],
        "physics": "화면 왼쪽 밖에서 뻗어 나온 팔이 기계 손의 무게를 지탱하고 있습니다."
       },
       {
        "label": "B",
        "direction": "베이지색 기계 손의 검지가 베이지색 로봇(찰리)의 가슴에 부착된 둥근 금속 링을 가리키고 있습니다.",
        "built_space": "창고 내부. 왼쪽에 문이 있으나, 이전 샷에서 오른쪽에 있던 파란색 로봇 부품들이 왼쪽 문의 바로 옆으로 이동하여 공간의 일관성이 깨졌습니다.",
        "entities": "찰리의 가슴에 부착된 금속 링의 형태는 정확하나, 지시하는 손이 짙은 회색이 아닌 베이지색으로 렌더링되어 B-200의 정체성과 불일치합니다.",
        "hard_violations": [],
        "physics": "화면 왼쪽 밖에서 뻗어 나온 팔이 기계 손을 안정적으로 지탱하고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "B-200의 짙은 회색 손과 찰리의 베이지색 몸체를 정확히 구분했으며, 이전 샷의 배경 요소(오른쪽의 상자와 새 등)를 훌륭하게 유지하여 지시문을 잘 충족했습니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "가슴의 링 형태는 명확하게 표현되었으나, B-200의 손을 찰리와 같은 베이지색으로 잘못 렌더링하고 배경 부품의 위치를 임의로 변경하여 아쉽습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "짙은 회색 기계 손의 검지가 베이지색 로봇(찰리)의 가슴 중앙에 있는 둥근 원형 구조물을 정확히 가리키고 있습니다.",
        "built_space": "창고 내부. 왼쪽에 열린 문이 있고, 오른쪽에는 나무 상자와 파란색 로봇 부품들이 있으며 상자 위에는 아기 새들이 놓여 있어 이전 샷의 배치와 일치합니다.",
        "entities": "화면으로 들어오는 손은 짙은 회색 장갑판을 지녀 B-200의 외형과 일치하며, 가리킴을 받는 로봇은 샌드 베이지색 몸체와 흰색 마스크를 지닌 찰리입니다. 가슴의 링은 뚫린 고리보다는 내장된 원형 패널에 가깝게 표현되었습니다.",
        "hard_violations": [],
        "physics": "화면 왼쪽 밖에서 뻗어 나온 팔이 기계 손의 무게를 지탱하고 있습니다."
       },
       {
        "label": "B",
        "direction": "베이지색 기계 손의 검지가 베이지색 로봇(찰리)의 가슴에 부착된 둥근 금속 링을 가리키고 있습니다.",
        "built_space": "창고 내부. 왼쪽에 문이 있으나, 이전 샷에서 오른쪽에 있던 파란색 로봇 부품들이 왼쪽 문의 바로 옆으로 이동하여 공간의 일관성이 깨졌습니다.",
        "entities": "찰리의 가슴에 부착된 금속 링의 형태는 정확하나, 지시하는 손이 짙은 회색이 아닌 베이지색으로 렌더링되어 B-200의 정체성과 불일치합니다.",
        "hard_violations": [],
        "physics": "화면 왼쪽 밖에서 뻗어 나온 팔이 기계 손을 안정적으로 지탱하고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "손끝이 손상된 가슴 링을 정확히 가리키는 클로즈업이지만, 지시하는 손과 팔이 찰리와 같은 베이지색 장갑으로 표현되어 B-200의 고정된 외형과 어긋납니다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "짙은 회색 B-200의 손이 찰리의 가슴 링을 가리키는 관계와 창고 배치를 더 충실히 유지하지만, 링의 손상과 개조된 포신 손은 불명확하고 숨겨져야 할 새들이 보입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽에서 들어온 기계 손의 굵은 검지가 오른쪽 찰리의 가슴 링 안쪽을 향하며, 손끝이 링의 왼쪽 안쪽 가장자리에 거의 닿습니다. 지시 대상은 분명히 가슴에 붙은 링입니다. 눈과 포구는 보이지 않습니다.",
        "built_space": "왼쪽에 낮빛이 들어오는 출입구 하나와 그 오른쪽 위 작은 창 하나, 천장 보와 매달린 조명 하나가 보입니다. 뒤쪽에는 선반과 폐기 로봇 더미가 있습니다. 참고의 낡은 창고 재질은 유지하지만 폐기물 더미는 참고의 오른쪽보다 왼쪽 배경으로 옮겨져 있습니다. 가슴 장갑은 비스듬히 보이며 링 주변 몸통이 충분히 포함됩니다. 좌석과 바닥 접점은 화면 밖입니다.",
        "entities": "베이지색 각진 가슴 장갑과 검은 통풍구는 찰리의 참고 외형에 부합합니다. 가슴에는 금이 가고 찌그러진 돌출형 금속 링 하나가 붙어 있습니다. 그러나 지시하는 손과 전완도 베이지색이며 찰리의 참고 손과 닮아, 짙은 회색 중장비형 B-200의 손으로 식별하기 어렵습니다. 보이는 손목과 손에는 개조 포신이 확인되지 않습니다. 얼굴, 새, 읽을 수 있는 글자는 보이지 않습니다.",
        "hard_violations": [],
        "physics": "검지는 관절을 통해 손바닥과 왼쪽 화면 밖으로 이어지는 전완에 연결되어 있어 지시 동작이 가능합니다. 링은 가슴 장갑에 부착되어 있으며 떠 있는 물체로 보이지 않습니다. 몸통은 아래쪽 골반으로 이어지고, 앉은 자세의 실제 지지점은 클로즈업 밖이라 확인할 수 없습니다."
       },
       {
        "label": "B",
        "direction": "왼쪽에서 뻗은 짙은 회색 기계 검지가 찰리 가슴의 원형 부품 중앙을 정확히 가리키고 거의 접촉합니다. 링 테두리보다 중앙 덮개를 직접 짚는 모습입니다. 찰리의 마스크 아래쪽만 보여 시선은 판별할 수 없고 포구도 보이지 않습니다.",
        "built_space": "왼쪽에 밝은 출입구 하나와 상부 창 하나, 천장 보가 보이고 오른쪽에는 선반 한 구역, 폐기 로봇 더미와 큰 나무 상자가 있습니다. 폐기물과 상자의 오른쪽 배치는 이전 장면과 더 가깝습니다. 찰리의 가슴은 사선으로 물러나며 손과 링이 함께 선명하게 보입니다. 하단의 굽힌 다리는 착석과 양립하지만 좌석 자체는 보이지 않습니다.",
        "entities": "지시하는 손과 전완은 짙은 회색 철제여서 B-200의 색상과 중량감에 부합합니다. 다만 보이는 부분은 관절 손가락이며 참고의 대형 포신 손 구조는 확인되지 않습니다. 찰리의 베이지색 장갑, 흰 마스크 하단, 흉부 통풍구와 긁힌 표면은 참고와 일치합니다. 가슴 부품은 중앙 덮개를 둘러싼 동심원형 링으로 표현되어 A보다 손상된 고리의 형태가 덜 분명합니다. 오른쪽 상자 옆에는 작은 새 여러 마리가 노출되어 은폐 유지 조건과 다릅니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "뻗은 검지는 손바닥과 전완에 관절로 연결되어 있고 가슴 부품을 짚는 자세가 기계적으로 가능합니다. 원형 부품은 장갑에 고정되어 있습니다. 찰리의 몸통은 골반과 굽힌 다리로 이어지며 공중에 떠 있다는 징후는 없습니다. 새들은 상자 옆 턱에 걸쳐 보이며, 지지 없이 떠 있는 물체는 확인되지 않습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "손끝이 손상된 가슴 링을 정확히 가리키는 클로즈업이지만, 지시하는 손과 팔이 찰리와 같은 베이지색 장갑으로 표현되어 B-200의 고정된 외형과 어긋납니다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "짙은 회색 B-200의 손이 찰리의 가슴 링을 가리키는 관계와 창고 배치를 더 충실히 유지하지만, 링의 손상과 개조된 포신 손은 불명확하고 숨겨져야 할 새들이 보입니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽에서 들어온 기계 손의 굵은 검지가 오른쪽 찰리의 가슴 링 안쪽을 향하며, 손끝이 링의 왼쪽 안쪽 가장자리에 거의 닿습니다. 지시 대상은 분명히 가슴에 붙은 링입니다. 눈과 포구는 보이지 않습니다.",
        "built_space": "왼쪽에 낮빛이 들어오는 출입구 하나와 그 오른쪽 위 작은 창 하나, 천장 보와 매달린 조명 하나가 보입니다. 뒤쪽에는 선반과 폐기 로봇 더미가 있습니다. 참고의 낡은 창고 재질은 유지하지만 폐기물 더미는 참고의 오른쪽보다 왼쪽 배경으로 옮겨져 있습니다. 가슴 장갑은 비스듬히 보이며 링 주변 몸통이 충분히 포함됩니다. 좌석과 바닥 접점은 화면 밖입니다.",
        "entities": "베이지색 각진 가슴 장갑과 검은 통풍구는 찰리의 참고 외형에 부합합니다. 가슴에는 금이 가고 찌그러진 돌출형 금속 링 하나가 붙어 있습니다. 그러나 지시하는 손과 전완도 베이지색이며 찰리의 참고 손과 닮아, 짙은 회색 중장비형 B-200의 손으로 식별하기 어렵습니다. 보이는 손목과 손에는 개조 포신이 확인되지 않습니다. 얼굴, 새, 읽을 수 있는 글자는 보이지 않습니다.",
        "hard_violations": [],
        "physics": "검지는 관절을 통해 손바닥과 왼쪽 화면 밖으로 이어지는 전완에 연결되어 있어 지시 동작이 가능합니다. 링은 가슴 장갑에 부착되어 있으며 떠 있는 물체로 보이지 않습니다. 몸통은 아래쪽 골반으로 이어지고, 앉은 자세의 실제 지지점은 클로즈업 밖이라 확인할 수 없습니다."
       },
       {
        "label": "A",
        "direction": "왼쪽에서 뻗은 짙은 회색 기계 검지가 찰리 가슴의 원형 부품 중앙을 정확히 가리키고 거의 접촉합니다. 링 테두리보다 중앙 덮개를 직접 짚는 모습입니다. 찰리의 마스크 아래쪽만 보여 시선은 판별할 수 없고 포구도 보이지 않습니다.",
        "built_space": "왼쪽에 밝은 출입구 하나와 상부 창 하나, 천장 보가 보이고 오른쪽에는 선반 한 구역, 폐기 로봇 더미와 큰 나무 상자가 있습니다. 폐기물과 상자의 오른쪽 배치는 이전 장면과 더 가깝습니다. 찰리의 가슴은 사선으로 물러나며 손과 링이 함께 선명하게 보입니다. 하단의 굽힌 다리는 착석과 양립하지만 좌석 자체는 보이지 않습니다.",
        "entities": "지시하는 손과 전완은 짙은 회색 철제여서 B-200의 색상과 중량감에 부합합니다. 다만 보이는 부분은 관절 손가락이며 참고의 대형 포신 손 구조는 확인되지 않습니다. 찰리의 베이지색 장갑, 흰 마스크 하단, 흉부 통풍구와 긁힌 표면은 참고와 일치합니다. 가슴 부품은 중앙 덮개를 둘러싼 동심원형 링으로 표현되어 A보다 손상된 고리의 형태가 덜 분명합니다. 오른쪽 상자 옆에는 작은 새 여러 마리가 노출되어 은폐 유지 조건과 다릅니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "뻗은 검지는 손바닥과 전완에 관절로 연결되어 있고 가슴 부품을 짚는 자세가 기계적으로 가능합니다. 원형 부품은 장갑에 고정되어 있습니다. 찰리의 몸통은 골반과 굽힌 다리로 이어지며 공중에 떠 있다는 징후는 없습니다. 새들은 상자 옆 턱에 걸쳐 보이며, 지지 없이 떠 있는 물체는 확인되지 않습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.375
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.375
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1375
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "B-200의 짙은 회색 손과 찰리의 베이지색 몸체를 정확히 구분했으며, 이전 샷의 배경 요소(오른쪽의 상자와 새 등)를 훌륭하게 유지하여 지시문을 잘 충족했습니다."
   },
   {
    "label": "B",
    "score": 1375,
    "verdict_ko": "가슴의 링 형태는 명확하게 표현되었으나, B-200의 손을 찰리와 같은 베이지색으로 잘못 렌더링하고 배경 부품의 위치를 임의로 변경하여 아쉽습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of B-200 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S64sh7_sel.png",
    "asset_id": "ac0fffb7-d7d6-4ce3-a286-139c152f3a01",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — B-200: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1339855>",
    "asset_id": "8091d94b-e8e7-407e-97d6-c030f55a73f9",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-9b77-72f7-811e-2c30b6674ae3",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S64sh7"
  }
 },
 "S64sh15::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:51:21.292005+00:00",
  "fingerprint": "49b4118ce579d83ed3c75ce2af3b6de9fbbfd67a5c91ca507de22bb6c6080247",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S64sh15_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S64sh15_sel.png",
  "source_sha256": "2c974930116e9a95238845e6344477abcfb9c2993d490edf7d11f0fde4d9df03",
  "file": "S64sh15_cine.png",
  "staged_sha256": "2b7f7924bd6c120454964d513a0a413ae6fa95109ba095284534cb2b0ebe9c71",
  "latency_ms": 9274
 },
 "S64sh24::signage": {
  "fp": "5f34007e464cce6a",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S64sh24": {
  "input_fingerprint": "fa4086ca9a1e7e99",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 아기 새(B-200이 돌보던 새)들을 내려다보며 안광이 부드럽게 반달 모양으로 휘어진 B-200의 따뜻한 금속 얼굴.\n\nLOCATION (lock): Beside a concealed box of nestlings in the warehouse's shadowed corner, with daylight beyond the half-open entrance. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Warehouse interior (Remains behind B-200 after 찰리 has left); used as A minimally resolved background keeps the face and the birds' small silhouettes readable without introducing new objects.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the warehouse illumination subdued and continuous, letting the softened crescent eye shapes convey warmth without adding an unsupported warm light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of B-200 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The shifted box exposes the baby birds in B-200's warehouse corner; the broken-robot heap and half-open door remain unchanged. B-200 retains his converted gun-hands, and Charlie has left the corner with his battle damage still unrepaired.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 아기 새(B-200이 돌보던 새)들을 내려다보며 안광이 부드럽게 반달 모양으로 휘어진 B-200의 따뜻한 금속 얼굴.\n\nLOCATION (lock): Beside a concealed box of nestlings in the warehouse's shadowed corner, with daylight beyond the half-open entrance. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Warehouse interior (Remains behind B-200 after 찰리 has left); used as A minimally resolved background keeps the face and the birds' small silhouettes readable without introducing new objects.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the warehouse illumination subdued and continuous, letting the softened crescent eye shapes convey warmth without adding an unsupported warm light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of B-200 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The shifted box exposes the baby birds in B-200's warehouse corner; the broken-robot heap and half-open door remain unchanged. B-200 retains his converted gun-hands, and Charlie has left the corner with his battle damage still unrepaired.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 아기 새(B-200이 돌보던 새)들을 내려다보며 안광이 부드럽게 반달 모양으로 휘어진 B-200의 따뜻한 금속 얼굴.\n\nLOCATION (lock): Beside a concealed box of nestlings in the warehouse's shadowed corner, with daylight beyond the half-open entrance. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Warehouse interior (Remains behind B-200 after 찰리 has left); used as A minimally resolved background keeps the face and the birds' small silhouettes readable without introducing new objects.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the warehouse illumination subdued and continuous, letting the softened crescent eye shapes convey warmth without adding an unsupported warm light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of B-200 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The shifted box exposes the baby birds in B-200's warehouse corner; the broken-robot heap and half-open door remain unchanged. B-200 retains his converted gun-hands, and Charlie has left the corner with his battle damage still unrepaired.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "로봇의 고개와 시선이 화면 왼쪽 하단에 놓인 새끼 새가 담긴 상자를 정확하게 향하고 있음.",
    "built_space": "창고 내부. 화면 왼쪽에 반쯤 열린 문과 야외의 빛이 보이며, 배경에는 흐릿하게 선반과 집기들이 배치되어 공간감이 자연스러움.",
    "entities": "B-200의 머리 형태, 회색 철제 장갑판, 양팔의 고사포 포신 등이 레퍼런스와 정확히 일치하며, 눈 부위에 흰색 반달 모양의 안광이 구현됨. 상자 안의 새끼 새들도 명확함.",
    "hard_violations": [],
    "physics": "로봇이 몸을 살짝 숙여 상자를 내려다보는 자세가 물리적으로 안정적이며, 상자는 바닥 또는 구조물 위에 제대로 놓여 있음."
   },
   {
    "label": "B",
    "direction": "로봇이 화면 하단 중앙에 위치한 상자와 새끼 새들을 향해 시선을 아래로 두고 있음.",
    "built_space": "창고 내부. 왼쪽에 열린 문, 오른쪽에 고철 더미가 있어 이전 샷의 배경 요소가 잘 반영됨.",
    "entities": "새끼 새들은 상자 안에 잘 표현되었으나, B-200의 머리 형태가 가로로 넓고 짐승의 코가 있는 형태로 심하게 변형되어 레퍼런스의 정체성을 상실함. 총구 형태의 손이 프레임에 없음.",
    "hard_violations": [],
    "physics": "로봇과 새들이 물리적으로 지탱되어 있으며, 부자연스럽게 떠 있거나 구조적으로 불가능한 부분은 없음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "지정된 B-200의 복잡한 외형과 무기 손을 레퍼런스와 동일하게 유지하면서, 요구된 반달 모양의 안광과 새들을 내려다보는 자연스러운 앵글을 성공적으로 연출했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "창고 배경 요소는 잘 구현했으나, 로봇의 머리 형태가 짐승처럼 크게 변형되어 레퍼런스와의 일치도가 매우 떨어지며 개조된 총구 손도 보이지 않습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "로봇의 고개와 시선이 화면 왼쪽 하단에 놓인 새끼 새가 담긴 상자를 정확하게 향하고 있음.",
        "built_space": "창고 내부. 화면 왼쪽에 반쯤 열린 문과 야외의 빛이 보이며, 배경에는 흐릿하게 선반과 집기들이 배치되어 공간감이 자연스러움.",
        "entities": "B-200의 머리 형태, 회색 철제 장갑판, 양팔의 고사포 포신 등이 레퍼런스와 정확히 일치하며, 눈 부위에 흰색 반달 모양의 안광이 구현됨. 상자 안의 새끼 새들도 명확함.",
        "hard_violations": [],
        "physics": "로봇이 몸을 살짝 숙여 상자를 내려다보는 자세가 물리적으로 안정적이며, 상자는 바닥 또는 구조물 위에 제대로 놓여 있음."
       },
       {
        "label": "B",
        "direction": "로봇이 화면 하단 중앙에 위치한 상자와 새끼 새들을 향해 시선을 아래로 두고 있음.",
        "built_space": "창고 내부. 왼쪽에 열린 문, 오른쪽에 고철 더미가 있어 이전 샷의 배경 요소가 잘 반영됨.",
        "entities": "새끼 새들은 상자 안에 잘 표현되었으나, B-200의 머리 형태가 가로로 넓고 짐승의 코가 있는 형태로 심하게 변형되어 레퍼런스의 정체성을 상실함. 총구 형태의 손이 프레임에 없음.",
        "hard_violations": [],
        "physics": "로봇과 새들이 물리적으로 지탱되어 있으며, 부자연스럽게 떠 있거나 구조적으로 불가능한 부분은 없음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "지정된 B-200의 복잡한 외형과 무기 손을 레퍼런스와 동일하게 유지하면서, 요구된 반달 모양의 안광과 새들을 내려다보는 자연스러운 앵글을 성공적으로 연출했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "창고 배경 요소는 잘 구현했으나, 로봇의 머리 형태가 짐승처럼 크게 변형되어 레퍼런스와의 일치도가 매우 떨어지며 개조된 총구 손도 보이지 않습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "로봇의 고개와 시선이 화면 왼쪽 하단에 놓인 새끼 새가 담긴 상자를 정확하게 향하고 있음.",
        "built_space": "창고 내부. 화면 왼쪽에 반쯤 열린 문과 야외의 빛이 보이며, 배경에는 흐릿하게 선반과 집기들이 배치되어 공간감이 자연스러움.",
        "entities": "B-200의 머리 형태, 회색 철제 장갑판, 양팔의 고사포 포신 등이 레퍼런스와 정확히 일치하며, 눈 부위에 흰색 반달 모양의 안광이 구현됨. 상자 안의 새끼 새들도 명확함.",
        "hard_violations": [],
        "physics": "로봇이 몸을 살짝 숙여 상자를 내려다보는 자세가 물리적으로 안정적이며, 상자는 바닥 또는 구조물 위에 제대로 놓여 있음."
       },
       {
        "label": "B",
        "direction": "로봇이 화면 하단 중앙에 위치한 상자와 새끼 새들을 향해 시선을 아래로 두고 있음.",
        "built_space": "창고 내부. 왼쪽에 열린 문, 오른쪽에 고철 더미가 있어 이전 샷의 배경 요소가 잘 반영됨.",
        "entities": "새끼 새들은 상자 안에 잘 표현되었으나, B-200의 머리 형태가 가로로 넓고 짐승의 코가 있는 형태로 심하게 변형되어 레퍼런스의 정체성을 상실함. 총구 형태의 손이 프레임에 없음.",
        "hard_violations": [],
        "physics": "로봇과 새들이 물리적으로 지탱되어 있으며, 부자연스럽게 떠 있거나 구조적으로 불가능한 부분은 없음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "얼굴 중심 클로즈업은 정확하지만, 얼굴 축이 카메라 정면에 가까워 바로 아래 아기 새들을 내려다보는 핵심 행동이 불명확하다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "상반신까지 보여 클로즈업보다 넓지만, 새들이 있는 왼쪽 아래로 숙인 얼굴과 부드러운 반달 안광이 핵심 순간을 더 명확히 구현한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "B-200의 얼굴은 거의 정면으로 카메라를 향한다. 동공 없는 반달 안광만으로 시선을 확정하기는 어렵지만, 턱 아래 전경의 새들을 향한 하향 고개 움직임은 뚜렷하지 않다. 전경의 새 두 마리는 머리를 위로 들어 로봇 쪽을 향한다. 포신은 프레임 밖이다.",
        "built_space": "왼쪽에 낮빛이 들어오는 출입구 하나와 높은 창 하나, 위쪽 철골 지붕, 오른쪽 선반과 고장 난 로봇 부품 더미 및 나무 상자가 보인다. 이전 장면의 좌우 공간 배치와 재질이 대체로 이어진다. 새 상자 하나가 로봇 바로 앞 하단에 걸쳐 있고 얼굴이 화면 대부분을 차지한다. 불가능한 반사나 중복된 고정 설비는 보이지 않는다.",
        "entities": "짙은 회색 장갑, 각진 금속 얼굴, 붉은 보조 렌즈와 굵은 목 기구를 가진 B-200 한 대가 보이며 캐릭터 참조의 기계 정체성을 따른다. 안광 두 개는 노란 반달 형태이고 주변 금속에도 약하게 빛이 번진다. 솜털 난 아기 새 두 마리가 뚜렷하고 하단에는 다른 새로 보이는 일부 형체가 겹친다. 찰리나 다른 사람은 없고 읽을 수 있는 글자도 보이지 않는다.",
        "hard_violations": [],
        "physics": "머리는 목의 기계 관절과 몸통에 연결되어 있다. 새들의 몸은 상자 안으로 이어져 바닥이나 둥지에 앉아 있는 것으로 보이며 비행 중인 개체는 없다. 상자의 하부 지지면과 로봇의 다리는 화면 밖이므로 접지 방식은 확인할 수 없지만, 공중에 떠 있다고 볼 증거는 없다."
       },
       {
        "label": "B",
        "direction": "B-200은 고개를 왼쪽 아래로 숙여 전경 상자 속 아기 새들을 향한다. 새들은 머리를 들어 로봇 얼굴 쪽을 바라본다. 하단에 보이는 포신들은 화면 왼쪽을 향하며 새들의 머리보다 낮은 높이를 지나고, 새들을 직접 겨냥하는 구도는 아니다.",
        "built_space": "왼쪽 출입구 하나와 높은 창 하나, 철골 지붕과 벽 기둥, 오른쪽 선반 일부가 보인다. 뒤쪽에는 작은 수납장 하나, 선반과 선풍기 하나가 드러나 A보다 배경 정보가 많다. 새 상자는 왼쪽 전경, 로봇은 오른쪽에 위치한다. 이전 장면의 창고 재질과 주광 방향은 이어지지만, 얼굴뿐 아니라 가슴과 포신까지 포함해 요구된 클로즈업보다 넓다. 부품 더미의 연속성은 가려져 확인하기 어렵다.",
        "entities": "B-200 한 대의 짙은 회색 중장갑, 각진 얼굴, 붉은 보조 센서, 굵은 관절과 손을 대체한 포신이 보인다. 참조의 주요 기계적 특징을 유지하며 눈은 부드러운 흰색 반달 발광부로 표현됐다. 상자에는 솜털 난 아기 새 네 마리가 보인다. 찰리나 추가 인물은 없다. 장갑 표면에 작은 문자 같은 흔적은 있으나 명확히 읽히지는 않는다.",
        "hard_violations": [],
        "physics": "숙인 머리는 목 관절로 몸통에 연결되고 포신도 팔 기구에 이어져 있어 지지가 확인된다. 새들의 몸은 상자 안에 놓여 있으며 떠 있는 개체는 없다. 상자의 밑면과 로봇의 하체는 잘렸으므로 구체적인 받침이나 접지는 확인할 수 없지만, 보이는 부분에 물리적으로 불가능한 지지나 동작은 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "얼굴 중심 클로즈업은 정확하지만, 얼굴 축이 카메라 정면에 가까워 바로 아래 아기 새들을 내려다보는 핵심 행동이 불명확하다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "상반신까지 보여 클로즈업보다 넓지만, 새들이 있는 왼쪽 아래로 숙인 얼굴과 부드러운 반달 안광이 핵심 순간을 더 명확히 구현한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "B-200의 얼굴은 거의 정면으로 카메라를 향한다. 동공 없는 반달 안광만으로 시선을 확정하기는 어렵지만, 턱 아래 전경의 새들을 향한 하향 고개 움직임은 뚜렷하지 않다. 전경의 새 두 마리는 머리를 위로 들어 로봇 쪽을 향한다. 포신은 프레임 밖이다.",
        "built_space": "왼쪽에 낮빛이 들어오는 출입구 하나와 높은 창 하나, 위쪽 철골 지붕, 오른쪽 선반과 고장 난 로봇 부품 더미 및 나무 상자가 보인다. 이전 장면의 좌우 공간 배치와 재질이 대체로 이어진다. 새 상자 하나가 로봇 바로 앞 하단에 걸쳐 있고 얼굴이 화면 대부분을 차지한다. 불가능한 반사나 중복된 고정 설비는 보이지 않는다.",
        "entities": "짙은 회색 장갑, 각진 금속 얼굴, 붉은 보조 렌즈와 굵은 목 기구를 가진 B-200 한 대가 보이며 캐릭터 참조의 기계 정체성을 따른다. 안광 두 개는 노란 반달 형태이고 주변 금속에도 약하게 빛이 번진다. 솜털 난 아기 새 두 마리가 뚜렷하고 하단에는 다른 새로 보이는 일부 형체가 겹친다. 찰리나 다른 사람은 없고 읽을 수 있는 글자도 보이지 않는다.",
        "hard_violations": [],
        "physics": "머리는 목의 기계 관절과 몸통에 연결되어 있다. 새들의 몸은 상자 안으로 이어져 바닥이나 둥지에 앉아 있는 것으로 보이며 비행 중인 개체는 없다. 상자의 하부 지지면과 로봇의 다리는 화면 밖이므로 접지 방식은 확인할 수 없지만, 공중에 떠 있다고 볼 증거는 없다."
       },
       {
        "label": "A",
        "direction": "B-200은 고개를 왼쪽 아래로 숙여 전경 상자 속 아기 새들을 향한다. 새들은 머리를 들어 로봇 얼굴 쪽을 바라본다. 하단에 보이는 포신들은 화면 왼쪽을 향하며 새들의 머리보다 낮은 높이를 지나고, 새들을 직접 겨냥하는 구도는 아니다.",
        "built_space": "왼쪽 출입구 하나와 높은 창 하나, 철골 지붕과 벽 기둥, 오른쪽 선반 일부가 보인다. 뒤쪽에는 작은 수납장 하나, 선반과 선풍기 하나가 드러나 A보다 배경 정보가 많다. 새 상자는 왼쪽 전경, 로봇은 오른쪽에 위치한다. 이전 장면의 창고 재질과 주광 방향은 이어지지만, 얼굴뿐 아니라 가슴과 포신까지 포함해 요구된 클로즈업보다 넓다. 부품 더미의 연속성은 가려져 확인하기 어렵다.",
        "entities": "B-200 한 대의 짙은 회색 중장갑, 각진 얼굴, 붉은 보조 센서, 굵은 관절과 손을 대체한 포신이 보인다. 참조의 주요 기계적 특징을 유지하며 눈은 부드러운 흰색 반달 발광부로 표현됐다. 상자에는 솜털 난 아기 새 네 마리가 보인다. 찰리나 추가 인물은 없다. 장갑 표면에 작은 문자 같은 흔적은 있으나 명확히 읽히지는 않는다.",
        "hard_violations": [],
        "physics": "숙인 머리는 목 관절로 몸통에 연결되고 포신도 팔 기구에 이어져 있어 지지가 확인된다. 새들의 몸은 상자 안에 놓여 있으며 떠 있는 개체는 없다. 상자의 밑면과 로봇의 하체는 잘렸으므로 구체적인 받침이나 접지는 확인할 수 없지만, 보이는 부분에 물리적으로 불가능한 지지나 동작은 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.125
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.125
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1125
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지정된 B-200의 복잡한 외형과 무기 손을 레퍼런스와 동일하게 유지하면서, 요구된 반달 모양의 안광과 새들을 내려다보는 자연스러운 앵글을 성공적으로 연출했습니다."
   },
   {
    "label": "B",
    "score": 1125,
    "verdict_ko": "창고 배경 요소는 잘 구현했으나, 로봇의 머리 형태가 짐승처럼 크게 변형되어 레퍼런스와의 일치도가 매우 떨어지며 개조된 총구 손도 보이지 않습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of B-200 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S64sh15_sel.png",
    "asset_id": "8a695140-800e-4b52-9785-84cb0048d521",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — B-200: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1339855>",
    "asset_id": "8091d94b-e8e7-407e-97d6-c030f55a73f9",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-9d2c-7881-983a-1316351cff59",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S64sh15"
  }
 },
 "S64sh24::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:52:23.074332+00:00",
  "fingerprint": "4437df52b334f7faf2cd402a9655398bbf1ddf1707ccff3d4105d0cd671d9d97",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S64sh24_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S64sh24_sel.png",
  "source_sha256": "1d711d3bf2853297b799e79068328d8cdffb18e8491daf39d5b837afea64a58d",
  "file": "S64sh24_cine.png",
  "staged_sha256": "082e642520575cf3fb8a8e18b1d6186867cd83cc6b632d8e853db32d2f5ed269",
  "latency_ms": 9205
 },
 "S65sh5::signage": {
  "fp": "3dfdfb28a150f0da",
  "inscriptions": [
   {
    "text_native": "0",
    "source": "scene_text_quoted",
    "reason_ko": "방사능 측정기 눈금이 가리키고 있는 숫자가 클로즈업 화면에 표시되어야 합니다.",
    "source_quote": "0"
   }
  ],
  "cues": [],
  "dropped": []
 },
 "S65sh5": {
  "input_fingerprint": "c314461c4486dc0e",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 방사능 측정기 눈금이 숫자 0에 정확히 멈춰 선 상태의 클로즈업.\n\nLOCATION (lock): On the overgrown approach to a partly collapsed department store, among moss, puddles, and uncontrolled vegetation. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Radiation meter dial (The needle has settled at the zero marking) — The marked face is directed toward the camera at a readable downward oblique angle, showing the needle and numeral 0; used as The primary focal detail, held within the supporting hands and environmental context rather than enlarged to fill the image; Moss-covered ground (Moss spreads across the approach to the collapsed department store); used as Soft surrounding context connects the apparently anomalous reading to the living environment.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight appropriate to the exterior keeps the dial markings readable and the surrounding moss subdued, without glare obscuring the zero reading.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The department store is half collapsed and overgrown with moss and trees, with standing pools of water. The radiation meter settles near zero, rather than establishing an exact zero reading.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nWORDS TO RENDER (authoritative — the scene itself calls for these; render each as period-real physical lettering in the native script, exactly as written; add no other readable text anywhere):\n- \"0\"\n\nThe WORDS TO RENDER above are the only readable writing in this image: render those words exactly as given, in the place and era's own language and script, and nothing else legible. Invent no other wording a viewer could read. No caption, subtitle, watermark, logo or overlay. Surfaces that would carry writing may still be present — stage any wording they would carry out of legibility: a hand across, an oblique angle, shallow focus.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 방사능 측정기 눈금이 숫자 0에 정확히 멈춰 선 상태의 클로즈업.\n\nLOCATION (lock): On the overgrown approach to a partly collapsed department store, among moss, puddles, and uncontrolled vegetation. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Radiation meter dial (The needle has settled at the zero marking) — The marked face is directed toward the camera at a readable downward oblique angle, showing the needle and numeral 0; used as The primary focal detail, held within the supporting hands and environmental context rather than enlarged to fill the image; Moss-covered ground (Moss spreads across the approach to the collapsed department store); used as Soft surrounding context connects the apparently anomalous reading to the living environment.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight appropriate to the exterior keeps the dial markings readable and the surrounding moss subdued, without glare obscuring the zero reading.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The department store is half collapsed and overgrown with moss and trees, with standing pools of water. The radiation meter settles near zero, rather than establishing an exact zero reading.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nWORDS TO RENDER (authoritative — the scene itself calls for these; render each as period-real physical lettering in the native script, exactly as written; add no other readable text anywhere):\n- \"0\"\n\nThe WORDS TO RENDER above are the only readable writing in this image: render those words exactly as given, in the place and era's own language and script, and nothing else legible. Invent no other wording a viewer could read. No caption, subtitle, watermark, logo or overlay. Surfaces that would carry writing may still be present — stage any wording they would carry out of legibility: a hand across, an oblique angle, shallow focus.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 방사능 측정기 눈금이 숫자 0에 정확히 멈춰 선 상태의 클로즈업.\n\nLOCATION (lock): On the overgrown approach to a partly collapsed department store, among moss, puddles, and uncontrolled vegetation. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Radiation meter dial (The needle has settled at the zero marking) — The marked face is directed toward the camera at a readable downward oblique angle, showing the needle and numeral 0; used as The primary focal detail, held within the supporting hands and environmental context rather than enlarged to fill the image; Moss-covered ground (Moss spreads across the approach to the collapsed department store); used as Soft surrounding context connects the apparently anomalous reading to the living environment.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight appropriate to the exterior keeps the dial markings readable and the surrounding moss subdued, without glare obscuring the zero reading.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The department store is half collapsed and overgrown with moss and trees, with standing pools of water. The radiation meter settles near zero, rather than establishing an exact zero reading.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nWORDS TO RENDER (authoritative — the scene itself calls for these; render each as period-real physical lettering in the native script, exactly as written; add no other readable text anywhere):\n- \"0\"\n\nThe WORDS TO RENDER above are the only readable writing in this image: render those words exactly as given, in the place and era's own language and script, and nothing else legible. Invent no other wording a viewer could read. No caption, subtitle, watermark, logo or overlay. Surfaces that would carry writing may still be present — stage any wording they would carry out of legibility: a hand across, an oblique angle, shallow focus.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "측정기 계기판이 카메라를 향하며, 바늘은 중앙 상단을 가리킴.",
    "built_space": "이끼 낀 바닥이 아닌 레퍼런스 사진의 원경(건물과 연못)을 그대로 배경으로 사용함.",
    "entities": "측정기를 쥔 양손. 오른손에 엄지손가락이 2개 존재함. 계기판에 '0'이 두 번 표기됨.",
    "hard_violations": [
     "[gemini-pro] 해부학적으로 불가능한 신체 구조 (오른손의 여분 엄지손가락)"
    ],
    "physics": "양손이 기기를 들고 지탱함."
   },
   {
    "label": "B",
    "direction": "측정기 계기판이 카메라를 비스듬히 향하며, 바늘이 정확히 0을 가리킴.",
    "built_space": "이끼 낀 바닥과 웅덩이가 배경을 채우며, 수면에 무너진 건물이 반사됨.",
    "entities": "측정기와 가죽 스트랩을 쥔 두 손. 계기판에 '0'이 하나만 명확히 표기됨.",
    "hard_violations": [],
    "physics": "양손이 기기를 안정적으로 쥐고 지탱함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지시된 클로즈업 앵글, 이끼 낀 바닥 배경, 0을 가리키는 바늘 등 프롬프트의 요구사항을 훌륭하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "해부학적으로 불가능한 손가락 기형이 발생했으며, 이끼 낀 바닥 대신 레퍼런스 이미지의 넓은 구도를 그대로 복사하여 감점되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "측정기 계기판이 카메라를 향하며, 바늘은 중앙 상단을 가리킴.",
        "built_space": "이끼 낀 바닥이 아닌 레퍼런스 사진의 원경(건물과 연못)을 그대로 배경으로 사용함.",
        "entities": "측정기를 쥔 양손. 오른손에 엄지손가락이 2개 존재함. 계기판에 '0'이 두 번 표기됨.",
        "hard_violations": [
         "해부학적으로 불가능한 신체 구조 (오른손의 여분 엄지손가락)"
        ],
        "physics": "양손이 기기를 들고 지탱함."
       },
       {
        "label": "B",
        "direction": "측정기 계기판이 카메라를 비스듬히 향하며, 바늘이 정확히 0을 가리킴.",
        "built_space": "이끼 낀 바닥과 웅덩이가 배경을 채우며, 수면에 무너진 건물이 반사됨.",
        "entities": "측정기와 가죽 스트랩을 쥔 두 손. 계기판에 '0'이 하나만 명확히 표기됨.",
        "hard_violations": [],
        "physics": "양손이 기기를 안정적으로 쥐고 지탱함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지시된 클로즈업 앵글, 이끼 낀 바닥 배경, 0을 가리키는 바늘 등 프롬프트의 요구사항을 훌륭하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "해부학적으로 불가능한 손가락 기형이 발생했으며, 이끼 낀 바닥 대신 레퍼런스 이미지의 넓은 구도를 그대로 복사하여 감점되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "측정기 계기판이 카메라를 향하며, 바늘은 중앙 상단을 가리킴.",
        "built_space": "이끼 낀 바닥이 아닌 레퍼런스 사진의 원경(건물과 연못)을 그대로 배경으로 사용함.",
        "entities": "측정기를 쥔 양손. 오른손에 엄지손가락이 2개 존재함. 계기판에 '0'이 두 번 표기됨.",
        "hard_violations": [
         "해부학적으로 불가능한 신체 구조 (오른손의 여분 엄지손가락)"
        ],
        "physics": "양손이 기기를 들고 지탱함."
       },
       {
        "label": "B",
        "direction": "측정기 계기판이 카메라를 비스듬히 향하며, 바늘이 정확히 0을 가리킴.",
        "built_space": "이끼 낀 바닥과 웅덩이가 배경을 채우며, 수면에 무너진 건물이 반사됨.",
        "entities": "측정기와 가죽 스트랩을 쥔 두 손. 계기판에 '0'이 하나만 명확히 표기됨.",
        "hard_violations": [],
        "physics": "양손이 기기를 안정적으로 쥐고 지탱함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "0의 시작 눈금에 놓인 바늘을 내려다보는 클로즈업으로, 손의 지지와 이끼 배경을 유지하면서 핵심 판독 상태를 가장 명료하게 구현했다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "바늘의 0 지시와 장소는 충실하지만, A보다 정면에 가까운 계기판 각도와 강조된 건물 배경, 중복된 0 표기가 핵심 클로즈업의 집중도를 낮춘다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "계기판은 카메라와 손을 뻗은 관찰자 쪽으로 기울어져 있으며, 위에서 비스듬히 내려다보는 방향으로 읽힌다. 바늘은 오른쪽 아래 회전축에서 왼쪽의 숫자 0 바로 위 시작 눈금을 향한다. 얼굴이나 시선은 보이지 않는다.",
        "built_space": "전경에는 이끼로 덮인 깨진 콘크리트와 작은 고인 물이 있고, 상단에는 큰 물웅덩이 하나가 보인다. 건물 본체 대신 물에 비친 콘크리트 외벽과 세로 창 구획이 나타난다. 맞은편 건물이 수면에 거꾸로 비치는 배치는 이 하향 시점에서 가능하다. 참고 장소의 재료와 식생은 맞지만 건물의 정확한 붕괴 형상은 직접 확인할 수 없다.",
        "entities": "낡은 금속 외함, 유리 덮개, 부채꼴 눈금과 바늘을 가진 휴대용 아날로그 측정기 한 대가 보인다. 기기 종류를 글자로 확인할 수는 없지만 방사능 측정기 소품으로 타당하다. 판독 가능한 문자는 숫자 0 하나이며 다른 문구는 없다. 흙 묻은 손 두 개와 팔 일부만 나오고 얼굴이나 추가 인물은 없다. 손만으로 민족·성별·정확한 나이를 판단할 수 없다. 이끼, 잡초, 물웅덩이와 잔해가 있으며 낮의 자연광으로 읽힌다.",
        "hard_violations": [],
        "physics": "왼손이 기기의 왼쪽과 아래를 받치고 오른손이 오른쪽 손잡이와 외함을 잡아 무게를 지지한다. 가죽 끈은 기기에 연결되고 손에 잡혀 있다. 바늘은 계기 내부 회전축에 연결되어 있으며, 떠 있는 물체나 불가능한 손 자세는 보이지 않는다. 유리의 반사는 있지만 0과 바늘 끝을 가리지 않는다."
       },
       {
        "label": "B",
        "direction": "계기판 앞면은 카메라 쪽으로 향하며 A보다 정면에 가깝게 보인다. 중앙 아래 회전축에서 뻗은 바늘은 위쪽의 작은 0에 해당하는 눈금을 가리킨다. 계기 중앙에도 큰 0이 있어 같은 숫자가 두 번 보인다. 얼굴이나 시선은 없다.",
        "built_space": "상단 중앙에 백화점 입구 한 곳과 그 앞 계단·접근로가 있고, 양쪽에는 무너진 콘크리트와 창 구획이 이어진다. 오른쪽에는 기둥으로 나뉜 창가 공간이 보이며, 접근로 주변에 식생과 여러 고인 물 구역이 있다. 참고 사진의 중앙 진입로와 붕괴된 입면 관계를 잘 유지한다. 인물은 건물 내부가 아니라 접근로에서 기기를 든 배치로 읽히며, 불가능한 반사나 명백히 중복된 고정 시설은 없다.",
        "entities": "낡은 직사각형 휴대용 아날로그 측정기 한 대, 이를 잡은 손 두 개, 회갈색 소매가 보인다. 방사능 측정기 소품으로 읽힐 수 있는 형태이며, 읽히는 문자는 두 곳의 0뿐이다. 추가 얼굴이나 전신 인물은 없다. 손만으로 민족·성별·정확한 나이를 확인할 수 없다. 이끼, 관목, 물웅덩이, 붕괴된 백화점이 낮의 자연광 아래 보인다.",
        "hard_violations": [],
        "physics": "왼손이 기기 왼쪽과 하단을 감싸 받치고 오른손이 오른쪽 외함을 잡고 있어 지지가 명확하다. 손목과 소매의 연결도 자연스럽다. 계기 바늘은 내부 축에서 뻗어 있고 잔해는 지면에 놓여 있다. 지지 없이 떠 있는 물체나 물리적으로 불가능한 자세는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "0의 시작 눈금에 놓인 바늘을 내려다보는 클로즈업으로, 손의 지지와 이끼 배경을 유지하면서 핵심 판독 상태를 가장 명료하게 구현했다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "바늘의 0 지시와 장소는 충실하지만, A보다 정면에 가까운 계기판 각도와 강조된 건물 배경, 중복된 0 표기가 핵심 클로즈업의 집중도를 낮춘다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "계기판은 카메라와 손을 뻗은 관찰자 쪽으로 기울어져 있으며, 위에서 비스듬히 내려다보는 방향으로 읽힌다. 바늘은 오른쪽 아래 회전축에서 왼쪽의 숫자 0 바로 위 시작 눈금을 향한다. 얼굴이나 시선은 보이지 않는다.",
        "built_space": "전경에는 이끼로 덮인 깨진 콘크리트와 작은 고인 물이 있고, 상단에는 큰 물웅덩이 하나가 보인다. 건물 본체 대신 물에 비친 콘크리트 외벽과 세로 창 구획이 나타난다. 맞은편 건물이 수면에 거꾸로 비치는 배치는 이 하향 시점에서 가능하다. 참고 장소의 재료와 식생은 맞지만 건물의 정확한 붕괴 형상은 직접 확인할 수 없다.",
        "entities": "낡은 금속 외함, 유리 덮개, 부채꼴 눈금과 바늘을 가진 휴대용 아날로그 측정기 한 대가 보인다. 기기 종류를 글자로 확인할 수는 없지만 방사능 측정기 소품으로 타당하다. 판독 가능한 문자는 숫자 0 하나이며 다른 문구는 없다. 흙 묻은 손 두 개와 팔 일부만 나오고 얼굴이나 추가 인물은 없다. 손만으로 민족·성별·정확한 나이를 판단할 수 없다. 이끼, 잡초, 물웅덩이와 잔해가 있으며 낮의 자연광으로 읽힌다.",
        "hard_violations": [],
        "physics": "왼손이 기기의 왼쪽과 아래를 받치고 오른손이 오른쪽 손잡이와 외함을 잡아 무게를 지지한다. 가죽 끈은 기기에 연결되고 손에 잡혀 있다. 바늘은 계기 내부 회전축에 연결되어 있으며, 떠 있는 물체나 불가능한 손 자세는 보이지 않는다. 유리의 반사는 있지만 0과 바늘 끝을 가리지 않는다."
       },
       {
        "label": "A",
        "direction": "계기판 앞면은 카메라 쪽으로 향하며 A보다 정면에 가깝게 보인다. 중앙 아래 회전축에서 뻗은 바늘은 위쪽의 작은 0에 해당하는 눈금을 가리킨다. 계기 중앙에도 큰 0이 있어 같은 숫자가 두 번 보인다. 얼굴이나 시선은 없다.",
        "built_space": "상단 중앙에 백화점 입구 한 곳과 그 앞 계단·접근로가 있고, 양쪽에는 무너진 콘크리트와 창 구획이 이어진다. 오른쪽에는 기둥으로 나뉜 창가 공간이 보이며, 접근로 주변에 식생과 여러 고인 물 구역이 있다. 참고 사진의 중앙 진입로와 붕괴된 입면 관계를 잘 유지한다. 인물은 건물 내부가 아니라 접근로에서 기기를 든 배치로 읽히며, 불가능한 반사나 명백히 중복된 고정 시설은 없다.",
        "entities": "낡은 직사각형 휴대용 아날로그 측정기 한 대, 이를 잡은 손 두 개, 회갈색 소매가 보인다. 방사능 측정기 소품으로 읽힐 수 있는 형태이며, 읽히는 문자는 두 곳의 0뿐이다. 추가 얼굴이나 전신 인물은 없다. 손만으로 민족·성별·정확한 나이를 확인할 수 없다. 이끼, 관목, 물웅덩이, 붕괴된 백화점이 낮의 자연광 아래 보인다.",
        "hard_violations": [],
        "physics": "왼손이 기기 왼쪽과 하단을 감싸 받치고 오른손이 오른쪽 외함을 잡고 있어 지지가 명확하다. 손목과 소매의 연결도 자연스럽다. 계기 바늘은 내부 축에서 뻗어 있고 잔해는 지면에 놓여 있다. 지지 없이 떠 있는 물체나 물리적으로 불가능한 자세는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.317,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.067,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 해부학적으로 불가능한 신체 구조 (오른손의 여분 엄지손가락)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1067
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "지시된 클로즈업 앵글, 이끼 낀 바닥 배경, 0을 가리키는 바늘 등 프롬프트의 요구사항을 훌륭하게 구현했습니다."
   },
   {
    "label": "A",
    "score": 1067,
    "verdict_ko": "해부학적으로 불가능한 손가락 기형이 발생했으며, 이끼 낀 바닥 대신 레퍼런스 이미지의 넓은 구도를 그대로 복사하여 감점되었습니다.  ★위반: [gemini-pro] 해부학적으로 불가능한 신체 구조 (오른손의 여분 엄지손가락)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L239B01.png",
    "asset_id": "a44a0231-5865-442f-9e00-425f802c380d",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-9ee4-7170-821e-c33fe5dbc424",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S65sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:00:05.873647+00:00",
  "fingerprint": "56311d2cdcddb5b0a15701247a0986dc765d95afde4a701e87ac9c0fbef3b0b1",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S65sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S65sh5_sel.png",
  "source_sha256": "a3678d2fa6db32d80064cefaf8e81ad98b919c32f9429d9286ae157c0d098045",
  "file": "S65sh5_cine.png",
  "staged_sha256": "9df9f5122f408eb2893788d36c3becefedbde972415b9e95c32f12e4656f1bd1",
  "latency_ms": 10042
 },
 "S65sh10::signage": {
  "fp": "b6225f47af2d9930",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S65sh10": {
  "input_fingerprint": "4d088aa96e10a240",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 맨얼굴로 두 눈을 감고 깊게 숨을 들이마시는 현우의 평온한 얼굴.\n\nLOCATION (lock): Among moss-covered rubble outside the ruined department store, on the approach before entering the building. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Collapsed department store (Partly collapsed by an earthquake) — A partial exterior section is visible obliquely behind 현우; used as A soft background edge maintains the dangerous setting against his newfound calm; Uncontrolled tree growth (Growing irregularly around the ruined department store); used as Provides a restrained background indication of returning life without crowding the facial close-up.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained daylight preserves natural skin detail and the supported color of the surrounding vegetation, allowing the peaceful expression to register without an added glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The half-collapsed department store remains surrounded by moss, trees, flowers and pooled water, with birds and a water deer present. The radiation meter continues to indicate a value near zero. 현우: He has removed his gas mask and breathes with his face uncovered. His previous fighting injuries remain.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 맨얼굴로 두 눈을 감고 깊게 숨을 들이마시는 현우의 평온한 얼굴.\n\nLOCATION (lock): Among moss-covered rubble outside the ruined department store, on the approach before entering the building. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Collapsed department store (Partly collapsed by an earthquake) — A partial exterior section is visible obliquely behind 현우; used as A soft background edge maintains the dangerous setting against his newfound calm; Uncontrolled tree growth (Growing irregularly around the ruined department store); used as Provides a restrained background indication of returning life without crowding the facial close-up.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained daylight preserves natural skin detail and the supported color of the surrounding vegetation, allowing the peaceful expression to register without an added glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The half-collapsed department store remains surrounded by moss, trees, flowers and pooled water, with birds and a water deer present. The radiation meter continues to indicate a value near zero. 현우: He has removed his gas mask and breathes with his face uncovered. His previous fighting injuries remain.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 맨얼굴로 두 눈을 감고 깊게 숨을 들이마시는 현우의 평온한 얼굴.\n\nLOCATION (lock): Among moss-covered rubble outside the ruined department store, on the approach before entering the building. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Collapsed department store (Partly collapsed by an earthquake) — A partial exterior section is visible obliquely behind 현우; used as A soft background edge maintains the dangerous setting against his newfound calm; Uncontrolled tree growth (Growing irregularly around the ruined department store); used as Provides a restrained background indication of returning life without crowding the facial close-up.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained daylight preserves natural skin detail and the supported color of the surrounding vegetation, allowing the peaceful expression to register without an added glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The half-collapsed department store remains surrounded by moss, trees, flowers and pooled water, with birds and a water deer present. The radiation meter continues to indicate a value near zero. 현우: He has removed his gas mask and breathes with his face uncovered. His previous fighting injuries remain.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "고개를 약간 들고 눈을 감은 채 정면을 향함.",
    "built_space": "이끼 낀 콘크리트 잔해, 물웅덩이, 붕괴된 건물 및 수면 반사가 레퍼런스와 일치하게 배치됨.",
    "entities": "현우(검은 머리, 남색 티셔츠)가 맨얼굴로 있음. 지시된 전투 상처는 보이지 않음.",
    "hard_violations": [],
    "physics": "특별한 지지대 없이 서서 얼굴을 들고 있는 자연스러운 자세."
   },
   {
    "label": "B",
    "direction": "고개를 위로 젖히고 눈을 감은 채 하늘 방향으로 깊게 숨을 들이마심.",
    "built_space": "이끼 낀 잔해와 물웅덩이, 크게 기울어진 채 붕괴된 폐건물 배경이 배치됨.",
    "entities": "현우(검은 머리, 남색 티셔츠)의 이마에 붉은 상처 흔적이 있으며 깊게 호흡 중임.",
    "hard_violations": [],
    "physics": "서 있는 상태로 고개를 뒤로 젖히며 호흡하는 근육의 움직임과 자세가 자연스러움."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "고개를 젖히고 깊게 숨을 들이마시는 역동성을 잘 살렸으며, 이마에 지시된 상처 흔적을 포함해 세부 요구사항을 가장 정확히 충족했습니다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "평온한 표정과 배경은 훌륭하나, 깊게 숨을 쉬는 동작이 덜 뚜렷하고 프롬프트에 명시된 얼굴의 상처가 누락되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "고개를 약간 들고 눈을 감은 채 정면을 향함.",
        "built_space": "이끼 낀 콘크리트 잔해, 물웅덩이, 붕괴된 건물 및 수면 반사가 레퍼런스와 일치하게 배치됨.",
        "entities": "현우(검은 머리, 남색 티셔츠)가 맨얼굴로 있음. 지시된 전투 상처는 보이지 않음.",
        "hard_violations": [],
        "physics": "특별한 지지대 없이 서서 얼굴을 들고 있는 자연스러운 자세."
       },
       {
        "label": "B",
        "direction": "고개를 위로 젖히고 눈을 감은 채 하늘 방향으로 깊게 숨을 들이마심.",
        "built_space": "이끼 낀 잔해와 물웅덩이, 크게 기울어진 채 붕괴된 폐건물 배경이 배치됨.",
        "entities": "현우(검은 머리, 남색 티셔츠)의 이마에 붉은 상처 흔적이 있으며 깊게 호흡 중임.",
        "hard_violations": [],
        "physics": "서 있는 상태로 고개를 뒤로 젖히며 호흡하는 근육의 움직임과 자세가 자연스러움."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "고개를 젖히고 깊게 숨을 들이마시는 역동성을 잘 살렸으며, 이마에 지시된 상처 흔적을 포함해 세부 요구사항을 가장 정확히 충족했습니다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "평온한 표정과 배경은 훌륭하나, 깊게 숨을 쉬는 동작이 덜 뚜렷하고 프롬프트에 명시된 얼굴의 상처가 누락되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "고개를 약간 들고 눈을 감은 채 정면을 향함.",
        "built_space": "이끼 낀 콘크리트 잔해, 물웅덩이, 붕괴된 건물 및 수면 반사가 레퍼런스와 일치하게 배치됨.",
        "entities": "현우(검은 머리, 남색 티셔츠)가 맨얼굴로 있음. 지시된 전투 상처는 보이지 않음.",
        "hard_violations": [],
        "physics": "특별한 지지대 없이 서서 얼굴을 들고 있는 자연스러운 자세."
       },
       {
        "label": "B",
        "direction": "고개를 위로 젖히고 눈을 감은 채 하늘 방향으로 깊게 숨을 들이마심.",
        "built_space": "이끼 낀 잔해와 물웅덩이, 크게 기울어진 채 붕괴된 폐건물 배경이 배치됨.",
        "entities": "현우(검은 머리, 남색 티셔츠)의 이마에 붉은 상처 흔적이 있으며 깊게 호흡 중임.",
        "hard_violations": [],
        "physics": "서 있는 상태로 고개를 뒤로 젖히며 호흡하는 근육의 움직임과 자세가 자연스러움."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "눈을 감고 턱을 들어 깊게 들이마시는 얼굴 클로즈업은 충실하지만, 붕괴 건물이 배경의 부드러운 가장자리보다 두드러지고 기존 부상 흔적은 불분명하다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "맨얼굴과 감긴 두 눈, 평온한 호흡을 얼굴 중심 클로즈업으로 담고, 부드러운 폐허 배경과 참조 인물의 얼굴·복장 및 뺨의 상처 흔적을 더 충실히 유지한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 얼굴을 카메라 쪽으로 두되 턱을 위로 들고 두 눈을 감았다. 특정 대상을 바라보지 않으며, 벌어진 입과 들린 턱은 깊게 숨을 들이마시는 순간으로 읽힌다. 무기나 방향을 확인할 휴대 물건은 보이지 않는다.",
        "built_space": "인물 뒤로 백화점 한 동의 무너진 왼쪽 외벽과 오른쪽 상층 보·기둥 열이 넓게 보인다. 하단 양쪽에는 이끼 낀 잔해, 오른쪽에는 고인 물과 건물 반사, 주변에는 나무와 풀이 있다. 인물은 건물 바깥 접근로에 있으며 위치 모순은 없다. 수면 반사는 가능한 배치다. 다만 건물의 붕괴 면과 식생이 상당히 선명하여 절제된 흐린 배경 가장자리라는 지시보다 존재감이 크다.",
        "entities": "젊은 동아시아계 남성 한 명만 보이며, 앳된 얼굴과 헝클어진 검은 머리는 현우 참조와 대체로 맞는다. 얼굴에는 방독면이 없고 하단에 남색 옷깃 일부가 보인다. 피부 질감은 자연스럽지만 전투 부상으로 명확히 구분되는 상처는 잘 보이지 않는다. 계측기·새·고라니·꽃은 이 얼굴 클로즈업에서 확인되지 않으며, 이를 보여주려고 구도를 넓힐 필요는 없다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리는 보이는 목과 어깨에 자연스럽게 연결되어 있으며 턱을 들어 숨 쉬는 자세는 가능하다. 발과 지면 접촉은 화면 밖이므로 확인할 수 없지만 공중에 떠 있는 모습은 아니다. 잔해는 지면이나 다른 잔해에 얹혀 있고 물은 낮은 지면에 고여 있다. 지지 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "현우는 거의 정면을 향하고 두 눈을 완전히 감았으며 입술을 살짝 벌렸다. 시선의 대상은 없고, 이완된 얼굴과 조금 들린 턱이 평온하게 숨을 들이마시는 연기를 이룬다. 방향을 따질 무기나 손에 든 물체는 없다.",
        "built_space": "화면 위쪽 뒤에 백화점 한 동의 밝은 콘크리트 외벽 일부, 창 개구부와 하부 기둥 사이의 어두운 공간이 흐리게 보인다. 인물 양옆에는 이끼 낀 콘크리트 잔해와 나무가 있고, 뒤쪽의 큰 물웅덩이에 외벽과 창이 반사된다. 건물 밖 접근로라는 위치와 참조의 재료·수면·식생 구성이 잘 이어진다. 수면 반사에도 뚜렷한 광학적 모순은 없다. 다만 반쯤 붕괴한 건물이라는 특징 자체는 A보다 약하게 드러난다.",
        "entities": "현우에 해당하는 젊은 동아시아계 남성 한 명만 있다. 얼굴 윤곽, 코와 입술, 흐트러진 검은 머리, 남색 둥근목 티셔츠가 인물 참조와 가깝다. 방독면 없이 얼굴이 드러나며 화면 오른쪽 뺨에는 작은 긁힌 상처 같은 흔적이 보인다. 계측기와 동물, 꽃은 프레임에서 확인되지 않지만 얼굴 클로즈업의 범위상 누락으로 감점할 사항은 아니다. 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "머리·목·어깨가 해부학적으로 자연스럽게 이어지고, 곧게 선 상체에서 편안히 호흡하는 자세다. 하체는 잘렸으므로 발의 지지는 보이지 않지만 부유를 암시하지 않는다. 옷은 어깨와 몸통을 따라 드리워져 있고 잔해와 물도 지면에 놓여 있다. 지지 없는 물체나 불가능한 동작은 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "눈을 감고 턱을 들어 깊게 들이마시는 얼굴 클로즈업은 충실하지만, 붕괴 건물이 배경의 부드러운 가장자리보다 두드러지고 기존 부상 흔적은 불분명하다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "맨얼굴과 감긴 두 눈, 평온한 호흡을 얼굴 중심 클로즈업으로 담고, 부드러운 폐허 배경과 참조 인물의 얼굴·복장 및 뺨의 상처 흔적을 더 충실히 유지한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 얼굴을 카메라 쪽으로 두되 턱을 위로 들고 두 눈을 감았다. 특정 대상을 바라보지 않으며, 벌어진 입과 들린 턱은 깊게 숨을 들이마시는 순간으로 읽힌다. 무기나 방향을 확인할 휴대 물건은 보이지 않는다.",
        "built_space": "인물 뒤로 백화점 한 동의 무너진 왼쪽 외벽과 오른쪽 상층 보·기둥 열이 넓게 보인다. 하단 양쪽에는 이끼 낀 잔해, 오른쪽에는 고인 물과 건물 반사, 주변에는 나무와 풀이 있다. 인물은 건물 바깥 접근로에 있으며 위치 모순은 없다. 수면 반사는 가능한 배치다. 다만 건물의 붕괴 면과 식생이 상당히 선명하여 절제된 흐린 배경 가장자리라는 지시보다 존재감이 크다.",
        "entities": "젊은 동아시아계 남성 한 명만 보이며, 앳된 얼굴과 헝클어진 검은 머리는 현우 참조와 대체로 맞는다. 얼굴에는 방독면이 없고 하단에 남색 옷깃 일부가 보인다. 피부 질감은 자연스럽지만 전투 부상으로 명확히 구분되는 상처는 잘 보이지 않는다. 계측기·새·고라니·꽃은 이 얼굴 클로즈업에서 확인되지 않으며, 이를 보여주려고 구도를 넓힐 필요는 없다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "머리는 보이는 목과 어깨에 자연스럽게 연결되어 있으며 턱을 들어 숨 쉬는 자세는 가능하다. 발과 지면 접촉은 화면 밖이므로 확인할 수 없지만 공중에 떠 있는 모습은 아니다. 잔해는 지면이나 다른 잔해에 얹혀 있고 물은 낮은 지면에 고여 있다. 지지 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "현우는 거의 정면을 향하고 두 눈을 완전히 감았으며 입술을 살짝 벌렸다. 시선의 대상은 없고, 이완된 얼굴과 조금 들린 턱이 평온하게 숨을 들이마시는 연기를 이룬다. 방향을 따질 무기나 손에 든 물체는 없다.",
        "built_space": "화면 위쪽 뒤에 백화점 한 동의 밝은 콘크리트 외벽 일부, 창 개구부와 하부 기둥 사이의 어두운 공간이 흐리게 보인다. 인물 양옆에는 이끼 낀 콘크리트 잔해와 나무가 있고, 뒤쪽의 큰 물웅덩이에 외벽과 창이 반사된다. 건물 밖 접근로라는 위치와 참조의 재료·수면·식생 구성이 잘 이어진다. 수면 반사에도 뚜렷한 광학적 모순은 없다. 다만 반쯤 붕괴한 건물이라는 특징 자체는 A보다 약하게 드러난다.",
        "entities": "현우에 해당하는 젊은 동아시아계 남성 한 명만 있다. 얼굴 윤곽, 코와 입술, 흐트러진 검은 머리, 남색 둥근목 티셔츠가 인물 참조와 가깝다. 방독면 없이 얼굴이 드러나며 화면 오른쪽 뺨에는 작은 긁힌 상처 같은 흔적이 보인다. 계측기와 동물, 꽃은 프레임에서 확인되지 않지만 얼굴 클로즈업의 범위상 누락으로 감점할 사항은 아니다. 읽을 수 있는 글자나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "머리·목·어깨가 해부학적으로 자연스럽게 이어지고, 곧게 선 상체에서 편안히 호흡하는 자세다. 하체는 잘렸으므로 발의 지지는 보이지 않지만 부유를 암시하지 않는다. 옷은 어깨와 몸통을 따라 드리워져 있고 잔해와 물도 지면에 놓여 있다. 지지 없는 물체나 불가능한 동작은 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.857,
    "B": 1.889
   },
   "adjusted": {
    "A": 1.857,
    "B": 1.889
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1889,
   "A": 1857
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1889,
    "verdict_ko": "고개를 젖히고 깊게 숨을 들이마시는 역동성을 잘 살렸으며, 이마에 지시된 상처 흔적을 포함해 세부 요구사항을 가장 정확히 충족했습니다."
   },
   {
    "label": "A",
    "score": 1857,
    "verdict_ko": "평온한 표정과 배경은 훌륭하나, 깊게 숨을 쉬는 동작이 덜 뚜렷하고 프롬프트에 명시된 얼굴의 상처가 누락되었습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S65sh5_sel.png",
    "asset_id": "e3e1a817-0d43-4ace-9d8f-51484a5940e9",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-a097-71c1-9488-ec6b52e4be8f",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S65sh5"
  }
 },
 "S65sh10::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:01:09.987530+00:00",
  "fingerprint": "93cd39494e76ab251d6a370180e8777ae10d497e49a4f7ffc15726ab460c3f51",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S65sh10_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S65sh10_sel.png",
  "source_sha256": "fc1eeb1dc9b321616d364ced2be686271bba8a357093cc1945fa4479ed09ccd5",
  "file": "S65sh10_cine.png",
  "staged_sha256": "173f73e08468d7677f07a516ee7f599bd4237da2a170df67e9e9b88bf84290fa",
  "latency_ms": 9914
 },
 "S65sh20::signage": {
  "fp": "51e81110944353f6",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S65sh20": {
  "input_fingerprint": "ce1d9afe45fd6e9f",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 높은 콘크리트 층 위에서 아래의 현우와 수빈을 몰래 내려다보는 낯선 인물의 어두운 뒷모습 전경.\n\nLOCATION (lock): On an elevated exposed concrete remnant of the collapsed department store, overlooking the overgrown approach below. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: High concrete floor edge beside the watcher in the lower-left of the frame, foreground; Lower opening being entered by the pair in the lower-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: High concrete floor (An elevated level within the partly collapsed building, occupied by the stranger) — Its near edge crosses the lower-left foreground, with the lower approach visible beyond it; used as Establishes the watcher's elevation and separates the foreground observer from the pair below; Opening into the collapsed department store (Open and being entered by 현우 and 수빈) — Seen steeply from above in the lower-right portion of the composition; used as The shared spatial anchor that makes the stranger's surveillance and the pair's inward route readable; Moss and irregular tree growth (Spread through the ruined approach around the building); used as Connects the distant lower space to the living environment established before the watcher reveal.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight separates the readable lower approach from the stranger's darker foreground back, preserving anonymity without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The half-collapsed, moss-covered department store remains overgrown with trees and surrounded by pools of water and wildlife. The radiation reading established outside is near zero. 한수: He watches from above; he has short hair and a sharp gaze.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 한수 (북한 출신, 성인 남성, 짧게 자른 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 높은 콘크리트 층 위에서 아래의 현우와 수빈을 몰래 내려다보는 낯선 인물의 어두운 뒷모습 전경.\n\nLOCATION (lock): On an elevated exposed concrete remnant of the collapsed department store, overlooking the overgrown approach below. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: High concrete floor edge beside the watcher in the lower-left of the frame, foreground; Lower opening being entered by the pair in the lower-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: High concrete floor (An elevated level within the partly collapsed building, occupied by the stranger) — Its near edge crosses the lower-left foreground, with the lower approach visible beyond it; used as Establishes the watcher's elevation and separates the foreground observer from the pair below; Opening into the collapsed department store (Open and being entered by 현우 and 수빈) — Seen steeply from above in the lower-right portion of the composition; used as The shared spatial anchor that makes the stranger's surveillance and the pair's inward route readable; Moss and irregular tree growth (Spread through the ruined approach around the building); used as Connects the distant lower space to the living environment established before the watcher reveal.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight separates the readable lower approach from the stranger's darker foreground back, preserving anonymity without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The half-collapsed, moss-covered department store remains overgrown with trees and surrounded by pools of water and wildlife. The radiation reading established outside is near zero. 한수: He watches from above; he has short hair and a sharp gaze.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 한수 (북한 출신, 성인 남성, 짧게 자른 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 높은 콘크리트 층 위에서 아래의 현우와 수빈을 몰래 내려다보는 낯선 인물의 어두운 뒷모습 전경.\n\nLOCATION (lock): On an elevated exposed concrete remnant of the collapsed department store, overlooking the overgrown approach below. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: High concrete floor edge beside the watcher in the lower-left of the frame, foreground; Lower opening being entered by the pair in the lower-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: High concrete floor (An elevated level within the partly collapsed building, occupied by the stranger) — Its near edge crosses the lower-left foreground, with the lower approach visible beyond it; used as Establishes the watcher's elevation and separates the foreground observer from the pair below; Opening into the collapsed department store (Open and being entered by 현우 and 수빈) — Seen steeply from above in the lower-right portion of the composition; used as The shared spatial anchor that makes the stranger's surveillance and the pair's inward route readable; Moss and irregular tree growth (Spread through the ruined approach around the building); used as Connects the distant lower space to the living environment established before the watcher reveal.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight separates the readable lower approach from the stranger's darker foreground back, preserving anonymity without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The half-collapsed, moss-covered department store remains overgrown with trees and surrounded by pools of water and wildlife. The radiation reading established outside is near zero. 한수: He watches from above; he has short hair and a sharp gaze.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 한수 (북한 출신, 성인 남성, 짧게 자른 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "전경 좌측의 남성이 아래쪽 우측 하단의 입구로 향하는 두 인물을 내려다보고 있습니다.",
    "built_space": "화면 좌측과 하단을 가로지르는 부서진 고층 콘크리트 바닥이 있고, 우측 아래에는 콘크리트 기둥과 보로 이루어진 입구가 보입니다. 건물은 심하게 파괴되었고 식물과 이끼로 덮여 있습니다.",
    "entities": "전경의 인물은 짧은 검은 머리를 한 성인 남성으로 한수의 뒷모습과 일치합니다. 아래에는 현우와 수빈을 나타내는 두 명의 작은 인물이 있습니다.",
    "hard_violations": [],
    "physics": "남성은 콘크리트 가장자리에 안정적으로 걸터앉아 있으며, 아래의 두 인물은 지면을 딛고 걷고 있습니다."
   },
   {
    "label": "B",
    "direction": "전경 좌측의 남성이 아래쪽 길을 걷고 있는 두 인물을 내려다보고 있습니다.",
    "built_space": "화면 좌측 절반을 차지하는 콘크리트 구조물이 있고, 우측 중앙 상단 쪽에 건물의 사각형 입구가 있습니다. 폐허가 된 지상에는 나무와 풀이 무성합니다.",
    "entities": "전경의 인물은 한수의 머리 스타일과 참조 이미지의 짙은 남색 재킷을 입고 있습니다. 아래에는 걷고 있는 두 명의 인물이 있습니다.",
    "hard_violations": [],
    "physics": "남성은 콘크리트 바닥에 안정적으로 서 있으며, 아래의 인물들도 길 위를 자연스럽게 딛고 걷고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "프롬프트가 지정한 구도(좌측 하단의 콘크리트 바닥 모서리와 우측 하단의 하층 입구)를 정확히 구현했으며, 가파르게 내려다보는 시점과 명암 대비도 성공적으로 연출했습니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "인물의 복장은 참조 이미지와 일치하나, 프롬프트에서 우측 하단에 배치하도록 지시한 하층 입구를 화면의 상단/중앙에 배치하여 프레임 레이아웃 조건을 어겼습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "전경 좌측의 남성이 아래쪽 우측 하단의 입구로 향하는 두 인물을 내려다보고 있습니다.",
        "built_space": "화면 좌측과 하단을 가로지르는 부서진 고층 콘크리트 바닥이 있고, 우측 아래에는 콘크리트 기둥과 보로 이루어진 입구가 보입니다. 건물은 심하게 파괴되었고 식물과 이끼로 덮여 있습니다.",
        "entities": "전경의 인물은 짧은 검은 머리를 한 성인 남성으로 한수의 뒷모습과 일치합니다. 아래에는 현우와 수빈을 나타내는 두 명의 작은 인물이 있습니다.",
        "hard_violations": [],
        "physics": "남성은 콘크리트 가장자리에 안정적으로 걸터앉아 있으며, 아래의 두 인물은 지면을 딛고 걷고 있습니다."
       },
       {
        "label": "B",
        "direction": "전경 좌측의 남성이 아래쪽 길을 걷고 있는 두 인물을 내려다보고 있습니다.",
        "built_space": "화면 좌측 절반을 차지하는 콘크리트 구조물이 있고, 우측 중앙 상단 쪽에 건물의 사각형 입구가 있습니다. 폐허가 된 지상에는 나무와 풀이 무성합니다.",
        "entities": "전경의 인물은 한수의 머리 스타일과 참조 이미지의 짙은 남색 재킷을 입고 있습니다. 아래에는 걷고 있는 두 명의 인물이 있습니다.",
        "hard_violations": [],
        "physics": "남성은 콘크리트 바닥에 안정적으로 서 있으며, 아래의 인물들도 길 위를 자연스럽게 딛고 걷고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "프롬프트가 지정한 구도(좌측 하단의 콘크리트 바닥 모서리와 우측 하단의 하층 입구)를 정확히 구현했으며, 가파르게 내려다보는 시점과 명암 대비도 성공적으로 연출했습니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "인물의 복장은 참조 이미지와 일치하나, 프롬프트에서 우측 하단에 배치하도록 지시한 하층 입구를 화면의 상단/중앙에 배치하여 프레임 레이아웃 조건을 어겼습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "전경 좌측의 남성이 아래쪽 우측 하단의 입구로 향하는 두 인물을 내려다보고 있습니다.",
        "built_space": "화면 좌측과 하단을 가로지르는 부서진 고층 콘크리트 바닥이 있고, 우측 아래에는 콘크리트 기둥과 보로 이루어진 입구가 보입니다. 건물은 심하게 파괴되었고 식물과 이끼로 덮여 있습니다.",
        "entities": "전경의 인물은 짧은 검은 머리를 한 성인 남성으로 한수의 뒷모습과 일치합니다. 아래에는 현우와 수빈을 나타내는 두 명의 작은 인물이 있습니다.",
        "hard_violations": [],
        "physics": "남성은 콘크리트 가장자리에 안정적으로 걸터앉아 있으며, 아래의 두 인물은 지면을 딛고 걷고 있습니다."
       },
       {
        "label": "B",
        "direction": "전경 좌측의 남성이 아래쪽 길을 걷고 있는 두 인물을 내려다보고 있습니다.",
        "built_space": "화면 좌측 절반을 차지하는 콘크리트 구조물이 있고, 우측 중앙 상단 쪽에 건물의 사각형 입구가 있습니다. 폐허가 된 지상에는 나무와 풀이 무성합니다.",
        "entities": "전경의 인물은 한수의 머리 스타일과 참조 이미지의 짙은 남색 재킷을 입고 있습니다. 아래에는 걷고 있는 두 명의 인물이 있습니다.",
        "hard_violations": [],
        "physics": "남성은 콘크리트 바닥에 안정적으로 서 있으며, 아래의 인물들도 길 위를 자연스럽게 딛고 걷고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "왼쪽의 어두운 관찰자와 오른쪽 아래 출입구, 참조의 붕괴 외벽을 가장 충실히 배치했지만 두 사람은 안으로 들어가기보다 밖으로 나오는 방향이다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "높은 층에서 내려다보는 뒷모습은 맞지만 출입구가 오른쪽 위에 있고, 두 사람도 입구를 등지고 걸어와 지정된 배치와 진입 동작을 놓쳤다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "관찰자는 등을 보이고 머리를 오른쪽으로 돌려 아래 접근로의 두 사람 쪽을 향한다. 눈은 보이지 않아 정확한 시선은 확인할 수 없다. 아래 두 사람은 정면이 카메라 쪽으로 보이며, 뒤쪽의 어두운 출입구에서 멀어지는 보행 방향이다. 요구된 백화점 안으로의 진입과 반대다. 무기나 방향성 있는 휴대 도구는 없다.",
        "built_space": "왼쪽 전경에 노출된 상층 콘크리트 바닥 하나가 있고, 그 가장자리 너머로 아래 접근로가 내려다보인다. 관찰자는 그 바닥 가장자리 안쪽에 서 있다. 뚜렷한 지상 출입구는 하나지만 오른쪽 위에 배치되어, 요구된 오른쪽 아래 출입구 구도를 충족하지 못한다. 오른쪽에는 물웅덩이 하나와 나무들, 접근로에는 이끼와 잔해가 있다. 재료와 식생은 참조와 유사하지만 참조의 크게 꺾여 무너진 외벽은 식별되지 않는다.",
        "entities": "총 세 사람이 보이며 관찰자 한 명과 현우·수빈 역할의 아래쪽 두 사람으로 읽힌다. 관찰자는 짧은 검은 머리의 성인 남성으로, 어두운 재킷과 푸른 셔츠 깃이 한수 참조에 부합한다. 얼굴은 가려져 정확한 동일인 여부와 날카로운 눈빛은 검증할 수 없다. 아래에는 검은 상의의 짧은 머리 남성과 밝은 겉옷·청바지를 입은 긴 머리 여성이 보이지만, 작은 크기여서 세부 정체성은 판단하기 어렵다. 콘크리트 폐허, 이끼, 나무, 물은 있으나 야생동물은 식별되지 않는다. 판독 가능한 글씨나 방사선 측정 도구는 없다.",
        "hard_violations": [],
        "physics": "관찰자의 발은 화면 밖이지만 몸이 상층 바닥 위에 수직으로 이어져 정상적으로 서 있는 배치이며, 공중에 뜬 증거는 없다. 아래 두 사람은 접근로에 발을 디디고 보행하고 있다. 나무는 지면에 뿌리내리고 잔해는 바닥에 놓여 있다. 물웅덩이의 반사는 주변 나무와 구조물의 위치에 비추어 불가능해 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "관찰자는 머리를 오른쪽 아래로 숙여 출입구 앞의 두 사람을 내려다본다. 눈 자체는 숨겨져 있지만 머리와 몸의 방향은 감시 대상과 맞는다. 아래 두 사람은 카메라 쪽 정면을 드러내며 어두운 내부에서 밝은 외부 접근로 쪽으로 걸어 나오는 모습이다. 따라서 두 사람의 안쪽 진입 방향은 충족하지 못한다. 무기나 조작 중인 도구는 없다.",
        "built_space": "왼쪽 아래의 상층 바닥 하나와 그 앞을 가로지르는 낮은 콘크리트 턱 하나가 관찰자의 높이를 명확히 만든다. 관찰자는 턱 안쪽 바닥에 앉아 있다. 오른쪽 아래에는 넓은 하층 개구부 하나가 있고 두 사람이 그 앞에 있어 지정된 화면 배치에 가깝다. 맞은편에는 참조처럼 비스듬히 붕괴한 대형 콘크리트 판들과 반복되는 외벽 창구획이 보인다. 이끼와 불규칙한 나무도 이어진다. 참조의 물웅덩이는 이 구도에서 뚜렷하지 않다.",
        "entities": "관찰자 한 명과 아래쪽 두 사람으로 총 세 명이다. 관찰자는 짧은 검은 머리의 성인 남성이며 어두운 재킷을 입어 한수의 보이는 신체·복장 특징과 대체로 맞는다. 뒷모습이므로 얼굴 동일성과 눈빛, 출신은 확인할 수 없다. 아래 두 사람은 짧은 머리의 검은 상의 남성과 긴 머리의 여성으로 읽히지만 세부 얼굴과 복장은 거리 때문에 불명확하다. 붕괴한 백화점, 콘크리트 층, 출입구, 이끼와 나무가 모두 보인다. 야생동물과 측정 도구는 식별되지 않으며 읽을 수 있는 글씨는 없다.",
        "hard_violations": [],
        "physics": "관찰자는 엉덩이를 상층 바닥에 대고 앉아 있으며 굽힌 다리 일부는 낮은 콘크리트 턱에 걸쳐 지지된다. 상체를 앞으로 기울여 아래를 보는 자세는 그 지지 관계에서 가능하다. 아래 두 사람의 발은 하층 바닥에 닿아 있고 보행 자세도 물리적으로 자연스럽다. 붕괴 판재들은 다른 잔해와 구조체에 기대어 있으며, 지지 없이 떠 있는 신체나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "왼쪽의 어두운 관찰자와 오른쪽 아래 출입구, 참조의 붕괴 외벽을 가장 충실히 배치했지만 두 사람은 안으로 들어가기보다 밖으로 나오는 방향이다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "높은 층에서 내려다보는 뒷모습은 맞지만 출입구가 오른쪽 위에 있고, 두 사람도 입구를 등지고 걸어와 지정된 배치와 진입 동작을 놓쳤다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "관찰자는 등을 보이고 머리를 오른쪽으로 돌려 아래 접근로의 두 사람 쪽을 향한다. 눈은 보이지 않아 정확한 시선은 확인할 수 없다. 아래 두 사람은 정면이 카메라 쪽으로 보이며, 뒤쪽의 어두운 출입구에서 멀어지는 보행 방향이다. 요구된 백화점 안으로의 진입과 반대다. 무기나 방향성 있는 휴대 도구는 없다.",
        "built_space": "왼쪽 전경에 노출된 상층 콘크리트 바닥 하나가 있고, 그 가장자리 너머로 아래 접근로가 내려다보인다. 관찰자는 그 바닥 가장자리 안쪽에 서 있다. 뚜렷한 지상 출입구는 하나지만 오른쪽 위에 배치되어, 요구된 오른쪽 아래 출입구 구도를 충족하지 못한다. 오른쪽에는 물웅덩이 하나와 나무들, 접근로에는 이끼와 잔해가 있다. 재료와 식생은 참조와 유사하지만 참조의 크게 꺾여 무너진 외벽은 식별되지 않는다.",
        "entities": "총 세 사람이 보이며 관찰자 한 명과 현우·수빈 역할의 아래쪽 두 사람으로 읽힌다. 관찰자는 짧은 검은 머리의 성인 남성으로, 어두운 재킷과 푸른 셔츠 깃이 한수 참조에 부합한다. 얼굴은 가려져 정확한 동일인 여부와 날카로운 눈빛은 검증할 수 없다. 아래에는 검은 상의의 짧은 머리 남성과 밝은 겉옷·청바지를 입은 긴 머리 여성이 보이지만, 작은 크기여서 세부 정체성은 판단하기 어렵다. 콘크리트 폐허, 이끼, 나무, 물은 있으나 야생동물은 식별되지 않는다. 판독 가능한 글씨나 방사선 측정 도구는 없다.",
        "hard_violations": [],
        "physics": "관찰자의 발은 화면 밖이지만 몸이 상층 바닥 위에 수직으로 이어져 정상적으로 서 있는 배치이며, 공중에 뜬 증거는 없다. 아래 두 사람은 접근로에 발을 디디고 보행하고 있다. 나무는 지면에 뿌리내리고 잔해는 바닥에 놓여 있다. 물웅덩이의 반사는 주변 나무와 구조물의 위치에 비추어 불가능해 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "관찰자는 머리를 오른쪽 아래로 숙여 출입구 앞의 두 사람을 내려다본다. 눈 자체는 숨겨져 있지만 머리와 몸의 방향은 감시 대상과 맞는다. 아래 두 사람은 카메라 쪽 정면을 드러내며 어두운 내부에서 밝은 외부 접근로 쪽으로 걸어 나오는 모습이다. 따라서 두 사람의 안쪽 진입 방향은 충족하지 못한다. 무기나 조작 중인 도구는 없다.",
        "built_space": "왼쪽 아래의 상층 바닥 하나와 그 앞을 가로지르는 낮은 콘크리트 턱 하나가 관찰자의 높이를 명확히 만든다. 관찰자는 턱 안쪽 바닥에 앉아 있다. 오른쪽 아래에는 넓은 하층 개구부 하나가 있고 두 사람이 그 앞에 있어 지정된 화면 배치에 가깝다. 맞은편에는 참조처럼 비스듬히 붕괴한 대형 콘크리트 판들과 반복되는 외벽 창구획이 보인다. 이끼와 불규칙한 나무도 이어진다. 참조의 물웅덩이는 이 구도에서 뚜렷하지 않다.",
        "entities": "관찰자 한 명과 아래쪽 두 사람으로 총 세 명이다. 관찰자는 짧은 검은 머리의 성인 남성이며 어두운 재킷을 입어 한수의 보이는 신체·복장 특징과 대체로 맞는다. 뒷모습이므로 얼굴 동일성과 눈빛, 출신은 확인할 수 없다. 아래 두 사람은 짧은 머리의 검은 상의 남성과 긴 머리의 여성으로 읽히지만 세부 얼굴과 복장은 거리 때문에 불명확하다. 붕괴한 백화점, 콘크리트 층, 출입구, 이끼와 나무가 모두 보인다. 야생동물과 측정 도구는 식별되지 않으며 읽을 수 있는 글씨는 없다.",
        "hard_violations": [],
        "physics": "관찰자는 엉덩이를 상층 바닥에 대고 앉아 있으며 굽힌 다리 일부는 낮은 콘크리트 턱에 걸쳐 지지된다. 상체를 앞으로 기울여 아래를 보는 자세는 그 지지 관계에서 가능하다. 아래 두 사람의 발은 하층 바닥에 닿아 있고 보행 자세도 물리적으로 자연스럽다. 붕괴 판재들은 다른 잔해와 구조체에 기대어 있으며, 지지 없이 떠 있는 신체나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.238
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.238
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1238
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "프롬프트가 지정한 구도(좌측 하단의 콘크리트 바닥 모서리와 우측 하단의 하층 입구)를 정확히 구현했으며, 가파르게 내려다보는 시점과 명암 대비도 성공적으로 연출했습니다."
   },
   {
    "label": "B",
    "score": 1238,
    "verdict_ko": "인물의 복장은 참조 이미지와 일치하나, 프롬프트에서 우측 하단에 배치하도록 지시한 하층 입구를 화면의 상단/중앙에 배치하여 프레임 레이아웃 조건을 어겼습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S65sh10_sel.png",
    "asset_id": "a55fd836-4cf7-43f1-b9eb-9c5b28a935f0",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 한수: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1103406>",
    "asset_id": "4b12902d-ff9d-48f0-a297-fcd5fb403477",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-a246-78fd-8358-63a21e16da0d",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S65sh10"
  },
  "lane_policy": "ab_select_bypass:prev"
 },
 "S65sh20::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:02:11.997505+00:00",
  "fingerprint": "9b177882fbe05ed7b2384879b3f1aabbd71f7b5887ba956fa9478222d1f6dcab",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S65sh20_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S65sh20_sel.png",
  "source_sha256": "894eaf0d5fd2be738d46d299cd61377920f8d5ec9c04ba894899b586be7eb6d4",
  "file": "S65sh20_cine.png",
  "staged_sha256": "a9200f2e07ba4584b0ed5933dbec4fb8fc80b93310db1f97a9fce397bf62a0da",
  "latency_ms": 9432
 },
 "S66sh14::signage": {
  "fp": "dea466e63a23a3f9",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S66sh14": {
  "input_fingerprint": "5ddf22fc3b9b7515",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 현우의 시점, 어둠 속 손전등 불빛 아래로 드러난 먼지 쌓인 진열대와 수많은 통조림 캔들이 널린 지하 식품 코너 전경.\n\nLOCATION (lock): Inside the collapsed department store's underground food section, where flashlights reveal dusty shelves and scattered canned goods. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Food-section displays (Dust-covered and still retaining their recognizable layout) — Shelf fronts and oblique ends remain visible along the lateral viewing direction; used as Establish depth and the unexpectedly extensive supply of food; Numerous canned foods (Accumulated throughout the dusty food section); used as Provide repeated small-scale discoveries across the composition without isolating a single oversized object.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The flashlight reveals localized patches of dusty stock within the dark department store, with controlled contrast preserving the edges of the surrounding displays.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A newly collapsed floor leaves a high opening above the underground food department. Flashlight beams reveal dusty wine, canned food, snacks and surviving food displays in the dark interior.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 현우의 시점, 어둠 속 손전등 불빛 아래로 드러난 먼지 쌓인 진열대와 수많은 통조림 캔들이 널린 지하 식품 코너 전경.\n\nLOCATION (lock): Inside the collapsed department store's underground food section, where flashlights reveal dusty shelves and scattered canned goods. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Food-section displays (Dust-covered and still retaining their recognizable layout) — Shelf fronts and oblique ends remain visible along the lateral viewing direction; used as Establish depth and the unexpectedly extensive supply of food; Numerous canned foods (Accumulated throughout the dusty food section); used as Provide repeated small-scale discoveries across the composition without isolating a single oversized object.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The flashlight reveals localized patches of dusty stock within the dark department store, with controlled contrast preserving the edges of the surrounding displays.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A newly collapsed floor leaves a high opening above the underground food department. Flashlight beams reveal dusty wine, canned food, snacks and surviving food displays in the dark interior.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 현우의 시점, 어둠 속 손전등 불빛 아래로 드러난 먼지 쌓인 진열대와 수많은 통조림 캔들이 널린 지하 식품 코너 전경.\n\nLOCATION (lock): Inside the collapsed department store's underground food section, where flashlights reveal dusty shelves and scattered canned goods. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Food-section displays (Dust-covered and still retaining their recognizable layout) — Shelf fronts and oblique ends remain visible along the lateral viewing direction; used as Establish depth and the unexpectedly extensive supply of food; Numerous canned foods (Accumulated throughout the dusty food section); used as Provide repeated small-scale discoveries across the composition without isolating a single oversized object.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The flashlight reveals localized patches of dusty stock within the dark department store, with controlled contrast preserving the edges of the surrounding displays.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A newly collapsed floor leaves a high opening above the underground food department. Flashlight beams reveal dusty wine, canned food, snacks and surviving food displays in the dark interior.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "우측 하단의 손전등 불빛은 진열대를, 좌측 상단 허공의 손전등은 중앙을 향함.",
    "built_space": "레퍼런스 사진의 구조(에스컬레이터, 기둥)가 전혀 반영되지 않은 임의의 마트 통로임.",
    "entities": "통조림 캔, 선반, 손전등을 쥔 손.",
    "hard_violations": [
     "[gemini-pro] 지지대 없이 허공에 떠 있는 물체 (좌측 상단의 손전등 헤드)"
    ],
    "physics": "우측 손전등은 손에 쥐어져 있으나, 좌측 상단의 손전등은 아무 지지대 없이 허공에 떠 있음."
   },
   {
    "label": "B",
    "direction": "화면 밖 시점에서 비추는 손전등 불빛이 중앙의 곡선형 진열대에 정확히 닿음.",
    "built_space": "에스컬레이터와 기둥의 배치 등 레퍼런스 사진의 공간 구조를 완벽하게 일치시킴.",
    "entities": "통조림 캔, 진열대, 잔해. 기둥에 '가공식품', '신선식품' 등 읽을 수 있는 글자가 남아있음.",
    "hard_violations": [
     "[gpt-high] 오른쪽 기둥 안내판에 ‘신선식품’과 ‘FRESH FOOD’라는 읽을 수 있는 글자가 노출됩니다."
    ],
    "physics": "모든 캔과 파편들이 바닥과 선반 위에 안정적으로 위치해 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "구도를 복사하지 말라는 지시와 텍스트를 읽을 수 없게 처리하라는 지시를 어겼으나, 지정된 장소의 구조와 조명 분위기를 충실히 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "레퍼런스 장소를 무시한 임의의 공간을 생성했으며, 허공에 떠 있는 손전등이 있어 물리적 오류(Hard Violation)로 실격입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "우측 하단의 손전등 불빛은 진열대를, 좌측 상단 허공의 손전등은 중앙을 향함.",
        "built_space": "레퍼런스 사진의 구조(에스컬레이터, 기둥)가 전혀 반영되지 않은 임의의 마트 통로임.",
        "entities": "통조림 캔, 선반, 손전등을 쥔 손.",
        "hard_violations": [
         "지지대 없이 허공에 떠 있는 물체 (좌측 상단의 손전등 헤드)"
        ],
        "physics": "우측 손전등은 손에 쥐어져 있으나, 좌측 상단의 손전등은 아무 지지대 없이 허공에 떠 있음."
       },
       {
        "label": "B",
        "direction": "화면 밖 시점에서 비추는 손전등 불빛이 중앙의 곡선형 진열대에 정확히 닿음.",
        "built_space": "에스컬레이터와 기둥의 배치 등 레퍼런스 사진의 공간 구조를 완벽하게 일치시킴.",
        "entities": "통조림 캔, 진열대, 잔해. 기둥에 '가공식품', '신선식품' 등 읽을 수 있는 글자가 남아있음.",
        "hard_violations": [],
        "physics": "모든 캔과 파편들이 바닥과 선반 위에 안정적으로 위치해 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "구도를 복사하지 말라는 지시와 텍스트를 읽을 수 없게 처리하라는 지시를 어겼으나, 지정된 장소의 구조와 조명 분위기를 충실히 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "레퍼런스 장소를 무시한 임의의 공간을 생성했으며, 허공에 떠 있는 손전등이 있어 물리적 오류(Hard Violation)로 실격입니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "우측 하단의 손전등 불빛은 진열대를, 좌측 상단 허공의 손전등은 중앙을 향함.",
        "built_space": "레퍼런스 사진의 구조(에스컬레이터, 기둥)가 전혀 반영되지 않은 임의의 마트 통로임.",
        "entities": "통조림 캔, 선반, 손전등을 쥔 손.",
        "hard_violations": [
         "지지대 없이 허공에 떠 있는 물체 (좌측 상단의 손전등 헤드)"
        ],
        "physics": "우측 손전등은 손에 쥐어져 있으나, 좌측 상단의 손전등은 아무 지지대 없이 허공에 떠 있음."
       },
       {
        "label": "B",
        "direction": "화면 밖 시점에서 비추는 손전등 불빛이 중앙의 곡선형 진열대에 정확히 닿음.",
        "built_space": "에스컬레이터와 기둥의 배치 등 레퍼런스 사진의 공간 구조를 완벽하게 일치시킴.",
        "entities": "통조림 캔, 진열대, 잔해. 기둥에 '가공식품', '신선식품' 등 읽을 수 있는 글자가 남아있음.",
        "hard_violations": [],
        "physics": "모든 캔과 파편들이 바닥과 선반 위에 안정적으로 위치해 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "장소의 진열대 배치는 잘 보존했지만, 읽을 수 있는 안내 글자가 금지 조건을 위반하고 참조 사진의 위층 내려다보기 구도를 거의 답습합니다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "지하 내부의 주관적 와이드 시점과 통조림을 향하는 손전등 빛을 충실히 구현했으나, 높은 선반 중심의 배치는 참조 장소의 낮은 독립 진열대 구성과 차이가 있습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 붕괴 가장자리에서 지하 매장을 비스듬히 내려다봅니다. 밝은 조명 영역은 앞쪽 통조림 진열대와 그 뒤 중앙 진열대에 닿지만, 광원이나 손전등의 조사 방향은 직접 보이지 않습니다. 사람이나 시선은 없습니다.",
        "built_space": "왼쪽 에스컬레이터 한 대, 앞쪽과 중앙의 독립 진열대 두 개, 뒤쪽 낮은 판매대 한 개, 오른쪽 가장자리에 일부 잘린 진열대 한 개와 후면 벽 선반들이 보입니다. 위쪽에는 부서진 슬래브와 늘어진 배선이 있고, 오른쪽에는 안내판이 붙은 기둥이 있습니다. 참조 장소의 주요 배치를 상당히 유지하지만 카메라 위치도 참조의 위층 가장자리 시점과 매우 유사합니다. 반사는 없습니다.",
        "entities": "먼지가 쌓인 진열대, 다수의 통조림, 바닥에 흩어진 캔과 콘크리트 잔해가 보입니다. 사람이나 얼굴은 없습니다. 손전등 자체는 보이지 않으며, 와인과 과자는 명확히 식별하기 어렵습니다. 오른쪽 기둥의 ‘신선식품’과 ‘FRESH FOOD’는 읽을 수 있어 무문자 조건에 어긋납니다.",
        "hard_violations": [
         "오른쪽 기둥 안내판에 ‘신선식품’과 ‘FRESH FOOD’라는 읽을 수 있는 글자가 노출됩니다."
        ],
        "physics": "진열대는 바닥에 놓여 있고 캔들은 선반이나 잔해가 있는 바닥에 지지됩니다. 늘어진 배선은 상부 구조에 연결되어 있습니다. 떠 있는 물체나 지지 없는 인체는 보이지 않습니다."
       },
       {
        "label": "B",
        "direction": "카메라는 지하 매장 통로를 따라 수평에 가깝게 바라봅니다. 오른손에 쥔 손전등은 앞쪽 오른편의 쏟아진 통조림 더미와 통로 안쪽을 향하고, 왼쪽 위에서 들어오는 손전등 빛은 왼쪽 선반 상부를 비춥니다. 빛이 실제 식품 재고에 닿아 발견의 순간을 만듭니다.",
        "built_space": "왼쪽 가장자리 선반 한 줄, 왼쪽 중앙의 긴 독립 선반 한 줄, 오른쪽 전경 선반 한 줄, 오른쪽 중경 선반과 후면 선반이 보입니다. 오른쪽 중경에는 큰 콘크리트 기둥이 있고, 머리 위에는 낮빛이 들어오는 붕괴 개구부 한 곳이 있습니다. 선반 정면과 비스듬한 끝면이 함께 보여 깊이가 형성됩니다. 다만 참조의 낮은 독립 진열대 중심 배치보다 높은 통로형 선반이 지배적입니다. 반사는 없습니다.",
        "entities": "다수의 통조림, 봉지 과자류, 병 제품, 먼지 낀 금속 선반과 붕괴 잔해가 보입니다. 병의 내용물이 와인인지는 확정하기 어렵습니다. 손전등을 조작하는 오른손만 나타나며 얼굴이나 전신 인물은 없습니다. 왼쪽 위에는 두 번째 손전등의 앞부분이 걸쳐 있습니다. 확실히 읽을 수 있는 문구나 자막은 보이지 않습니다.",
        "hard_violations": [],
        "physics": "오른쪽 손전등은 손가락과 엄지가 몸통을 감싸 지지합니다. 왼쪽 위 손전등은 몸통이 화면 밖으로 이어져 파지 부위가 잘린 것으로 보이며 공중에 독립적으로 떠 있지는 않습니다. 캔 더미는 바닥과 서로 맞닿은 캔들에 지지되고, 기울어진 선반 판은 바닥과 주변 잔해에 걸쳐 있습니다. 매달린 철근과 배선은 부서진 천장에 연결되어 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "장소의 진열대 배치는 잘 보존했지만, 읽을 수 있는 안내 글자가 금지 조건을 위반하고 참조 사진의 위층 내려다보기 구도를 거의 답습합니다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "지하 내부의 주관적 와이드 시점과 통조림을 향하는 손전등 빛을 충실히 구현했으나, 높은 선반 중심의 배치는 참조 장소의 낮은 독립 진열대 구성과 차이가 있습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "카메라는 붕괴 가장자리에서 지하 매장을 비스듬히 내려다봅니다. 밝은 조명 영역은 앞쪽 통조림 진열대와 그 뒤 중앙 진열대에 닿지만, 광원이나 손전등의 조사 방향은 직접 보이지 않습니다. 사람이나 시선은 없습니다.",
        "built_space": "왼쪽 에스컬레이터 한 대, 앞쪽과 중앙의 독립 진열대 두 개, 뒤쪽 낮은 판매대 한 개, 오른쪽 가장자리에 일부 잘린 진열대 한 개와 후면 벽 선반들이 보입니다. 위쪽에는 부서진 슬래브와 늘어진 배선이 있고, 오른쪽에는 안내판이 붙은 기둥이 있습니다. 참조 장소의 주요 배치를 상당히 유지하지만 카메라 위치도 참조의 위층 가장자리 시점과 매우 유사합니다. 반사는 없습니다.",
        "entities": "먼지가 쌓인 진열대, 다수의 통조림, 바닥에 흩어진 캔과 콘크리트 잔해가 보입니다. 사람이나 얼굴은 없습니다. 손전등 자체는 보이지 않으며, 와인과 과자는 명확히 식별하기 어렵습니다. 오른쪽 기둥의 ‘신선식품’과 ‘FRESH FOOD’는 읽을 수 있어 무문자 조건에 어긋납니다.",
        "hard_violations": [
         "오른쪽 기둥 안내판에 ‘신선식품’과 ‘FRESH FOOD’라는 읽을 수 있는 글자가 노출됩니다."
        ],
        "physics": "진열대는 바닥에 놓여 있고 캔들은 선반이나 잔해가 있는 바닥에 지지됩니다. 늘어진 배선은 상부 구조에 연결되어 있습니다. 떠 있는 물체나 지지 없는 인체는 보이지 않습니다."
       },
       {
        "label": "A",
        "direction": "카메라는 지하 매장 통로를 따라 수평에 가깝게 바라봅니다. 오른손에 쥔 손전등은 앞쪽 오른편의 쏟아진 통조림 더미와 통로 안쪽을 향하고, 왼쪽 위에서 들어오는 손전등 빛은 왼쪽 선반 상부를 비춥니다. 빛이 실제 식품 재고에 닿아 발견의 순간을 만듭니다.",
        "built_space": "왼쪽 가장자리 선반 한 줄, 왼쪽 중앙의 긴 독립 선반 한 줄, 오른쪽 전경 선반 한 줄, 오른쪽 중경 선반과 후면 선반이 보입니다. 오른쪽 중경에는 큰 콘크리트 기둥이 있고, 머리 위에는 낮빛이 들어오는 붕괴 개구부 한 곳이 있습니다. 선반 정면과 비스듬한 끝면이 함께 보여 깊이가 형성됩니다. 다만 참조의 낮은 독립 진열대 중심 배치보다 높은 통로형 선반이 지배적입니다. 반사는 없습니다.",
        "entities": "다수의 통조림, 봉지 과자류, 병 제품, 먼지 낀 금속 선반과 붕괴 잔해가 보입니다. 병의 내용물이 와인인지는 확정하기 어렵습니다. 손전등을 조작하는 오른손만 나타나며 얼굴이나 전신 인물은 없습니다. 왼쪽 위에는 두 번째 손전등의 앞부분이 걸쳐 있습니다. 확실히 읽을 수 있는 문구나 자막은 보이지 않습니다.",
        "hard_violations": [],
        "physics": "오른쪽 손전등은 손가락과 엄지가 몸통을 감싸 지지합니다. 왼쪽 위 손전등은 몸통이 화면 밖으로 이어져 파지 부위가 잘린 것으로 보이며 공중에 독립적으로 떠 있지는 않습니다. 캔 더미는 바닥과 서로 맞닿은 캔들에 지지되고, 기울어진 선반 판은 바닥과 주변 잔해에 걸쳐 있습니다. 매달린 철근과 배선은 부서진 천장에 연결되어 있습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.5,
    "B": 1.375
   },
   "adjusted": {
    "A": 1.25,
    "B": 1.125
   },
   "violations": {
    "A": [
     "[gemini-pro] 지지대 없이 허공에 떠 있는 물체 (좌측 상단의 손전등 헤드)"
    ],
    "B": [
     "[gpt-high] 오른쪽 기둥 안내판에 ‘신선식품’과 ‘FRESH FOOD’라는 읽을 수 있는 글자가 노출됩니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1125,
   "A": 1250
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1125,
    "verdict_ko": "구도를 복사하지 말라는 지시와 텍스트를 읽을 수 없게 처리하라는 지시를 어겼으나, 지정된 장소의 구조와 조명 분위기를 충실히 구현했습니다.  ★위반: [gpt-high] 오른쪽 기둥 안내판에 ‘신선식품’과 ‘FRESH FOOD’라는 읽을 수 있는 글자가 노출됩니다."
   },
   {
    "label": "A",
    "score": 1250,
    "verdict_ko": "레퍼런스 장소를 무시한 임의의 공간을 생성했으며, 허공에 떠 있는 손전등이 있어 물리적 오류(Hard Violation)로 실격입니다.  ★위반: [gemini-pro] 지지대 없이 허공에 떠 있는 물체 (좌측 상단의 손전등 헤드)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L240B01.png",
    "asset_id": "75e19812-b609-48b1-a66c-c19637b5c909",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-a3f9-7b13-a337-d3ed7fac1d12",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S66sh14::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:03:23.677285+00:00",
  "fingerprint": "5295fc4c9948dcf1348f2dcebff707529e3642c60d6729b2dd79a057918a7e55",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S66sh14_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S66sh14_sel.png",
  "source_sha256": "cfd90d10256a5961bf57364a31a095ca60a7d66b589c4a7d4ee202aa6510d56a",
  "file": "S66sh14_cine.png",
  "staged_sha256": "a557ac0eb78611863af0fb71b9fa2b29ef6e67c765bbf717a26c188b6d288d93",
  "latency_ms": 9984
 },
 "S66sh26::signage": {
  "fp": "c85fb77f81b4efbf",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::9d63210811f5f1fe": {
  "subjects": [],
  "subject_text": "무너진 백화점 지상층과 붕괴 구멍 주변\n무너진 벽과 바닥 틈이 이어지는 어두운 상업시설 지상층. 먼지 쌓인 층별 안내판과 아래층으로 향하는 에스컬레이터가 보인다.",
  "identity": "canonical",
  "scope_id": "L240",
  "scope_role": "location_interior",
  "scope_sha": "cf1e843ecc3e18c5"
 },
 "S66sh26::bgfirst_bg": {
  "input_fingerprint": "a9ea5ff6bc1cbb2b",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 천장 붕괴 구멍 위쪽 가장자리로 소리 없이 다가선 낯선 남자의 짙은 실루엣 전경.\n\nLOCATION (lock): Inside the ruined department store's upper sales level, at the edge of the floor collapse above the dark food section.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Collapsed floor opening (Open following the collapse that dropped 현우 and 수빈 below) — The near upper-floor rim is seen obliquely from behind the approaching man; used as Connect the concealed observer to the lower-floor space without revealing its occupants.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the department-store interior dim, separating the dark rear silhouette from the opening through restrained tonal contrast without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 천장 붕괴 구멍 위쪽 가장자리로 소리 없이 다가선 낯선 남자의 짙은 실루엣 전경.\n\nLOCATION (lock): Inside the ruined department store's upper sales level, at the edge of the floor collapse above the dark food section.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Collapsed floor opening (Open following the collapse that dropped 현우 and 수빈 below) — The near upper-floor rim is seen obliquely from behind the approaching man; used as Connect the concealed observer to the lower-floor space without revealing its occupants.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the department-store interior dim, separating the dark rear silhouette from the opening through restrained tonal contrast without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S66sh26__bgfirst_bg.png",
  "asset_id": "46131388-6cc7-4fff-87df-cd5907e8b98a",
  "input_asset_ids": [
   "a6150871-4cda-4118-b344-d3f6a8f9efc5",
   "75e19812-b609-48b1-a66c-c19637b5c909"
  ]
 },
 "S66sh26": {
  "input_fingerprint": "6f13ce6f114ee4fe",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 천장 붕괴 구멍 위쪽 가장자리로 소리 없이 다가선 낯선 남자의 짙은 실루엣 전경.\n\nLOCATION (lock): Inside the ruined department store's upper sales level, at the edge of the floor collapse above the dark food section. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Collapsed floor opening (Open following the collapse that dropped 현우 and 수빈 below) — The near upper-floor rim is seen obliquely from behind the approaching man; used as Connect the concealed observer to the lower-floor space without revealing its occupants.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the department-store interior dim, separating the dark rear silhouette from the opening through restrained tonal contrast without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The collapse opening remains far above the food department, and the other exits are blocked. Dusty food stock remains in place, with canned food opened during the wait. 한수: He stands near the upper edge of the collapse opening, with short hair and a sharp gaze.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 한수 (북한 출신, 성인 남성, 짧게 자른 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 천장 붕괴 구멍 위쪽 가장자리로 소리 없이 다가선 낯선 남자의 짙은 실루엣 전경.\n\nLOCATION (lock): Inside the ruined department store's upper sales level, at the edge of the floor collapse above the dark food section. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Collapsed floor opening (Open following the collapse that dropped 현우 and 수빈 below) — The near upper-floor rim is seen obliquely from behind the approaching man; used as Connect the concealed observer to the lower-floor space without revealing its occupants.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the department-store interior dim, separating the dark rear silhouette from the opening through restrained tonal contrast without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The collapse opening remains far above the food department, and the other exits are blocked. Dusty food stock remains in place, with canned food opened during the wait. 한수: He stands near the upper edge of the collapse opening, with short hair and a sharp gaze.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 한수 (북한 출신, 성인 남성, 짧게 자른 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 천장 붕괴 구멍 위쪽 가장자리로 소리 없이 다가선 낯선 남자의 짙은 실루엣 전경.\n\nLOCATION (lock): Inside the ruined department store's upper sales level, at the edge of the floor collapse above the dark food section. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Collapsed floor opening (Open following the collapse that dropped 현우 and 수빈 below) — The near upper-floor rim is seen obliquely from behind the approaching man; used as Connect the concealed observer to the lower-floor space without revealing its occupants.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the department-store interior dim, separating the dark rear silhouette from the opening through restrained tonal contrast without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The collapse opening remains far above the food department, and the other exits are blocked. Dusty food stock remains in place, with canned food opened during the wait. 한수: He stands near the upper edge of the collapse opening, with short hair and a sharp gaze.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 한수 (북한 출신, 성인 남성, 짧게 자른 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S66sh26__bgfirst_bg.png",
     "asset_id": "46131388-6cc7-4fff-87df-cd5907e8b98a",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S66sh26.png",
     "asset_id": "a6150871-4cda-4118-b344-d3f6a8f9efc5",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 한수: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1103406>",
     "asset_id": "4b12902d-ff9d-48f0-a297-fcd5fb403477",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L240B01.png",
     "asset_id": "75e19812-b609-48b1-a66c-c19637b5c909",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 한수: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1103406>",
     "asset_id": "4b12902d-ff9d-48f0-a297-fcd5fb403477",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "남자가 앞쪽의 붕괴된 층 구멍을 내려다보며 서 있음.",
    "built_space": "무너진 2층 바닥 가장자리와 에스컬레이터, 아래층 식료품 매장이 레퍼런스와 유사하게 배치되어 있음.",
    "entities": "지시된 레퍼런스와 일치하는 짧은 머리의 성인 남성(한수)이 어두운 옷을 입고 있음.",
    "hard_violations": [
     "[gpt-high] 장소 참조에 하나인 에스컬레이터를 별도 난간과 계단열을 가진 두 줄로 늘려 고정 시설을 중복시켰다."
    ],
    "physics": "남자의 두 발이 2층 바닥 가장자리에 단단히 딛고 있어 체중이 정상적으로 지탱됨."
   },
   {
    "label": "B",
    "direction": "인물의 시선 방향을 알 수 없으며, 붕괴 구멍 위쪽 공중에 떠 있음.",
    "built_space": "레퍼런스 사진의 백화점 내부 구조와 간판이 거의 동일하게 복제됨.",
    "entities": "인물이 구체적인 인체 묘사 없이 완전히 2D 평면의 검은 단색 실루엣으로 처리됨.",
    "hard_violations": [
     "[gemini-pro] 물리적 지지대 없이 허공에 떠 있는 인물",
     "[gemini-pro] 식별 가능한 텍스트(간판 글씨) 노출",
     "[gemini-pro] 인체를 완전한 평면 실루엣(검은 덩어리)으로 축소함",
     "[gpt-high] 한수가 지정된 상층 가장자리가 아닌 붕괴 구멍 내부에 배치되어 있다.",
     "[gpt-high] 남자의 발에 닿는 바닥이나 받침이 보이지 않으며, 상층 가장자리와 떨어진 공중의 몸을 지지하는 것이 없다.",
     "[gpt-high] 안내판과 벽 광고에 읽을 수 있는 글자가 남아 있다."
    ],
    "physics": "인물의 발이 바닥에 닿지 않고 무너진 구멍 위 허공에 아무런 지지 없이 떠 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "레퍼런스의 공간을 적절히 재구성하여 텍스트 노출을 방지했으며, 지시된 가장자리에 서 있는 인물을 물리적으로 자연스럽게 구현하여 프롬프트를 충실히 따름."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "인물이 물리적 지지 없이 허공에 떠 있고 완전히 평면적인 검은 덩어리로 묘사되었으며, 식별 가능한 텍스트가 그대로 노출되어 다수의 치명적 위반이 발생함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "남자가 앞쪽의 붕괴된 층 구멍을 내려다보며 서 있음.",
        "built_space": "무너진 2층 바닥 가장자리와 에스컬레이터, 아래층 식료품 매장이 레퍼런스와 유사하게 배치되어 있음.",
        "entities": "지시된 레퍼런스와 일치하는 짧은 머리의 성인 남성(한수)이 어두운 옷을 입고 있음.",
        "hard_violations": [],
        "physics": "남자의 두 발이 2층 바닥 가장자리에 단단히 딛고 있어 체중이 정상적으로 지탱됨."
       },
       {
        "label": "B",
        "direction": "인물의 시선 방향을 알 수 없으며, 붕괴 구멍 위쪽 공중에 떠 있음.",
        "built_space": "레퍼런스 사진의 백화점 내부 구조와 간판이 거의 동일하게 복제됨.",
        "entities": "인물이 구체적인 인체 묘사 없이 완전히 2D 평면의 검은 단색 실루엣으로 처리됨.",
        "hard_violations": [
         "물리적 지지대 없이 허공에 떠 있는 인물",
         "식별 가능한 텍스트(간판 글씨) 노출",
         "인체를 완전한 평면 실루엣(검은 덩어리)으로 축소함"
        ],
        "physics": "인물의 발이 바닥에 닿지 않고 무너진 구멍 위 허공에 아무런 지지 없이 떠 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "레퍼런스의 공간을 적절히 재구성하여 텍스트 노출을 방지했으며, 지시된 가장자리에 서 있는 인물을 물리적으로 자연스럽게 구현하여 프롬프트를 충실히 따름."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "인물이 물리적 지지 없이 허공에 떠 있고 완전히 평면적인 검은 덩어리로 묘사되었으며, 식별 가능한 텍스트가 그대로 노출되어 다수의 치명적 위반이 발생함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "남자가 앞쪽의 붕괴된 층 구멍을 내려다보며 서 있음.",
        "built_space": "무너진 2층 바닥 가장자리와 에스컬레이터, 아래층 식료품 매장이 레퍼런스와 유사하게 배치되어 있음.",
        "entities": "지시된 레퍼런스와 일치하는 짧은 머리의 성인 남성(한수)이 어두운 옷을 입고 있음.",
        "hard_violations": [],
        "physics": "남자의 두 발이 2층 바닥 가장자리에 단단히 딛고 있어 체중이 정상적으로 지탱됨."
       },
       {
        "label": "B",
        "direction": "인물의 시선 방향을 알 수 없으며, 붕괴 구멍 위쪽 공중에 떠 있음.",
        "built_space": "레퍼런스 사진의 백화점 내부 구조와 간판이 거의 동일하게 복제됨.",
        "entities": "인물이 구체적인 인체 묘사 없이 완전히 2D 평면의 검은 단색 실루엣으로 처리됨.",
        "hard_violations": [
         "물리적 지지대 없이 허공에 떠 있는 인물",
         "식별 가능한 텍스트(간판 글씨) 노출",
         "인체를 완전한 평면 실루엣(검은 덩어리)으로 축소함"
        ],
        "physics": "인물의 발이 바닥에 닿지 않고 무너진 구멍 위 허공에 아무런 지지 없이 떠 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 1,
        "verdict_ko": "한수가 상층 가장자리 전경이 아니라 구멍 안에 지지 없이 떠 있는 검은 형상으로 배치되었고, 읽을 수 있는 간판 문구까지 남아 핵심 연출을 위반한다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "상층 바닥에 선 한수의 어두운 뒷모습과 아래를 살피는 와이드 구도는 맞지만, 참조에 하나뿐인 에스컬레이터를 두 줄로 만들어 장소의 고정 구조를 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "남자의 머리는 화면 왼쪽 아래, 에스컬레이터 하단 부근을 향한다. 얼굴과 눈은 검게 묻혀 정확한 시선은 확인할 수 없다. 가까운 상층 가장자리에서 하층을 내려다보는 관찰자의 방향 관계가 아니라, 구멍 안쪽에 옆모습으로 놓인 인물처럼 보인다. 무기나 손에 든 물건은 없다.",
        "built_space": "왼쪽 에스컬레이터 한 줄, 중앙의 큰 붕괴 개구부 하나, 왼쪽 독립 안내판 하나, 아래층의 상품 진열대들과 기둥들이 보이며 건축 배치는 참조와 가깝다. 그러나 가까운 상층 바닥은 비어 있고 남자는 개구부 안쪽 하층 공간과 겹쳐 있다. 요구된 상층 가장자리의 전경 인물 배치가 아니다.",
        "entities": "사람은 한 명이며 짧은 머리의 성인 남성 윤곽은 보인다. 몸과 옷이 거의 단색 검정으로 뭉쳐 한수의 얼굴, 체격 세부, 남색 재킷과 파란 셔츠는 확인되지 않는다. 식품 선반과 통조림은 보이고 다른 사람은 없다. 왼쪽 안내판의 층수와 품목, 상단 벽 광고의 한국어가 읽혀 무문자 조건에 어긋난다.",
        "hard_violations": [
         "한수가 지정된 상층 가장자리가 아닌 붕괴 구멍 내부에 배치되어 있다.",
         "남자의 발에 닿는 바닥이나 받침이 보이지 않으며, 상층 가장자리와 떨어진 공중의 몸을 지지하는 것이 없다.",
         "안내판과 벽 광고에 읽을 수 있는 글자가 남아 있다."
        ],
        "physics": "남자의 발은 아래층 후면 진열 공간 위에 겹쳐 보이지만 발바닥과 지면의 접촉이 확인되지 않는다. 상층 슬래브는 인물 앞쪽과 양옆에 떨어져 있고, 손으로 잡은 지지물도 없다. 도약이나 착지 동작도 아니므로 몸을 떠받치는 근거가 없다. 진열대는 아래층 바닥에 놓여 있고 늘어진 선들은 파손된 구조체에 연결되어 있다."
       },
       {
        "label": "B",
        "direction": "남자는 카메라에 등을 비스듬히 보이고 머리를 오른쪽 아래로 숙여 붕괴 구멍 너머의 어두운 식품 매장을 살핀다. 정확한 눈동자는 보이지 않지만 머리와 몸의 방향은 요구된 관찰 대상을 향한다. 카메라를 보지 않으며 손에 든 물건은 없다.",
        "built_space": "카메라는 상층 바닥에서 남자의 뒤쪽을 보고, 남자는 왼쪽 전경의 가까운 붕괴 가장자리에 서 있다. 중앙과 오른쪽으로 큰 개구부 하나와 아래층 선반들이 펼쳐진다. 다만 남자 뒤·왼쪽의 계단열과 그 오른쪽의 별도 계단열에 각각 난간이 있어 에스컬레이터가 두 줄로 보인다. 참조의 한 줄짜리 에스컬레이터 구조와 다르다. 창, 기둥, 파손된 콘크리트와 노출 배선은 같은 장소의 주요 재료와 분위기를 따른다.",
        "entities": "짧은 검은 머리의 동아시아계 성인 남성 한 명이 보이며, 다른 인물은 드러나지 않는다. 얼굴은 측후면으로 가려져 참조 인물과의 정확한 얼굴 일치는 확인하기 어렵다. 어두운 바지와 신발을 착용했지만 상의는 참조의 남색 정장 재킷보다 짧은 작업용 재킷에 가깝다. 아래층에는 식품 선반과 상자가 있고, 통조림을 열어 둔 상태는 이 거리에서 확인되지 않는다. 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [
         "장소 참조에 하나인 에스컬레이터를 별도 난간과 계단열을 가진 두 줄로 늘려 고정 시설을 중복시켰다."
        ],
        "physics": "남자의 두 신발은 무너지지 않은 상층 타일 바닥에 닿아 몸을 지지한다. 발을 약간 벌리고 가장자리에서 멈춘 자세로, 공중 부양이나 불가능한 관절은 없다. 손은 몸 옆에 자연스럽게 놓여 있고 재킷의 주름도 입체적으로 보인다. 아래층 선반과 상자는 바닥에 놓이며 늘어진 배선은 위쪽 구조체에 매달려 있다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "한수가 상층 가장자리 전경이 아니라 구멍 안에 지지 없이 떠 있는 검은 형상으로 배치되었고, 읽을 수 있는 간판 문구까지 남아 핵심 연출을 위반한다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "상층 바닥에 선 한수의 어두운 뒷모습과 아래를 살피는 와이드 구도는 맞지만, 참조에 하나뿐인 에스컬레이터를 두 줄로 만들어 장소의 고정 구조를 위반한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "남자의 머리는 화면 왼쪽 아래, 에스컬레이터 하단 부근을 향한다. 얼굴과 눈은 검게 묻혀 정확한 시선은 확인할 수 없다. 가까운 상층 가장자리에서 하층을 내려다보는 관찰자의 방향 관계가 아니라, 구멍 안쪽에 옆모습으로 놓인 인물처럼 보인다. 무기나 손에 든 물건은 없다.",
        "built_space": "왼쪽 에스컬레이터 한 줄, 중앙의 큰 붕괴 개구부 하나, 왼쪽 독립 안내판 하나, 아래층의 상품 진열대들과 기둥들이 보이며 건축 배치는 참조와 가깝다. 그러나 가까운 상층 바닥은 비어 있고 남자는 개구부 안쪽 하층 공간과 겹쳐 있다. 요구된 상층 가장자리의 전경 인물 배치가 아니다.",
        "entities": "사람은 한 명이며 짧은 머리의 성인 남성 윤곽은 보인다. 몸과 옷이 거의 단색 검정으로 뭉쳐 한수의 얼굴, 체격 세부, 남색 재킷과 파란 셔츠는 확인되지 않는다. 식품 선반과 통조림은 보이고 다른 사람은 없다. 왼쪽 안내판의 층수와 품목, 상단 벽 광고의 한국어가 읽혀 무문자 조건에 어긋난다.",
        "hard_violations": [
         "한수가 지정된 상층 가장자리가 아닌 붕괴 구멍 내부에 배치되어 있다.",
         "남자의 발에 닿는 바닥이나 받침이 보이지 않으며, 상층 가장자리와 떨어진 공중의 몸을 지지하는 것이 없다.",
         "안내판과 벽 광고에 읽을 수 있는 글자가 남아 있다."
        ],
        "physics": "남자의 발은 아래층 후면 진열 공간 위에 겹쳐 보이지만 발바닥과 지면의 접촉이 확인되지 않는다. 상층 슬래브는 인물 앞쪽과 양옆에 떨어져 있고, 손으로 잡은 지지물도 없다. 도약이나 착지 동작도 아니므로 몸을 떠받치는 근거가 없다. 진열대는 아래층 바닥에 놓여 있고 늘어진 선들은 파손된 구조체에 연결되어 있다."
       },
       {
        "label": "A",
        "direction": "남자는 카메라에 등을 비스듬히 보이고 머리를 오른쪽 아래로 숙여 붕괴 구멍 너머의 어두운 식품 매장을 살핀다. 정확한 눈동자는 보이지 않지만 머리와 몸의 방향은 요구된 관찰 대상을 향한다. 카메라를 보지 않으며 손에 든 물건은 없다.",
        "built_space": "카메라는 상층 바닥에서 남자의 뒤쪽을 보고, 남자는 왼쪽 전경의 가까운 붕괴 가장자리에 서 있다. 중앙과 오른쪽으로 큰 개구부 하나와 아래층 선반들이 펼쳐진다. 다만 남자 뒤·왼쪽의 계단열과 그 오른쪽의 별도 계단열에 각각 난간이 있어 에스컬레이터가 두 줄로 보인다. 참조의 한 줄짜리 에스컬레이터 구조와 다르다. 창, 기둥, 파손된 콘크리트와 노출 배선은 같은 장소의 주요 재료와 분위기를 따른다.",
        "entities": "짧은 검은 머리의 동아시아계 성인 남성 한 명이 보이며, 다른 인물은 드러나지 않는다. 얼굴은 측후면으로 가려져 참조 인물과의 정확한 얼굴 일치는 확인하기 어렵다. 어두운 바지와 신발을 착용했지만 상의는 참조의 남색 정장 재킷보다 짧은 작업용 재킷에 가깝다. 아래층에는 식품 선반과 상자가 있고, 통조림을 열어 둔 상태는 이 거리에서 확인되지 않는다. 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [
         "장소 참조에 하나인 에스컬레이터를 별도 난간과 계단열을 가진 두 줄로 늘려 고정 시설을 중복시켰다."
        ],
        "physics": "남자의 두 신발은 무너지지 않은 상층 타일 바닥에 닿아 몸을 지지한다. 발을 약간 벌리고 가장자리에서 멈춘 자세로, 공중 부양이나 불가능한 관절은 없다. 손은 몸 옆에 자연스럽게 놓여 있고 재킷의 주름도 입체적으로 보인다. 아래층 선반과 상자는 바닥에 놓이며 늘어진 배선은 위쪽 구조체에 매달려 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.679
   },
   "adjusted": {
    "A": 1.75,
    "B": 0.429
   },
   "violations": {
    "B": [
     "[gemini-pro] 물리적 지지대 없이 허공에 떠 있는 인물",
     "[gemini-pro] 식별 가능한 텍스트(간판 글씨) 노출",
     "[gemini-pro] 인체를 완전한 평면 실루엣(검은 덩어리)으로 축소함",
     "[gpt-high] 한수가 지정된 상층 가장자리가 아닌 붕괴 구멍 내부에 배치되어 있다.",
     "[gpt-high] 남자의 발에 닿는 바닥이나 받침이 보이지 않으며, 상층 가장자리와 떨어진 공중의 몸을 지지하는 것이 없다.",
     "[gpt-high] 안내판과 벽 광고에 읽을 수 있는 글자가 남아 있다."
    ],
    "A": [
     "[gpt-high] 장소 참조에 하나인 에스컬레이터를 별도 난간과 계단열을 가진 두 줄로 늘려 고정 시설을 중복시켰다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 429
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "레퍼런스의 공간을 적절히 재구성하여 텍스트 노출을 방지했으며, 지시된 가장자리에 서 있는 인물을 물리적으로 자연스럽게 구현하여 프롬프트를 충실히 따름.  ★위반: [gpt-high] 장소 참조에 하나인 에스컬레이터를 별도 난간과 계단열을 가진 두 줄로 늘려 고정 시설을 중복시켰다."
   },
   {
    "label": "B",
    "score": 429,
    "verdict_ko": "인물이 물리적 지지 없이 허공에 떠 있고 완전히 평면적인 검은 덩어리로 묘사되었으며, 식별 가능한 텍스트가 그대로 노출되어 다수의 치명적 위반이 발생함.  ★위반: [gemini-pro] 물리적 지지대 없이 허공에 떠 있는 인물 / [gemini-pro] 식별 가능한 텍스트(간판 글씨) 노출 / [gemini-pro] 인체를 완전한 평면 실루엣(검은 덩어리)으로 축소함 / [gpt-high] 한수가 지정된 상층 가장자리가 아닌 붕괴 구멍 내부에 배치되어 있다. / [gpt-high] 남자의 발에 닿는 바닥이나 받침이 보이지 않으며, 상층 가장자리와 떨어진 공중의 몸을 지지하는 것이 없다. / [gpt-high] 안내판과 벽 광고에 읽을 수 있는 글자가 남아 있다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L240B01.png",
    "asset_id": "75e19812-b609-48b1-a66c-c19637b5c909",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 한수: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1103406>",
    "asset_id": "4b12902d-ff9d-48f0-a297-fcd5fb403477",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-a5a4-7f97-bca9-d02c4f362bf8",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S66sh26__bgfirst_bg.png",
   "bg_asset_id": "46131388-6cc7-4fff-87df-cd5907e8b98a",
   "bg_record_key": "S66sh26::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S66sh26::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:05:08.148470+00:00",
  "fingerprint": "b19d2f3e9843bf362804a37169ba8613887293a4a337c4e9b61fabb33a3afdbe",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S66sh26_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S66sh26_sel.png",
  "source_sha256": "e0e5bd0021301ca49d1e157a326d8a357d84bf0c6dc42132e14f9fede0e29a7b",
  "file": "S66sh26_cine.png",
  "staged_sha256": "64113a3bca3700ebf6cb1f25eb2f0b9843153293a46863d878a569f8d9cf272d",
  "latency_ms": 8415
 },
 "S66sh36::signage": {
  "fp": "58aa5cc5a108d959",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S66sh36": {
  "input_fingerprint": "e3b4255fb3f4a592",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 현우의 시점, 날카로운 눈빛에 짧은 머리를 한 한수가 밧줄을 쥔 채 내려다보는 굳은 전신.\n\nLOCATION (lock): At the upper edge of the department store's collapse shaft, inside the dim ruined sales floor where the rescue rope is anchored. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Rescue rope (Held by 한수 and extending into the opening) — Runs obliquely from his hand toward the lower-left edge; used as Link the revealed rescuer to the unseen climber without showing the POV character; Upper-floor opening rim (Collapsed and open) — A narrow portion of the upper edge is visible at the bottom of the low viewpoint; used as Ground the emerging subjective perspective and establish separation from 한수.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the dim interior ambience while allowing enough facial separation to reveal 한수's sharp expression without adding a new source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 한수 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the broken concrete floor, rubble, and edges of the collapse opening. Exclude the basement food shelves and canned goods as furnishings of this upper-level location.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A rescue rope hangs through the collapse opening into the food department. The opened parasol remains on the lower level, where rain has dripped through from above. 한수: He stands at the upper rescue point, with short hair and a sharp gaze.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 한수 right now, so 한수's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 한수: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 한수 (북한 출신, 성인 남성, 짧게 자른 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 현우의 시점, 날카로운 눈빛에 짧은 머리를 한 한수가 밧줄을 쥔 채 내려다보는 굳은 전신.\n\nLOCATION (lock): At the upper edge of the department store's collapse shaft, inside the dim ruined sales floor where the rescue rope is anchored. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Rescue rope (Held by 한수 and extending into the opening) — Runs obliquely from his hand toward the lower-left edge; used as Link the revealed rescuer to the unseen climber without showing the POV character; Upper-floor opening rim (Collapsed and open) — A narrow portion of the upper edge is visible at the bottom of the low viewpoint; used as Ground the emerging subjective perspective and establish separation from 한수.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the dim interior ambience while allowing enough facial separation to reveal 한수's sharp expression without adding a new source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 한수 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the broken concrete floor, rubble, and edges of the collapse opening. Exclude the basement food shelves and canned goods as furnishings of this upper-level location.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A rescue rope hangs through the collapse opening into the food department. The opened parasol remains on the lower level, where rain has dripped through from above. 한수: He stands at the upper rescue point, with short hair and a sharp gaze.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 한수 right now, so 한수's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 한수: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 한수 (북한 출신, 성인 남성, 짧게 자른 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 현우의 시점, 날카로운 눈빛에 짧은 머리를 한 한수가 밧줄을 쥔 채 내려다보는 굳은 전신.\n\nLOCATION (lock): At the upper edge of the department store's collapse shaft, inside the dim ruined sales floor where the rescue rope is anchored. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Rescue rope (Held by 한수 and extending into the opening) — Runs obliquely from his hand toward the lower-left edge; used as Link the revealed rescuer to the unseen climber without showing the POV character; Upper-floor opening rim (Collapsed and open) — A narrow portion of the upper edge is visible at the bottom of the low viewpoint; used as Ground the emerging subjective perspective and establish separation from 한수.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the dim interior ambience while allowing enough facial separation to reveal 한수's sharp expression without adding a new source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 한수 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the broken concrete floor, rubble, and edges of the collapse opening. Exclude the basement food shelves and canned goods as furnishings of this upper-level location.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A rescue rope hangs through the collapse opening into the food department. The opened parasol remains on the lower level, where rain has dripped through from above. 한수: He stands at the upper rescue point, with short hair and a sharp gaze.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 한수 right now, so 한수's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 한수: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 한수 (북한 출신, 성인 남성, 짧게 자른 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "한수는 붕괴된 구멍 아래쪽을 내려다보고 있으며, 밧줄은 그의 손에서 좌측 하단으로 비스듬히 향한다.",
    "built_space": "카메라는 위층에 위치해 구멍 건너편을 바라보며, 프레임 하단에 부서진 바닥의 좁은 테두리가 명확히 보인다. 레퍼런스와 일치하는 타일 바닥과 기둥이 존재한다.",
    "entities": "짧은 머리와 날카로운 눈빛을 가진 한수가 레퍼런스와 정확히 일치하는 의상을 입고 구조 밧줄을 쥐고 있다.",
    "hard_violations": [
     "[gpt-high] 구멍에서 올라오는 현우의 낮은 시점 대신 한수와 상층 바닥을 위에서 내려다보는 시점으로 촬영하여, 명시된 구조물 내 카메라 위치를 뒤집었다."
    ],
    "physics": "한수의 두 발이 바닥을 단단히 딛고 체중을 지탱하며, 두 손은 밧줄을 현실적으로 쥐고 있다. 팽팽한 밧줄은 구멍으로 이어지고 느슨한 부분은 뒤쪽 바닥에 놓여 있다."
   },
   {
    "label": "B",
    "direction": "한수는 카메라 렌즈를 향해 아래를 내려다보고 있으며, 밧줄은 그의 손에서 좌측 하단으로 향한다.",
    "built_space": "카메라는 위층 아래 구멍 내부에 위치해 위를 올려다보는 구도이며, 부서진 슬래브가 프레임 하단이 아닌 중앙을 크게 가로지르고 있다.",
    "entities": "한수가 레퍼런스와 일치하는 의상을 입고 밧줄을 쥐고 있다.",
    "hard_violations": [
     "[gemini-pro] 카메라 위치 및 공간 위반: 프롬프트에 지정된 위층 매장 내부가 아닌 붕괴된 구멍 아래 허공에서 위를 올려다보는 시점으로 렌더링됨.",
     "[gemini-pro] 프레이밍 위반: 바닥 가장자리가 화면 하단에 좁게 위치해야 한다는 지시를 무시하고 화면 중앙을 가로지르도록 배치됨."
    ],
    "physics": "오른발이 슬래브를 딛고 체중을 지탱하며 밧줄은 아래로 늘어져 있으나, 밧줄을 쥔 손의 형태가 명확하지 않고 밧줄과 뭉개져 있다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 프레이밍과 카메라 위치를 정확히 구현했으며, 인물 묘사와 공간의 물리적 사실성이 뛰어납니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "카메라를 구멍 아래에 배치하여 공간 위치 잠금을 위반했으며, 지정된 프레이밍 구성에도 어긋납니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "한수는 붕괴된 구멍 아래쪽을 내려다보고 있으며, 밧줄은 그의 손에서 좌측 하단으로 비스듬히 향한다.",
        "built_space": "카메라는 위층에 위치해 구멍 건너편을 바라보며, 프레임 하단에 부서진 바닥의 좁은 테두리가 명확히 보인다. 레퍼런스와 일치하는 타일 바닥과 기둥이 존재한다.",
        "entities": "짧은 머리와 날카로운 눈빛을 가진 한수가 레퍼런스와 정확히 일치하는 의상을 입고 구조 밧줄을 쥐고 있다.",
        "hard_violations": [],
        "physics": "한수의 두 발이 바닥을 단단히 딛고 체중을 지탱하며, 두 손은 밧줄을 현실적으로 쥐고 있다. 팽팽한 밧줄은 구멍으로 이어지고 느슨한 부분은 뒤쪽 바닥에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "한수는 카메라 렌즈를 향해 아래를 내려다보고 있으며, 밧줄은 그의 손에서 좌측 하단으로 향한다.",
        "built_space": "카메라는 위층 아래 구멍 내부에 위치해 위를 올려다보는 구도이며, 부서진 슬래브가 프레임 하단이 아닌 중앙을 크게 가로지르고 있다.",
        "entities": "한수가 레퍼런스와 일치하는 의상을 입고 밧줄을 쥐고 있다.",
        "hard_violations": [
         "카메라 위치 및 공간 위반: 프롬프트에 지정된 위층 매장 내부가 아닌 붕괴된 구멍 아래 허공에서 위를 올려다보는 시점으로 렌더링됨.",
         "프레이밍 위반: 바닥 가장자리가 화면 하단에 좁게 위치해야 한다는 지시를 무시하고 화면 중앙을 가로지르도록 배치됨."
        ],
        "physics": "오른발이 슬래브를 딛고 체중을 지탱하며 밧줄은 아래로 늘어져 있으나, 밧줄을 쥔 손의 형태가 명확하지 않고 밧줄과 뭉개져 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지정된 프레이밍과 카메라 위치를 정확히 구현했으며, 인물 묘사와 공간의 물리적 사실성이 뛰어납니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "카메라를 구멍 아래에 배치하여 공간 위치 잠금을 위반했으며, 지정된 프레이밍 구성에도 어긋납니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "한수는 붕괴된 구멍 아래쪽을 내려다보고 있으며, 밧줄은 그의 손에서 좌측 하단으로 비스듬히 향한다.",
        "built_space": "카메라는 위층에 위치해 구멍 건너편을 바라보며, 프레임 하단에 부서진 바닥의 좁은 테두리가 명확히 보인다. 레퍼런스와 일치하는 타일 바닥과 기둥이 존재한다.",
        "entities": "짧은 머리와 날카로운 눈빛을 가진 한수가 레퍼런스와 정확히 일치하는 의상을 입고 구조 밧줄을 쥐고 있다.",
        "hard_violations": [],
        "physics": "한수의 두 발이 바닥을 단단히 딛고 체중을 지탱하며, 두 손은 밧줄을 현실적으로 쥐고 있다. 팽팽한 밧줄은 구멍으로 이어지고 느슨한 부분은 뒤쪽 바닥에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "한수는 카메라 렌즈를 향해 아래를 내려다보고 있으며, 밧줄은 그의 손에서 좌측 하단으로 향한다.",
        "built_space": "카메라는 위층 아래 구멍 내부에 위치해 위를 올려다보는 구도이며, 부서진 슬래브가 프레임 하단이 아닌 중앙을 크게 가로지르고 있다.",
        "entities": "한수가 레퍼런스와 일치하는 의상을 입고 밧줄을 쥐고 있다.",
        "hard_violations": [
         "카메라 위치 및 공간 위반: 프롬프트에 지정된 위층 매장 내부가 아닌 붕괴된 구멍 아래 허공에서 위를 올려다보는 시점으로 렌더링됨.",
         "프레이밍 위반: 바닥 가장자리가 화면 하단에 좁게 위치해야 한다는 지시를 무시하고 화면 중앙을 가로지르도록 배치됨."
        ],
        "physics": "오른발이 슬래브를 딛고 체중을 지탱하며 밧줄은 아래로 늘어져 있으나, 밧줄을 쥔 손의 형태가 명확하지 않고 밧줄과 뭉개져 있다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "아래의 현우를 내려다보는 한수와 왼쪽 아래로 이어지는 밧줄은 정확하지만, 카메라가 너무 깊어 붕괴 가장자리가 화면 하단의 좁은 부분이 아니라 중앙을 크게 차지한다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "전신과 밧줄을 잡은 손은 분명하지만, 카메라가 한수보다 높고 한수가 위를 올려다보므로 현우의 낮은 주관 시점과 내려다보는 행동을 반대로 구현했다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "한수는 고개와 눈을 아래쪽 카메라, 즉 구멍 속 현우가 있을 방향으로 향한다. 손에 잡힌 밧줄은 왼쪽 아래 화면 끝으로 비스듬히 이어져 보이지 않는 등반자와 연결되는 방향이 맞는다.",
        "built_space": "한수는 붕괴된 상층 콘크리트 바닥 끝에 서 있다. 중앙 뒤와 오른쪽 창가에 굵은 기둥이 각각 하나씩 보이고, 오른쪽 위에는 창 구획들, 오른쪽 아래에는 에스컬레이터의 난간 일부가 보인다. 파손된 슬래브와 철근, 잔해는 참고 장소와 어울리며 식품 선반은 없다. 다만 상층 가장자리가 화면 중앙을 가로지르고 그 아래 벽체가 넓게 보여, 하단에 가장자리를 조금만 드러내라는 구도보다 훨씬 아래에서 촬영한 모습이다.",
        "entities": "인물은 성인 동아시아계 남성 한 명뿐이며, 짧은 검은 머리와 어두운 작업 재킷·바지 차림은 이전 장면의 한수와 대체로 일치한다. 얼굴은 작게 보이지만 굳은 표정과 아래를 향한 눈빛은 읽힌다. 구조 밧줄과 무너진 콘크리트, 잔해가 있으며 현우의 신체나 다른 인물은 없다. 하층 파라솔은 프레임에 없고 읽을 수 있는 글자도 없다.",
        "hard_violations": [],
        "physics": "한수의 두 다리는 상층 바닥까지 이어지며 발 접점은 가장자리 잔해에 일부 가려져 있다. 몸을 조금 앞으로 기울인 상태로 밧줄을 손으로 잡아 지지한다. 밧줄은 손에서 화면 밖 아래로 팽팽하게 이어지고, 손 아래의 다른 구간은 수직으로 처져 있다. 별도 고정점은 보이지 않지만 떠 있는 신체나 지지 없는 물체는 없다."
       },
       {
        "label": "B",
        "direction": "한수는 상체를 숙였지만 눈은 자신보다 높은 카메라를 올려다본다. 따라서 아래의 현우를 내려다보는 시선이 아니다. 손에서 뻗은 밧줄은 화면 왼쪽 아래로 향하지만, 높은 전경 가장자리 쪽으로 이어져 낮은 구멍 속 등반자와의 연결감이 약하다.",
        "built_space": "상층 바닥과 붕괴 구멍, 하단 전경의 반대쪽 가장자리가 보인다. 왼쪽 전경과 뒤쪽 중앙, 오른쪽 창가와 화면 오른쪽 끝에 기둥 또는 벽체가 보이며, 오른쪽 뒤에는 여러 창 구획이 있다. 식품 선반은 없고 콘크리트와 잔해의 재질은 참고와 유사하다. 그러나 카메라가 상층 바닥과 한수의 정수리를 내려다보는 높이에 있어 지정된 낮은 주관 시점과 반대다.",
        "entities": "짧은 검은 머리의 성인 동아시아계 남성 한 명이 등장하며, 얼굴과 체격은 한수 참고에 대체로 부합한다. 어두운 재킷과 바지, 검은 신발은 이전 장면의 복장을 따른다. 두 손에 구조 밧줄이 있고 여분의 줄은 왼쪽 바닥에 놓여 있다. 추가 인물이나 읽을 수 있는 글자는 없으며 하층 파라솔은 보이지 않는다.",
        "hard_violations": [
         "구멍에서 올라오는 현우의 낮은 시점 대신 한수와 상층 바닥을 위에서 내려다보는 시점으로 촬영하여, 명시된 구조물 내 카메라 위치를 뒤집었다."
        ],
        "physics": "양발이 붕괴 가장자리의 콘크리트에 닿아 있고 다리를 벌려 상체의 전방 기울기를 받친다. 두 손이 밧줄을 잡고 있으며 여분의 줄은 바닥에 지지된다. 가장자리에서의 자세는 위험해 보이지만 물리적으로 불가능하지는 않다. 공중에 뜬 인물이나 지지 없는 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "아래의 현우를 내려다보는 한수와 왼쪽 아래로 이어지는 밧줄은 정확하지만, 카메라가 너무 깊어 붕괴 가장자리가 화면 하단의 좁은 부분이 아니라 중앙을 크게 차지한다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "전신과 밧줄을 잡은 손은 분명하지만, 카메라가 한수보다 높고 한수가 위를 올려다보므로 현우의 낮은 주관 시점과 내려다보는 행동을 반대로 구현했다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "한수는 고개와 눈을 아래쪽 카메라, 즉 구멍 속 현우가 있을 방향으로 향한다. 손에 잡힌 밧줄은 왼쪽 아래 화면 끝으로 비스듬히 이어져 보이지 않는 등반자와 연결되는 방향이 맞는다.",
        "built_space": "한수는 붕괴된 상층 콘크리트 바닥 끝에 서 있다. 중앙 뒤와 오른쪽 창가에 굵은 기둥이 각각 하나씩 보이고, 오른쪽 위에는 창 구획들, 오른쪽 아래에는 에스컬레이터의 난간 일부가 보인다. 파손된 슬래브와 철근, 잔해는 참고 장소와 어울리며 식품 선반은 없다. 다만 상층 가장자리가 화면 중앙을 가로지르고 그 아래 벽체가 넓게 보여, 하단에 가장자리를 조금만 드러내라는 구도보다 훨씬 아래에서 촬영한 모습이다.",
        "entities": "인물은 성인 동아시아계 남성 한 명뿐이며, 짧은 검은 머리와 어두운 작업 재킷·바지 차림은 이전 장면의 한수와 대체로 일치한다. 얼굴은 작게 보이지만 굳은 표정과 아래를 향한 눈빛은 읽힌다. 구조 밧줄과 무너진 콘크리트, 잔해가 있으며 현우의 신체나 다른 인물은 없다. 하층 파라솔은 프레임에 없고 읽을 수 있는 글자도 없다.",
        "hard_violations": [],
        "physics": "한수의 두 다리는 상층 바닥까지 이어지며 발 접점은 가장자리 잔해에 일부 가려져 있다. 몸을 조금 앞으로 기울인 상태로 밧줄을 손으로 잡아 지지한다. 밧줄은 손에서 화면 밖 아래로 팽팽하게 이어지고, 손 아래의 다른 구간은 수직으로 처져 있다. 별도 고정점은 보이지 않지만 떠 있는 신체나 지지 없는 물체는 없다."
       },
       {
        "label": "A",
        "direction": "한수는 상체를 숙였지만 눈은 자신보다 높은 카메라를 올려다본다. 따라서 아래의 현우를 내려다보는 시선이 아니다. 손에서 뻗은 밧줄은 화면 왼쪽 아래로 향하지만, 높은 전경 가장자리 쪽으로 이어져 낮은 구멍 속 등반자와의 연결감이 약하다.",
        "built_space": "상층 바닥과 붕괴 구멍, 하단 전경의 반대쪽 가장자리가 보인다. 왼쪽 전경과 뒤쪽 중앙, 오른쪽 창가와 화면 오른쪽 끝에 기둥 또는 벽체가 보이며, 오른쪽 뒤에는 여러 창 구획이 있다. 식품 선반은 없고 콘크리트와 잔해의 재질은 참고와 유사하다. 그러나 카메라가 상층 바닥과 한수의 정수리를 내려다보는 높이에 있어 지정된 낮은 주관 시점과 반대다.",
        "entities": "짧은 검은 머리의 성인 동아시아계 남성 한 명이 등장하며, 얼굴과 체격은 한수 참고에 대체로 부합한다. 어두운 재킷과 바지, 검은 신발은 이전 장면의 복장을 따른다. 두 손에 구조 밧줄이 있고 여분의 줄은 왼쪽 바닥에 놓여 있다. 추가 인물이나 읽을 수 있는 글자는 없으며 하층 파라솔은 보이지 않는다.",
        "hard_violations": [
         "구멍에서 올라오는 현우의 낮은 시점 대신 한수와 상층 바닥을 위에서 내려다보는 시점으로 촬영하여, 명시된 구조물 내 카메라 위치를 뒤집었다."
        ],
        "physics": "양발이 붕괴 가장자리의 콘크리트에 닿아 있고 다리를 벌려 상체의 전방 기울기를 받친다. 두 손이 밧줄을 잡고 있으며 여분의 줄은 바닥에 지지된다. 가장자리에서의 자세는 위험해 보이지만 물리적으로 불가능하지는 않다. 공중에 뜬 인물이나 지지 없는 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.429,
    "B": 1.429
   },
   "adjusted": {
    "A": 1.179,
    "B": 1.179
   },
   "violations": {
    "B": [
     "[gemini-pro] 카메라 위치 및 공간 위반: 프롬프트에 지정된 위층 매장 내부가 아닌 붕괴된 구멍 아래 허공에서 위를 올려다보는 시점으로 렌더링됨.",
     "[gemini-pro] 프레이밍 위반: 바닥 가장자리가 화면 하단에 좁게 위치해야 한다는 지시를 무시하고 화면 중앙을 가로지르도록 배치됨."
    ],
    "A": [
     "[gpt-high] 구멍에서 올라오는 현우의 낮은 시점 대신 한수와 상층 바닥을 위에서 내려다보는 시점으로 촬영하여, 명시된 구조물 내 카메라 위치를 뒤집었다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1179,
   "B": 1179
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1179,
    "verdict_ko": "지정된 프레이밍과 카메라 위치를 정확히 구현했으며, 인물 묘사와 공간의 물리적 사실성이 뛰어납니다.  ★위반: [gpt-high] 구멍에서 올라오는 현우의 낮은 시점 대신 한수와 상층 바닥을 위에서 내려다보는 시점으로 촬영하여, 명시된 구조물 내 카메라 위치를 뒤집었다."
   },
   {
    "label": "B",
    "score": 1179,
    "verdict_ko": "카메라를 구멍 아래에 배치하여 공간 위치 잠금을 위반했으며, 지정된 프레이밍 구성에도 어긋납니다.  ★위반: [gemini-pro] 카메라 위치 및 공간 위반: 프롬프트에 지정된 위층 매장 내부가 아닌 붕괴된 구멍 아래 허공에서 위를 올려다보는 시점으로 렌더링됨. / [gemini-pro] 프레이밍 위반: 바닥 가장자리가 화면 하단에 좁게 위치해야 한다는 지시를 무시하고 화면 중앙을 가로지르도록 배치됨."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 한수 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S66sh26_sel.png",
    "asset_id": "3b33003c-806e-44fa-a9bd-549fde572e63",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 한수: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1103406>",
    "asset_id": "4b12902d-ff9d-48f0-a297-fcd5fb403477",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-a8e5-7109-ad52-412d7e7a19c1",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S66sh26"
  }
 },
 "S66sh36::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:07:22.312840+00:00",
  "fingerprint": "03899fde169fc4241b624d555cfaad5f792767479def54f8da7c6c49a379ed1d",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S66sh36_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S66sh36_sel.png",
  "source_sha256": "c38307323cf4a5bdc2dadc18238f7c3e884cb61e887eee462129548e97f4e10e",
  "file": "S66sh36_cine.png",
  "staged_sha256": "6b297e03685348842a905300a5f78f5752f76b87577f08739a194f9d2bcf8ef5",
  "latency_ms": 9284
 },
 "S67sh40::signage": {
  "fp": "01d87f1ac80ec18e",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "groupbg::village_lodging_yard": {
  "input_fingerprint": "36b68fc8253e9dfe",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "village_lodging_yard",
    "tags": [
     "S67sh40"
    ]
   },
   "context_sig": "de29d4e7b39271bd"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: In the village's outdoor residential gathering area at night, under military vehicle headlights during the roundup.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n익산 한옥마을 거리와 공터: 황폐화된 전통 건축물 사이로 불길이 일고 낡은 공을 차는 흙바닥 넓은 터. (특징: 부서진 기와와 낡은 목조 한옥 잔해들; 밤을 밝히는 드럼통 모닥불과 횃불; 창, 도끼, 몽둥이를 든 하회탈 무리; 흙먼지 날리는 공터 바닥과 낡은 축구공)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 찰리와 B-200. 숙소를 내려다보는데... 여기저기 터지는 군용 헤드라이트 불빛!\n- 이때 어디선가 끌려 나오는 현우와 앰버.\n\nTIME OF DAY (lock): night, bright moonlight.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: In the village's outdoor residential gathering area at night, under military vehicle headlights during the roundup.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n익산 한옥마을 거리와 공터: 황폐화된 전통 건축물 사이로 불길이 일고 낡은 공을 차는 흙바닥 넓은 터. (특징: 부서진 기와와 낡은 목조 한옥 잔해들; 밤을 밝히는 드럼통 모닥불과 횃불; 창, 도끼, 몽둥이를 든 하회탈 무리; 흙먼지 날리는 공터 바닥과 낡은 축구공)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 찰리와 B-200. 숙소를 내려다보는데... 여기저기 터지는 군용 헤드라이트 불빛!\n- 이때 어디선가 끌려 나오는 현우와 앰버.\n\nTIME OF DAY (lock): night, bright moonlight.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_village_lodging_yard_928b09.png",
  "asset_id": "69fb2d3b-78b7-46d7-89c8-296f2c05b2be",
  "input_asset_ids": [
   "a32fd8a9-4876-47cb-ad74-e319feef74f4"
  ],
  "origin_tag": "S67sh40",
  "place_text": "In the village's outdoor residential gathering area at night, under military vehicle headlights during the roundup.",
  "origin_inputs": {
   "place_text": "In the village's outdoor residential gathering area at night, under military vehicle headlights during the roundup.",
   "time_of_day_en": "night, bright moonlight",
   "conti_asset_id": "a32fd8a9-4876-47cb-ad74-e319feef74f4"
  }
 },
 "S67sh40::bgfirst_bg": {
  "input_fingerprint": "94e73786571ee72e",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 현우의 관자놀이에 차가운 권총 총구를 바짝 들이댄 박철진의 살벌한 밀착 상체.\n\nLOCATION (lock): In the village's outdoor residential gathering area at night, under military vehicle headlights during the roundup.\n\nTIME OF DAY (lock): night, bright moonlight.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Pistol (Muzzle pressed against 현우's temple) — The side of the weapon is visible between 박철진's hand and 현우's head; used as Make the physical threat readable within the shared close framing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Bright full-moon ambience and the established military headlight spill create controlled hard-edged contrast across the hostage pair without obscuring the muzzle contact.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 현우의 관자놀이에 차가운 권총 총구를 바짝 들이댄 박철진의 살벌한 밀착 상체.\n\nLOCATION (lock): In the village's outdoor residential gathering area at night, under military vehicle headlights during the roundup.\n\nTIME OF DAY (lock): night, bright moonlight.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Pistol (Muzzle pressed against 현우's temple) — The side of the weapon is visible between 박철진's hand and 현우's head; used as Make the physical threat readable within the shared close framing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Bright full-moon ambience and the established military headlight spill create controlled hard-edged contrast across the hostage pair without obscuring the muzzle contact.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S67sh40__bgfirst_bg.png",
  "asset_id": "6d5cb97b-e1e7-40f2-8bdc-880bfadc45a7",
  "input_asset_ids": [
   "a32fd8a9-4876-47cb-ad74-e319feef74f4",
   "69fb2d3b-78b7-46d7-89c8-296f2c05b2be"
  ]
 },
 "S67sh40": {
  "input_fingerprint": "fba53143aa2143d7",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, bright moonlight.\n\nSHOT TEXT (authoritative, Korean): 현우의 관자놀이에 차가운 권총 총구를 바짝 들이댄 박철진의 살벌한 밀착 상체.\n\nLOCATION (lock): In the village's outdoor residential gathering area at night, under military vehicle headlights during the roundup. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Pistol (Muzzle pressed against 현우's temple) — The side of the weapon is visible between 박철진's hand and 현우's head; used as Make the physical threat readable within the shared close framing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Bright full-moon ambience and the established military headlight spill create controlled hard-edged contrast across the hostage pair without obscuring the muzzle contact.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The village is lit by a bright full moon and military headlights, and the dump truck used for the infiltration remains present. Charlie stands with his previously dented, punctured body and impaired systems; B-200's gun-hands are still intact. 박철진: He holds a raised pistol in a hostage-threatening position. 현우: He is held captive in the village and retains his earlier fighting injuries.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 박철진 right now, so 박철진's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 박철진: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, bright moonlight.\n\nSHOT TEXT (authoritative, Korean): 현우의 관자놀이에 차가운 권총 총구를 바짝 들이댄 박철진의 살벌한 밀착 상체.\n\nLOCATION (lock): In the village's outdoor residential gathering area at night, under military vehicle headlights during the roundup. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Pistol (Muzzle pressed against 현우's temple) — The side of the weapon is visible between 박철진's hand and 현우's head; used as Make the physical threat readable within the shared close framing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Bright full-moon ambience and the established military headlight spill create controlled hard-edged contrast across the hostage pair without obscuring the muzzle contact.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The village is lit by a bright full moon and military headlights, and the dump truck used for the infiltration remains present. Charlie stands with his previously dented, punctured body and impaired systems; B-200's gun-hands are still intact. 박철진: He holds a raised pistol in a hostage-threatening position. 현우: He is held captive in the village and retains his earlier fighting injuries.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 박철진 right now, so 박철진's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 박철진: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, bright moonlight.\n\nSHOT TEXT (authoritative, Korean): 현우의 관자놀이에 차가운 권총 총구를 바짝 들이댄 박철진의 살벌한 밀착 상체.\n\nLOCATION (lock): In the village's outdoor residential gathering area at night, under military vehicle headlights during the roundup. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Pistol (Muzzle pressed against 현우's temple) — The side of the weapon is visible between 박철진's hand and 현우's head; used as Make the physical threat readable within the shared close framing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Bright full-moon ambience and the established military headlight spill create controlled hard-edged contrast across the hostage pair without obscuring the muzzle contact.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The village is lit by a bright full moon and military headlights, and the dump truck used for the infiltration remains present. Charlie stands with his previously dented, punctured body and impaired systems; B-200's gun-hands are still intact. 박철진: He holds a raised pistol in a hostage-threatening position. 현우: He is held captive in the village and retains his earlier fighting injuries.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 박철진 right now, so 박철진's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 박철진: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S67sh40__bgfirst_bg.png",
     "asset_id": "6d5cb97b-e1e7-40f2-8bdc-880bfadc45a7",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S67sh40.png",
     "asset_id": "a32fd8a9-4876-47cb-ad74-e319feef74f4",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1401722>",
     "asset_id": "fee7383c-fb61-4b3a-ba7c-79f2555de00b",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:842741>",
     "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
     "role": "prop_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_village_lodging_yard_928b09.png",
     "asset_id": "69fb2d3b-78b7-46d7-89c8-296f2c05b2be",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1401722>",
     "asset_id": "fee7383c-fb61-4b3a-ba7c-79f2555de00b",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:842741>",
     "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
     "role": "prop_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "박철진의 시선은 화면 밖 우측을 향하고 현우는 아래를 내려다봄. 총구는 현우의 관자놀이에 밀착됨.",
    "built_space": "야간의 폐허 마을과 헤드라이트를 켠 트럭이 배경에 묘사됨. 인물 배치에 따른 공간적 오류는 없음.",
    "entities": "인물들의 외형은 레퍼런스와 부합하나, 박철진이 쥐고 있는 권총 하단부(손잡이)가 완전히 생략되어 있음.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 소품 및 신체 구조 (권총 손잡이 누락 및 공중을 쥐고 있는 기형적인 손)"
    ],
    "physics": "박철진의 손이 총을 쥐고 있으나 쥐어야 할 권총 손잡이가 존재하지 않아 손가락이 허공에 둥글게 말려 있는 물리적 오류가 관찰됨."
   },
   {
    "label": "B",
    "direction": "박철진과 현우가 서로 시선을 교환하며 대치함. 권총의 총구는 현우의 이마 측면에 닿아 있음.",
    "built_space": "헤드라이트를 켠 군용 트럭들과 파괴된 한옥 마을이 배경에 배치됨. 트럭 뒤쪽으로 프롬프트에 없는 군인 2명이 서 있음.",
    "entities": "현우와 박철진의 외형 및 상처, 권총의 질감은 레퍼런스와 잘 일치함. 지시되지 않은 추가 인물이 묘사됨.",
    "hard_violations": [
     "[gemini-pro] 지시문에 없는 인물 임의 추가 (배경의 군인 2명)",
     "[gpt-high] 허용된 두 사람 외에 배경 인물 두 명이 추가되었다.",
     "[gpt-high] 권총 측면에 읽을 수 있는 구경 각인이 있어 글자 금지 조건을 위반한다."
    ],
    "physics": "박철진의 손이 권총의 손잡이와 방아쇠 울 주변을 정상적으로 파지하고 있으며, 총구가 머리에 닿는 접촉면의 지지가 명확함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "두 인물 간의 물리적 밀착과 긴장감 묘사는 우수하나, 프롬프트에서 엄격히 금지한 추가 인물(배경의 군인)이 등장하여 심각한 위반입니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "총기의 손잡이가 완전히 누락되어 손과 융합되는 등 물리적으로 불가능한 소품 묘사가 발생한 치명적인 오류가 있습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "박철진과 현우가 서로 시선을 교환하며 대치함. 권총의 총구는 현우의 이마 측면에 닿아 있음.",
        "built_space": "헤드라이트를 켠 군용 트럭들과 파괴된 한옥 마을이 배경에 배치됨. 트럭 뒤쪽으로 프롬프트에 없는 군인 2명이 서 있음.",
        "entities": "현우와 박철진의 외형 및 상처, 권총의 질감은 레퍼런스와 잘 일치함. 지시되지 않은 추가 인물이 묘사됨.",
        "hard_violations": [
         "지시문에 없는 인물 임의 추가 (배경의 군인 2명)"
        ],
        "physics": "박철진의 손이 권총의 손잡이와 방아쇠 울 주변을 정상적으로 파지하고 있으며, 총구가 머리에 닿는 접촉면의 지지가 명확함."
       },
       {
        "label": "A",
        "direction": "박철진의 시선은 화면 밖 우측을 향하고 현우는 아래를 내려다봄. 총구는 현우의 관자놀이에 밀착됨.",
        "built_space": "야간의 폐허 마을과 헤드라이트를 켠 트럭이 배경에 묘사됨. 인물 배치에 따른 공간적 오류는 없음.",
        "entities": "인물들의 외형은 레퍼런스와 부합하나, 박철진이 쥐고 있는 권총 하단부(손잡이)가 완전히 생략되어 있음.",
        "hard_violations": [
         "물리적으로 불가능한 소품 및 신체 구조 (권총 손잡이 누락 및 공중을 쥐고 있는 기형적인 손)"
        ],
        "physics": "박철진의 손이 총을 쥐고 있으나 쥐어야 할 권총 손잡이가 존재하지 않아 손가락이 허공에 둥글게 말려 있는 물리적 오류가 관찰됨."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "두 인물 간의 물리적 밀착과 긴장감 묘사는 우수하나, 프롬프트에서 엄격히 금지한 추가 인물(배경의 군인)이 등장하여 심각한 위반입니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "총기의 손잡이가 완전히 누락되어 손과 융합되는 등 물리적으로 불가능한 소품 묘사가 발생한 치명적인 오류가 있습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "박철진과 현우가 서로 시선을 교환하며 대치함. 권총의 총구는 현우의 이마 측면에 닿아 있음.",
        "built_space": "헤드라이트를 켠 군용 트럭들과 파괴된 한옥 마을이 배경에 배치됨. 트럭 뒤쪽으로 프롬프트에 없는 군인 2명이 서 있음.",
        "entities": "현우와 박철진의 외형 및 상처, 권총의 질감은 레퍼런스와 잘 일치함. 지시되지 않은 추가 인물이 묘사됨.",
        "hard_violations": [
         "지시문에 없는 인물 임의 추가 (배경의 군인 2명)"
        ],
        "physics": "박철진의 손이 권총의 손잡이와 방아쇠 울 주변을 정상적으로 파지하고 있으며, 총구가 머리에 닿는 접촉면의 지지가 명확함."
       },
       {
        "label": "A",
        "direction": "박철진의 시선은 화면 밖 우측을 향하고 현우는 아래를 내려다봄. 총구는 현우의 관자놀이에 밀착됨.",
        "built_space": "야간의 폐허 마을과 헤드라이트를 켠 트럭이 배경에 묘사됨. 인물 배치에 따른 공간적 오류는 없음.",
        "entities": "인물들의 외형은 레퍼런스와 부합하나, 박철진이 쥐고 있는 권총 하단부(손잡이)가 완전히 생략되어 있음.",
        "hard_violations": [
         "물리적으로 불가능한 소품 및 신체 구조 (권총 손잡이 누락 및 공중을 쥐고 있는 기형적인 손)"
        ],
        "physics": "박철진의 손이 총을 쥐고 있으나 쥐어야 할 권총 손잡이가 존재하지 않아 손가락이 허공에 둥글게 말려 있는 물리적 오류가 관찰됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "밀착 위협과 총구 접촉은 보이지만, 금지된 배경 인물 두 명과 읽히는 총기 각인이 있으며 박철진의 전술복도 참조와 다르다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "두 사람의 밀착 클로즈업, 현우의 관자놀이에 닿은 총구, 박철진의 정장과 장소를 충실히 구현했으나 현우의 티셔츠 색과 개머리판 없는 총기는 참조와 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "박철진은 오른쪽의 현우 얼굴을 바라보고 현우는 왼쪽 박철진 쪽으로 눈을 돌린다. 권총은 화면 오른쪽으로 향하며 총구가 현우의 눈썹 바깥 위쪽 머리선에 닿는다. 머리를 향한 위협은 분명하지만 관자놀이 옆면보다 이마 가장자리에 가까운 접촉이다.",
        "built_space": "왼쪽 목조 처마와 기둥, 중앙 기와집과 돌담, 오른쪽 위 전봇대가 참조 마을의 재료와 형태를 따른다. 트럭은 왼쪽 가까이에 한 대, 중앙 뒤쪽에 한 대로 두 대가 보인다. 전경 두 사람 외에 중앙과 총 아래 배경에 각각 한 명씩 서 있다. 가까운 왼쪽 트럭은 참조의 중앙 원거리 차량 배치와 다르다.",
        "entities": "전경에는 짧게 빗어 넘긴 검은 머리의 중년 남성과 헝클어진 검은 머리의 앳된 남성이 있으며 두 참조 얼굴과 대체로 닮았다. 현우의 얼굴에는 찰과상과 피가 있고 어두운 티셔츠를 입었다. 박철진은 참조의 정장 대신 검은 전술복과 조끼를 입었다. 권총의 금속 측면은 보이지만 참조의 길게 연결된 개머리판은 없으며 총기 측면의 구경 표기가 읽힌다. 배경 남성 두 명은 허용된 등장인물이 아니다.",
        "hard_violations": [
         "허용된 두 사람 외에 배경 인물 두 명이 추가되었다.",
         "권총 측면에 읽을 수 있는 구경 각인이 있어 글자 금지 조건을 위반한다."
        ],
        "physics": "권총 손잡이는 박철진의 손에 잡혀 있고 손목과 소매, 굽힌 팔이 연결되어 총의 무게를 지탱한다. 반대쪽 팔은 현우의 어깨와 목 뒤로 이어져 밀착 제압 자세를 만든다. 두 사람의 하체는 프레임 밖이며, 보이는 상체에서 부유나 불가능한 관절 연결은 발견되지 않는다."
       },
       {
        "label": "B",
        "direction": "현우와 박철진은 모두 화면 오른쪽의 프레임 밖을 경계하듯 바라본다. 박철진의 권총은 왼쪽으로 향하고 총구 끝이 현우의 바깥 눈꼬리 위 관자놀이에 밀착한다. 손과 머리 사이에 총의 측면이 드러나 목표와 접촉 지점이 명확하다.",
        "built_space": "두 사람은 마을 길 전경에 나란히 밀착하고 박철진이 현우보다 조금 뒤에 있다. 왼쪽 목조 처마와 돌담, 오른쪽 기와 담장, 중앙 뒤 트럭 한 대와 콘크리트 차단물, 오른쪽 상자 더미와 불붙은 드럼통 한 개가 보인다. 오른쪽에는 가까운 전봇대 하나와 먼 전봇대 하나가 보이며 참조의 마을 공간 구성을 유지한다. 추가 인물이나 불가능한 반사는 없다.",
        "entities": "현우는 참조와 닮은 앳된 얼굴과 헝클어진 검은 머리이며 볼과 턱에 전투 상처가 남아 있다. 다만 티셔츠는 참조의 남색 대신 회갈색이다. 박철진은 중년 얼굴 윤곽과 짧게 정돈한 검은 머리, 남색 정장과 흰 셔츠, 넥타이로 참조에 가깝다. 권총 한 정의 금속 질감과 기본 형태는 유사하지만 참조의 개머리판이 없다. 보름달과 군용 트럭 전조등이 보이며 두 사람 외 인물은 없다. 총기 각인은 있으나 선명하게 읽히는 문구는 확인되지 않는다.",
        "hard_violations": [],
        "physics": "박철진의 손이 권총 손잡이를 감싸고 검지가 방아쇠울 안에 놓여 있다. 손목은 흰 셔츠 소매와 정장 팔로 자연스럽게 이어져 총을 지지한다. 팔을 굽혀 옆에 있는 현우의 관자놀이에 총구를 대는 자세는 가능하다. 하체는 클로즈업 밖에 있으며 상체나 소품이 지지 없이 떠 있는 부분은 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "밀착 위협과 총구 접촉은 보이지만, 금지된 배경 인물 두 명과 읽히는 총기 각인이 있으며 박철진의 전술복도 참조와 다르다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "두 사람의 밀착 클로즈업, 현우의 관자놀이에 닿은 총구, 박철진의 정장과 장소를 충실히 구현했으나 현우의 티셔츠 색과 개머리판 없는 총기는 참조와 다르다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "박철진은 오른쪽의 현우 얼굴을 바라보고 현우는 왼쪽 박철진 쪽으로 눈을 돌린다. 권총은 화면 오른쪽으로 향하며 총구가 현우의 눈썹 바깥 위쪽 머리선에 닿는다. 머리를 향한 위협은 분명하지만 관자놀이 옆면보다 이마 가장자리에 가까운 접촉이다.",
        "built_space": "왼쪽 목조 처마와 기둥, 중앙 기와집과 돌담, 오른쪽 위 전봇대가 참조 마을의 재료와 형태를 따른다. 트럭은 왼쪽 가까이에 한 대, 중앙 뒤쪽에 한 대로 두 대가 보인다. 전경 두 사람 외에 중앙과 총 아래 배경에 각각 한 명씩 서 있다. 가까운 왼쪽 트럭은 참조의 중앙 원거리 차량 배치와 다르다.",
        "entities": "전경에는 짧게 빗어 넘긴 검은 머리의 중년 남성과 헝클어진 검은 머리의 앳된 남성이 있으며 두 참조 얼굴과 대체로 닮았다. 현우의 얼굴에는 찰과상과 피가 있고 어두운 티셔츠를 입었다. 박철진은 참조의 정장 대신 검은 전술복과 조끼를 입었다. 권총의 금속 측면은 보이지만 참조의 길게 연결된 개머리판은 없으며 총기 측면의 구경 표기가 읽힌다. 배경 남성 두 명은 허용된 등장인물이 아니다.",
        "hard_violations": [
         "허용된 두 사람 외에 배경 인물 두 명이 추가되었다.",
         "권총 측면에 읽을 수 있는 구경 각인이 있어 글자 금지 조건을 위반한다."
        ],
        "physics": "권총 손잡이는 박철진의 손에 잡혀 있고 손목과 소매, 굽힌 팔이 연결되어 총의 무게를 지탱한다. 반대쪽 팔은 현우의 어깨와 목 뒤로 이어져 밀착 제압 자세를 만든다. 두 사람의 하체는 프레임 밖이며, 보이는 상체에서 부유나 불가능한 관절 연결은 발견되지 않는다."
       },
       {
        "label": "A",
        "direction": "현우와 박철진은 모두 화면 오른쪽의 프레임 밖을 경계하듯 바라본다. 박철진의 권총은 왼쪽으로 향하고 총구 끝이 현우의 바깥 눈꼬리 위 관자놀이에 밀착한다. 손과 머리 사이에 총의 측면이 드러나 목표와 접촉 지점이 명확하다.",
        "built_space": "두 사람은 마을 길 전경에 나란히 밀착하고 박철진이 현우보다 조금 뒤에 있다. 왼쪽 목조 처마와 돌담, 오른쪽 기와 담장, 중앙 뒤 트럭 한 대와 콘크리트 차단물, 오른쪽 상자 더미와 불붙은 드럼통 한 개가 보인다. 오른쪽에는 가까운 전봇대 하나와 먼 전봇대 하나가 보이며 참조의 마을 공간 구성을 유지한다. 추가 인물이나 불가능한 반사는 없다.",
        "entities": "현우는 참조와 닮은 앳된 얼굴과 헝클어진 검은 머리이며 볼과 턱에 전투 상처가 남아 있다. 다만 티셔츠는 참조의 남색 대신 회갈색이다. 박철진은 중년 얼굴 윤곽과 짧게 정돈한 검은 머리, 남색 정장과 흰 셔츠, 넥타이로 참조에 가깝다. 권총 한 정의 금속 질감과 기본 형태는 유사하지만 참조의 개머리판이 없다. 보름달과 군용 트럭 전조등이 보이며 두 사람 외 인물은 없다. 총기 각인은 있으나 선명하게 읽히는 문구는 확인되지 않는다.",
        "hard_violations": [],
        "physics": "박철진의 손이 권총 손잡이를 감싸고 검지가 방아쇠울 안에 놓여 있다. 손목은 흰 셔츠 소매와 정장 팔로 자연스럽게 이어져 총을 지지한다. 팔을 굽혀 옆에 있는 현우의 관자놀이에 총구를 대는 자세는 가능하다. 하체는 클로즈업 밖에 있으며 상체나 소품이 지지 없이 떠 있는 부분은 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.75,
    "B": 1.375
   },
   "adjusted": {
    "A": 1.5,
    "B": 1.125
   },
   "violations": {
    "B": [
     "[gemini-pro] 지시문에 없는 인물 임의 추가 (배경의 군인 2명)",
     "[gpt-high] 허용된 두 사람 외에 배경 인물 두 명이 추가되었다.",
     "[gpt-high] 권총 측면에 읽을 수 있는 구경 각인이 있어 글자 금지 조건을 위반한다."
    ],
    "A": [
     "[gemini-pro] 물리적으로 불가능한 소품 및 신체 구조 (권총 손잡이 누락 및 공중을 쥐고 있는 기형적인 손)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1125,
   "A": 1500
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1125,
    "verdict_ko": "두 인물 간의 물리적 밀착과 긴장감 묘사는 우수하나, 프롬프트에서 엄격히 금지한 추가 인물(배경의 군인)이 등장하여 심각한 위반입니다.  ★위반: [gemini-pro] 지시문에 없는 인물 임의 추가 (배경의 군인 2명) / [gpt-high] 허용된 두 사람 외에 배경 인물 두 명이 추가되었다. / [gpt-high] 권총 측면에 읽을 수 있는 구경 각인이 있어 글자 금지 조건을 위반한다."
   },
   {
    "label": "A",
    "score": 1500,
    "verdict_ko": "총기의 손잡이가 완전히 누락되어 손과 융합되는 등 물리적으로 불가능한 소품 묘사가 발생한 치명적인 오류가 있습니다.  ★위반: [gemini-pro] 물리적으로 불가능한 소품 및 신체 구조 (권총 손잡이 누락 및 공중을 쥐고 있는 기형적인 손)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_village_lodging_yard_928b09.png",
    "asset_id": "69fb2d3b-78b7-46d7-89c8-296f2c05b2be",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1401722>",
    "asset_id": "fee7383c-fb61-4b3a-ba7c-79f2555de00b",
    "role": "character_ref"
   },
   {
    "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
    "path": "<bytes:842741>",
    "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
    "role": "prop_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-aae0-7f4d-b659-74145624c87d",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S67sh40__bgfirst_bg.png",
   "bg_asset_id": "6d5cb97b-e1e7-40f2-8bdc-880bfadc45a7",
   "bg_record_key": "S67sh40::bgfirst_bg",
   "chain_winner": true,
   "authority": "groupbg",
   "group_key": "village_lodging_yard",
   "groupbg_asset_id": "69fb2d3b-78b7-46d7-89c8-296f2c05b2be"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S67sh40::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:10:00.000103+00:00",
  "fingerprint": "675bb0623eec98c15555782d277096de9f61497a08f1fbdabfba96b5b5c3f4bc",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S67sh40_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S67sh40_sel.png",
  "source_sha256": "2bf703b41b461f9bd8e042217db1d865d62755ed59e8dc3acc6d723a2ea316ce",
  "file": "S67sh40_cine.png",
  "staged_sha256": "f5e57d6e6e9223332ba9ec28e75116efa1a885f7ca3f5188ecd4491ecf70e973",
  "latency_ms": 10028
 },
 "S67sh49::signage": {
  "fp": "9e3dbed7e7b2d35a",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S67sh49": {
  "input_fingerprint": "ffe843da5c273202",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, bright moonlight.\n\nSHOT TEXT (authoritative, Korean): 달리는 트럭들 앞을 우뚝 가로막고 선 채 거대한 고사포를 정조준한 B-200의 위협적인 전신.\n\nLOCATION (lock): On the road leading out of the village, directly ahead of the departing militia truck convoy at night. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Trucks' forward route (Blocked by B-200 before the convoy overturns) — The approach runs from the lower-left foreground toward B-200 in the middle distance; used as Preserve the opposing movement and aiming directions without placing the camera on the firing line.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established bright moonlit night with restrained tonal separation, withholding muzzle flash because B-200 has not yet fired.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The militia trucks are still upright and traveling away; Charlie is confined in the transport cage with heavy restraints around his neck and both hands, retaining his earlier battle damage. B-200 blocks the route with intact gun-hands aimed forward.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by B-200 right now, so B-200's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to B-200: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, bright moonlight.\n\nSHOT TEXT (authoritative, Korean): 달리는 트럭들 앞을 우뚝 가로막고 선 채 거대한 고사포를 정조준한 B-200의 위협적인 전신.\n\nLOCATION (lock): On the road leading out of the village, directly ahead of the departing militia truck convoy at night. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Trucks' forward route (Blocked by B-200 before the convoy overturns) — The approach runs from the lower-left foreground toward B-200 in the middle distance; used as Preserve the opposing movement and aiming directions without placing the camera on the firing line.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established bright moonlit night with restrained tonal separation, withholding muzzle flash because B-200 has not yet fired.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The militia trucks are still upright and traveling away; Charlie is confined in the transport cage with heavy restraints around his neck and both hands, retaining his earlier battle damage. B-200 blocks the route with intact gun-hands aimed forward.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by B-200 right now, so B-200's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to B-200: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, bright moonlight.\n\nSHOT TEXT (authoritative, Korean): 달리는 트럭들 앞을 우뚝 가로막고 선 채 거대한 고사포를 정조준한 B-200의 위협적인 전신.\n\nLOCATION (lock): On the road leading out of the village, directly ahead of the departing militia truck convoy at night. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Trucks' forward route (Blocked by B-200 before the convoy overturns) — The approach runs from the lower-left foreground toward B-200 in the middle distance; used as Preserve the opposing movement and aiming directions without placing the camera on the firing line.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established bright moonlit night with restrained tonal separation, withholding muzzle flash because B-200 has not yet fired.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The militia trucks are still upright and traveling away; Charlie is confined in the transport cage with heavy restraints around his neck and both hands, retaining his earlier battle damage. B-200 blocks the route with intact gun-hands aimed forward.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by B-200 right now, so B-200's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to B-200: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: B-200 (거대한 중장비형 실루엣, 육중한 몸체, 짙은 회색 철제 장갑판, 굵은 기계 관절, 양손의 대형 고사포 포신) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S67sh49__bgfirst_bg.png",
     "asset_id": "5b89bd25-951b-4371-9526-4d76774415ae",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S67sh49.png",
     "asset_id": "f64f2986-6e3d-4497-9926-988ba8db7762",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — B-200: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1339855>",
     "asset_id": "8091d94b-e8e7-407e-97d6-c030f55a73f9",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:842741>",
     "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
     "role": "prop_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_village_truck_ambush_bde2df.png",
     "asset_id": "bb89f6c6-a1e1-4ea1-b9c4-7696c210aebc",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — B-200: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1339855>",
     "asset_id": "8091d94b-e8e7-407e-97d6-c030f55a73f9",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:842741>",
     "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
     "role": "prop_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "B-200은 도로 우측에서 왼쪽 허공을 겨누며, 트럭의 정면 진행 경로를 가로막지 않습니다.",
    "built_space": "폐허가 된 마을의 흙길이며, 왼쪽 편에 멀어지는 방향의 트럭 1대가 있습니다.",
    "entities": "B-200의 양팔이 권총 레퍼런스로 기괴하게 변형되었습니다. 찰리는 트럭 뒤칸에 손이 묶인 채 서 있습니다.",
    "hard_violations": [
     "[gemini-pro] 로봇의 양팔이 권총 레퍼런스로 대체되어 방아쇠와 손잡이가 허공에 매달린 물리적 불가능 상태",
     "[gemini-pro] 트럭의 정면을 가로막지 않고 도로 우측에 서서 측면을 겨누는 동선 및 연출 오류",
     "[gpt-high] B-200 외에는 인물을 보이지 말라는 제한에도 철창 속 남성과 다음 트럭의 민병대원을 노출했다."
    ],
    "physics": "합성된 거대한 권총 팔의 부품들이 물리적 지지 없이 매달려 있습니다."
   },
   {
    "label": "B",
    "direction": "B-200은 도로 중앙에서 정면을 겨눕니다. 오른쪽 트럭은 로봇을 향해, 왼쪽 트럭은 카메라를 향해 달립니다.",
    "built_space": "달빛이 비치는 마을 흙길에 2대의 트럭이 배치되어 있습니다.",
    "entities": "B-200의 형태는 캐릭터 레퍼런스와 일치합니다. 찰리는 목과 손이 묶인 채 우측 트럭에 앉아 있습니다. 총기 레퍼런스는 누락되었습니다.",
    "hard_violations": [
     "[gemini-pro] 출발하는 수송대(convoy)라는 설정과 달리 트럭 두 대가 서로 반대 방향으로 주행하는 심각한 동선 및 연출 위반",
     "[gpt-high] 가시 인물을 B-200으로 한정했는데 민병대원·운전자와 철창 속 남성 등 추가 인물을 노출했다."
    ],
    "physics": "모든 차량과 로봇의 몸체는 지면에 안정적으로 닿아 지지받고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "B-200의 외형과 인물 구속 상태는 잘 표현되었으나, 수송대 트럭들이 서로 반대 방향으로 달리는 치명적인 동선 오류가 발생했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "로봇의 팔에 권총 레퍼런스가 그대로 합성되는 심각한 형태 오류가 있으며, 트럭의 정면을 가로막지 못했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "B-200은 도로 우측에서 왼쪽 허공을 겨누며, 트럭의 정면 진행 경로를 가로막지 않습니다.",
        "built_space": "폐허가 된 마을의 흙길이며, 왼쪽 편에 멀어지는 방향의 트럭 1대가 있습니다.",
        "entities": "B-200의 양팔이 권총 레퍼런스로 기괴하게 변형되었습니다. 찰리는 트럭 뒤칸에 손이 묶인 채 서 있습니다.",
        "hard_violations": [
         "로봇의 양팔이 권총 레퍼런스로 대체되어 방아쇠와 손잡이가 허공에 매달린 물리적 불가능 상태",
         "트럭의 정면을 가로막지 않고 도로 우측에 서서 측면을 겨누는 동선 및 연출 오류"
        ],
        "physics": "합성된 거대한 권총 팔의 부품들이 물리적 지지 없이 매달려 있습니다."
       },
       {
        "label": "B",
        "direction": "B-200은 도로 중앙에서 정면을 겨눕니다. 오른쪽 트럭은 로봇을 향해, 왼쪽 트럭은 카메라를 향해 달립니다.",
        "built_space": "달빛이 비치는 마을 흙길에 2대의 트럭이 배치되어 있습니다.",
        "entities": "B-200의 형태는 캐릭터 레퍼런스와 일치합니다. 찰리는 목과 손이 묶인 채 우측 트럭에 앉아 있습니다. 총기 레퍼런스는 누락되었습니다.",
        "hard_violations": [
         "출발하는 수송대(convoy)라는 설정과 달리 트럭 두 대가 서로 반대 방향으로 주행하는 심각한 동선 및 연출 위반"
        ],
        "physics": "모든 차량과 로봇의 몸체는 지면에 안정적으로 닿아 지지받고 있습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "B-200의 외형과 인물 구속 상태는 잘 표현되었으나, 수송대 트럭들이 서로 반대 방향으로 달리는 치명적인 동선 오류가 발생했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "로봇의 팔에 권총 레퍼런스가 그대로 합성되는 심각한 형태 오류가 있으며, 트럭의 정면을 가로막지 못했습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "B-200은 도로 우측에서 왼쪽 허공을 겨누며, 트럭의 정면 진행 경로를 가로막지 않습니다.",
        "built_space": "폐허가 된 마을의 흙길이며, 왼쪽 편에 멀어지는 방향의 트럭 1대가 있습니다.",
        "entities": "B-200의 양팔이 권총 레퍼런스로 기괴하게 변형되었습니다. 찰리는 트럭 뒤칸에 손이 묶인 채 서 있습니다.",
        "hard_violations": [
         "로봇의 양팔이 권총 레퍼런스로 대체되어 방아쇠와 손잡이가 허공에 매달린 물리적 불가능 상태",
         "트럭의 정면을 가로막지 않고 도로 우측에 서서 측면을 겨누는 동선 및 연출 오류"
        ],
        "physics": "합성된 거대한 권총 팔의 부품들이 물리적 지지 없이 매달려 있습니다."
       },
       {
        "label": "B",
        "direction": "B-200은 도로 중앙에서 정면을 겨눕니다. 오른쪽 트럭은 로봇을 향해, 왼쪽 트럭은 카메라를 향해 달립니다.",
        "built_space": "달빛이 비치는 마을 흙길에 2대의 트럭이 배치되어 있습니다.",
        "entities": "B-200의 형태는 캐릭터 레퍼런스와 일치합니다. 찰리는 목과 손이 묶인 채 우측 트럭에 앉아 있습니다. 총기 레퍼런스는 누락되었습니다.",
        "hard_violations": [
         "출발하는 수송대(convoy)라는 설정과 달리 트럭 두 대가 서로 반대 방향으로 주행하는 심각한 동선 및 연출 위반"
        ],
        "physics": "모든 차량과 로봇의 몸체는 지면에 안정적으로 닿아 지지받고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "장소와 B-200의 장갑·포신은 가깝지만, 포구가 트럭을 겨누지 않고 차량의 대향 진행도 성립하지 않으며 금지된 인물들이 추가되었다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "전신 와이드와 트럭 쪽으로 향한 조준은 상대적으로 낫지만, 행렬 앞 차단 구도가 불명확하고 추가 인물 및 고사포 대신 확대된 권총형 무기가 요구를 어긴다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "B-200의 머리와 양팔 포구는 화면 오른쪽을 향한다. 왼쪽 전경 트럭이나 오른쪽 아래 수송차를 정조준하는 축이 아니다. 왼쪽 트럭은 전면과 전조등을 카메라 쪽으로 드러내며 B-200에서 멀어지는 방향이고, 오른쪽 수송차는 후면을 보이며 왼쪽 안쪽을 향한다. 두 차량이 왼쪽 아래에서 중경의 B-200을 향해 함께 접근하는 관계가 성립하지 않는다.",
        "built_space": "흙길 양쪽의 파손된 목조 기와집, 오른쪽 돌담, 왼쪽 전경 불통 하나, 오른쪽 높은 횃불 바구니 하나와 안쪽의 여러 작은 화점이 장소 참조와 가깝다. 트럭 두 대는 전경 양쪽에 있고 B-200은 그 뒤 도로 중앙에 있다. 전신은 보이지만, 적어도 왼쪽 트럭의 진행 앞을 막는 배치는 아니다.",
        "entities": "B-200 한 기의 짙은 회색 장갑판, 육중한 관절, 양팔의 다연장 포신은 캐릭터 참조와 가깝다. 별도 소품 참조의 개머리판 달린 권총은 보이지 않는다. 왼쪽 트럭에 적어도 적재함 인물 네 명과 운전자 한 명, 오른쪽 철창에 성인 남성 한 명이 보여 B-200만 보이라는 제한을 어긴다. 철창 남성은 목과 손 주변의 사슬이 보이나 기존 전투 손상은 확정하기 어렵다. 차량은 모두 직립해 있고, 밝은 달밤이며 발포 섬광은 없다.",
        "hard_violations": [
         "가시 인물을 B-200으로 한정했는데 민병대원·운전자와 철창 속 남성 등 추가 인물을 노출했다."
        ],
        "physics": "B-200은 양발을 흙길에 딛고 있고 포신은 팔의 기계 구조에 연결되어 지지된다. 트럭은 타이어로 노면에 서 있으며 먼지가 이동을 암시한다. 사람들의 하체는 차체에 가려지지만 적재함·운전석·철창 바닥이 지지할 수 있는 위치다. 지지 없이 떠 있는 몸이나 무기는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "B-200의 머리와 두 무기의 총구는 화면 왼쪽 트럭들 쪽을 향한다. 가까운 철창 수송차 쪽으로 겨누는 방향은 A보다 분명하고 카메라도 사선에서 벗어나 있다. 다만 트럭들은 후면을 보인 채 왼쪽 도로 안쪽으로 이어져, 오른쪽의 B-200을 향해 접근하기보다는 그 옆을 지나 멀어지는 행렬처럼 보인다.",
        "built_space": "B-200의 전신은 도로 오른쪽 중경에, 가장 가까운 수송차는 왼쪽 전경에 배치된다. 왼쪽으로 최소 네 대의 차량이 이어진다. 파손된 기와집과 돌담·산 능선은 있으나, 참조의 중앙 마을길과 달리 오른쪽에 긴 금속 가드레일과 전선 달린 전신주가 두드러진다. 행렬의 주행로는 B-200 왼쪽으로 열려 있어 도로를 정면 차단하는 관계가 약하다.",
        "entities": "B-200의 회색 장갑과 굵은 기계 다리는 유지되지만, 양손 고사포가 아니라 개머리판·권총 손잡이가 있는 거대한 총 두 정이 보인다. 이는 소품 참조의 형태에는 가깝지만 캐릭터 참조의 다연장 고사포 손과 다르다. 가까운 철창에 목과 손이 구속된 남성 한 명, 다음 트럭에 헬멧 쓴 인물 한 명이 보여 인물 제한을 위반한다. 기존 전투 손상은 명확하지 않다. 달빛 아래 차량들은 직립해 있고 발포 섬광이나 뚜렷하게 읽히는 문자는 없다.",
        "hard_violations": [
         "B-200 외에는 인물을 보이지 말라는 제한에도 철창 속 남성과 다음 트럭의 민병대원을 노출했다."
        ],
        "physics": "B-200의 벌린 두 발은 노면에 확실히 닿아 체중을 지지한다. 총들은 팔 앞의 기계 연결부에 걸쳐 있어 완전히 떠 있다고 단정할 수는 없지만, 아래로 노출된 권총 손잡이를 손으로 잡아 조작하는 모습은 없다. 차량 바퀴는 지면에 닿아 있고 두 인물도 차량 내부 바닥이 지지할 수 있는 위치다. 트럭의 실제 주행을 보여주는 동작 단서는 약하다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "장소와 B-200의 장갑·포신은 가깝지만, 포구가 트럭을 겨누지 않고 차량의 대향 진행도 성립하지 않으며 금지된 인물들이 추가되었다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "전신 와이드와 트럭 쪽으로 향한 조준은 상대적으로 낫지만, 행렬 앞 차단 구도가 불명확하고 추가 인물 및 고사포 대신 확대된 권총형 무기가 요구를 어긴다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "B-200의 머리와 양팔 포구는 화면 오른쪽을 향한다. 왼쪽 전경 트럭이나 오른쪽 아래 수송차를 정조준하는 축이 아니다. 왼쪽 트럭은 전면과 전조등을 카메라 쪽으로 드러내며 B-200에서 멀어지는 방향이고, 오른쪽 수송차는 후면을 보이며 왼쪽 안쪽을 향한다. 두 차량이 왼쪽 아래에서 중경의 B-200을 향해 함께 접근하는 관계가 성립하지 않는다.",
        "built_space": "흙길 양쪽의 파손된 목조 기와집, 오른쪽 돌담, 왼쪽 전경 불통 하나, 오른쪽 높은 횃불 바구니 하나와 안쪽의 여러 작은 화점이 장소 참조와 가깝다. 트럭 두 대는 전경 양쪽에 있고 B-200은 그 뒤 도로 중앙에 있다. 전신은 보이지만, 적어도 왼쪽 트럭의 진행 앞을 막는 배치는 아니다.",
        "entities": "B-200 한 기의 짙은 회색 장갑판, 육중한 관절, 양팔의 다연장 포신은 캐릭터 참조와 가깝다. 별도 소품 참조의 개머리판 달린 권총은 보이지 않는다. 왼쪽 트럭에 적어도 적재함 인물 네 명과 운전자 한 명, 오른쪽 철창에 성인 남성 한 명이 보여 B-200만 보이라는 제한을 어긴다. 철창 남성은 목과 손 주변의 사슬이 보이나 기존 전투 손상은 확정하기 어렵다. 차량은 모두 직립해 있고, 밝은 달밤이며 발포 섬광은 없다.",
        "hard_violations": [
         "가시 인물을 B-200으로 한정했는데 민병대원·운전자와 철창 속 남성 등 추가 인물을 노출했다."
        ],
        "physics": "B-200은 양발을 흙길에 딛고 있고 포신은 팔의 기계 구조에 연결되어 지지된다. 트럭은 타이어로 노면에 서 있으며 먼지가 이동을 암시한다. 사람들의 하체는 차체에 가려지지만 적재함·운전석·철창 바닥이 지지할 수 있는 위치다. 지지 없이 떠 있는 몸이나 무기는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "B-200의 머리와 두 무기의 총구는 화면 왼쪽 트럭들 쪽을 향한다. 가까운 철창 수송차 쪽으로 겨누는 방향은 A보다 분명하고 카메라도 사선에서 벗어나 있다. 다만 트럭들은 후면을 보인 채 왼쪽 도로 안쪽으로 이어져, 오른쪽의 B-200을 향해 접근하기보다는 그 옆을 지나 멀어지는 행렬처럼 보인다.",
        "built_space": "B-200의 전신은 도로 오른쪽 중경에, 가장 가까운 수송차는 왼쪽 전경에 배치된다. 왼쪽으로 최소 네 대의 차량이 이어진다. 파손된 기와집과 돌담·산 능선은 있으나, 참조의 중앙 마을길과 달리 오른쪽에 긴 금속 가드레일과 전선 달린 전신주가 두드러진다. 행렬의 주행로는 B-200 왼쪽으로 열려 있어 도로를 정면 차단하는 관계가 약하다.",
        "entities": "B-200의 회색 장갑과 굵은 기계 다리는 유지되지만, 양손 고사포가 아니라 개머리판·권총 손잡이가 있는 거대한 총 두 정이 보인다. 이는 소품 참조의 형태에는 가깝지만 캐릭터 참조의 다연장 고사포 손과 다르다. 가까운 철창에 목과 손이 구속된 남성 한 명, 다음 트럭에 헬멧 쓴 인물 한 명이 보여 인물 제한을 위반한다. 기존 전투 손상은 명확하지 않다. 달빛 아래 차량들은 직립해 있고 발포 섬광이나 뚜렷하게 읽히는 문자는 없다.",
        "hard_violations": [
         "B-200 외에는 인물을 보이지 말라는 제한에도 철창 속 남성과 다음 트럭의 민병대원을 노출했다."
        ],
        "physics": "B-200의 벌린 두 발은 노면에 확실히 닿아 체중을 지지한다. 총들은 팔 앞의 기계 연결부에 걸쳐 있어 완전히 떠 있다고 단정할 수는 없지만, 아래로 노출된 권총 손잡이를 손으로 잡아 조작하는 모습은 없다. 차량 바퀴는 지면에 닿아 있고 두 인물도 차량 내부 바닥이 지지할 수 있는 위치다. 트럭의 실제 주행을 보여주는 동작 단서는 약하다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.75,
    "B": 1.667
   },
   "adjusted": {
    "A": 1.5,
    "B": 1.417
   },
   "violations": {
    "A": [
     "[gemini-pro] 로봇의 양팔이 권총 레퍼런스로 대체되어 방아쇠와 손잡이가 허공에 매달린 물리적 불가능 상태",
     "[gemini-pro] 트럭의 정면을 가로막지 않고 도로 우측에 서서 측면을 겨누는 동선 및 연출 오류",
     "[gpt-high] B-200 외에는 인물을 보이지 말라는 제한에도 철창 속 남성과 다음 트럭의 민병대원을 노출했다."
    ],
    "B": [
     "[gemini-pro] 출발하는 수송대(convoy)라는 설정과 달리 트럭 두 대가 서로 반대 방향으로 주행하는 심각한 동선 및 연출 위반",
     "[gpt-high] 가시 인물을 B-200으로 한정했는데 민병대원·운전자와 철창 속 남성 등 추가 인물을 노출했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1417,
   "A": 1500
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1417,
    "verdict_ko": "B-200의 외형과 인물 구속 상태는 잘 표현되었으나, 수송대 트럭들이 서로 반대 방향으로 달리는 치명적인 동선 오류가 발생했습니다.  ★위반: [gemini-pro] 출발하는 수송대(convoy)라는 설정과 달리 트럭 두 대가 서로 반대 방향으로 주행하는 심각한 동선 및 연출 위반 / [gpt-high] 가시 인물을 B-200으로 한정했는데 민병대원·운전자와 철창 속 남성 등 추가 인물을 노출했다."
   },
   {
    "label": "A",
    "score": 1500,
    "verdict_ko": "로봇의 팔에 권총 레퍼런스가 그대로 합성되는 심각한 형태 오류가 있으며, 트럭의 정면을 가로막지 못했습니다.  ★위반: [gemini-pro] 로봇의 양팔이 권총 레퍼런스로 대체되어 방아쇠와 손잡이가 허공에 매달린 물리적 불가능 상태 / [gemini-pro] 트럭의 정면을 가로막지 않고 도로 우측에 서서 측면을 겨누는 동선 및 연출 오류 / [gpt-high] B-200 외에는 인물을 보이지 말라는 제한에도 철창 속 남성과 다음 트럭의 민병대원을 노출했다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_village_truck_ambush_bde2df.png",
    "asset_id": "bb89f6c6-a1e1-4ea1-b9c4-7696c210aebc",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — B-200: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1339855>",
    "asset_id": "8091d94b-e8e7-407e-97d6-c030f55a73f9",
    "role": "character_ref"
   },
   {
    "label": "PROP REFERENCE — 총: the exact object appearing in this shot; match its look, material and wear exactly.",
    "path": "<bytes:842741>",
    "asset_id": "aa00d3c0-9fbe-454f-83b9-7a4bd62b8cd3",
    "role": "prop_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-afbe-71b8-a84e-538a05e6fb17",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S67sh49__bgfirst_bg.png",
   "bg_asset_id": "5b89bd25-951b-4371-9526-4d76774415ae",
   "bg_record_key": "S67sh49::bgfirst_bg",
   "chain_winner": true,
   "authority": "groupbg",
   "group_key": "village_truck_ambush",
   "groupbg_asset_id": "bb89f6c6-a1e1-4ea1-b9c4-7696c210aebc"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S67sh49::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:54:59.307848+00:00",
  "fingerprint": "f617e70ed8a71cdaf21f220dcb357f5b03cf4da9fd968663dcab6b10bfb7f8fd",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S67sh49_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S67sh49_sel.png",
  "source_sha256": "7d90931d09c2346a987f7446f4a3f8b97db1462586b95c10e3fb0ad103f0e07f",
  "file": "S67sh49_cine.png",
  "staged_sha256": "b87c1738e6edd7ec528856af8f077542f15aeab2848b198de64135cb9196ae4e",
  "latency_ms": 10614
 },
 "S67sh70::signage": {
  "fp": "7726b4ee80764f48",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S67sh70": {
  "input_fingerprint": "cc5c38bc4b4ad491",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, bright moonlight.\n\nSHOT TEXT (authoritative, Korean): B-200의 부품을 가슴에 깊이 품은 채 두 눈의 디지털 불빛이 완전히 꺼진 찰리의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): At the blast site on the village departure road, among wrecked trucks and robot debris after the nighttime battle. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: B-200's remaining chest component (Retrieved after the explosion and held against 찰리's chest); used as Remain partly visible at the lower edge as the tangible object of grief; no intact B-200 figure appears.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The established moonlit ambience softly separates 찰리's damaged face from the subdued background, while both eyes remain completely unlit.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): B-200 has been destroyed, with only a component from his chest remaining, held against Charlie's chest. No articulated head, torso or limbs remain to form a whole-body pose.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Overturned trucks and the blast wreckage remain after the smoke clears; B-200 is destroyed, leaving the recovered chest component. Charlie holds that component with his newly shot and previously battered body, closing his eyes without an established permanent loss of power.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, bright moonlight.\n\nSHOT TEXT (authoritative, Korean): B-200의 부품을 가슴에 깊이 품은 채 두 눈의 디지털 불빛이 완전히 꺼진 찰리의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): At the blast site on the village departure road, among wrecked trucks and robot debris after the nighttime battle. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: B-200's remaining chest component (Retrieved after the explosion and held against 찰리's chest); used as Remain partly visible at the lower edge as the tangible object of grief; no intact B-200 figure appears.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The established moonlit ambience softly separates 찰리's damaged face from the subdued background, while both eyes remain completely unlit.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): B-200 has been destroyed, with only a component from his chest remaining, held against Charlie's chest. No articulated head, torso or limbs remain to form a whole-body pose.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Overturned trucks and the blast wreckage remain after the smoke clears; B-200 is destroyed, leaving the recovered chest component. Charlie holds that component with his newly shot and previously battered body, closing his eyes without an established permanent loss of power.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night, bright moonlight.\n\nSHOT TEXT (authoritative, Korean): B-200의 부품을 가슴에 깊이 품은 채 두 눈의 디지털 불빛이 완전히 꺼진 찰리의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): At the blast site on the village departure road, among wrecked trucks and robot debris after the nighttime battle. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: B-200's remaining chest component (Retrieved after the explosion and held against 찰리's chest); used as Remain partly visible at the lower edge as the tangible object of grief; no intact B-200 figure appears.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The established moonlit ambience softly separates 찰리's damaged face from the subdued background, while both eyes remain completely unlit.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): B-200 has been destroyed, with only a component from his chest remaining, held against Charlie's chest. No articulated head, torso or limbs remain to form a whole-body pose.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Overturned trucks and the blast wreckage remain after the smoke clears; B-200 is destroyed, leaving the recovered chest component. Charlie holds that component with his newly shot and previously battered body, closing his eyes without an established permanent loss of power.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S67sh70__bgfirst_bg.png",
     "asset_id": "24c611fd-00da-42ad-afe0-c4324cb61919",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S67sh70.png",
     "asset_id": "0909ba75-540e-4c7a-b173-2342ee20e201",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_village_truck_ambush_bde2df.png",
     "asset_id": "bb89f6c6-a1e1-4ea1-b9c4-7696c210aebc",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "찰리의 시선은 아래로 향해 가슴에 품은 부품 쪽에 머물고 있음.",
    "built_space": "달빛이 비치는 야간의 마을 폐허 도로이며, 주변에 부서진 구조물과 파편들이 널려 있음.",
    "entities": "샌드 베이지색 장갑과 흰색 마스크형 얼굴을 한 찰리가 등장하며, 두 눈의 불빛이 완전히 꺼져 있음. 가슴 앞에는 기계 부품이 위치함.",
    "hard_violations": [],
    "physics": "찰리의 양팔과 가슴이 기계 부품을 안정적으로 받치고 있음."
   },
   {
    "label": "B",
    "direction": "찰리의 시선은 정면 카메라 쪽을 향하고 있음.",
    "built_space": "달빛이 비치는 야간의 폐허로, 배경에 부서진 트럭 잔해와 파편들이 위치함.",
    "entities": "샌드 베이지색 장갑의 찰리가 등장하나, 두 눈에 붉은 불빛이 켜져 있음. 무릎 부근에서 기계 부품을 잡고 있음.",
    "hard_violations": [],
    "physics": "찰리의 양손이 기계 부품을 쥐고 다리 위에서 지탱하고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "눈의 불빛이 완전히 꺼진 상태와 부품을 가슴에 품은 클로즈업 구도를 프롬프트에 맞게 훌륭히 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "눈에 붉은 불빛이 켜져 있어 핵심 지시를 어겼으며, 구도가 지정된 클로즈업보다 지나치게 넓음."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 시선은 아래로 향해 가슴에 품은 부품 쪽에 머물고 있음.",
        "built_space": "달빛이 비치는 야간의 마을 폐허 도로이며, 주변에 부서진 구조물과 파편들이 널려 있음.",
        "entities": "샌드 베이지색 장갑과 흰색 마스크형 얼굴을 한 찰리가 등장하며, 두 눈의 불빛이 완전히 꺼져 있음. 가슴 앞에는 기계 부품이 위치함.",
        "hard_violations": [],
        "physics": "찰리의 양팔과 가슴이 기계 부품을 안정적으로 받치고 있음."
       },
       {
        "label": "B",
        "direction": "찰리의 시선은 정면 카메라 쪽을 향하고 있음.",
        "built_space": "달빛이 비치는 야간의 폐허로, 배경에 부서진 트럭 잔해와 파편들이 위치함.",
        "entities": "샌드 베이지색 장갑의 찰리가 등장하나, 두 눈에 붉은 불빛이 켜져 있음. 무릎 부근에서 기계 부품을 잡고 있음.",
        "hard_violations": [],
        "physics": "찰리의 양손이 기계 부품을 쥐고 다리 위에서 지탱하고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "눈의 불빛이 완전히 꺼진 상태와 부품을 가슴에 품은 클로즈업 구도를 프롬프트에 맞게 훌륭히 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "눈에 붉은 불빛이 켜져 있어 핵심 지시를 어겼으며, 구도가 지정된 클로즈업보다 지나치게 넓음."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 시선은 아래로 향해 가슴에 품은 부품 쪽에 머물고 있음.",
        "built_space": "달빛이 비치는 야간의 마을 폐허 도로이며, 주변에 부서진 구조물과 파편들이 널려 있음.",
        "entities": "샌드 베이지색 장갑과 흰색 마스크형 얼굴을 한 찰리가 등장하며, 두 눈의 불빛이 완전히 꺼져 있음. 가슴 앞에는 기계 부품이 위치함.",
        "hard_violations": [],
        "physics": "찰리의 양팔과 가슴이 기계 부품을 안정적으로 받치고 있음."
       },
       {
        "label": "B",
        "direction": "찰리의 시선은 정면 카메라 쪽을 향하고 있음.",
        "built_space": "달빛이 비치는 야간의 폐허로, 배경에 부서진 트럭 잔해와 파편들이 위치함.",
        "entities": "샌드 베이지색 장갑의 찰리가 등장하나, 두 눈에 붉은 불빛이 켜져 있음. 무릎 부근에서 기계 부품을 잡고 있음.",
        "hard_violations": [],
        "physics": "찰리의 양손이 기계 부품을 쥐고 다리 위에서 지탱하고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "꺼진 두 눈과 양손으로 받쳐 품은 흉부 잔해는 맞지만, 얼굴 클로즈업 대신 상체와 잔해를 크게 보여 주어 핵심 구도를 놓쳤다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "두 눈에 붉은 디지털 불빛이 남아 있고 무릎까지 드러나는 넓은 구도여서, 소등된 얼굴 클로즈업과 가슴 깊이 품는 동작 모두 어긋난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴은 거의 정면이며 눈은 카메라 쪽을 향한다. 양쪽 눈구멍 안에 붉은 발광점이 뚜렷하여 완전 소등 조건에 어긋난다. 두 손은 아래의 흉부 잔해를 향하지만, 잔해는 가슴보다 무릎 부근에 놓여 있다.",
        "built_space": "야외 도로 양옆으로 파손된 기와 건물과 잔해가 보인다. 왼쪽에는 똑바로 선 트럭 한 대, 오른쪽에는 기울어진 차량 잔해가 보인다. 찰리는 중앙에서 무릎을 벌리고 낮게 앉은 형태다. 참조의 폐허 마을 재료는 반영했지만, 화면을 넓혀 장소와 하체까지 보여 주므로 지정된 얼굴 클로즈업이 아니다.",
        "entities": "등장 인물은 로봇 찰리 하나다. 흰 각진 마스크, 두 안테나, 샌드 베이지 장갑과 육중한 팔은 참조와 대체로 일치한다. 표면에는 긁힘과 오염이 있지만 새 총상의 위치는 분명하지 않다. 하단에 파손된 기계 흉부 부품이 있으며 온전한 B-200은 없다. 달과 전투 잔해가 보이고 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "찰리의 양손이 잔해 윗부분과 옆부분에 닿고 잔해는 벌어진 무릎 사이에 걸쳐 있어 손과 무릎이 무게를 받는 형태다. 팔은 팔꿈치를 굽혀 몸통에 연결되어 있으며, 지지 없이 떠 있는 물체는 보이지 않는다. 다만 이는 부품을 가슴 깊이 끌어안기보다 무릎에 내려놓은 자세다."
       },
       {
        "label": "B",
        "direction": "찰리는 고개를 약간 숙이고 두 눈을 감은 듯한 비발광 상태다. 눈에서 디지털 불빛은 보이지 않는다. 양손은 가슴 아래의 흉부 잔해 양옆을 감싸며 안쪽으로 모여 있어 A보다 품는 동작에 가깝다.",
        "built_space": "달빛이 비치는 야외 폭발 현장으로, 왼쪽과 오른쪽 뒤에 옆으로 전복된 트럭이 각각 한 대씩 보인다. 왼쪽 위의 기와지붕 일부, 지면의 돌과 금속 잔해, 오른쪽의 도로 가드레일이 보인다. 참조 장소의 기와 건축과 돌담 배치는 이 화면만으로 확정하기 어렵다. 얼굴뿐 아니라 양어깨·가슴·팔 대부분이 포함되어 요청보다 넓은 상반신 구도다.",
        "entities": "찰리 하나만 등장하며 흰 마스크형 얼굴, 베이지 장갑판, 안테나와 굵은 기계 팔은 참조 정체성을 유지한다. 얼굴과 장갑에 마모가 있지만 새 총상은 뚜렷하지 않다. 양손에 든 물체는 머리와 팔다리가 없는 흉부 장갑 부품으로 읽히며 온전한 B-200은 없다. 다만 부품이 하단에 조금만 걸치는 것이 아니라 화면 아래쪽을 크게 차지한다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "흉부 잔해는 찰리의 양손과 굽힌 팔로 아래와 옆에서 받쳐져 있고 몸통 가까이에 놓여 있다. 손과 물체의 접촉이 보이므로 부유하는 잔해가 아니다. 하체 지지는 프레임 밖이라 판정할 수 없으며, 보이는 상체에 물리적으로 불가능한 자세는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "꺼진 두 눈과 양손으로 받쳐 품은 흉부 잔해는 맞지만, 얼굴 클로즈업 대신 상체와 잔해를 크게 보여 주어 핵심 구도를 놓쳤다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "두 눈에 붉은 디지털 불빛이 남아 있고 무릎까지 드러나는 넓은 구도여서, 소등된 얼굴 클로즈업과 가슴 깊이 품는 동작 모두 어긋난다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴은 거의 정면이며 눈은 카메라 쪽을 향한다. 양쪽 눈구멍 안에 붉은 발광점이 뚜렷하여 완전 소등 조건에 어긋난다. 두 손은 아래의 흉부 잔해를 향하지만, 잔해는 가슴보다 무릎 부근에 놓여 있다.",
        "built_space": "야외 도로 양옆으로 파손된 기와 건물과 잔해가 보인다. 왼쪽에는 똑바로 선 트럭 한 대, 오른쪽에는 기울어진 차량 잔해가 보인다. 찰리는 중앙에서 무릎을 벌리고 낮게 앉은 형태다. 참조의 폐허 마을 재료는 반영했지만, 화면을 넓혀 장소와 하체까지 보여 주므로 지정된 얼굴 클로즈업이 아니다.",
        "entities": "등장 인물은 로봇 찰리 하나다. 흰 각진 마스크, 두 안테나, 샌드 베이지 장갑과 육중한 팔은 참조와 대체로 일치한다. 표면에는 긁힘과 오염이 있지만 새 총상의 위치는 분명하지 않다. 하단에 파손된 기계 흉부 부품이 있으며 온전한 B-200은 없다. 달과 전투 잔해가 보이고 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "찰리의 양손이 잔해 윗부분과 옆부분에 닿고 잔해는 벌어진 무릎 사이에 걸쳐 있어 손과 무릎이 무게를 받는 형태다. 팔은 팔꿈치를 굽혀 몸통에 연결되어 있으며, 지지 없이 떠 있는 물체는 보이지 않는다. 다만 이는 부품을 가슴 깊이 끌어안기보다 무릎에 내려놓은 자세다."
       },
       {
        "label": "A",
        "direction": "찰리는 고개를 약간 숙이고 두 눈을 감은 듯한 비발광 상태다. 눈에서 디지털 불빛은 보이지 않는다. 양손은 가슴 아래의 흉부 잔해 양옆을 감싸며 안쪽으로 모여 있어 A보다 품는 동작에 가깝다.",
        "built_space": "달빛이 비치는 야외 폭발 현장으로, 왼쪽과 오른쪽 뒤에 옆으로 전복된 트럭이 각각 한 대씩 보인다. 왼쪽 위의 기와지붕 일부, 지면의 돌과 금속 잔해, 오른쪽의 도로 가드레일이 보인다. 참조 장소의 기와 건축과 돌담 배치는 이 화면만으로 확정하기 어렵다. 얼굴뿐 아니라 양어깨·가슴·팔 대부분이 포함되어 요청보다 넓은 상반신 구도다.",
        "entities": "찰리 하나만 등장하며 흰 마스크형 얼굴, 베이지 장갑판, 안테나와 굵은 기계 팔은 참조 정체성을 유지한다. 얼굴과 장갑에 마모가 있지만 새 총상은 뚜렷하지 않다. 양손에 든 물체는 머리와 팔다리가 없는 흉부 장갑 부품으로 읽히며 온전한 B-200은 없다. 다만 부품이 하단에 조금만 걸치는 것이 아니라 화면 아래쪽을 크게 차지한다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "흉부 잔해는 찰리의 양손과 굽힌 팔로 아래와 옆에서 받쳐져 있고 몸통 가까이에 놓여 있다. 손과 물체의 접촉이 보이므로 부유하는 잔해가 아니다. 하체 지지는 프레임 밖이라 판정할 수 없으며, 보이는 상체에 물리적으로 불가능한 자세는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.929
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.929
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 929
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "눈의 불빛이 완전히 꺼진 상태와 부품을 가슴에 품은 클로즈업 구도를 프롬프트에 맞게 훌륭히 구현함."
   },
   {
    "label": "B",
    "score": 929,
    "verdict_ko": "눈에 붉은 불빛이 켜져 있어 핵심 지시를 어겼으며, 구도가 지정된 클로즈업보다 지나치게 넓음."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_village_truck_ambush_bde2df.png",
    "asset_id": "bb89f6c6-a1e1-4ea1-b9c4-7696c210aebc",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-b48f-7366-beb6-b6e69eddf2bf",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S67sh70__bgfirst_bg.png",
   "bg_asset_id": "24c611fd-00da-42ad-afe0-c4324cb61919",
   "bg_record_key": "S67sh70::bgfirst_bg",
   "chain_winner": true,
   "authority": "groupbg",
   "group_key": "village_truck_ambush",
   "groupbg_asset_id": "bb89f6c6-a1e1-4ea1-b9c4-7696c210aebc"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S67sh70::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:56:40.182053+00:00",
  "fingerprint": "df485ac9cd3fc71169dacf56a5a2fda7df3f49c408c793f692bf69423b369e1b",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S67sh70_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S67sh70_sel.png",
  "source_sha256": "ce5dbf8a3e9a26a9fc8a21d294d4ad50d74fe98f2a2ea992bfd240426ec717cd",
  "file": "S67sh70_cine.png",
  "staged_sha256": "4c92c2048f270bd2b09e43064fd97dbf28ce77092c43a0c0b5ad9ad5dcf0913b",
  "latency_ms": 9744
 },
 "S68sh3::signage": {
  "fp": "4a83faa4cbfa3b76",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S68sh3": {
  "input_fingerprint": "052f65896b07a6fc",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 양 갈래 흙길에서 서로 다른 방향을 향해 달려가며 바퀴 뒤로 흙먼지를 일으키고 있는 두 낡은 트럭의 뒷모습 풀샷.\n\nLOCATION (lock): At a fork in the rural dirt road outside the village, where the two truck groups separate in daylight. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: 현우's truck on the left branch in the upper-left of the frame, midground, moves toward left branch extending away from the fork; 수빈's truck on the right branch in the upper-right of the frame, midground, moves toward right branch followed by the other vehicles.\n- KEY BACKGROUND ELEMENTS: 현우's truck (Old, roofless, and moving away along one branch) — Its rear and right side are visible as it recedes toward the upper left; used as Establish the branch the camera will subsequently follow; 수빈's truck and following vehicles (수빈's old roofless truck separates onto the other branch, with vehicles following it) — Rear profiles recede toward the upper right along the same branch; used as Make the separation of the two traveling groups spatially unambiguous; Forked dirt road (Both branches are being used, with dust rising behind the trucks' wheels) — The shared approach enters from the lower center and divides toward the upper corners; used as Provide a readable ground plan for the farewell and subsequent camera pursuit.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight keeps both branching routes and the departing trucks legible with restrained contrast and no added atmospheric treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Two old roofless trucks separate at the fork, with more vehicles following one branch. Charlie rides in the cargo bed with his battle damage unrepaired and the recovered B-200 chest component in his possession.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 양 갈래 흙길에서 서로 다른 방향을 향해 달려가며 바퀴 뒤로 흙먼지를 일으키고 있는 두 낡은 트럭의 뒷모습 풀샷.\n\nLOCATION (lock): At a fork in the rural dirt road outside the village, where the two truck groups separate in daylight. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: 현우's truck on the left branch in the upper-left of the frame, midground, moves toward left branch extending away from the fork; 수빈's truck on the right branch in the upper-right of the frame, midground, moves toward right branch followed by the other vehicles.\n- KEY BACKGROUND ELEMENTS: 현우's truck (Old, roofless, and moving away along one branch) — Its rear and right side are visible as it recedes toward the upper left; used as Establish the branch the camera will subsequently follow; 수빈's truck and following vehicles (수빈's old roofless truck separates onto the other branch, with vehicles following it) — Rear profiles recede toward the upper right along the same branch; used as Make the separation of the two traveling groups spatially unambiguous; Forked dirt road (Both branches are being used, with dust rising behind the trucks' wheels) — The shared approach enters from the lower center and divides toward the upper corners; used as Provide a readable ground plan for the farewell and subsequent camera pursuit.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight keeps both branching routes and the departing trucks legible with restrained contrast and no added atmospheric treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Two old roofless trucks separate at the fork, with more vehicles following one branch. Charlie rides in the cargo bed with his battle damage unrepaired and the recovered B-200 chest component in his possession.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 양 갈래 흙길에서 서로 다른 방향을 향해 달려가며 바퀴 뒤로 흙먼지를 일으키고 있는 두 낡은 트럭의 뒷모습 풀샷.\n\nLOCATION (lock): At a fork in the rural dirt road outside the village, where the two truck groups separate in daylight. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: 현우's truck on the left branch in the upper-left of the frame, midground, moves toward left branch extending away from the fork; 수빈's truck on the right branch in the upper-right of the frame, midground, moves toward right branch followed by the other vehicles.\n- KEY BACKGROUND ELEMENTS: 현우's truck (Old, roofless, and moving away along one branch) — Its rear and right side are visible as it recedes toward the upper left; used as Establish the branch the camera will subsequently follow; 수빈's truck and following vehicles (수빈's old roofless truck separates onto the other branch, with vehicles following it) — Rear profiles recede toward the upper right along the same branch; used as Make the separation of the two traveling groups spatially unambiguous; Forked dirt road (Both branches are being used, with dust rising behind the trucks' wheels) — The shared approach enters from the lower center and divides toward the upper corners; used as Provide a readable ground plan for the farewell and subsequent camera pursuit.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight keeps both branching routes and the departing trucks legible with restrained contrast and no added atmospheric treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Two old roofless trucks separate at the fork, with more vehicles following one branch. Charlie rides in the cargo bed with his battle damage unrepaired and the recovered B-200 chest component in his possession.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "두 트럭과 차량들이 갈래길을 따라 각각 멀어지는 방향으로 이동 중임.",
    "built_space": "화면 하단 중앙에서 시작되어 좌상단과 우상단으로 갈라지는 흙길이 명확하게 구현됨.",
    "entities": "왼쪽 트럭 짐칸에 짐이 덮여 있으나 찰리의 모습은 보이지 않음. 오른쪽 길의 오토바이에 사람이 탑승해 있음.",
    "hard_violations": [
     "[gemini-pro] 화면에 사람을 묘사하지 말라는 규칙을 위반하고 오토바이 탑승자(사람)를 생성함",
     "[gpt-high] 사람이 전혀 등장하지 않아야 하는 장면에 오토바이 탑승자가 명확히 등장한다."
    ],
    "physics": "차량들이 땅에 닿아 주행하며 바퀴 뒤로 흙먼지가 자연스럽게 발생함."
   },
   {
    "label": "B",
    "direction": "두 트럭과 오토바이 무리가 갈래길을 따라 양쪽으로 멀어지는 방향을 향하고 있음.",
    "built_space": "흙길이 중앙에서 좌우로 뻗어나가는 구조가 프레임에 맞게 배치됨.",
    "entities": "왼쪽 트럭의 짐칸이 완전히 비어 있어 찰리가 누락됨. 오른쪽 길을 따르는 다수의 오토바이에 사람들이 탑승해 있음.",
    "hard_violations": [
     "[gemini-pro] 화면에 사람을 묘사하지 말라는 규칙을 위반하고 다수의 오토바이 탑승자(사람)를 생성함",
     "[gpt-high] 사람이 전혀 등장하지 않아야 하는 장면에 오토바이 탑승자 세 명과 운전석의 사람 형상이 등장한다."
    ],
    "physics": "차량과 오토바이들이 노면 위를 주행하며 흙먼지를 일으키고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "사람을 포함하지 말라는 지침을 어기고 오토바이 탑승자를 생성했으며, 짐칸에 명시된 찰리(Charlie)가 누락되어 치명적 감점 요소가 됨."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "사람을 묘사하지 말라는 명시적 지침을 위반하고 다수의 오토바이 탑승자를 화면에 포함시켰으며, 짐칸이 비어 있어 실패함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 트럭과 차량들이 갈래길을 따라 각각 멀어지는 방향으로 이동 중임.",
        "built_space": "화면 하단 중앙에서 시작되어 좌상단과 우상단으로 갈라지는 흙길이 명확하게 구현됨.",
        "entities": "왼쪽 트럭 짐칸에 짐이 덮여 있으나 찰리의 모습은 보이지 않음. 오른쪽 길의 오토바이에 사람이 탑승해 있음.",
        "hard_violations": [
         "화면에 사람을 묘사하지 말라는 규칙을 위반하고 오토바이 탑승자(사람)를 생성함"
        ],
        "physics": "차량들이 땅에 닿아 주행하며 바퀴 뒤로 흙먼지가 자연스럽게 발생함."
       },
       {
        "label": "B",
        "direction": "두 트럭과 오토바이 무리가 갈래길을 따라 양쪽으로 멀어지는 방향을 향하고 있음.",
        "built_space": "흙길이 중앙에서 좌우로 뻗어나가는 구조가 프레임에 맞게 배치됨.",
        "entities": "왼쪽 트럭의 짐칸이 완전히 비어 있어 찰리가 누락됨. 오른쪽 길을 따르는 다수의 오토바이에 사람들이 탑승해 있음.",
        "hard_violations": [
         "화면에 사람을 묘사하지 말라는 규칙을 위반하고 다수의 오토바이 탑승자(사람)를 생성함"
        ],
        "physics": "차량과 오토바이들이 노면 위를 주행하며 흙먼지를 일으키고 있음."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "사람을 포함하지 말라는 지침을 어기고 오토바이 탑승자를 생성했으며, 짐칸에 명시된 찰리(Charlie)가 누락되어 치명적 감점 요소가 됨."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "사람을 묘사하지 말라는 명시적 지침을 위반하고 다수의 오토바이 탑승자를 화면에 포함시켰으며, 짐칸이 비어 있어 실패함."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "두 트럭과 차량들이 갈래길을 따라 각각 멀어지는 방향으로 이동 중임.",
        "built_space": "화면 하단 중앙에서 시작되어 좌상단과 우상단으로 갈라지는 흙길이 명확하게 구현됨.",
        "entities": "왼쪽 트럭 짐칸에 짐이 덮여 있으나 찰리의 모습은 보이지 않음. 오른쪽 길의 오토바이에 사람이 탑승해 있음.",
        "hard_violations": [
         "화면에 사람을 묘사하지 말라는 규칙을 위반하고 오토바이 탑승자(사람)를 생성함"
        ],
        "physics": "차량들이 땅에 닿아 주행하며 바퀴 뒤로 흙먼지가 자연스럽게 발생함."
       },
       {
        "label": "B",
        "direction": "두 트럭과 오토바이 무리가 갈래길을 따라 양쪽으로 멀어지는 방향을 향하고 있음.",
        "built_space": "흙길이 중앙에서 좌우로 뻗어나가는 구조가 프레임에 맞게 배치됨.",
        "entities": "왼쪽 트럭의 짐칸이 완전히 비어 있어 찰리가 누락됨. 오른쪽 길을 따르는 다수의 오토바이에 사람들이 탑승해 있음.",
        "hard_violations": [
         "화면에 사람을 묘사하지 말라는 규칙을 위반하고 다수의 오토바이 탑승자(사람)를 생성함"
        ],
        "physics": "차량과 오토바이들이 노면 위를 주행하며 흙먼지를 일으키고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "두 트럭의 좌우 상단 배치와 갈림길 풀샷은 더 정확하지만, 금지된 사람들이 등장하고 오른쪽 추가 차량들이 수빈의 트럭을 뒤따르지 않아 부적합하다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "갈라지는 주행과 흙먼지는 구현했지만, 사람 등장 금지를 위반하고 트럭들이 화면 중앙에 내려와 있으며 추가 차량도 선행한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 트럭은 왼쪽 위 갈래로, 오른쪽 트럭은 오른쪽 위 갈래로 카메라에서 멀어진다. 두 차량의 후면이 보이고 왼쪽 트럭의 측면도 드러난다. 오른쪽 도로의 추가 트럭들과 오토바이들은 같은 방향으로 향하지만, 주된 오른쪽 트럭보다 앞에 있어 '뒤따르는 차량' 관계와 반대다. 인물들의 정확한 시선은 판독하기 어렵다.",
        "built_space": "아래 중앙의 공통 흙길 하나가 좌우 두 갈래로 나뉘며, 두 주된 트럭은 각각 화면 왼쪽 위와 오른쪽 위의 중경에 놓인다. 오른쪽 갈래에 추가 트럭 두 대와 오토바이 세 대가 보인다. 주변은 경작지와 낮은 산이며 실내 설비나 반사는 없다. 낮의 농촌 갈림길이라는 공간 조건과 와이드 구도는 잘 읽힌다.",
        "entities": "낡고 녹슨 화물 트럭 두 대, 양 갈래 흙길, 바퀴 뒤 흙먼지가 보인다. 적재함은 개방되어 있고 운전석 상부는 프레임이 두드러지지만 지붕 제거 여부는 완전히 명확하지 않다. 오토바이 탑승자 세 명과 오른쪽 주 트럭 운전석의 사람 형상이 보여 무인 화면 조건에 어긋난다. 인물의 국적·나이·성별은 이 크기에서 확인할 수 없다. 찰리의 손상 상태와 B-200 부품은 식별되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "사람이 전혀 등장하지 않아야 하는 장면에 오토바이 탑승자 세 명과 운전석의 사람 형상이 등장한다."
        ],
        "physics": "트럭과 오토바이는 바퀴로 흙길에 지지되어 있고, 탑승자들은 좌석에 앉아 있다. 먼지는 바퀴 부근에서 뒤쪽으로 퍼져 주행 동작과 부합한다. 지지 없이 떠 있는 차체나 인물은 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "왼쪽 트럭은 왼쪽 위로 굽은 도로를 따라, 오른쪽 트럭은 오른쪽 위 도로를 따라 멀어진다. 두 트럭의 후면과 측면이 보인다. 오른쪽 갈래의 승용차, 오토바이, 먼 트럭은 모두 주된 오른쪽 트럭보다 앞에 있으므로 요구된 후속 차량 배치가 아니다. 오토바이 탑승자는 진행 방향을 향한다.",
        "built_space": "아래 중앙에서 들어오는 흙길 하나가 수풀을 사이에 두고 두 갈래로 나뉜다. 주된 트럭 두 대는 각 갈래에 있지만 화면 상단보다는 중앙 높이에 크게 놓인다. 오른쪽 갈래에는 승용차 한 대, 오토바이 한 대, 추가 트럭 한 대가 보인다. 배경에는 수풀, 산, 작은 건물과 전신주가 있으며 반사나 고정 설비의 중복 문제는 없다.",
        "entities": "낡은 주 트럭 두 대와 갈림길, 바퀴 뒤 흙먼지가 보인다. 적재함은 열려 있지만 운전석에는 지붕 외피가 남아 보여 '지붕 없는 트럭'과 맞지 않는다. 왼쪽 적재함에는 천과 여러 짐이 실려 있으나 찰리나 B-200으로 식별할 수 없다. 오른쪽 오토바이에는 최소 한 명의 사람이 명확히 보이며, 그 사람의 국적·나이·성별은 판독할 수 없다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "사람이 전혀 등장하지 않아야 하는 장면에 오토바이 탑승자가 명확히 등장한다."
        ],
        "physics": "차량들은 바퀴로 노면에 지지되고, 오토바이 탑승자는 좌석에 앉아 있다. 왼쪽 트럭의 짐과 천은 적재함 바닥과 난간에 놓여 지지된다. 바퀴 뒤 먼지와 진행 방향은 자연스럽고, 지지 없이 떠 있는 물체나 인물은 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "두 트럭의 좌우 상단 배치와 갈림길 풀샷은 더 정확하지만, 금지된 사람들이 등장하고 오른쪽 추가 차량들이 수빈의 트럭을 뒤따르지 않아 부적합하다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "갈라지는 주행과 흙먼지는 구현했지만, 사람 등장 금지를 위반하고 트럭들이 화면 중앙에 내려와 있으며 추가 차량도 선행한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽 트럭은 왼쪽 위 갈래로, 오른쪽 트럭은 오른쪽 위 갈래로 카메라에서 멀어진다. 두 차량의 후면이 보이고 왼쪽 트럭의 측면도 드러난다. 오른쪽 도로의 추가 트럭들과 오토바이들은 같은 방향으로 향하지만, 주된 오른쪽 트럭보다 앞에 있어 '뒤따르는 차량' 관계와 반대다. 인물들의 정확한 시선은 판독하기 어렵다.",
        "built_space": "아래 중앙의 공통 흙길 하나가 좌우 두 갈래로 나뉘며, 두 주된 트럭은 각각 화면 왼쪽 위와 오른쪽 위의 중경에 놓인다. 오른쪽 갈래에 추가 트럭 두 대와 오토바이 세 대가 보인다. 주변은 경작지와 낮은 산이며 실내 설비나 반사는 없다. 낮의 농촌 갈림길이라는 공간 조건과 와이드 구도는 잘 읽힌다.",
        "entities": "낡고 녹슨 화물 트럭 두 대, 양 갈래 흙길, 바퀴 뒤 흙먼지가 보인다. 적재함은 개방되어 있고 운전석 상부는 프레임이 두드러지지만 지붕 제거 여부는 완전히 명확하지 않다. 오토바이 탑승자 세 명과 오른쪽 주 트럭 운전석의 사람 형상이 보여 무인 화면 조건에 어긋난다. 인물의 국적·나이·성별은 이 크기에서 확인할 수 없다. 찰리의 손상 상태와 B-200 부품은 식별되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "사람이 전혀 등장하지 않아야 하는 장면에 오토바이 탑승자 세 명과 운전석의 사람 형상이 등장한다."
        ],
        "physics": "트럭과 오토바이는 바퀴로 흙길에 지지되어 있고, 탑승자들은 좌석에 앉아 있다. 먼지는 바퀴 부근에서 뒤쪽으로 퍼져 주행 동작과 부합한다. 지지 없이 떠 있는 차체나 인물은 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "왼쪽 트럭은 왼쪽 위로 굽은 도로를 따라, 오른쪽 트럭은 오른쪽 위 도로를 따라 멀어진다. 두 트럭의 후면과 측면이 보인다. 오른쪽 갈래의 승용차, 오토바이, 먼 트럭은 모두 주된 오른쪽 트럭보다 앞에 있으므로 요구된 후속 차량 배치가 아니다. 오토바이 탑승자는 진행 방향을 향한다.",
        "built_space": "아래 중앙에서 들어오는 흙길 하나가 수풀을 사이에 두고 두 갈래로 나뉜다. 주된 트럭 두 대는 각 갈래에 있지만 화면 상단보다는 중앙 높이에 크게 놓인다. 오른쪽 갈래에는 승용차 한 대, 오토바이 한 대, 추가 트럭 한 대가 보인다. 배경에는 수풀, 산, 작은 건물과 전신주가 있으며 반사나 고정 설비의 중복 문제는 없다.",
        "entities": "낡은 주 트럭 두 대와 갈림길, 바퀴 뒤 흙먼지가 보인다. 적재함은 열려 있지만 운전석에는 지붕 외피가 남아 보여 '지붕 없는 트럭'과 맞지 않는다. 왼쪽 적재함에는 천과 여러 짐이 실려 있으나 찰리나 B-200으로 식별할 수 없다. 오른쪽 오토바이에는 최소 한 명의 사람이 명확히 보이며, 그 사람의 국적·나이·성별은 판독할 수 없다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "사람이 전혀 등장하지 않아야 하는 장면에 오토바이 탑승자가 명확히 등장한다."
        ],
        "physics": "차량들은 바퀴로 노면에 지지되고, 오토바이 탑승자는 좌석에 앉아 있다. 왼쪽 트럭의 짐과 천은 적재함 바닥과 난간에 놓여 지지된다. 바퀴 뒤 먼지와 진행 방향은 자연스럽고, 지지 없이 떠 있는 물체나 인물은 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.667,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.417,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 화면에 사람을 묘사하지 말라는 규칙을 위반하고 오토바이 탑승자(사람)를 생성함",
     "[gpt-high] 사람이 전혀 등장하지 않아야 하는 장면에 오토바이 탑승자가 명확히 등장한다."
    ],
    "B": [
     "[gemini-pro] 화면에 사람을 묘사하지 말라는 규칙을 위반하고 다수의 오토바이 탑승자(사람)를 생성함",
     "[gpt-high] 사람이 전혀 등장하지 않아야 하는 장면에 오토바이 탑승자 세 명과 운전석의 사람 형상이 등장한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1417,
   "B": 1750
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1417,
    "verdict_ko": "사람을 포함하지 말라는 지침을 어기고 오토바이 탑승자를 생성했으며, 짐칸에 명시된 찰리(Charlie)가 누락되어 치명적 감점 요소가 됨.  ★위반: [gemini-pro] 화면에 사람을 묘사하지 말라는 규칙을 위반하고 오토바이 탑승자(사람)를 생성함 / [gpt-high] 사람이 전혀 등장하지 않아야 하는 장면에 오토바이 탑승자가 명확히 등장한다."
   },
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "사람을 묘사하지 말라는 명시적 지침을 위반하고 다수의 오토바이 탑승자를 화면에 포함시켰으며, 짐칸이 비어 있어 실패함.  ★위반: [gemini-pro] 화면에 사람을 묘사하지 말라는 규칙을 위반하고 다수의 오토바이 탑승자(사람)를 생성함 / [gpt-high] 사람이 전혀 등장하지 않아야 하는 장면에 오토바이 탑승자 세 명과 운전석의 사람 형상이 등장한다."
   }
  ],
  "refs": [],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-b957-744f-b143-c62728240ca9",
  "ref_mode": "lane(map_marker): 스케치",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S68sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:13:06.518654+00:00",
  "fingerprint": "0214ba487f2d4ea1b6711587d3826031db40e3a15cf7c280b253d8e87b068db3",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S68sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S68sh3_sel.png",
  "source_sha256": "f7f360c88b06c2faaba0f604c26c7f8e244f6e5dc1c7e6b172d44e8ce62639be",
  "file": "S68sh3_cine.png",
  "staged_sha256": "48556c0842a3d5a93b9b4c69c15c26b18f13a740db6cb6f092f0a59119e1753b",
  "latency_ms": 10034
 },
 "S68sh7::signage": {
  "fp": "28960e5541b3014d",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S68sh7": {
  "input_fingerprint": "43f93f215e52084f",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 손가락 위의 아기 새(B-200이 돌보던 새)를 가만히 내려다보며 눈을 지그시 감은 찰리의 낡은 금속 얼굴 클로즈업.\n\nLOCATION (lock): In the open rear cargo bed of a moving old truck on the rural road, with unobstructed daylight above. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Truck cargo bed (Carrying the seated 찰리 during the journey) — Only a narrow oblique portion remains visible behind his lower shoulder; used as Anchor the farewell in the moving truck before the dream transition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight reveals the worn metal face and the small bird with gentle contrast, without introducing dream coloration before the dissolve.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie sits in the roofless truck's cargo bed with the recovered B-200 chest component and his unrepaired battle damage. A baby bird formerly sheltered in the warehouse now perches on his finger.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 아기 새(B-200이 돌보던 새) (손가락 위에 앉는 작은 크기, 둥근 몸통, 작은 부리, 짧은 날개, 가느다란 발) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 손가락 위의 아기 새(B-200이 돌보던 새)를 가만히 내려다보며 눈을 지그시 감은 찰리의 낡은 금속 얼굴 클로즈업.\n\nLOCATION (lock): In the open rear cargo bed of a moving old truck on the rural road, with unobstructed daylight above. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Truck cargo bed (Carrying the seated 찰리 during the journey) — Only a narrow oblique portion remains visible behind his lower shoulder; used as Anchor the farewell in the moving truck before the dream transition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight reveals the worn metal face and the small bird with gentle contrast, without introducing dream coloration before the dissolve.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie sits in the roofless truck's cargo bed with the recovered B-200 chest component and his unrepaired battle damage. A baby bird formerly sheltered in the warehouse now perches on his finger.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 아기 새(B-200이 돌보던 새) (손가락 위에 앉는 작은 크기, 둥근 몸통, 작은 부리, 짧은 날개, 가느다란 발) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 손가락 위의 아기 새(B-200이 돌보던 새)를 가만히 내려다보며 눈을 지그시 감은 찰리의 낡은 금속 얼굴 클로즈업.\n\nLOCATION (lock): In the open rear cargo bed of a moving old truck on the rural road, with unobstructed daylight above. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Truck cargo bed (Carrying the seated 찰리 during the journey) — Only a narrow oblique portion remains visible behind his lower shoulder; used as Anchor the farewell in the moving truck before the dream transition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight reveals the worn metal face and the small bird with gentle contrast, without introducing dream coloration before the dissolve.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie sits in the roofless truck's cargo bed with the recovered B-200 chest component and his unrepaired battle damage. A baby bird formerly sheltered in the warehouse now perches on his finger.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 아기 새(B-200이 돌보던 새) (손가락 위에 앉는 작은 크기, 둥근 몸통, 작은 부리, 짧은 날개, 가느다란 발) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S68sh7__bgfirst_bg.png",
     "asset_id": "16f6020b-ef17-4edc-b91e-0f09780d6056",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S68sh7.png",
     "asset_id": "aa8d3a60-5597-4fad-886f-90bc9f48a604",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 아기 새(B-200이 돌보던 새): the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:833571>",
     "asset_id": "0eaf7589-8c85-4da6-aebd-27ac974cb20f",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_open_truck_bed_9f9871.png",
     "asset_id": "f8675ac4-ad27-42ab-a598-758f4a775a49",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 아기 새(B-200이 돌보던 새): the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:833571>",
     "asset_id": "0eaf7589-8c85-4da6-aebd-27ac974cb20f",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "찰리가 손가락 위에 있는 아기 새를 향해 고개를 숙이고 있으며 눈을 감고 있음.",
    "built_space": "트럭 적재함 내부에 위치하며, 왼쪽 배경으로 레퍼런스의 시골길이 보임. 적재함 벽면이 왼쪽에 다소 넓게 배치됨.",
    "entities": "찰리의 마스크와 낡은 장갑판, 아기 새의 형태가 레퍼런스와 일치함.",
    "hard_violations": [
     "[gemini-pro] 프롬프트 지시문 텍스트('손가락 위의 아기 새...')가 트럭 적재함 벽면에 그대로 유출되어 적혀 있음 (읽을 수 있는 글씨 금지 위반).",
     "[gpt-high] 적재함 뒤쪽 문에 '손가락 위의 아기 새' 등 프롬프트의 촬영 지시문이 읽을 수 있는 글자로 노출되어, 지시문 유출 및 화면 내 가독성 있는 문자 금지 조건을 위반한다."
    ],
    "physics": "아기 새가 찰리의 기계 손가락 위에 물리적으로 안정감 있게 앉아 있음."
   },
   {
    "label": "B",
    "direction": "찰리가 손가락 위에 앉은 아기 새 쪽으로 고개를 숙이고 눈을 지그시 감고 있음.",
    "built_space": "야외 시골길을 달리는 트럭 적재함 내부. 찰리의 어깨 뒤로 적재함 벽면 일부와 풍경이 보임.",
    "entities": "찰리의 외형(얼굴, 장갑판)과 아기 새 모두 레퍼런스의 특징을 정확하게 구현함.",
    "hard_violations": [],
    "physics": "아기 새가 찰리의 구부린 기계 손가락 위에 자연스럽게 지탱되어 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "요청된 찰리와 아기 새의 디자인을 잘 반영했으며, 달리는 트럭의 배경과 인물의 자세 및 감정 표현이 프롬프트와 일치합니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "프롬프트의 지시문 텍스트가 배경 트럭 적재함에 그대로 노출되는 치명적인 오류가 발생하여 실격입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리가 손가락 위에 있는 아기 새를 향해 고개를 숙이고 있으며 눈을 감고 있음.",
        "built_space": "트럭 적재함 내부에 위치하며, 왼쪽 배경으로 레퍼런스의 시골길이 보임. 적재함 벽면이 왼쪽에 다소 넓게 배치됨.",
        "entities": "찰리의 마스크와 낡은 장갑판, 아기 새의 형태가 레퍼런스와 일치함.",
        "hard_violations": [
         "프롬프트 지시문 텍스트('손가락 위의 아기 새...')가 트럭 적재함 벽면에 그대로 유출되어 적혀 있음 (읽을 수 있는 글씨 금지 위반)."
        ],
        "physics": "아기 새가 찰리의 기계 손가락 위에 물리적으로 안정감 있게 앉아 있음."
       },
       {
        "label": "B",
        "direction": "찰리가 손가락 위에 앉은 아기 새 쪽으로 고개를 숙이고 눈을 지그시 감고 있음.",
        "built_space": "야외 시골길을 달리는 트럭 적재함 내부. 찰리의 어깨 뒤로 적재함 벽면 일부와 풍경이 보임.",
        "entities": "찰리의 외형(얼굴, 장갑판)과 아기 새 모두 레퍼런스의 특징을 정확하게 구현함.",
        "hard_violations": [],
        "physics": "아기 새가 찰리의 구부린 기계 손가락 위에 자연스럽게 지탱되어 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "요청된 찰리와 아기 새의 디자인을 잘 반영했으며, 달리는 트럭의 배경과 인물의 자세 및 감정 표현이 프롬프트와 일치합니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "프롬프트의 지시문 텍스트가 배경 트럭 적재함에 그대로 노출되는 치명적인 오류가 발생하여 실격입니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리가 손가락 위에 있는 아기 새를 향해 고개를 숙이고 있으며 눈을 감고 있음.",
        "built_space": "트럭 적재함 내부에 위치하며, 왼쪽 배경으로 레퍼런스의 시골길이 보임. 적재함 벽면이 왼쪽에 다소 넓게 배치됨.",
        "entities": "찰리의 마스크와 낡은 장갑판, 아기 새의 형태가 레퍼런스와 일치함.",
        "hard_violations": [
         "프롬프트 지시문 텍스트('손가락 위의 아기 새...')가 트럭 적재함 벽면에 그대로 유출되어 적혀 있음 (읽을 수 있는 글씨 금지 위반)."
        ],
        "physics": "아기 새가 찰리의 기계 손가락 위에 물리적으로 안정감 있게 앉아 있음."
       },
       {
        "label": "B",
        "direction": "찰리가 손가락 위에 앉은 아기 새 쪽으로 고개를 숙이고 눈을 지그시 감고 있음.",
        "built_space": "야외 시골길을 달리는 트럭 적재함 내부. 찰리의 어깨 뒤로 적재함 벽면 일부와 풍경이 보임.",
        "entities": "찰리의 외형(얼굴, 장갑판)과 아기 새 모두 레퍼런스의 특징을 정확하게 구현함.",
        "hard_violations": [],
        "physics": "아기 새가 찰리의 구부린 기계 손가락 위에 자연스럽게 지탱되어 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "눈을 감고 손가락 위 새를 향해 고개를 숙인 행동과 찰리의 정체성이 잘 맞지만, 얼굴 클로즈업보다 넓어 상체와 적재함이 과하게 보인다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "새를 향해 숙인 얼굴과 감은 눈은 맞지만, 적재함에 지시문이 읽히도록 새겨진 치명적 위반이 있고 배경도 지정된 좁은 조각보다 훨씬 넓다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 눈을 감고 고개를 아래로 숙였으며, 얼굴 바로 아래 손가락에 앉은 새를 향한다. 새는 대체로 카메라 쪽을 바라본다. 찰리가 새를 내려다보다 눈을 감은 순간으로 읽힌다.",
        "built_space": "화면 오른쪽에 녹슨 적재함 측판 한 면이 보이고 왼쪽 가장자리에도 반대편 측판 일부가 보인다. 지붕 없이 낮 하늘이 열려 있으며 논밭과 전신주가 장소 참조와 부합한다. 중복된 설비는 없다. 다만 양어깨와 가슴, 팔까지 포함한 구도여서 얼굴 중심 클로즈업보다 넓고, 오른쪽 측판도 어깨 아래의 좁은 사선 조각 이상으로 드러난다. 앉은 자리와 바닥 접촉은 화면 밖이다.",
        "entities": "찰리 한 개체와 아기 새 한 마리만 보인다. 찰리의 샌드 베이지 각진 장갑판, 흰 분할 마스크, 검은 목 기구와 육중한 팔은 참조에 가깝다. 얼굴과 장갑에 마모가 있다. 새는 황갈색의 둥근 몸통, 작은 부리, 짧은 접힌 날개와 가는 발을 갖춰 참조와 부합한다. 회수한 B-200 흉부 부품은 식별되지 않지만 하단이 잘려 있어 소지 여부를 단정할 수 없다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "새의 발이 수평으로 뻗은 금속 손가락에 닿아 몸을 지탱한다. 손은 손목과 굽힌 팔에 연결되어 가슴 앞에 들려 있고 머리는 목 기구로 지지된다. 하체가 잘려 좌면 접촉은 확인되지 않으나, 공중에 떠 있거나 지지 없이 분리된 신체는 보이지 않는다. 트럭의 이동 여부는 정지 화면만으로 확정하기 어렵다."
       },
       {
        "label": "B",
        "direction": "찰리는 화면 왼쪽 아래의 새 쪽으로 얼굴을 숙이고 눈을 감고 있다. 새의 부리와 머리는 오른쪽의 찰리를 향한다. 두 대상의 방향 관계는 요청한 조용한 작별 순간과 맞는다.",
        "built_space": "뒤쪽 적재함 문 한 면, 왼쪽 측판 한 면, 그 사이 바닥이 넓게 보인다. 녹슨 밝은 금속과 낡은 바닥, 지붕 없는 구조는 장소 참조와 부합하며 설비 중복은 없다. 그러나 적재함 전체 모서리와 도로까지 크게 펼쳐져, 어깨 아래에 좁은 사선 부분만 남기라는 배경 제한을 벗어난다. 찰리의 좌면과 하체는 화면 밖이다.",
        "entities": "찰리 한 개체와 새 한 마리가 보인다. 베이지 장갑과 흰 금속 얼굴은 맞지만 얼굴 판의 분할, 이마 점 배열과 눈 주변 형태는 참조와 차이가 난다. 새의 황갈색 깃털, 둥근 몸통, 작은 부리와 가는 발은 대체로 맞는다. 얼굴과 장갑의 긁힘은 보이며, 별도로 회수한 흉부 부품인지는 확인되지 않는다. 적재함 뒤쪽 문에는 촬영 지시문을 옮긴 한글 문장이 선명하게 보인다.",
        "hard_violations": [
         "적재함 뒤쪽 문에 '손가락 위의 아기 새' 등 프롬프트의 촬영 지시문이 읽을 수 있는 글자로 노출되어, 지시문 유출 및 화면 내 가독성 있는 문자 금지 조건을 위반한다."
        ],
        "physics": "새는 두 발로 금속 손가락을 딛고 있으며 발가락이 표면을 감싸 체중을 지지한다. 손가락은 화면 아래의 손 기구에 연결되고 머리는 목의 기계 구조에 연결된다. 보이는 범위에 지지 없는 부유 물체나 불가능한 신체 배치는 없다. 앉은 자세의 하체 접촉과 트럭의 실제 이동은 이 화면에서 확인할 수 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "눈을 감고 손가락 위 새를 향해 고개를 숙인 행동과 찰리의 정체성이 잘 맞지만, 얼굴 클로즈업보다 넓어 상체와 적재함이 과하게 보인다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "새를 향해 숙인 얼굴과 감은 눈은 맞지만, 적재함에 지시문이 읽히도록 새겨진 치명적 위반이 있고 배경도 지정된 좁은 조각보다 훨씬 넓다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "찰리는 눈을 감고 고개를 아래로 숙였으며, 얼굴 바로 아래 손가락에 앉은 새를 향한다. 새는 대체로 카메라 쪽을 바라본다. 찰리가 새를 내려다보다 눈을 감은 순간으로 읽힌다.",
        "built_space": "화면 오른쪽에 녹슨 적재함 측판 한 면이 보이고 왼쪽 가장자리에도 반대편 측판 일부가 보인다. 지붕 없이 낮 하늘이 열려 있으며 논밭과 전신주가 장소 참조와 부합한다. 중복된 설비는 없다. 다만 양어깨와 가슴, 팔까지 포함한 구도여서 얼굴 중심 클로즈업보다 넓고, 오른쪽 측판도 어깨 아래의 좁은 사선 조각 이상으로 드러난다. 앉은 자리와 바닥 접촉은 화면 밖이다.",
        "entities": "찰리 한 개체와 아기 새 한 마리만 보인다. 찰리의 샌드 베이지 각진 장갑판, 흰 분할 마스크, 검은 목 기구와 육중한 팔은 참조에 가깝다. 얼굴과 장갑에 마모가 있다. 새는 황갈색의 둥근 몸통, 작은 부리, 짧은 접힌 날개와 가는 발을 갖춰 참조와 부합한다. 회수한 B-200 흉부 부품은 식별되지 않지만 하단이 잘려 있어 소지 여부를 단정할 수 없다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "새의 발이 수평으로 뻗은 금속 손가락에 닿아 몸을 지탱한다. 손은 손목과 굽힌 팔에 연결되어 가슴 앞에 들려 있고 머리는 목 기구로 지지된다. 하체가 잘려 좌면 접촉은 확인되지 않으나, 공중에 떠 있거나 지지 없이 분리된 신체는 보이지 않는다. 트럭의 이동 여부는 정지 화면만으로 확정하기 어렵다."
       },
       {
        "label": "A",
        "direction": "찰리는 화면 왼쪽 아래의 새 쪽으로 얼굴을 숙이고 눈을 감고 있다. 새의 부리와 머리는 오른쪽의 찰리를 향한다. 두 대상의 방향 관계는 요청한 조용한 작별 순간과 맞는다.",
        "built_space": "뒤쪽 적재함 문 한 면, 왼쪽 측판 한 면, 그 사이 바닥이 넓게 보인다. 녹슨 밝은 금속과 낡은 바닥, 지붕 없는 구조는 장소 참조와 부합하며 설비 중복은 없다. 그러나 적재함 전체 모서리와 도로까지 크게 펼쳐져, 어깨 아래에 좁은 사선 부분만 남기라는 배경 제한을 벗어난다. 찰리의 좌면과 하체는 화면 밖이다.",
        "entities": "찰리 한 개체와 새 한 마리가 보인다. 베이지 장갑과 흰 금속 얼굴은 맞지만 얼굴 판의 분할, 이마 점 배열과 눈 주변 형태는 참조와 차이가 난다. 새의 황갈색 깃털, 둥근 몸통, 작은 부리와 가는 발은 대체로 맞는다. 얼굴과 장갑의 긁힘은 보이며, 별도로 회수한 흉부 부품인지는 확인되지 않는다. 적재함 뒤쪽 문에는 촬영 지시문을 옮긴 한글 문장이 선명하게 보인다.",
        "hard_violations": [
         "적재함 뒤쪽 문에 '손가락 위의 아기 새' 등 프롬프트의 촬영 지시문이 읽을 수 있는 글자로 노출되어, 지시문 유출 및 화면 내 가독성 있는 문자 금지 조건을 위반한다."
        ],
        "physics": "새는 두 발로 금속 손가락을 딛고 있으며 발가락이 표면을 감싸 체중을 지지한다. 손가락은 화면 아래의 손 기구에 연결되고 머리는 목의 기계 구조에 연결된다. 보이는 범위에 지지 없는 부유 물체나 불가능한 신체 배치는 없다. 앉은 자세의 하체 접촉과 트럭의 실제 이동은 이 화면에서 확인할 수 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.714,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.464,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 프롬프트 지시문 텍스트('손가락 위의 아기 새...')가 트럭 적재함 벽면에 그대로 유출되어 적혀 있음 (읽을 수 있는 글씨 금지 위반).",
     "[gpt-high] 적재함 뒤쪽 문에 '손가락 위의 아기 새' 등 프롬프트의 촬영 지시문이 읽을 수 있는 글자로 노출되어, 지시문 유출 및 화면 내 가독성 있는 문자 금지 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 464
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "요청된 찰리와 아기 새의 디자인을 잘 반영했으며, 달리는 트럭의 배경과 인물의 자세 및 감정 표현이 프롬프트와 일치합니다."
   },
   {
    "label": "A",
    "score": 464,
    "verdict_ko": "프롬프트의 지시문 텍스트가 배경 트럭 적재함에 그대로 노출되는 치명적인 오류가 발생하여 실격입니다.  ★위반: [gemini-pro] 프롬프트 지시문 텍스트('손가락 위의 아기 새...')가 트럭 적재함 벽면에 그대로 유출되어 적혀 있음 (읽을 수 있는 글씨 금지 위반). / [gpt-high] 적재함 뒤쪽 문에 '손가락 위의 아기 새' 등 프롬프트의 촬영 지시문이 읽을 수 있는 글자로 노출되어, 지시문 유출 및 화면 내 가독성 있는 문자 금지 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_open_truck_bed_9f9871.png",
    "asset_id": "f8675ac4-ad27-42ab-a598-758f4a775a49",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 아기 새(B-200이 돌보던 새): the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:833571>",
    "asset_id": "0eaf7589-8c85-4da6-aebd-27ac974cb20f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-baff-7562-86b5-1ba8a8d8b4eb",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S68sh7__bgfirst_bg.png",
   "bg_asset_id": "16f6020b-ef17-4edc-b91e-0f09780d6056",
   "bg_record_key": "S68sh7::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "open_truck_bed",
   "groupbg_asset_id": "f8675ac4-ad27-42ab-a598-758f4a775a49"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S68sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:58:33.134334+00:00",
  "fingerprint": "663f0aa4c563e53257efa48a065b6e497d5049e355175a99c1c95c10b014a1cf",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S68sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S68sh7_sel.png",
  "source_sha256": "ccb0ccb08d5c4525245ed0013a5452def375dbb6088bfcefa60273445d418608",
  "file": "S68sh7_cine.png",
  "staged_sha256": "4578db3dfea59fdec84b8c5e90bdf30effa239d3b636c16dbfcc884940bb5220",
  "latency_ms": 13019
 },
 "S69sh3::signage": {
  "fp": "12dedf02e3cfd39a",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::b540692c4ba06c93": {
  "subjects": [],
  "subject_text": "현우의 트럭 적재함\n지붕 없이 위로 열린 낡은 적재 공간. 낮은 측벽이 바닥을 둘러싸고, 한쪽에는 옮겨 펼칠 수 있는 그늘막이 놓여 있다.",
  "identity": "canonical",
  "scope_id": "L233",
  "scope_role": "location_exterior",
  "scope_sha": "df9b8f1b489c5ca2"
 },
 "S69sh3::bgfirst_bg": {
  "input_fingerprint": "20bd23281715627f",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 잠든 일행의 위로 커다란 그늘 막 천을 양손으로 넓게 펼쳐 든 찰리의 역동적인 상체.\n\nLOCATION (lock): Over the open rear cargo bed of a truck parked beside a shaded rock formation, where a cloth canopy shields the sleepers.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: shade cloth stretched above the sleeping group in the upper-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Shade cloth (Being spread and repositioned by 찰리 to shield the sleepers) — Its underside extends obliquely across the upper third, above rather than in front of the sleeping faces; used as Connect both raised hands and visibly shelter the group without obscuring 찰리's face; Parked truck cargo bed (Stationary, with 현우, 앰버, and 라울 asleep head-to-head) — Seen from outside the same side established during 찰리's approach; used as Provide the shared spatial base beneath the raised cloth; Rock formations beside the truck (The truck is parked in their shade); used as Provide a restrained background boundary locating the resting place.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight is softened by the existing rock shade and the cloth being moved over the sleeping group, making the protective reduction of light the emotional center.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 잠든 일행의 위로 커다란 그늘 막 천을 양손으로 넓게 펼쳐 든 찰리의 역동적인 상체.\n\nLOCATION (lock): Over the open rear cargo bed of a truck parked beside a shaded rock formation, where a cloth canopy shields the sleepers.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: shade cloth stretched above the sleeping group in the upper-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Shade cloth (Being spread and repositioned by 찰리 to shield the sleepers) — Its underside extends obliquely across the upper third, above rather than in front of the sleeping faces; used as Connect both raised hands and visibly shelter the group without obscuring 찰리's face; Parked truck cargo bed (Stationary, with 현우, 앰버, and 라울 asleep head-to-head) — Seen from outside the same side established during 찰리's approach; used as Provide the shared spatial base beneath the raised cloth; Rock formations beside the truck (The truck is parked in their shade); used as Provide a restrained background boundary locating the resting place.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight is softened by the existing rock shade and the cloth being moved over the sleeping group, making the protective reduction of light the emotional center.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S69sh3__bgfirst_bg.png",
  "asset_id": "ccf56b84-5714-4267-b1b6-cfb4a6ff27b2",
  "input_asset_ids": [
   "acde8000-023e-40d0-b3ff-468a1735a66a",
   "c94c53e3-76a2-4f78-b4ae-95204318d3c2"
  ]
 },
 "S69sh3": {
  "input_fingerprint": "8276c2232dea64a8",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 잠든 일행의 위로 커다란 그늘 막 천을 양손으로 넓게 펼쳐 든 찰리의 역동적인 상체.\n\nLOCATION (lock): Over the open rear cargo bed of a truck parked beside a shaded rock formation, where a cloth canopy shields the sleepers. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: shade cloth stretched above the sleeping group in the upper-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Shade cloth (Being spread and repositioned by 찰리 to shield the sleepers) — Its underside extends obliquely across the upper third, above rather than in front of the sleeping faces; used as Connect both raised hands and visibly shelter the group without obscuring 찰리's face; Parked truck cargo bed (Stationary, with 현우, 앰버, and 라울 asleep head-to-head) — Seen from outside the same side established during 찰리's approach; used as Provide the shared spatial base beneath the raised cloth; Rock formations beside the truck (The truck is parked in their shade); used as Provide a restrained background boundary locating the resting place.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight is softened by the existing rock shade and the cloth being moved over the sleeping group, making the protective reduction of light the emotional center.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Hyunwoo is asleep in the truck's rear cargo bed with his head grouped closely beside Amber's and Raul's heads beneath the shade Charlie provides. The source does not specify his torso's orientation or the positions of his arms and legs.\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Amber is asleep in the truck's rear cargo bed with her head grouped closely beside Hyunwoo's and Raul's heads beneath the shade Charlie provides. The source does not specify her torso's orientation or the positions of her arms and legs.\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Raul is asleep in the truck's rear cargo bed with his head grouped closely beside Hyunwoo's and Amber's heads beneath the shade Charlie provides. The source does not specify his torso's orientation or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The roofless truck is parked beside rocks for shade, with a movable shade cloth over the cargo-bed resting area. Charlie adjusts the cloth while retaining his unrepaired battle damage and the recovered B-200 chest component in his possession. 현우: He is asleep in the rear cargo bed, still bearing his recent injuries. 앰버: She is asleep in the rear cargo bed; no treatment of her recent head injury has been established. 라울: He is asleep in the rear cargo bed.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 잠든 일행의 위로 커다란 그늘 막 천을 양손으로 넓게 펼쳐 든 찰리의 역동적인 상체.\n\nLOCATION (lock): Over the open rear cargo bed of a truck parked beside a shaded rock formation, where a cloth canopy shields the sleepers. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: shade cloth stretched above the sleeping group in the upper-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Shade cloth (Being spread and repositioned by 찰리 to shield the sleepers) — Its underside extends obliquely across the upper third, above rather than in front of the sleeping faces; used as Connect both raised hands and visibly shelter the group without obscuring 찰리's face; Parked truck cargo bed (Stationary, with 현우, 앰버, and 라울 asleep head-to-head) — Seen from outside the same side established during 찰리's approach; used as Provide the shared spatial base beneath the raised cloth; Rock formations beside the truck (The truck is parked in their shade); used as Provide a restrained background boundary locating the resting place.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight is softened by the existing rock shade and the cloth being moved over the sleeping group, making the protective reduction of light the emotional center.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Hyunwoo is asleep in the truck's rear cargo bed with his head grouped closely beside Amber's and Raul's heads beneath the shade Charlie provides. The source does not specify his torso's orientation or the positions of his arms and legs.\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Amber is asleep in the truck's rear cargo bed with her head grouped closely beside Hyunwoo's and Raul's heads beneath the shade Charlie provides. The source does not specify her torso's orientation or the positions of her arms and legs.\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Raul is asleep in the truck's rear cargo bed with his head grouped closely beside Hyunwoo's and Amber's heads beneath the shade Charlie provides. The source does not specify his torso's orientation or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The roofless truck is parked beside rocks for shade, with a movable shade cloth over the cargo-bed resting area. Charlie adjusts the cloth while retaining his unrepaired battle damage and the recovered B-200 chest component in his possession. 현우: He is asleep in the rear cargo bed, still bearing his recent injuries. 앰버: She is asleep in the rear cargo bed; no treatment of her recent head injury has been established. 라울: He is asleep in the rear cargo bed.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 잠든 일행의 위로 커다란 그늘 막 천을 양손으로 넓게 펼쳐 든 찰리의 역동적인 상체.\n\nLOCATION (lock): Over the open rear cargo bed of a truck parked beside a shaded rock formation, where a cloth canopy shields the sleepers. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: shade cloth stretched above the sleeping group in the upper-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Shade cloth (Being spread and repositioned by 찰리 to shield the sleepers) — Its underside extends obliquely across the upper third, above rather than in front of the sleeping faces; used as Connect both raised hands and visibly shelter the group without obscuring 찰리's face; Parked truck cargo bed (Stationary, with 현우, 앰버, and 라울 asleep head-to-head) — Seen from outside the same side established during 찰리's approach; used as Provide the shared spatial base beneath the raised cloth; Rock formations beside the truck (The truck is parked in their shade); used as Provide a restrained background boundary locating the resting place.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight is softened by the existing rock shade and the cloth being moved over the sleeping group, making the protective reduction of light the emotional center.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Hyunwoo is asleep in the truck's rear cargo bed with his head grouped closely beside Amber's and Raul's heads beneath the shade Charlie provides. The source does not specify his torso's orientation or the positions of his arms and legs.\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Amber is asleep in the truck's rear cargo bed with her head grouped closely beside Hyunwoo's and Raul's heads beneath the shade Charlie provides. The source does not specify her torso's orientation or the positions of her arms and legs.\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Raul is asleep in the truck's rear cargo bed with his head grouped closely beside Hyunwoo's and Amber's heads beneath the shade Charlie provides. The source does not specify his torso's orientation or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The roofless truck is parked beside rocks for shade, with a movable shade cloth over the cargo-bed resting area. Charlie adjusts the cloth while retaining his unrepaired battle damage and the recovered B-200 chest component in his possession. 현우: He is asleep in the rear cargo bed, still bearing his recent injuries. 앰버: She is asleep in the rear cargo bed; no treatment of her recent head injury has been established. 라울: He is asleep in the rear cargo bed.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S69sh3__bgfirst_bg.png",
     "asset_id": "ccf56b84-5714-4267-b1b6-cfb4a6ff27b2",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S69sh3.png",
     "asset_id": "acde8000-023e-40d0-b3ff-468a1735a66a",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163202>",
     "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_open_roof_truck_sel.png",
     "asset_id": "c94c53e3-76a2-4f78-b4ae-95204318d3c2",
     "role": "location_seed_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163202>",
     "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "찰리는 정면을 향해 시선을 두고 양손으로 천을 위로 뻗어 들고 있습니다.",
    "built_space": "트럭 짐칸이 전경에 있고, 배경에는 지시된 좁은 바위 틈 대신 일반적인 암벽과 산길이 보입니다.",
    "entities": "현우, 앰버(이마에 밴드 부착), 라울이 짐칸에 누워 있으며, 찰리의 얼굴 마스크는 레퍼런스와 다른 격자무늬 패턴을 가집니다.",
    "hard_violations": [],
    "physics": "찰리의 양손이 천을 쥐고 형태를 지탱하며, 잠든 아이들의 몸은 트럭 짐칸 바닥에 자연스럽게 닿아 있습니다."
   },
   {
    "label": "B",
    "direction": "찰리는 아래쪽 짐칸에 잠든 일행을 향해 시선을 내리깔고 양손으로 천을 넓게 쥐고 있습니다.",
    "built_space": "트럭 짐칸 안쪽에 아이들이 누워 있고, 배경은 로케이션 레퍼런스와 정확히 일치하는 형태의 좁은 바위 틈이 자리 잡고 있습니다.",
    "entities": "현우, 앰버(미치료된 핏자국 상처), 라울이 짐칸에 나란히 누워 있고, 찰리는 레퍼런스의 안광과 마스크 디자인을 잘 유지하고 있습니다.",
    "hard_violations": [],
    "physics": "찰리의 손가락이 천 가장자리를 단단히 쥐고 지탱하며, 누워있는 세 명은 트럭 바닥에 체중을 싣고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지정된 로케이션 레퍼런스의 바위 틈을 완벽하게 재현했으며, 앰버의 미치료 상처와 찰리의 마스크 등 디테일이 프롬프트와 잘 일치합니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "지정된 배경 레퍼런스와 전혀 다른 장소를 묘사했으며, 앰버의 이마에 밴드가 있어 '미치료 상태' 지시를 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 정면을 향해 시선을 두고 양손으로 천을 위로 뻗어 들고 있습니다.",
        "built_space": "트럭 짐칸이 전경에 있고, 배경에는 지시된 좁은 바위 틈 대신 일반적인 암벽과 산길이 보입니다.",
        "entities": "현우, 앰버(이마에 밴드 부착), 라울이 짐칸에 누워 있으며, 찰리의 얼굴 마스크는 레퍼런스와 다른 격자무늬 패턴을 가집니다.",
        "hard_violations": [],
        "physics": "찰리의 양손이 천을 쥐고 형태를 지탱하며, 잠든 아이들의 몸은 트럭 짐칸 바닥에 자연스럽게 닿아 있습니다."
       },
       {
        "label": "B",
        "direction": "찰리는 아래쪽 짐칸에 잠든 일행을 향해 시선을 내리깔고 양손으로 천을 넓게 쥐고 있습니다.",
        "built_space": "트럭 짐칸 안쪽에 아이들이 누워 있고, 배경은 로케이션 레퍼런스와 정확히 일치하는 형태의 좁은 바위 틈이 자리 잡고 있습니다.",
        "entities": "현우, 앰버(미치료된 핏자국 상처), 라울이 짐칸에 나란히 누워 있고, 찰리는 레퍼런스의 안광과 마스크 디자인을 잘 유지하고 있습니다.",
        "hard_violations": [],
        "physics": "찰리의 손가락이 천 가장자리를 단단히 쥐고 지탱하며, 누워있는 세 명은 트럭 바닥에 체중을 싣고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지정된 로케이션 레퍼런스의 바위 틈을 완벽하게 재현했으며, 앰버의 미치료 상처와 찰리의 마스크 등 디테일이 프롬프트와 잘 일치합니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "지정된 배경 레퍼런스와 전혀 다른 장소를 묘사했으며, 앰버의 이마에 밴드가 있어 '미치료 상태' 지시를 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 정면을 향해 시선을 두고 양손으로 천을 위로 뻗어 들고 있습니다.",
        "built_space": "트럭 짐칸이 전경에 있고, 배경에는 지시된 좁은 바위 틈 대신 일반적인 암벽과 산길이 보입니다.",
        "entities": "현우, 앰버(이마에 밴드 부착), 라울이 짐칸에 누워 있으며, 찰리의 얼굴 마스크는 레퍼런스와 다른 격자무늬 패턴을 가집니다.",
        "hard_violations": [],
        "physics": "찰리의 양손이 천을 쥐고 형태를 지탱하며, 잠든 아이들의 몸은 트럭 짐칸 바닥에 자연스럽게 닿아 있습니다."
       },
       {
        "label": "B",
        "direction": "찰리는 아래쪽 짐칸에 잠든 일행을 향해 시선을 내리깔고 양손으로 천을 넓게 쥐고 있습니다.",
        "built_space": "트럭 짐칸 안쪽에 아이들이 누워 있고, 배경은 로케이션 레퍼런스와 정확히 일치하는 형태의 좁은 바위 틈이 자리 잡고 있습니다.",
        "entities": "현우, 앰버(미치료된 핏자국 상처), 라울이 짐칸에 나란히 누워 있고, 찰리는 레퍼런스의 안광과 마스크 디자인을 잘 유지하고 있습니다.",
        "hard_violations": [],
        "physics": "찰리의 손가락이 천 가장자리를 단단히 쥐고 지탱하며, 누워있는 세 명은 트럭 바닥에 체중을 싣고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "찰리와 잠든 세 사람의 외형은 잘 맞지만, 천이 머리 위 상단이 아니라 상체 앞에 낮게 처져 핵심인 차양 펼치기 구도를 놓쳤다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "양손을 높이 벌린 찰리의 상체와 세 사람 위를 덮는 천의 아랫면을 정확히 배치했으나, 앰버의 반창고와 찰리 얼굴 세부는 설정·참조와 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 얼굴은 정면에서 약간 아래의 적재함 쪽을 향한다. 양팔은 좌우로 벌어져 각각 천의 모서리를 잡지만 손은 머리보다 낮다. 천은 잠든 사람들의 머리 위로 펼쳐지기보다 찰리의 몸 앞에서 깊게 처져 있다. 세 사람은 눈을 감고 있다.",
        "built_space": "열린 적재함 하나와 앞쪽 운전실 일부, 운전실 뒤 보호 프레임 하나가 보인다. 세 사람은 적재함 바닥에 누워 머리를 가운데로 모았고, 찰리는 반대편 측벽 너머에 있다. 뒤편의 큰 회색 바위와 좁은 틈은 장소 참조에 가깝다. 카메라는 적재함 밖에서 내부를 바라보지만 천이 화면 상단 3분의 1이 아닌 중앙 상당 부분을 차지한다.",
        "entities": "찰리는 긴 기계 팔, 모래색 장갑판, 흰 마스크와 주황색 눈을 지녀 참조와 가깝다. 검은 헝클어진 머리의 청년 현우, 금발 여자아이 앰버, 피부가 짙고 머리를 뒤로 넘긴 남자아이 라울로 구별되는 세 사람만 보이며 남색 상의도 맞는다. 현우와 앰버에게 상처가 보이고 앰버의 이마에는 처치되지 않은 상처가 있다. 찰리 장갑의 마모는 보이지만 회수한 가슴 부품의 별도 소지 여부는 확인되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "천 양끝은 찰리의 두 손에 잡혀 있고 중앙이 중력에 따라 깊게 늘어져 있어 지지는 명확하다. 다만 이 형태는 넓은 수평 차양보다 앞에 늘어뜨린 천에 가깝다. 세 사람의 머리는 겹친 침구나 옆 사람 가까이에 놓이고 몸과 팔도 적재함 침구에 받쳐져 있다. 찰리의 하체는 가려져 있으므로 발 접지는 확인할 수 없지만 공중에 떠 있다고 볼 근거는 없다."
       },
       {
        "label": "B",
        "direction": "찰리의 얼굴은 화면 왼쪽 아래, 잠든 일행 방향으로 돌아가 있다. 두 팔은 위쪽 양옆으로 뻗고 양손이 천의 서로 떨어진 부분을 잡는다. 천의 아랫면이 세 사람의 머리 위를 가로질러 실제로 그들을 덮는 방향이며 찰리의 얼굴을 가리지 않는다. 세 사람은 눈을 감고 서로 가까이 머리를 모았다.",
        "built_space": "적재함 하나, 그 앞의 운전실 하나, 바깥 거울 하나가 보인다. 카메라는 적재함 바깥의 뒤쪽 사선에서 보고 찰리는 반대편 측벽 너머에 서 있다. 천은 화면 상단을 비스듬히 덮고 찰리의 상체가 중심을 이룬다. 회색 바위의 세로 틈과 끼인 큰 돌들이 참조 장소의 특징을 유지하지만 오른쪽에는 더 열린 산 배경이 보인다. 운전실에는 지붕이 있어 지붕 없는 트럭이라는 문구에는 완전히 맞지 않는다.",
        "entities": "모래색 장갑과 육중한 긴 팔은 찰리의 정체성에 맞으나, 흰 얼굴의 눈 부분은 참조의 주황색 광학 눈 대신 점무늬처럼 표현된다. 현우는 검은 머리의 앳된 청년으로, 앰버는 금발 여자아이로, 라울은 뒤로 묶은 머리의 피부가 짙은 남자아이로 표현되어 지정된 인물 구성과 대체로 맞는다. 세 사람의 남색 옷도 참조에 부합한다. 현우의 얼굴과 팔에 상처가 있으나 앰버의 이마에는 치료가 설정되지 않았음에도 작은 반창고가 붙어 있다. 찰리에게 마모 흔적은 있지만 별도 가슴 부품의 소지는 확인되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "천은 양손의 실제 접촉으로 지지되고, 팽팽한 가장자리와 완만한 중앙 처짐이 넓게 펼치는 동작에 맞는다. 찰리의 팔꿈치와 어깨는 천을 들어 올리는 자세로 연결된다. 하체는 적재함 벽에 가려져 있지만 부유를 나타내지는 않는다. 잠든 세 사람의 머리와 몸은 침구에 놓여 있고 손과 팔도 침구나 자신의 몸 위에 내려놓여 있어 무의식적인 수면 자세의 중력 조건을 충족한다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "찰리와 잠든 세 사람의 외형은 잘 맞지만, 천이 머리 위 상단이 아니라 상체 앞에 낮게 처져 핵심인 차양 펼치기 구도를 놓쳤다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "양손을 높이 벌린 찰리의 상체와 세 사람 위를 덮는 천의 아랫면을 정확히 배치했으나, 앰버의 반창고와 찰리 얼굴 세부는 설정·참조와 다르다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 얼굴은 정면에서 약간 아래의 적재함 쪽을 향한다. 양팔은 좌우로 벌어져 각각 천의 모서리를 잡지만 손은 머리보다 낮다. 천은 잠든 사람들의 머리 위로 펼쳐지기보다 찰리의 몸 앞에서 깊게 처져 있다. 세 사람은 눈을 감고 있다.",
        "built_space": "열린 적재함 하나와 앞쪽 운전실 일부, 운전실 뒤 보호 프레임 하나가 보인다. 세 사람은 적재함 바닥에 누워 머리를 가운데로 모았고, 찰리는 반대편 측벽 너머에 있다. 뒤편의 큰 회색 바위와 좁은 틈은 장소 참조에 가깝다. 카메라는 적재함 밖에서 내부를 바라보지만 천이 화면 상단 3분의 1이 아닌 중앙 상당 부분을 차지한다.",
        "entities": "찰리는 긴 기계 팔, 모래색 장갑판, 흰 마스크와 주황색 눈을 지녀 참조와 가깝다. 검은 헝클어진 머리의 청년 현우, 금발 여자아이 앰버, 피부가 짙고 머리를 뒤로 넘긴 남자아이 라울로 구별되는 세 사람만 보이며 남색 상의도 맞는다. 현우와 앰버에게 상처가 보이고 앰버의 이마에는 처치되지 않은 상처가 있다. 찰리 장갑의 마모는 보이지만 회수한 가슴 부품의 별도 소지 여부는 확인되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "천 양끝은 찰리의 두 손에 잡혀 있고 중앙이 중력에 따라 깊게 늘어져 있어 지지는 명확하다. 다만 이 형태는 넓은 수평 차양보다 앞에 늘어뜨린 천에 가깝다. 세 사람의 머리는 겹친 침구나 옆 사람 가까이에 놓이고 몸과 팔도 적재함 침구에 받쳐져 있다. 찰리의 하체는 가려져 있으므로 발 접지는 확인할 수 없지만 공중에 떠 있다고 볼 근거는 없다."
       },
       {
        "label": "A",
        "direction": "찰리의 얼굴은 화면 왼쪽 아래, 잠든 일행 방향으로 돌아가 있다. 두 팔은 위쪽 양옆으로 뻗고 양손이 천의 서로 떨어진 부분을 잡는다. 천의 아랫면이 세 사람의 머리 위를 가로질러 실제로 그들을 덮는 방향이며 찰리의 얼굴을 가리지 않는다. 세 사람은 눈을 감고 서로 가까이 머리를 모았다.",
        "built_space": "적재함 하나, 그 앞의 운전실 하나, 바깥 거울 하나가 보인다. 카메라는 적재함 바깥의 뒤쪽 사선에서 보고 찰리는 반대편 측벽 너머에 서 있다. 천은 화면 상단을 비스듬히 덮고 찰리의 상체가 중심을 이룬다. 회색 바위의 세로 틈과 끼인 큰 돌들이 참조 장소의 특징을 유지하지만 오른쪽에는 더 열린 산 배경이 보인다. 운전실에는 지붕이 있어 지붕 없는 트럭이라는 문구에는 완전히 맞지 않는다.",
        "entities": "모래색 장갑과 육중한 긴 팔은 찰리의 정체성에 맞으나, 흰 얼굴의 눈 부분은 참조의 주황색 광학 눈 대신 점무늬처럼 표현된다. 현우는 검은 머리의 앳된 청년으로, 앰버는 금발 여자아이로, 라울은 뒤로 묶은 머리의 피부가 짙은 남자아이로 표현되어 지정된 인물 구성과 대체로 맞는다. 세 사람의 남색 옷도 참조에 부합한다. 현우의 얼굴과 팔에 상처가 있으나 앰버의 이마에는 치료가 설정되지 않았음에도 작은 반창고가 붙어 있다. 찰리에게 마모 흔적은 있지만 별도 가슴 부품의 소지는 확인되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "천은 양손의 실제 접촉으로 지지되고, 팽팽한 가장자리와 완만한 중앙 처짐이 넓게 펼치는 동작에 맞는다. 찰리의 팔꿈치와 어깨는 천을 들어 올리는 자세로 연결된다. 하체는 적재함 벽에 가려져 있지만 부유를 나타내지는 않는다. 잠든 세 사람의 머리와 몸은 침구에 놓여 있고 손과 팔도 침구나 자신의 몸 위에 내려놓여 있어 무의식적인 수면 자세의 중력 조건을 충족한다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.571,
    "B": 1.75
   },
   "adjusted": {
    "A": 1.571,
    "B": 1.75
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1750,
   "A": 1571
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "지정된 로케이션 레퍼런스의 바위 틈을 완벽하게 재현했으며, 앰버의 미치료 상처와 찰리의 마스크 등 디테일이 프롬프트와 잘 일치합니다."
   },
   {
    "label": "A",
    "score": 1571,
    "verdict_ko": "지정된 배경 레퍼런스와 전혀 다른 장소를 묘사했으며, 앰버의 이마에 밴드가 있어 '미치료 상태' 지시를 위반했습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_open_roof_truck_sel.png",
    "asset_id": "c94c53e3-76a2-4f78-b4ae-95204318d3c2",
    "role": "location_seed_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163202>",
    "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-bfc2-7d49-8edf-553dd6c93f32",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S69sh3__bgfirst_bg.png",
   "bg_asset_id": "ccf56b84-5714-4267-b1b6-cfb4a6ff27b2",
   "bg_record_key": "S69sh3::bgfirst_bg",
   "chain_winner": false,
   "authority": "seed_bg"
  },
  "ref_mode": "seed-bg+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  },
  "lane_policy": "ab_select_ready"
 },
 "S69sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:15:55.249415+00:00",
  "fingerprint": "61809060e189c61861b826f04d8b8367a7939b2714c57839bd63b772c5ee3c7f",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S69sh3_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S69sh3_sel.png",
  "source_sha256": "e2b647e7fc37f5759244763761c49e6461d35686871885e65da51c8598a06fb8",
  "file": "S69sh3_cine.png",
  "staged_sha256": "bb3e6d8f9b6269a41dcf4455cfc7e9e668a4139cc0be9a50d3354044af04672a",
  "latency_ms": 11193
 },
 "S70sh5::signage": {
  "fp": "6185ff30a3ef5644",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::5210e323d6019fba": {
  "subjects": [],
  "subject_text": "군 병원 복도\n여러 병실의 출입문이 이어지는 실내 복도. 천장 조명 아래로 긴 통로가 뻗어 있고 병실 안쪽을 볼 수 있는 출입부가 있다.",
  "identity": "canonical",
  "scope_id": "L244",
  "scope_role": "location_interior",
  "scope_sha": "d5e2db056cd48c88"
 },
 "groupbg::military_hospital_room": {
  "input_fingerprint": "746b93870f51bf8b",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "military_hospital_room",
    "tags": [
     "S70sh5",
     "S70sh9"
    ]
   },
   "context_sig": "3be9ca213d18297f"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: Inside a military-hospital patient room, beside the injured commander's bed under nighttime ward lighting.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n군 병원 복도: 차갑고 통제된 군 의료 시설의 하얀 통로. (특징: 깨끗하지만 차가운 색감의 병원 복도; 제복을 입은 장교들의 걸음; 환자들이 있는 병실 문들)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 병실에 누워있는 사람은... 한 쪽 눈, 팔과 다리 깁스를 하고 치료 중인 박철진이다.\n\nTIME OF DAY (lock): night.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: Inside a military-hospital patient room, beside the injured commander's bed under nighttime ward lighting.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n군 병원 복도: 차갑고 통제된 군 의료 시설의 하얀 통로. (특징: 깨끗하지만 차가운 색감의 병원 복도; 제복을 입은 장교들의 걸음; 환자들이 있는 병실 문들)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 병실에 누워있는 사람은... 한 쪽 눈, 팔과 다리 깁스를 하고 치료 중인 박철진이다.\n\nTIME OF DAY (lock): night.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_military_hospital_room_03987e.png",
  "asset_id": "83fc438f-e6e6-45b5-9493-51fe5f5a3faa",
  "input_asset_ids": [
   "e18af667-e876-4f68-a303-2e9caf00a411"
  ],
  "origin_tag": "S70sh5",
  "place_text": "Inside a military-hospital patient room, beside the injured commander's bed under nighttime ward lighting.",
  "origin_inputs": {
   "place_text": "Inside a military-hospital patient room, beside the injured commander's bed under nighttime ward lighting.",
   "time_of_day_en": "night",
   "conti_asset_id": "e18af667-e876-4f68-a303-2e9caf00a411"
  }
 },
 "S70sh5::bgfirst_bg": {
  "input_fingerprint": "a2a594571d2e0ef7",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 병상에 누운 박철진을 차가운 눈빛으로 멸시하듯 내려다보는 국방장관의 상체.\n\nLOCATION (lock): Inside a military-hospital patient room, beside the injured commander's bed under nighttime ward lighting.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Hospital bed (Occupied by 박철진 during treatment for his injuries) — The shoulder end and near side are seen from the bedside camera position; used as Anchor the patient's low foreground position beneath the minister's upper body.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral hospital interior illumination preserves the minister's downward expression and the patient's injured face with restrained tonal contrast and no invented fixture or color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 병상에 누운 박철진을 차가운 눈빛으로 멸시하듯 내려다보는 국방장관의 상체.\n\nLOCATION (lock): Inside a military-hospital patient room, beside the injured commander's bed under nighttime ward lighting.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Hospital bed (Occupied by 박철진 during treatment for his injuries) — The shoulder end and near side are seen from the bedside camera position; used as Anchor the patient's low foreground position beneath the minister's upper body.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral hospital interior illumination preserves the minister's downward expression and the patient's injured face with restrained tonal contrast and no invented fixture or color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S70sh5__bgfirst_bg.png",
  "asset_id": "a0d10e4e-ddb6-42d2-9c3e-f1455cd58d3a",
  "input_asset_ids": [
   "e18af667-e876-4f68-a303-2e9caf00a411",
   "83fc438f-e6e6-45b5-9493-51fe5f5a3faa"
  ]
 },
 "S70sh5": {
  "input_fingerprint": "23a5d68fc9466909",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 병상에 누운 박철진을 차가운 눈빛으로 멸시하듯 내려다보는 국방장관의 상체.\n\nLOCATION (lock): Inside a military-hospital patient room, beside the injured commander's bed under nighttime ward lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Hospital bed (Occupied by 박철진 during treatment for his injuries) — The shoulder end and near side are seen from the bedside camera position; used as Anchor the patient's low foreground position beneath the minister's upper body.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral hospital interior illumination preserves the minister's downward expression and the patient's injured face with restrained tonal contrast and no invented fixture or color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Park Cheoljin is reclined on a hospital bed, with one eye covered and an arm and a leg immobilized in casts, looking upward with his uncovered eye. The source does not identify the affected sides or specify his limbs' exact placement and his head's angle.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): It is night in a military hospital ward containing beds for wounded patients. 박철진: He is reclined on a hospital bed under treatment for severe injuries, with one eye covered and an arm and a leg immobilized in casts. He is only faintly conscious. 국방장관: He has reached the ward and stands looking into it.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리); 국방장관 (한국인, 성인 남성, 단정한 짧은 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 병상에 누운 박철진을 차가운 눈빛으로 멸시하듯 내려다보는 국방장관의 상체.\n\nLOCATION (lock): Inside a military-hospital patient room, beside the injured commander's bed under nighttime ward lighting. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Hospital bed (Occupied by 박철진 during treatment for his injuries) — The shoulder end and near side are seen from the bedside camera position; used as Anchor the patient's low foreground position beneath the minister's upper body.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral hospital interior illumination preserves the minister's downward expression and the patient's injured face with restrained tonal contrast and no invented fixture or color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Park Cheoljin is reclined on a hospital bed, with one eye covered and an arm and a leg immobilized in casts, looking upward with his uncovered eye. The source does not identify the affected sides or specify his limbs' exact placement and his head's angle.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): It is night in a military hospital ward containing beds for wounded patients. 박철진: He is reclined on a hospital bed under treatment for severe injuries, with one eye covered and an arm and a leg immobilized in casts. He is only faintly conscious. 국방장관: He has reached the ward and stands looking into it.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리); 국방장관 (한국인, 성인 남성, 단정한 짧은 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 병상에 누운 박철진을 차가운 눈빛으로 멸시하듯 내려다보는 국방장관의 상체.\n\nLOCATION (lock): Inside a military-hospital patient room, beside the injured commander's bed under nighttime ward lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Hospital bed (Occupied by 박철진 during treatment for his injuries) — The shoulder end and near side are seen from the bedside camera position; used as Anchor the patient's low foreground position beneath the minister's upper body.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral hospital interior illumination preserves the minister's downward expression and the patient's injured face with restrained tonal contrast and no invented fixture or color cast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Park Cheoljin is reclined on a hospital bed, with one eye covered and an arm and a leg immobilized in casts, looking upward with his uncovered eye. The source does not identify the affected sides or specify his limbs' exact placement and his head's angle.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): It is night in a military hospital ward containing beds for wounded patients. 박철진: He is reclined on a hospital bed under treatment for severe injuries, with one eye covered and an arm and a leg immobilized in casts. He is only faintly conscious. 국방장관: He has reached the ward and stands looking into it.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리); 국방장관 (한국인, 성인 남성, 단정한 짧은 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S70sh5__bgfirst_bg.png",
     "asset_id": "a0d10e4e-ddb6-42d2-9c3e-f1455cd58d3a",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S70sh5.png",
     "asset_id": "e18af667-e876-4f68-a303-2e9caf00a411",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1401722>",
     "asset_id": "fee7383c-fb61-4b3a-ba7c-79f2555de00b",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 국방장관: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1183281>",
     "asset_id": "65f118fd-34f5-434a-8c3f-1378f0e8a758",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_military_hospital_room_03987e.png",
     "asset_id": "83fc438f-e6e6-45b5-9493-51fe5f5a3faa",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1401722>",
     "asset_id": "fee7383c-fb61-4b3a-ba7c-79f2555de00b",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 국방장관: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1183281>",
     "asset_id": "65f118fd-34f5-434a-8c3f-1378f0e8a758",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "서 있는 인물이 누운 인물을 내려다보고, 누운 인물은 위를 올려다보며 시선이 교차함.",
    "built_space": "야간 병실 분위기와 침대, 커튼, 문 등 주요 구조물이 레퍼런스와 일치함.",
    "entities": "역할 오류 발생. 서 있는 인물에 환자인 박철진의 얼굴(귀걸이 포함)이 적용되었고, 누운 인물은 식별이 모호함.",
    "hard_violations": [
     "[gemini-pro] 잘못된 캐릭터 배치 (환자인 박철진이 서 있음)"
    ],
    "physics": "누운 인물은 침대에, 서 있는 인물은 바닥에 자연스럽게 지지됨."
   },
   {
    "label": "B",
    "direction": "국방장관이 침대에 누운 박철진을 내려다보고, 박철진은 장관을 향해 시선을 둠.",
    "built_space": "지정된 병실 레퍼런스의 구조와 야간 조명을 정확히 반영하여 배치됨.",
    "entities": "국방장관과 박철진의 신원이 올바르게 배역에 맞게 적용됨. 박철진은 안대와 팔다리 깁스를 착용한 상태임.",
    "hard_violations": [],
    "physics": "박철진의 몸은 침대에 완전히 밀착되어 있고, 장관은 안정적인 자세로 서 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "캐릭터의 신원과 요구된 부상 상태(안대, 깁스)를 정확히 구현했으나, 지시된 상체 중심의 미디엄 샷보다 프레임이 다소 넓게 잡혔습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "서 있는 장관과 누워 있는 환자의 캐릭터 신원이 서로 뒤바뀌어 샷의 기본 설정을 크게 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "서 있는 인물이 누운 인물을 내려다보고, 누운 인물은 위를 올려다보며 시선이 교차함.",
        "built_space": "야간 병실 분위기와 침대, 커튼, 문 등 주요 구조물이 레퍼런스와 일치함.",
        "entities": "역할 오류 발생. 서 있는 인물에 환자인 박철진의 얼굴(귀걸이 포함)이 적용되었고, 누운 인물은 식별이 모호함.",
        "hard_violations": [
         "잘못된 캐릭터 배치 (환자인 박철진이 서 있음)"
        ],
        "physics": "누운 인물은 침대에, 서 있는 인물은 바닥에 자연스럽게 지지됨."
       },
       {
        "label": "B",
        "direction": "국방장관이 침대에 누운 박철진을 내려다보고, 박철진은 장관을 향해 시선을 둠.",
        "built_space": "지정된 병실 레퍼런스의 구조와 야간 조명을 정확히 반영하여 배치됨.",
        "entities": "국방장관과 박철진의 신원이 올바르게 배역에 맞게 적용됨. 박철진은 안대와 팔다리 깁스를 착용한 상태임.",
        "hard_violations": [],
        "physics": "박철진의 몸은 침대에 완전히 밀착되어 있고, 장관은 안정적인 자세로 서 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "캐릭터의 신원과 요구된 부상 상태(안대, 깁스)를 정확히 구현했으나, 지시된 상체 중심의 미디엄 샷보다 프레임이 다소 넓게 잡혔습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "서 있는 장관과 누워 있는 환자의 캐릭터 신원이 서로 뒤바뀌어 샷의 기본 설정을 크게 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "서 있는 인물이 누운 인물을 내려다보고, 누운 인물은 위를 올려다보며 시선이 교차함.",
        "built_space": "야간 병실 분위기와 침대, 커튼, 문 등 주요 구조물이 레퍼런스와 일치함.",
        "entities": "역할 오류 발생. 서 있는 인물에 환자인 박철진의 얼굴(귀걸이 포함)이 적용되었고, 누운 인물은 식별이 모호함.",
        "hard_violations": [
         "잘못된 캐릭터 배치 (환자인 박철진이 서 있음)"
        ],
        "physics": "누운 인물은 침대에, 서 있는 인물은 바닥에 자연스럽게 지지됨."
       },
       {
        "label": "B",
        "direction": "국방장관이 침대에 누운 박철진을 내려다보고, 박철진은 장관을 향해 시선을 둠.",
        "built_space": "지정된 병실 레퍼런스의 구조와 야간 조명을 정확히 반영하여 배치됨.",
        "entities": "국방장관과 박철진의 신원이 올바르게 배역에 맞게 적용됨. 박철진은 안대와 팔다리 깁스를 착용한 상태임.",
        "hard_violations": [],
        "physics": "박철진의 몸은 침대에 완전히 밀착되어 있고, 장관은 안정적인 자세로 서 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "내려다보는 시선과 병실은 맞지만, 환자의 발끝까지 보여주는 넓은 구도가 장관 상체 중심의 미디엄 숏을 약화하고 장관 의상도 참조와 다르다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "장관 상체 아래 낮은 전경에 환자를 두는 미디엄 구도가 더 충실하지만, 장관의 얼굴과 정장이 박철진 참조를 닮아 인물 정체성에 뚜렷한 혼선이 있다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "장관은 고개와 눈을 화면 왼쪽 아래로 내려 병상의 박철진 얼굴을 바라본다. 환자의 드러난 눈은 오른쪽 위 장관 쪽을 향한다. 상호 시선의 방향은 맞지만 장관의 표정은 강한 멸시보다는 절제된 냉담함에 가깝다. 무기나 방향을 확인할 휴대 물체는 없다.",
        "built_space": "전경 병상 한 개와 뒤쪽 빈 병상 한 개가 보인다. 전경 침대의 머리판, 가까운 측면 난간, 발치 난간까지 포함된다. 왼쪽 수액대·수액 용기·펌프가 각각 하나, 벽 의료 설비판과 그 아래 조명이 하나씩 있으며, 회녹색 커튼, 미닫이문 하나, 뒤쪽 두 칸 창 하나, 천장등 두 개, 환기구 하나, 오른쪽 서랍 카트 하나가 참조 병실과 대응한다. 장관은 침대의 반대쪽 측면에 서 있어 침대와 충돌하지 않는다. 다만 환자의 거의 전신과 병실 깊이를 보여주어 요구한 상체 중심 구도보다 넓다.",
        "entities": "인물은 두 명뿐이며 모두 짧은 검은 머리의 한국인 성인 남성이라는 설정에 부합하는 외관이다. 누운 환자는 중년 얼굴로 박철진 참조와 대체로 연결되고, 한쪽 눈 가림과 한 팔 및 한 다리의 고정이 보인다. 장관 얼굴은 장관 참조에 비교적 가깝지만 남색 셔츠 대신 정장과 넥타이를 입었다. 창밖은 밤이며 병실의 중성 조명도 유지된다. 뚜렷하게 읽히는 글이나 덧씌운 표시는 없다.",
        "hard_violations": [],
        "physics": "환자의 머리는 베개에, 등과 다리는 침대에 받쳐져 있다. 고정된 팔은 매트리스 위에 놓이고 반대 손은 복부의 이불에 얹혀 있다. 발끝이 위를 향하지만 다리와 발뒤꿈치는 침대에 지지되어 공중에 떠 있는 자세가 아니다. 장관의 발은 화면 밖이지만 몸통은 침대 옆에서 자연스럽게 선 자세로 이어진다. 수액 용기는 수액대 고리에 걸리고 펌프도 기둥에 고정되어 있다."
       },
       {
        "label": "B",
        "direction": "장관의 눈과 숙인 고개가 왼쪽 아래 환자의 얼굴을 향하고, 좁힌 눈과 굳은 입매가 냉정하게 내려다보는 연기를 만든다. 환자의 얼굴은 장관 쪽 위로 향하지만 가리지 않은 눈은 측면 각도 때문에 시선까지 확실히 판별하기 어렵다. 카메라를 응시하는 인물은 없다.",
        "built_space": "전경 병상 한 개의 어깨 쪽과 가까운 측면 난간이 보이고, 뒤쪽 병상 일부는 장관에게 가려져 있다. 왼쪽 수액대·수액 용기·펌프 각 하나, 의료 설비판과 하부 조명 각 하나, 회녹색 커튼, 미닫이문 하나, 뒤쪽 두 칸 창 하나, 천장등 두 개, 환기구 하나, 오른쪽 카트 하나가 참조 공간과 대응한다. 장관은 침대 옆 통로에 서 있고 환자는 화면 아래에 놓인다. 장관을 머리부터 허리 아래 정도까지 크게 담아 A보다 요구된 미디엄 숏과 상하 배치에 가깝다.",
        "entities": "두 사람만 등장하며 한국인 성인 남성 설정에 맞는 외관이다. 그러나 서 있는 장관은 중년의 얼굴 윤곽, 귀 장식, 정장과 사선 넥타이까지 박철진 참조를 강하게 닮아 장관 참조의 얼굴 및 남색 셔츠와 어긋난다. 누운 환자는 측면과 붕대 때문에 얼굴 일치 여부가 제한적으로만 확인된다. 환자의 한쪽 눈 가림과 팔 고정은 분명하다. 다리는 이불과 화면 밖에 가려져 고정 여부를 판단할 수 없으며, 이를 누락으로 보지는 않는다. 야간 창밖과 병원 조명은 맞고 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "환자의 뒤통수와 등은 베개 및 올려진 침대 등받이에 지지되고, 고정된 팔과 손은 침구 위에 놓여 있다. 머리나 팔이 지지 없이 들려 있지 않다. 장관은 침대 옆에서 곧게 서 있으며 보이는 손도 난간 부근에 자연스럽게 내려와 있다. 하체는 프레임 밖이지만 공중 부양을 암시하는 자세는 없다. 수액 용기와 펌프에는 각각 고리와 기둥 고정 장치가 보인다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "내려다보는 시선과 병실은 맞지만, 환자의 발끝까지 보여주는 넓은 구도가 장관 상체 중심의 미디엄 숏을 약화하고 장관 의상도 참조와 다르다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "장관 상체 아래 낮은 전경에 환자를 두는 미디엄 구도가 더 충실하지만, 장관의 얼굴과 정장이 박철진 참조를 닮아 인물 정체성에 뚜렷한 혼선이 있다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "장관은 고개와 눈을 화면 왼쪽 아래로 내려 병상의 박철진 얼굴을 바라본다. 환자의 드러난 눈은 오른쪽 위 장관 쪽을 향한다. 상호 시선의 방향은 맞지만 장관의 표정은 강한 멸시보다는 절제된 냉담함에 가깝다. 무기나 방향을 확인할 휴대 물체는 없다.",
        "built_space": "전경 병상 한 개와 뒤쪽 빈 병상 한 개가 보인다. 전경 침대의 머리판, 가까운 측면 난간, 발치 난간까지 포함된다. 왼쪽 수액대·수액 용기·펌프가 각각 하나, 벽 의료 설비판과 그 아래 조명이 하나씩 있으며, 회녹색 커튼, 미닫이문 하나, 뒤쪽 두 칸 창 하나, 천장등 두 개, 환기구 하나, 오른쪽 서랍 카트 하나가 참조 병실과 대응한다. 장관은 침대의 반대쪽 측면에 서 있어 침대와 충돌하지 않는다. 다만 환자의 거의 전신과 병실 깊이를 보여주어 요구한 상체 중심 구도보다 넓다.",
        "entities": "인물은 두 명뿐이며 모두 짧은 검은 머리의 한국인 성인 남성이라는 설정에 부합하는 외관이다. 누운 환자는 중년 얼굴로 박철진 참조와 대체로 연결되고, 한쪽 눈 가림과 한 팔 및 한 다리의 고정이 보인다. 장관 얼굴은 장관 참조에 비교적 가깝지만 남색 셔츠 대신 정장과 넥타이를 입었다. 창밖은 밤이며 병실의 중성 조명도 유지된다. 뚜렷하게 읽히는 글이나 덧씌운 표시는 없다.",
        "hard_violations": [],
        "physics": "환자의 머리는 베개에, 등과 다리는 침대에 받쳐져 있다. 고정된 팔은 매트리스 위에 놓이고 반대 손은 복부의 이불에 얹혀 있다. 발끝이 위를 향하지만 다리와 발뒤꿈치는 침대에 지지되어 공중에 떠 있는 자세가 아니다. 장관의 발은 화면 밖이지만 몸통은 침대 옆에서 자연스럽게 선 자세로 이어진다. 수액 용기는 수액대 고리에 걸리고 펌프도 기둥에 고정되어 있다."
       },
       {
        "label": "A",
        "direction": "장관의 눈과 숙인 고개가 왼쪽 아래 환자의 얼굴을 향하고, 좁힌 눈과 굳은 입매가 냉정하게 내려다보는 연기를 만든다. 환자의 얼굴은 장관 쪽 위로 향하지만 가리지 않은 눈은 측면 각도 때문에 시선까지 확실히 판별하기 어렵다. 카메라를 응시하는 인물은 없다.",
        "built_space": "전경 병상 한 개의 어깨 쪽과 가까운 측면 난간이 보이고, 뒤쪽 병상 일부는 장관에게 가려져 있다. 왼쪽 수액대·수액 용기·펌프 각 하나, 의료 설비판과 하부 조명 각 하나, 회녹색 커튼, 미닫이문 하나, 뒤쪽 두 칸 창 하나, 천장등 두 개, 환기구 하나, 오른쪽 카트 하나가 참조 공간과 대응한다. 장관은 침대 옆 통로에 서 있고 환자는 화면 아래에 놓인다. 장관을 머리부터 허리 아래 정도까지 크게 담아 A보다 요구된 미디엄 숏과 상하 배치에 가깝다.",
        "entities": "두 사람만 등장하며 한국인 성인 남성 설정에 맞는 외관이다. 그러나 서 있는 장관은 중년의 얼굴 윤곽, 귀 장식, 정장과 사선 넥타이까지 박철진 참조를 강하게 닮아 장관 참조의 얼굴 및 남색 셔츠와 어긋난다. 누운 환자는 측면과 붕대 때문에 얼굴 일치 여부가 제한적으로만 확인된다. 환자의 한쪽 눈 가림과 팔 고정은 분명하다. 다리는 이불과 화면 밖에 가려져 고정 여부를 판단할 수 없으며, 이를 누락으로 보지는 않는다. 야간 창밖과 병원 조명은 맞고 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "환자의 뒤통수와 등은 베개 및 올려진 침대 등받이에 지지되고, 고정된 팔과 손은 침구 위에 놓여 있다. 머리나 팔이 지지 없이 들려 있지 않다. 장관은 침대 옆에서 곧게 서 있으며 보이는 손도 난간 부근에 자연스럽게 내려와 있다. 하체는 프레임 밖이지만 공중 부양을 암시하는 자세는 없다. 수액 용기와 펌프에는 각각 고리와 기둥 고정 장치가 보인다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.429,
    "B": 1.857
   },
   "adjusted": {
    "A": 1.179,
    "B": 1.857
   },
   "violations": {
    "A": [
     "[gemini-pro] 잘못된 캐릭터 배치 (환자인 박철진이 서 있음)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1857,
   "A": 1179
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1857,
    "verdict_ko": "캐릭터의 신원과 요구된 부상 상태(안대, 깁스)를 정확히 구현했으나, 지시된 상체 중심의 미디엄 샷보다 프레임이 다소 넓게 잡혔습니다."
   },
   {
    "label": "A",
    "score": 1179,
    "verdict_ko": "서 있는 장관과 누워 있는 환자의 캐릭터 신원이 서로 뒤바뀌어 샷의 기본 설정을 크게 위반했습니다.  ★위반: [gemini-pro] 잘못된 캐릭터 배치 (환자인 박철진이 서 있음)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_military_hospital_room_03987e.png",
    "asset_id": "83fc438f-e6e6-45b5-9493-51fe5f5a3faa",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 박철진: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1401722>",
    "asset_id": "fee7383c-fb61-4b3a-ba7c-79f2555de00b",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 국방장관: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1183281>",
    "asset_id": "65f118fd-34f5-434a-8c3f-1378f0e8a758",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-c34c-760c-ac14-f6f41d63ac68",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S70sh5__bgfirst_bg.png",
   "bg_asset_id": "a0d10e4e-ddb6-42d2-9c3e-f1455cd58d3a",
   "bg_record_key": "S70sh5::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "military_hospital_room",
   "groupbg_asset_id": "83fc438f-e6e6-45b5-9493-51fe5f5a3faa"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S70sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:18:03.445799+00:00",
  "fingerprint": "8321c2e0bd6df40bab5bf3cbcaab2d3f32428d8b595ea641fd5b6720d339b91e",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S70sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S70sh5_sel.png",
  "source_sha256": "ccc56c398b57b2acff7cdd6b9bed7d078d46403d843661ef870ec49b533966c4",
  "file": "S70sh5_cine.png",
  "staged_sha256": "50fa5aca74b38f963a1bd0753f58e7109c1f385a01b05b624ee97c58c02c64c9",
  "latency_ms": 9424
 },
 "S70sh9::signage": {
  "fp": "ba8a0135d3930382",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S70sh9": {
  "input_fingerprint": "462aabb63dc73dbe",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 붉게 충혈된 눈에서 뺨을 타고 눈물 한 방울이 흘러내리고 있는 찰나의 박철진의 억울하고 처절한 얼굴.\n\nLOCATION (lock): On the patient's bed inside the military-hospital room, under the ward's nighttime lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Hospital bed (Occupied by 박철진 during treatment) — Only a narrow portion beside his head remains visible; used as A subdued edge establishes his recumbent position without distracting from the tear.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient hospital illumination preserves the reddened eye and the small highlight on the tear without specifying an unsupported fixture or light color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Park Cheoljin is reclined on a hospital bed, with one eye covered and an arm and a leg immobilized in casts, looking upward with his uncovered eye. The source does not identify the affected sides or specify his limbs' exact placement and his head's angle.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The setting remains a military hospital ward at night. 박철진: He remains bedridden with severe injuries to his face, arm and leg, with one eye covered and his arm and leg immobilized in casts. His remaining eye is bloodshot and filling with tears.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 붉게 충혈된 눈에서 뺨을 타고 눈물 한 방울이 흘러내리고 있는 찰나의 박철진의 억울하고 처절한 얼굴.\n\nLOCATION (lock): On the patient's bed inside the military-hospital room, under the ward's nighttime lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Hospital bed (Occupied by 박철진 during treatment) — Only a narrow portion beside his head remains visible; used as A subdued edge establishes his recumbent position without distracting from the tear.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient hospital illumination preserves the reddened eye and the small highlight on the tear without specifying an unsupported fixture or light color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Park Cheoljin is reclined on a hospital bed, with one eye covered and an arm and a leg immobilized in casts, looking upward with his uncovered eye. The source does not identify the affected sides or specify his limbs' exact placement and his head's angle.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The setting remains a military hospital ward at night. 박철진: He remains bedridden with severe injuries to his face, arm and leg, with one eye covered and his arm and leg immobilized in casts. His remaining eye is bloodshot and filling with tears.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 붉게 충혈된 눈에서 뺨을 타고 눈물 한 방울이 흘러내리고 있는 찰나의 박철진의 억울하고 처절한 얼굴.\n\nLOCATION (lock): On the patient's bed inside the military-hospital room, under the ward's nighttime lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Hospital bed (Occupied by 박철진 during treatment) — Only a narrow portion beside his head remains visible; used as A subdued edge establishes his recumbent position without distracting from the tear.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient hospital illumination preserves the reddened eye and the small highlight on the tear without specifying an unsupported fixture or light color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Park Cheoljin is reclined on a hospital bed, with one eye covered and an arm and a leg immobilized in casts, looking upward with his uncovered eye. The source does not identify the affected sides or specify his limbs' exact placement and his head's angle.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The setting remains a military hospital ward at night. 박철진: He remains bedridden with severe injuries to his face, arm and leg, with one eye covered and his arm and leg immobilized in casts. His remaining eye is bloodshot and filling with tears.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 박철진 (한국인 남성, 44세, 중년의 얼굴 윤곽, 짧은 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "안대를 하지 않은 왼쪽 눈의 시선이 위쪽을 향하고 있음.",
    "built_space": "병원 침대 위이며, 머리 뒤쪽과 옆으로 베개 및 침대 난간이 올바르게 위치함.",
    "entities": "레퍼런스와 일치하는 외형의 박철진이 회색 환자복을 입고 오른쪽 눈에 안대를 착용함.",
    "hard_violations": [],
    "physics": "머리는 베개에 안정적으로 지지되어 있으나, 뺨을 타고 흐르는 눈물이 일정한 간격의 물방울들로 끊어져 있어 물리적으로 약간 어색함."
   },
   {
    "label": "B",
    "direction": "안대를 하지 않은 왼쪽 눈의 시선이 위쪽을 향하고 있음.",
    "built_space": "병원 침대 위이며, 화면 가장자리에 베개와 침대 난간 일부가 적절하게 배치됨.",
    "entities": "레퍼런스와 일치하는 외형의 박철진이 회색 환자복을 입고 오른쪽 눈에 안대를 착용함.",
    "hard_violations": [],
    "physics": "머리는 베개에 자연스럽게 얹혀 있으며, 눈물이 중력에 따라 뺨을 타고 흐르는 궤적과 맺힘이 현실적임."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "지시된 클로즈업 구도와 안대, 충혈된 눈 등 주요 설정은 잘 따랐으나, 눈물 자국이 점선처럼 다소 인위적으로 표현되었습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "요구된 클로즈업 샷 안에서 충혈된 눈과 현실적으로 흐르는 눈물, 레퍼런스와 일치하는 인물의 외형 및 감정 표현을 매우 훌륭하게 구현했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "안대를 하지 않은 왼쪽 눈의 시선이 위쪽을 향하고 있음.",
        "built_space": "병원 침대 위이며, 머리 뒤쪽과 옆으로 베개 및 침대 난간이 올바르게 위치함.",
        "entities": "레퍼런스와 일치하는 외형의 박철진이 회색 환자복을 입고 오른쪽 눈에 안대를 착용함.",
        "hard_violations": [],
        "physics": "머리는 베개에 안정적으로 지지되어 있으나, 뺨을 타고 흐르는 눈물이 일정한 간격의 물방울들로 끊어져 있어 물리적으로 약간 어색함."
       },
       {
        "label": "B",
        "direction": "안대를 하지 않은 왼쪽 눈의 시선이 위쪽을 향하고 있음.",
        "built_space": "병원 침대 위이며, 화면 가장자리에 베개와 침대 난간 일부가 적절하게 배치됨.",
        "entities": "레퍼런스와 일치하는 외형의 박철진이 회색 환자복을 입고 오른쪽 눈에 안대를 착용함.",
        "hard_violations": [],
        "physics": "머리는 베개에 자연스럽게 얹혀 있으며, 눈물이 중력에 따라 뺨을 타고 흐르는 궤적과 맺힘이 현실적임."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "지시된 클로즈업 구도와 안대, 충혈된 눈 등 주요 설정은 잘 따랐으나, 눈물 자국이 점선처럼 다소 인위적으로 표현되었습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "요구된 클로즈업 샷 안에서 충혈된 눈과 현실적으로 흐르는 눈물, 레퍼런스와 일치하는 인물의 외형 및 감정 표현을 매우 훌륭하게 구현했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "안대를 하지 않은 왼쪽 눈의 시선이 위쪽을 향하고 있음.",
        "built_space": "병원 침대 위이며, 머리 뒤쪽과 옆으로 베개 및 침대 난간이 올바르게 위치함.",
        "entities": "레퍼런스와 일치하는 외형의 박철진이 회색 환자복을 입고 오른쪽 눈에 안대를 착용함.",
        "hard_violations": [],
        "physics": "머리는 베개에 안정적으로 지지되어 있으나, 뺨을 타고 흐르는 눈물이 일정한 간격의 물방울들로 끊어져 있어 물리적으로 약간 어색함."
       },
       {
        "label": "B",
        "direction": "안대를 하지 않은 왼쪽 눈의 시선이 위쪽을 향하고 있음.",
        "built_space": "병원 침대 위이며, 화면 가장자리에 베개와 침대 난간 일부가 적절하게 배치됨.",
        "entities": "레퍼런스와 일치하는 외형의 박철진이 회색 환자복을 입고 오른쪽 눈에 안대를 착용함.",
        "hard_violations": [],
        "physics": "머리는 베개에 자연스럽게 얹혀 있으며, 눈물이 중력에 따라 뺨을 타고 흐르는 궤적과 맺힘이 현실적임."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "얼굴 중심의 클로즈업에서 더 뚜렷한 충혈과 뺨을 흐르는 한 방울, 벌어진 입의 처절한 표정이 핵심 순간을 더 충실하게 구현한다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "누운 자세와 위쪽 시선은 맞지만, 눈물방울이 두 개로 읽히고 충혈과 처절한 표정이 A보다 약하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "가리지 않은 눈은 렌즈가 아니라 화면 위쪽의 프레임 밖을 향한다. 특정 대상은 보이지 않지만 누운 채 위를 보는 지정 시선에 부합한다. 눈물 자국은 아래 눈꺼풀에서 뺨 아래쪽으로 이어진다.",
        "built_space": "흰 베개 하나가 머리 뒤를 받치며, 왼쪽에는 침대 난간 일부 하나, 오른쪽 위에는 머리판 일부가 보인다. 뒤의 회녹색 커튼과 절제된 조명은 참조 병실과 양립한다. 얼굴이 화면 대부분을 차지하고 침대는 주변부에 머문다. 중복 설비나 반사는 보이지 않는다.",
        "entities": "짧은 검은 머리의 중년 동아시아계 남성 한 명으로, 참조 속 박철진의 얼굴과 회색 상의, 흰 안대가 대체로 유지된다. 드러난 눈의 충혈과 젖은 아래 눈꺼풀, 뺨의 눈물 자국 끝에 맺힌 한 방울이 선명하다. 눈썹의 긴장과 조금 벌어진 입이 고통과 절박함을 드러낸다. 팔·다리의 깁스는 올바른 클로즈업 범위 밖이다. 다른 사람이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "뒤통수와 목 뒤가 베개에 지지되고 몸통은 침대에 누워 있다. 공중에 떠 있는 신체나 소품은 없다. 안대는 머리를 두른 끈으로 고정되어 있으며, 눈물은 피부에 붙어 광택을 내고 뺨을 따라 흘러 물리적으로 자연스럽다."
       },
       {
        "label": "B",
        "direction": "가리지 않은 눈이 화면 위쪽의 프레임 밖을 바라보며 렌즈를 직접 보지 않는다. 특정 대상은 화면에 없고, 지정된 위쪽 시선은 유지된다. 눈물은 눈 아래에서 뺨 아래쪽으로 흐른다.",
        "built_space": "머리를 받치는 흰 베개 하나, 왼쪽 침대 난간 일부 하나, 오른쪽 위 머리판 일부와 배경 커튼이 보인다. 참조 병실의 재질과 조명 분위기에 부합하며 중복 설비나 불가능한 반사는 없다. 얼굴 클로즈업이지만 A보다 상의와 왼쪽 난간이 조금 더 드러난다.",
        "entities": "참조 인물과 대체로 일치하는 짧은 검은 머리의 중년 동아시아계 남성 한 명이며, 회색 상의와 흰 안대를 유지한다. 드러난 눈은 젖어 있고 붉지만 A보다 충혈이 약하다. 뺨의 눈물 자국에 분리된 두 방울이 보여 '눈물 한 방울'과 덜 정확히 맞는다. 다문 입과 긴장된 눈썹은 슬픔을 표현하지만 처절함은 상대적으로 절제되어 있다. 깁스 부위는 프레임 밖이며 다른 사람이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "뒤통수와 목은 베개에 기대고 어깨와 몸통은 침대에 지지된다. 지지 없이 떠 있는 부위는 없다. 안대 끈은 머리에 걸려 있고 눈물방울은 뺨 표면에 붙어 있어, 자세와 액체 표현 모두 물리적으로 가능하다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "얼굴 중심의 클로즈업에서 더 뚜렷한 충혈과 뺨을 흐르는 한 방울, 벌어진 입의 처절한 표정이 핵심 순간을 더 충실하게 구현한다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "누운 자세와 위쪽 시선은 맞지만, 눈물방울이 두 개로 읽히고 충혈과 처절한 표정이 A보다 약하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "가리지 않은 눈은 렌즈가 아니라 화면 위쪽의 프레임 밖을 향한다. 특정 대상은 보이지 않지만 누운 채 위를 보는 지정 시선에 부합한다. 눈물 자국은 아래 눈꺼풀에서 뺨 아래쪽으로 이어진다.",
        "built_space": "흰 베개 하나가 머리 뒤를 받치며, 왼쪽에는 침대 난간 일부 하나, 오른쪽 위에는 머리판 일부가 보인다. 뒤의 회녹색 커튼과 절제된 조명은 참조 병실과 양립한다. 얼굴이 화면 대부분을 차지하고 침대는 주변부에 머문다. 중복 설비나 반사는 보이지 않는다.",
        "entities": "짧은 검은 머리의 중년 동아시아계 남성 한 명으로, 참조 속 박철진의 얼굴과 회색 상의, 흰 안대가 대체로 유지된다. 드러난 눈의 충혈과 젖은 아래 눈꺼풀, 뺨의 눈물 자국 끝에 맺힌 한 방울이 선명하다. 눈썹의 긴장과 조금 벌어진 입이 고통과 절박함을 드러낸다. 팔·다리의 깁스는 올바른 클로즈업 범위 밖이다. 다른 사람이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "뒤통수와 목 뒤가 베개에 지지되고 몸통은 침대에 누워 있다. 공중에 떠 있는 신체나 소품은 없다. 안대는 머리를 두른 끈으로 고정되어 있으며, 눈물은 피부에 붙어 광택을 내고 뺨을 따라 흘러 물리적으로 자연스럽다."
       },
       {
        "label": "A",
        "direction": "가리지 않은 눈이 화면 위쪽의 프레임 밖을 바라보며 렌즈를 직접 보지 않는다. 특정 대상은 화면에 없고, 지정된 위쪽 시선은 유지된다. 눈물은 눈 아래에서 뺨 아래쪽으로 흐른다.",
        "built_space": "머리를 받치는 흰 베개 하나, 왼쪽 침대 난간 일부 하나, 오른쪽 위 머리판 일부와 배경 커튼이 보인다. 참조 병실의 재질과 조명 분위기에 부합하며 중복 설비나 불가능한 반사는 없다. 얼굴 클로즈업이지만 A보다 상의와 왼쪽 난간이 조금 더 드러난다.",
        "entities": "참조 인물과 대체로 일치하는 짧은 검은 머리의 중년 동아시아계 남성 한 명이며, 회색 상의와 흰 안대를 유지한다. 드러난 눈은 젖어 있고 붉지만 A보다 충혈이 약하다. 뺨의 눈물 자국에 분리된 두 방울이 보여 '눈물 한 방울'과 덜 정확히 맞는다. 다문 입과 긴장된 눈썹은 슬픔을 표현하지만 처절함은 상대적으로 절제되어 있다. 깁스 부위는 프레임 밖이며 다른 사람이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "뒤통수와 목은 베개에 기대고 어깨와 몸통은 침대에 지지된다. 지지 없이 떠 있는 부위는 없다. 안대 끈은 머리에 걸려 있고 눈물방울은 뺨 표면에 붙어 있어, 자세와 액체 표현 모두 물리적으로 가능하다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.746,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.746,
    "B": 2.0
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "A": 1746,
   "B": 2000
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1746,
    "verdict_ko": "지시된 클로즈업 구도와 안대, 충혈된 눈 등 주요 설정은 잘 따랐으나, 눈물 자국이 점선처럼 다소 인위적으로 표현되었습니다."
   },
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "요구된 클로즈업 샷 안에서 충혈된 눈과 현실적으로 흐르는 눈물, 레퍼런스와 일치하는 인물의 외형 및 감정 표현을 매우 훌륭하게 구현했습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 박철진 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S70sh5_sel.png",
    "asset_id": "72b73958-43a2-4bac-b427-60a4856e24c0",
    "role": "prev_still"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-c81a-7932-982d-e121563e77ac",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S70sh5"
  },
  "locked_char_refs_excluded": [
   "박철진(C16)"
  ]
 },
 "S70sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:18:54.227156+00:00",
  "fingerprint": "7908c909f602170a5d50d75b9df3643d88a0a0364d6146b1ec3787f5792d403a",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S70sh9_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S70sh9_sel.png",
  "source_sha256": "563164dbd9a32fde20fed3c84e54bf92e020feef233bedbe4813da2b1707bd91",
  "file": "S70sh9_cine.png",
  "staged_sha256": "55142076d76c57556a42028b9b9fb4c7e2db8c2d0c131cc9ff58cddce00d5458",
  "latency_ms": 10131
 },
 "S71sh4::signage": {
  "fp": "21b31e69af2b58e2",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "groupbg::open_truck_cab": {
  "input_fingerprint": "b375775d773c387c",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "open_truck_cab",
    "tags": [
     "S63sh17",
     "S71sh17",
     "S71sh4",
     "S72sh43",
     "S72sh58"
    ]
   },
   "context_sig": "db8201d83d85b000"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At the passenger area adjoining the old truck's open cargo bed during the predawn drive, with faint dawn light outside.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n해남 지동현의 연구소 1층: 오래 방치되어 먼지가 쌓인 텅 빈 로비와 나선형 구조의 실내. (특징: 끼익 소리를 내며 열리는 낡은 거대 철문; 비어 있는 1층 실내 공간과 뽀얀 바닥 먼지; 바닥에 떨어져 유리가 깨진 액자; 위층으로 이어지는 나선형 철제 계단) / 현우와 수빈이 사용하는 낡은 트럭 운전석: 지붕 덮개가 떨어져 나가 하늘이 개방된 오래된 트럭의 앞좌석. (특징: 지붕 패널이 없는 오픈형 트럭 실내; 낡은 운전대와 대시보드; 마주 보는 운전자와 조수석 인물의 상반신)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 홀로 트럭을 운전하는 현우. 보조석에는 수빈이 타고 있고.\n- 지붕 하나 없는 낡은 트럭들이다.\n- 71. 도로 위, 현우의 트럭 – 새벽\n- / 도로 위. 차 안 - N\n- 드르르르르륵-! 문이 열리면-!\n\nTIME OF DAY (lock): dawn.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At the passenger area adjoining the old truck's open cargo bed during the predawn drive, with faint dawn light outside.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n해남 지동현의 연구소 1층: 오래 방치되어 먼지가 쌓인 텅 빈 로비와 나선형 구조의 실내. (특징: 끼익 소리를 내며 열리는 낡은 거대 철문; 비어 있는 1층 실내 공간과 뽀얀 바닥 먼지; 바닥에 떨어져 유리가 깨진 액자; 위층으로 이어지는 나선형 철제 계단) / 현우와 수빈이 사용하는 낡은 트럭 운전석: 지붕 덮개가 떨어져 나가 하늘이 개방된 오래된 트럭의 앞좌석. (특징: 지붕 패널이 없는 오픈형 트럭 실내; 낡은 운전대와 대시보드; 마주 보는 운전자와 조수석 인물의 상반신)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 홀로 트럭을 운전하는 현우. 보조석에는 수빈이 타고 있고.\n- 지붕 하나 없는 낡은 트럭들이다.\n- 71. 도로 위, 현우의 트럭 – 새벽\n- / 도로 위. 차 안 - N\n- 드르르르르륵-! 문이 열리면-!\n\nTIME OF DAY (lock): dawn.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_open_truck_cab_cebf8e.png",
  "asset_id": "cf509bfd-c4f2-4e5c-850c-88fea08ed824",
  "input_asset_ids": [
   "2d015eed-4450-4662-acb3-9b08a2be5734"
  ],
  "origin_tag": "S71sh4",
  "place_text": "At the passenger area adjoining the old truck's open cargo bed during the predawn drive, with faint dawn light outside.",
  "origin_inputs": {
   "place_text": "At the passenger area adjoining the old truck's open cargo bed during the predawn drive, with faint dawn light outside.",
   "time_of_day_en": "dawn",
   "conti_asset_id": "2d015eed-4450-4662-acb3-9b08a2be5734"
  }
 },
 "S71sh4::bgfirst_bg": {
  "input_fingerprint": "5a5a48233e11ce51",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리가 내민 삐뚤빼뚤한 꽃 그림 종이를 쥔 채 어리둥절한 표정을 짓는 앰버의 상체.\n\nLOCATION (lock): At the passenger area adjoining the old truck's open cargo bed during the predawn drive, with faint dawn light outside.\n\nTIME OF DAY (lock): dawn.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Flower drawing paper (Held by 앰버 after 찰리 hands it over) — The reverse side bearing the uneven flower drawing faces obliquely toward the camera; no pigment color is specified; used as Connects her puzzled expression to the inadequate clue in her hands; Truck passenger-side dashboard (Inside the traveling truck) — A shallow oblique edge appears along the lower foreground; used as Locates the camera inside the cab without blocking the paper.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued dawn ambience keeps 앰버's expression and the uneven drawing legible with restrained contrast and no added colored light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리가 내민 삐뚤빼뚤한 꽃 그림 종이를 쥔 채 어리둥절한 표정을 짓는 앰버의 상체.\n\nLOCATION (lock): At the passenger area adjoining the old truck's open cargo bed during the predawn drive, with faint dawn light outside.\n\nTIME OF DAY (lock): dawn.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Flower drawing paper (Held by 앰버 after 찰리 hands it over) — The reverse side bearing the uneven flower drawing faces obliquely toward the camera; no pigment color is specified; used as Connects her puzzled expression to the inadequate clue in her hands; Truck passenger-side dashboard (Inside the traveling truck) — A shallow oblique edge appears along the lower foreground; used as Locates the camera inside the cab without blocking the paper.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued dawn ambience keeps 앰버's expression and the uneven drawing legible with restrained contrast and no added colored light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S71sh4__bgfirst_bg.png",
  "asset_id": "1e25c307-2725-4133-9fb5-567004d17fb4",
  "input_asset_ids": [
   "2d015eed-4450-4662-acb3-9b08a2be5734",
   "cf509bfd-c4f2-4e5c-850c-88fea08ed824"
  ]
 },
 "S71sh4": {
  "input_fingerprint": "6f067055ae26613f",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 찰리가 내민 삐뚤빼뚤한 꽃 그림 종이를 쥔 채 어리둥절한 표정을 짓는 앰버의 상체.\n\nLOCATION (lock): At the passenger area adjoining the old truck's open cargo bed during the predawn drive, with faint dawn light outside. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Flower drawing paper (Held by 앰버 after 찰리 hands it over) — The reverse side bearing the uneven flower drawing faces obliquely toward the camera; no pigment color is specified; used as Connects her puzzled expression to the inadequate clue in her hands; Truck passenger-side dashboard (Inside the traveling truck) — A shallow oblique edge appears along the lower foreground; used as Locates the camera inside the cab without blocking the paper.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued dawn ambience keeps 앰버's expression and the uneven drawing legible with restrained contrast and no added colored light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old truck is traveling south at dawn; the violet sketch is drawn on the reverse of a sheet of paper. Charlie remains damaged, with punctured and dented bodywork, an exposed chest opening and malfunctioning sensors, and retains B-200's chest component. 앰버: She holds the sheet bearing the flower drawing. The injury to the back of her head from the rifle-butt blow has not been treated on screen.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 앰버 right now, so 앰버's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 앰버: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 찰리가 내민 삐뚤빼뚤한 꽃 그림 종이를 쥔 채 어리둥절한 표정을 짓는 앰버의 상체.\n\nLOCATION (lock): At the passenger area adjoining the old truck's open cargo bed during the predawn drive, with faint dawn light outside. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Flower drawing paper (Held by 앰버 after 찰리 hands it over) — The reverse side bearing the uneven flower drawing faces obliquely toward the camera; no pigment color is specified; used as Connects her puzzled expression to the inadequate clue in her hands; Truck passenger-side dashboard (Inside the traveling truck) — A shallow oblique edge appears along the lower foreground; used as Locates the camera inside the cab without blocking the paper.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued dawn ambience keeps 앰버's expression and the uneven drawing legible with restrained contrast and no added colored light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old truck is traveling south at dawn; the violet sketch is drawn on the reverse of a sheet of paper. Charlie remains damaged, with punctured and dented bodywork, an exposed chest opening and malfunctioning sensors, and retains B-200's chest component. 앰버: She holds the sheet bearing the flower drawing. The injury to the back of her head from the rifle-butt blow has not been treated on screen.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 앰버 right now, so 앰버's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 앰버: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 찰리가 내민 삐뚤빼뚤한 꽃 그림 종이를 쥔 채 어리둥절한 표정을 짓는 앰버의 상체.\n\nLOCATION (lock): At the passenger area adjoining the old truck's open cargo bed during the predawn drive, with faint dawn light outside. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Flower drawing paper (Held by 앰버 after 찰리 hands it over) — The reverse side bearing the uneven flower drawing faces obliquely toward the camera; no pigment color is specified; used as Connects her puzzled expression to the inadequate clue in her hands; Truck passenger-side dashboard (Inside the traveling truck) — A shallow oblique edge appears along the lower foreground; used as Locates the camera inside the cab without blocking the paper.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued dawn ambience keeps 앰버's expression and the uneven drawing legible with restrained contrast and no added colored light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old truck is traveling south at dawn; the violet sketch is drawn on the reverse of a sheet of paper. Charlie remains damaged, with punctured and dented bodywork, an exposed chest opening and malfunctioning sensors, and retains B-200's chest component. 앰버: She holds the sheet bearing the flower drawing. The injury to the back of her head from the rifle-butt blow has not been treated on screen.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 앰버 right now, so 앰버's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 앰버: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S71sh4__bgfirst_bg.png",
     "asset_id": "1e25c307-2725-4133-9fb5-567004d17fb4",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S71sh4.png",
     "asset_id": "2d015eed-4450-4662-acb3-9b08a2be5734",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 찰리의 꽃 그림: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:1316405>",
     "asset_id": "108fb164-c05c-4d1c-9144-293628b1d420",
     "role": "prop_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_open_truck_cab_cebf8e.png",
     "asset_id": "cf509bfd-c4f2-4e5c-850c-88fea08ed824",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 찰리의 꽃 그림: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:1316405>",
     "asset_id": "108fb164-c05c-4d1c-9144-293628b1d420",
     "role": "prop_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "B",
    "direction": "앰버는 눈앞에 든 종이 쪽으로 시선을 내리고 미간을 찌푸립니다. 꽃이 그려진 면은 카메라를 향해 비스듬히 드러나며, 이는 그림 면을 관객에게 보여주라는 명시적 배치에 맞습니다. 도로는 왼쪽 소실점으로 이어지지만 정지 화면만으로 트럭의 남행 여부는 확인할 수 없습니다.",
    "built_space": "검은 좌석 등받이 두 개, 왼쪽 문과 창틀 한 세트, 사이드미러 한 개, 중앙 세로 지지대가 있는 뒤쪽 금속 프레임과 열린 적재함이 보입니다. 앰버는 왼쪽 좌석 앞에 앉아 있고 오른쪽 좌석은 비어 있습니다. 녹슨 차체와 낡은 직물은 장소 참조에 가깝습니다. 대시보드는 하단에 비스듬하게 걸리며 그림을 가리지 않습니다. 거울 속 하늘과 도로 일부에 명백한 광학적 모순은 없습니다.",
    "entities": "보이는 인물은 금발의 약 10세 여자아이 한 명뿐이며, 둥근 얼굴과 큰 눈, 남색 반소매 상의가 앰버 참조에 가깝습니다. 혼혈 배경 자체는 외관만으로 확정할 수 없습니다. 두 손의 크기와 피부는 아이에게 자연스럽게 연결됩니다. 종이에는 보라색 꽃 한 송이와 녹색 줄기·잎이 그려져 있어 삐뚤빼뚤한 꽃 그림이라는 요구에는 맞지만, 참조의 여러 송이 제비꽃과 거친 종이 가장자리는 재현하지 못했습니다. 읽을 수 있는 글자는 없습니다. 찰리와 뒤통수 상처는 이 구도에서 보이지 않습니다.",
    "hard_violations": [],
    "physics": "앰버의 몸은 좌석에 앉은 자세로 지지되며, 하체는 프레임 밖입니다. 양손이 종이의 양쪽 아래 가장자리를 실제로 잡고 있어 종이가 떠 있지 않습니다. 손목과 팔의 연결도 자연스럽고 종이는 약간 휘어진 물성으로 보입니다."
   },
   {
    "label": "A",
    "direction": "앰버의 시선은 손에 든 종이가 아니라 화면 왼쪽 바깥을 향합니다. 시선의 대상은 보이지 않아 찰리를 보는 것인지는 확인할 수 없습니다. 꽃 그림 면은 카메라 쪽으로 드러나지만 A보다 정면에 가깝습니다. 도로의 진행 축은 보이나 남행 여부는 확인할 수 없습니다.",
    "built_space": "검은 좌석 등받이 두 개, 왼쪽 문과 창틀 한 세트, 사이드미러 한 개, 뒤쪽 금속 프레임의 중앙 지지대와 열린 적재함이 보입니다. 앰버는 왼쪽 좌석에 앉고 오른쪽 좌석은 비어 있어 기본 공간 배치는 참조와 맞습니다. 대각선 안전벨트도 보입니다. 다만 인물 위와 주변의 공간이 많이 포함되고 대시보드가 하단을 넓게 차지해, 요구한 상체 중심 미디엄 숏과 얕은 전경보다 넓은 구도입니다. 거울 반사에 명백한 모순은 없습니다.",
    "entities": "금발의 어린 여자아이 한 명만 보이며 얼굴과 체격은 앰버 참조에 가깝습니다. 다만 남색 상의가 참조의 반소매가 아닌 긴소매입니다. 손의 크기와 피부는 아이에게 맞습니다. 종이는 양손에 들려 있으나 그림은 단순한 무채색 꽃 한 송이로, 참조의 제비꽃 묶음과 다릅니다. 무채색 자체는 색을 지정하지 않은 촬영 지시의 위반이 아닙니다. 읽을 수 있는 글자는 없고 찰리와 뒤통수 상처도 보이지 않습니다.",
    "hard_violations": [],
    "physics": "앰버는 등받이를 뒤로 둔 정상적인 착석 자세이며 좌석이 몸을 지지합니다. 양손은 종이 양옆을 잡고 있고 팔과 손목도 자연스럽게 이어집니다. 안전벨트는 차체 쪽 고정부에서 몸 앞으로 이어져 보입니다. 지지 없이 떠 있는 신체나 물체는 없습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": null,
     "normalized": null,
     "ok": false
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "상체 중심의 미디엄 숏과 종이를 향한 의아한 표정, 얕은 대시보드 전경이 지시에 더 충실하지만 꽃 그림은 참조의 제비꽃 묶음과 다릅니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "인물을 작게 잡고 대시보드를 넓게 보여 상체 중심 구도가 약해졌으며, 종이 밖을 보는 시선과 긴소매 의상도 A보다 덜 충실합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버는 눈앞에 든 종이 쪽으로 시선을 내리고 미간을 찌푸립니다. 꽃이 그려진 면은 카메라를 향해 비스듬히 드러나며, 이는 그림 면을 관객에게 보여주라는 명시적 배치에 맞습니다. 도로는 왼쪽 소실점으로 이어지지만 정지 화면만으로 트럭의 남행 여부는 확인할 수 없습니다.",
        "built_space": "검은 좌석 등받이 두 개, 왼쪽 문과 창틀 한 세트, 사이드미러 한 개, 중앙 세로 지지대가 있는 뒤쪽 금속 프레임과 열린 적재함이 보입니다. 앰버는 왼쪽 좌석 앞에 앉아 있고 오른쪽 좌석은 비어 있습니다. 녹슨 차체와 낡은 직물은 장소 참조에 가깝습니다. 대시보드는 하단에 비스듬하게 걸리며 그림을 가리지 않습니다. 거울 속 하늘과 도로 일부에 명백한 광학적 모순은 없습니다.",
        "entities": "보이는 인물은 금발의 약 10세 여자아이 한 명뿐이며, 둥근 얼굴과 큰 눈, 남색 반소매 상의가 앰버 참조에 가깝습니다. 혼혈 배경 자체는 외관만으로 확정할 수 없습니다. 두 손의 크기와 피부는 아이에게 자연스럽게 연결됩니다. 종이에는 보라색 꽃 한 송이와 녹색 줄기·잎이 그려져 있어 삐뚤빼뚤한 꽃 그림이라는 요구에는 맞지만, 참조의 여러 송이 제비꽃과 거친 종이 가장자리는 재현하지 못했습니다. 읽을 수 있는 글자는 없습니다. 찰리와 뒤통수 상처는 이 구도에서 보이지 않습니다.",
        "hard_violations": [],
        "physics": "앰버의 몸은 좌석에 앉은 자세로 지지되며, 하체는 프레임 밖입니다. 양손이 종이의 양쪽 아래 가장자리를 실제로 잡고 있어 종이가 떠 있지 않습니다. 손목과 팔의 연결도 자연스럽고 종이는 약간 휘어진 물성으로 보입니다."
       },
       {
        "label": "B",
        "direction": "앰버의 시선은 손에 든 종이가 아니라 화면 왼쪽 바깥을 향합니다. 시선의 대상은 보이지 않아 찰리를 보는 것인지는 확인할 수 없습니다. 꽃 그림 면은 카메라 쪽으로 드러나지만 A보다 정면에 가깝습니다. 도로의 진행 축은 보이나 남행 여부는 확인할 수 없습니다.",
        "built_space": "검은 좌석 등받이 두 개, 왼쪽 문과 창틀 한 세트, 사이드미러 한 개, 뒤쪽 금속 프레임의 중앙 지지대와 열린 적재함이 보입니다. 앰버는 왼쪽 좌석에 앉고 오른쪽 좌석은 비어 있어 기본 공간 배치는 참조와 맞습니다. 대각선 안전벨트도 보입니다. 다만 인물 위와 주변의 공간이 많이 포함되고 대시보드가 하단을 넓게 차지해, 요구한 상체 중심 미디엄 숏과 얕은 전경보다 넓은 구도입니다. 거울 반사에 명백한 모순은 없습니다.",
        "entities": "금발의 어린 여자아이 한 명만 보이며 얼굴과 체격은 앰버 참조에 가깝습니다. 다만 남색 상의가 참조의 반소매가 아닌 긴소매입니다. 손의 크기와 피부는 아이에게 맞습니다. 종이는 양손에 들려 있으나 그림은 단순한 무채색 꽃 한 송이로, 참조의 제비꽃 묶음과 다릅니다. 무채색 자체는 색을 지정하지 않은 촬영 지시의 위반이 아닙니다. 읽을 수 있는 글자는 없고 찰리와 뒤통수 상처도 보이지 않습니다.",
        "hard_violations": [],
        "physics": "앰버는 등받이를 뒤로 둔 정상적인 착석 자세이며 좌석이 몸을 지지합니다. 양손은 종이 양옆을 잡고 있고 팔과 손목도 자연스럽게 이어집니다. 안전벨트는 차체 쪽 고정부에서 몸 앞으로 이어져 보입니다. 지지 없이 떠 있는 신체나 물체는 없습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "상체 중심의 미디엄 숏과 종이를 향한 의아한 표정, 얕은 대시보드 전경이 지시에 더 충실하지만 꽃 그림은 참조의 제비꽃 묶음과 다릅니다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "인물을 작게 잡고 대시보드를 넓게 보여 상체 중심 구도가 약해졌으며, 종이 밖을 보는 시선과 긴소매 의상도 A보다 덜 충실합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "앰버는 눈앞에 든 종이 쪽으로 시선을 내리고 미간을 찌푸립니다. 꽃이 그려진 면은 카메라를 향해 비스듬히 드러나며, 이는 그림 면을 관객에게 보여주라는 명시적 배치에 맞습니다. 도로는 왼쪽 소실점으로 이어지지만 정지 화면만으로 트럭의 남행 여부는 확인할 수 없습니다.",
        "built_space": "검은 좌석 등받이 두 개, 왼쪽 문과 창틀 한 세트, 사이드미러 한 개, 중앙 세로 지지대가 있는 뒤쪽 금속 프레임과 열린 적재함이 보입니다. 앰버는 왼쪽 좌석 앞에 앉아 있고 오른쪽 좌석은 비어 있습니다. 녹슨 차체와 낡은 직물은 장소 참조에 가깝습니다. 대시보드는 하단에 비스듬하게 걸리며 그림을 가리지 않습니다. 거울 속 하늘과 도로 일부에 명백한 광학적 모순은 없습니다.",
        "entities": "보이는 인물은 금발의 약 10세 여자아이 한 명뿐이며, 둥근 얼굴과 큰 눈, 남색 반소매 상의가 앰버 참조에 가깝습니다. 혼혈 배경 자체는 외관만으로 확정할 수 없습니다. 두 손의 크기와 피부는 아이에게 자연스럽게 연결됩니다. 종이에는 보라색 꽃 한 송이와 녹색 줄기·잎이 그려져 있어 삐뚤빼뚤한 꽃 그림이라는 요구에는 맞지만, 참조의 여러 송이 제비꽃과 거친 종이 가장자리는 재현하지 못했습니다. 읽을 수 있는 글자는 없습니다. 찰리와 뒤통수 상처는 이 구도에서 보이지 않습니다.",
        "hard_violations": [],
        "physics": "앰버의 몸은 좌석에 앉은 자세로 지지되며, 하체는 프레임 밖입니다. 양손이 종이의 양쪽 아래 가장자리를 실제로 잡고 있어 종이가 떠 있지 않습니다. 손목과 팔의 연결도 자연스럽고 종이는 약간 휘어진 물성으로 보입니다."
       },
       {
        "label": "A",
        "direction": "앰버의 시선은 손에 든 종이가 아니라 화면 왼쪽 바깥을 향합니다. 시선의 대상은 보이지 않아 찰리를 보는 것인지는 확인할 수 없습니다. 꽃 그림 면은 카메라 쪽으로 드러나지만 A보다 정면에 가깝습니다. 도로의 진행 축은 보이나 남행 여부는 확인할 수 없습니다.",
        "built_space": "검은 좌석 등받이 두 개, 왼쪽 문과 창틀 한 세트, 사이드미러 한 개, 뒤쪽 금속 프레임의 중앙 지지대와 열린 적재함이 보입니다. 앰버는 왼쪽 좌석에 앉고 오른쪽 좌석은 비어 있어 기본 공간 배치는 참조와 맞습니다. 대각선 안전벨트도 보입니다. 다만 인물 위와 주변의 공간이 많이 포함되고 대시보드가 하단을 넓게 차지해, 요구한 상체 중심 미디엄 숏과 얕은 전경보다 넓은 구도입니다. 거울 반사에 명백한 모순은 없습니다.",
        "entities": "금발의 어린 여자아이 한 명만 보이며 얼굴과 체격은 앰버 참조에 가깝습니다. 다만 남색 상의가 참조의 반소매가 아닌 긴소매입니다. 손의 크기와 피부는 아이에게 맞습니다. 종이는 양손에 들려 있으나 그림은 단순한 무채색 꽃 한 송이로, 참조의 제비꽃 묶음과 다릅니다. 무채색 자체는 색을 지정하지 않은 촬영 지시의 위반이 아닙니다. 읽을 수 있는 글자는 없고 찰리와 뒤통수 상처도 보이지 않습니다.",
        "hard_violations": [],
        "physics": "앰버는 등받이를 뒤로 둔 정상적인 착석 자세이며 좌석이 몸을 지지합니다. 양손은 종이 양옆을 잡고 있고 팔과 손목도 자연스럽게 이어집니다. 안전벨트는 차체 쪽 고정부에서 몸 앞으로 이어져 보입니다. 지지 없이 떠 있는 신체나 물체는 없습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gemini-pro"
   ],
   "route": "single_reverse"
  },
  "totals": {
   "B": 8,
   "A": 5
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 8,
    "verdict_ko": "상체 중심의 미디엄 숏과 종이를 향한 의아한 표정, 얕은 대시보드 전경이 지시에 더 충실하지만 꽃 그림은 참조의 제비꽃 묶음과 다릅니다."
   },
   {
    "label": "A",
    "score": 5,
    "verdict_ko": "인물을 작게 잡고 대시보드를 넓게 보여 상체 중심 구도가 약해졌으며, 종이 밖을 보는 시선과 긴소매 의상도 A보다 덜 충실합니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_open_truck_cab_cebf8e.png",
    "asset_id": "cf509bfd-c4f2-4e5c-850c-88fea08ed824",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   },
   {
    "label": "PROP REFERENCE — 찰리의 꽃 그림: the exact object appearing in this shot; match its look, material and wear exactly.",
    "path": "<bytes:1316405>",
    "asset_id": "108fb164-c05c-4d1c-9144-293628b1d420",
    "role": "prop_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-c9c6-71f9-9810-191cd9c314c1",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S71sh4__bgfirst_bg.png",
   "bg_asset_id": "1e25c307-2725-4133-9fb5-567004d17fb4",
   "bg_record_key": "S71sh4::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "open_truck_cab",
   "groupbg_asset_id": "cf509bfd-c4f2-4e5c-850c-88fea08ed824"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S71sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:20:47.176404+00:00",
  "fingerprint": "e99c1d64ca153a834a27ef979e3c8223826728a8f58ec48bc9ba32743022c019",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S71sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S71sh4_sel.png",
  "source_sha256": "397cb6dec4dfeb23d8e871c245aa909884b1172d05e69534fe098f7cb6180aca",
  "file": "S71sh4_cine.png",
  "staged_sha256": "76d569045c182eefe2d7442725c909eb33481eba0f50b1c62b299b1a58c26731",
  "latency_ms": 9065
 },
 "S71sh7::signage": {
  "fp": "9c51bf276a36526b",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "groupbg::coastal_search_hill": {
  "input_fingerprint": "7a41bf6375577d7c",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "coastal_search_hill",
    "tags": [
     "S71sh7"
    ]
   },
   "context_sig": "81a953d4a0d2dc9c"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On the coastal roadside hill where the group has stopped to search the ridges for violet flowers.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n현우와 수빈이 사용하는 낡은 트럭 운전석: 지붕 덮개가 떨어져 나가 하늘이 개방된 오래된 트럭의 앞좌석. (특징: 지붕 패널이 없는 오픈형 트럭 실내; 낡은 운전대와 대시보드; 마주 보는 운전자와 조수석 인물의 상반신)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 언덕 위에 서서 능선을 쳐다보는 현우.\n- 현우가 능선 너머로 제비꽃을 찾는 사이, 앰버와 찰리, 라울은 흙으로 장난을 치고 있다.\n\nTIME OF DAY (lock): dawn.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On the coastal roadside hill where the group has stopped to search the ridges for violet flowers.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n현우와 수빈이 사용하는 낡은 트럭 운전석: 지붕 덮개가 떨어져 나가 하늘이 개방된 오래된 트럭의 앞좌석. (특징: 지붕 패널이 없는 오픈형 트럭 실내; 낡은 운전대와 대시보드; 마주 보는 운전자와 조수석 인물의 상반신)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 언덕 위에 서서 능선을 쳐다보는 현우.\n- 현우가 능선 너머로 제비꽃을 찾는 사이, 앰버와 찰리, 라울은 흙으로 장난을 치고 있다.\n\nTIME OF DAY (lock): dawn.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_coastal_search_hill_7092b9.png",
  "asset_id": "ab08fa93-ee50-4572-91f7-b41dbd747f21",
  "input_asset_ids": [
   "6eebd1f8-91eb-42c9-ac4e-716f7b521fe1"
  ],
  "origin_tag": "S71sh7",
  "place_text": "On the coastal roadside hill where the group has stopped to search the ridges for violet flowers.",
  "origin_inputs": {
   "place_text": "On the coastal roadside hill where the group has stopped to search the ridges for violet flowers.",
   "time_of_day_en": "dawn",
   "conti_asset_id": "6eebd1f8-91eb-42c9-ac4e-716f7b521fe1"
  }
 },
 "S71sh7::bgfirst_bg": {
  "input_fingerprint": "6dca32eb5c7fcfc8",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 이별을 직감한 듯 눈물이 맺힌 눈으로 찰리를 올려다보는 앰버의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): On the coastal roadside hill where the group has stopped to search the ridges for violet flowers.\n\nTIME OF DAY (lock): dawn.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Hillside ground (Where the companions have been playing with soil); used as A softly resolved lower background preserves the outdoor setting without competing with her expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight renders the moisture in 앰버's eyes subtly, with controlled contrast and no dreamlike treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 이별을 직감한 듯 눈물이 맺힌 눈으로 찰리를 올려다보는 앰버의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): On the coastal roadside hill where the group has stopped to search the ridges for violet flowers.\n\nTIME OF DAY (lock): dawn.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Hillside ground (Where the companions have been playing with soil); used as A softly resolved lower background preserves the outdoor setting without competing with her expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight renders the moisture in 앰버's eyes subtly, with controlled contrast and no dreamlike treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S71sh7__bgfirst_bg.png",
  "asset_id": "0cca7227-2e69-4db8-ab19-8acea477fee5",
  "input_asset_ids": [
   "6eebd1f8-91eb-42c9-ac4e-716f7b521fe1",
   "ab08fa93-ee50-4572-91f7-b41dbd747f21"
  ]
 },
 "S71sh7": {
  "input_fingerprint": "3093f6d532756527",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 이별을 직감한 듯 눈물이 맺힌 눈으로 찰리를 올려다보는 앰버의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): On the coastal roadside hill where the group has stopped to search the ridges for violet flowers. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Hillside ground (Where the companions have been playing with soil); used as A softly resolved lower background preserves the outdoor setting without competing with her expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight renders the moisture in 앰버's eyes subtly, with controlled contrast and no dreamlike treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The truck has stopped near the coastal hills, and the flower sketch remains available without any stated disposal. Charlie's damaged bodywork, exposed chest opening and B-200 chest component persist. 앰버: She has tearful eyes and is wiping away tears; her earlier head injury remains unresolved.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 이별을 직감한 듯 눈물이 맺힌 눈으로 찰리를 올려다보는 앰버의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): On the coastal roadside hill where the group has stopped to search the ridges for violet flowers. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Hillside ground (Where the companions have been playing with soil); used as A softly resolved lower background preserves the outdoor setting without competing with her expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight renders the moisture in 앰버's eyes subtly, with controlled contrast and no dreamlike treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The truck has stopped near the coastal hills, and the flower sketch remains available without any stated disposal. Charlie's damaged bodywork, exposed chest opening and B-200 chest component persist. 앰버: She has tearful eyes and is wiping away tears; her earlier head injury remains unresolved.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 이별을 직감한 듯 눈물이 맺힌 눈으로 찰리를 올려다보는 앰버의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): On the coastal roadside hill where the group has stopped to search the ridges for violet flowers. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Hillside ground (Where the companions have been playing with soil); used as A softly resolved lower background preserves the outdoor setting without competing with her expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight renders the moisture in 앰버's eyes subtly, with controlled contrast and no dreamlike treatment.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The truck has stopped near the coastal hills, and the flower sketch remains available without any stated disposal. Charlie's damaged bodywork, exposed chest opening and B-200 chest component persist. 앰버: She has tearful eyes and is wiping away tears; her earlier head injury remains unresolved.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S71sh7__bgfirst_bg.png",
     "asset_id": "0cca7227-2e69-4db8-ab19-8acea477fee5",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S71sh7.png",
     "asset_id": "6eebd1f8-91eb-42c9-ac4e-716f7b521fe1",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_coastal_search_hill_7092b9.png",
     "asset_id": "ab08fa93-ee50-4572-91f7-b41dbd747f21",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "앰버의 시선이 화면 우측에 있는 찰리(로봇)의 얼굴 쪽을 향해 위로 향해 있습니다.",
    "built_space": "배경의 해안가 도로에 레퍼런스와 동일한 형태와 방향의 흰색 트럭이 주차되어 있습니다.",
    "entities": "앰버의 외모와 이마의 상처가 잘 나타나 있으며, 화면 우측의 찰리는 손상된 가슴 외장과 B-200 컴포넌트를 부착하고 있습니다.",
    "hard_violations": [
     "[gpt-high] 가시 인물을 앰버로 한정한 조건과 달리, 찰리의 얼굴 일부와 상체를 전경에 추가했다.",
     "[gpt-high] 오른쪽 몸체에 읽을 수 있는 'B-200' 표기가 있어 판독 가능한 글자 금지 조건을 위반한다."
    ],
    "physics": "앰버와 로봇 모두 지면에 안정적으로 서 있는 자세를 취하고 있습니다."
   },
   {
    "label": "B",
    "direction": "앰버의 시선이 특정 대상 없이 화면 밖 하늘을 향하고 있습니다.",
    "built_space": "흐릿한 해안가 언덕 배경만 보이며, 트럭 등의 구조물은 묘사되지 않았습니다.",
    "entities": "앰버의 얼굴과 눈물이 묘사되었으나 명시된 이마의 상처가 누락되었으며, 찰리도 프레임에 존재하지 않습니다.",
    "hard_violations": [],
    "physics": "앰버의 손이 뺨에 얹혀 있으며 눈물이 뺨을 타고 자연스럽게 맺혀 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "레퍼런스의 배경(트럭 포함), 앰버의 머리 상처, 찰리의 손상된 가슴 파츠를 매우 사실적으로 구현했으나, 눈물을 닦는 동작이 없고 로봇에 읽을 수 있는 글자(B-200)가 포함된 점이 감점 요인입니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "눈물이 맺힌 얼굴과 뺨에 손을 올린 모습은 표현되었으나, 필수 유지 조건인 앰버의 머리 상처가 누락되었고 찰리와 배경의 트럭 등 핵심 상황적 디테일이 생략되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버의 시선이 화면 우측에 있는 찰리(로봇)의 얼굴 쪽을 향해 위로 향해 있습니다.",
        "built_space": "배경의 해안가 도로에 레퍼런스와 동일한 형태와 방향의 흰색 트럭이 주차되어 있습니다.",
        "entities": "앰버의 외모와 이마의 상처가 잘 나타나 있으며, 화면 우측의 찰리는 손상된 가슴 외장과 B-200 컴포넌트를 부착하고 있습니다.",
        "hard_violations": [],
        "physics": "앰버와 로봇 모두 지면에 안정적으로 서 있는 자세를 취하고 있습니다."
       },
       {
        "label": "B",
        "direction": "앰버의 시선이 특정 대상 없이 화면 밖 하늘을 향하고 있습니다.",
        "built_space": "흐릿한 해안가 언덕 배경만 보이며, 트럭 등의 구조물은 묘사되지 않았습니다.",
        "entities": "앰버의 얼굴과 눈물이 묘사되었으나 명시된 이마의 상처가 누락되었으며, 찰리도 프레임에 존재하지 않습니다.",
        "hard_violations": [],
        "physics": "앰버의 손이 뺨에 얹혀 있으며 눈물이 뺨을 타고 자연스럽게 맺혀 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "레퍼런스의 배경(트럭 포함), 앰버의 머리 상처, 찰리의 손상된 가슴 파츠를 매우 사실적으로 구현했으나, 눈물을 닦는 동작이 없고 로봇에 읽을 수 있는 글자(B-200)가 포함된 점이 감점 요인입니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "눈물이 맺힌 얼굴과 뺨에 손을 올린 모습은 표현되었으나, 필수 유지 조건인 앰버의 머리 상처가 누락되었고 찰리와 배경의 트럭 등 핵심 상황적 디테일이 생략되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "앰버의 시선이 화면 우측에 있는 찰리(로봇)의 얼굴 쪽을 향해 위로 향해 있습니다.",
        "built_space": "배경의 해안가 도로에 레퍼런스와 동일한 형태와 방향의 흰색 트럭이 주차되어 있습니다.",
        "entities": "앰버의 외모와 이마의 상처가 잘 나타나 있으며, 화면 우측의 찰리는 손상된 가슴 외장과 B-200 컴포넌트를 부착하고 있습니다.",
        "hard_violations": [],
        "physics": "앰버와 로봇 모두 지면에 안정적으로 서 있는 자세를 취하고 있습니다."
       },
       {
        "label": "B",
        "direction": "앰버의 시선이 특정 대상 없이 화면 밖 하늘을 향하고 있습니다.",
        "built_space": "흐릿한 해안가 언덕 배경만 보이며, 트럭 등의 구조물은 묘사되지 않았습니다.",
        "entities": "앰버의 얼굴과 눈물이 묘사되었으나 명시된 이마의 상처가 누락되었으며, 찰리도 프레임에 존재하지 않습니다.",
        "hard_violations": [],
        "physics": "앰버의 손이 뺨에 얹혀 있으며 눈물이 뺨을 타고 자연스럽게 맺혀 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "눈물 어린 굳은 얼굴의 클로즈업, 화면 밖 위쪽을 향한 시선과 눈물을 닦는 손동작이 정확하지만, 드러난 이마에 지속되어야 할 부상 흔적이 없다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "찰리를 올려다보는 시선과 이마 상처는 맞지만, 금지된 추가 인물과 읽히는 표기가 등장하며 앰버의 얼굴 클로즈업을 두 인물의 넓은 구도로 바꿨다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버의 두 눈은 렌즈가 아니라 화면 왼쪽 위의 화면 밖 지점을 향한다. 찰리는 보이지 않아 대상의 실제 위치를 확인할 수 없지만, 자신보다 높은 곳의 찰리를 올려다보는 설정과 일치하는 시선이다. 손끝은 자신의 볼에 닿아 있다.",
        "built_space": "얼굴 뒤 아래쪽에는 흐릿한 흙과 돌, 풀과 작은 보라색 꽃이 있고, 왼쪽 위에는 도로와 가드레일 한 줄, 멀리 해안 능선과 바다가 보인다. 참조 장소의 지형과 재료가 맞으며, 얼굴이 크게 잡힌 구도에서 트럭이 보이지 않는 것은 결함이 아니다. 중복 시설이나 불가능한 반사는 없다.",
        "entities": "보이는 인물은 앰버 한 명이다. 약 10세 여자아이의 둥근 얼굴, 큰 눈, 금발과 참조의 얼굴 특징이 대체로 맞으며, 혼혈 배경 자체는 외형만으로 확정할 수 없다. 눈가와 볼에 실제 액체처럼 보이는 눈물이 있고 입매는 굳어 있다. 화면 아래에 남색 옷 일부가 보인다. 드러난 이마에는 부상 흔적이 없다. 읽히는 글자나 추가 소품은 없다.",
        "hard_violations": [],
        "physics": "머리는 목에 자연스럽게 연결되고, 볼에 댄 손은 화면 아래로 이어지는 손목과 팔로 지지된다. 손끝으로 볼의 눈물을 닦는 순간으로 가능한 자세다. 발과 몸의 지지면은 클로즈업 밖이므로 판단 대상이 아니며, 떠 있는 신체나 물체는 없다."
       },
       {
        "label": "B",
        "direction": "앰버는 화면 오른쪽 위에 보이는 찰리의 얼굴 쪽을 올려다본다. 시선의 목표가 실제로 프레임 안에 있어 관계가 명확하다. 찰리의 눈은 잘려 있어 그 시선은 확인할 수 없다.",
        "built_space": "왼쪽 뒤 도로에 낡은 픽업트럭 한 대와 가드레일 한 줄이 있고, 주변에는 돌과 흙, 풀과 보라색 꽃이 펼쳐진다. 뒤의 해안 능선과 바다도 참조 장소에 부합한다. 다만 배경이 비교적 선명하고 넓게 보이며, 오른쪽 전경의 찰리 상체가 큰 면적을 차지해 요청된 앰버 얼굴 중심 클로즈업과 다르다.",
        "entities": "앰버의 금발, 큰 눈, 둥근 얼굴과 남색 상의는 참조에 대체로 부합한다. 눈에 눈물이 맺혀 있고 이마에 상처가 있다. 오른쪽에는 찰리로 읽히는 별도 인물의 얼굴 일부와 손상된 기계식 몸체, 열린 가슴이 추가로 보인다. 몸체의 'B-200' 표기는 읽을 수 있어 무문자 조건에 어긋난다. 눈물을 닦는 손은 화면에 보이지 않는다.",
        "hard_violations": [
         "가시 인물을 앰버로 한정한 조건과 달리, 찰리의 얼굴 일부와 상체를 전경에 추가했다.",
         "오른쪽 몸체에 읽을 수 있는 'B-200' 표기가 있어 판독 가능한 글자 금지 조건을 위반한다."
        ],
        "physics": "앰버의 머리와 어깨, 화면 아래로 이어지는 몸통은 자연스럽게 연결되어 있다. 찰리의 열린 가슴과 외장도 몸통에 부착되어 있으며 떠 있는 부품은 보이지 않는다. 두 인물의 발과 지지면은 프레임 밖이므로 지면 접촉은 확인할 수 없지만, 보이는 부분에 불가능한 부유나 해부학적 문제는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "눈물 어린 굳은 얼굴의 클로즈업, 화면 밖 위쪽을 향한 시선과 눈물을 닦는 손동작이 정확하지만, 드러난 이마에 지속되어야 할 부상 흔적이 없다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "찰리를 올려다보는 시선과 이마 상처는 맞지만, 금지된 추가 인물과 읽히는 표기가 등장하며 앰버의 얼굴 클로즈업을 두 인물의 넓은 구도로 바꿨다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "앰버의 두 눈은 렌즈가 아니라 화면 왼쪽 위의 화면 밖 지점을 향한다. 찰리는 보이지 않아 대상의 실제 위치를 확인할 수 없지만, 자신보다 높은 곳의 찰리를 올려다보는 설정과 일치하는 시선이다. 손끝은 자신의 볼에 닿아 있다.",
        "built_space": "얼굴 뒤 아래쪽에는 흐릿한 흙과 돌, 풀과 작은 보라색 꽃이 있고, 왼쪽 위에는 도로와 가드레일 한 줄, 멀리 해안 능선과 바다가 보인다. 참조 장소의 지형과 재료가 맞으며, 얼굴이 크게 잡힌 구도에서 트럭이 보이지 않는 것은 결함이 아니다. 중복 시설이나 불가능한 반사는 없다.",
        "entities": "보이는 인물은 앰버 한 명이다. 약 10세 여자아이의 둥근 얼굴, 큰 눈, 금발과 참조의 얼굴 특징이 대체로 맞으며, 혼혈 배경 자체는 외형만으로 확정할 수 없다. 눈가와 볼에 실제 액체처럼 보이는 눈물이 있고 입매는 굳어 있다. 화면 아래에 남색 옷 일부가 보인다. 드러난 이마에는 부상 흔적이 없다. 읽히는 글자나 추가 소품은 없다.",
        "hard_violations": [],
        "physics": "머리는 목에 자연스럽게 연결되고, 볼에 댄 손은 화면 아래로 이어지는 손목과 팔로 지지된다. 손끝으로 볼의 눈물을 닦는 순간으로 가능한 자세다. 발과 몸의 지지면은 클로즈업 밖이므로 판단 대상이 아니며, 떠 있는 신체나 물체는 없다."
       },
       {
        "label": "A",
        "direction": "앰버는 화면 오른쪽 위에 보이는 찰리의 얼굴 쪽을 올려다본다. 시선의 목표가 실제로 프레임 안에 있어 관계가 명확하다. 찰리의 눈은 잘려 있어 그 시선은 확인할 수 없다.",
        "built_space": "왼쪽 뒤 도로에 낡은 픽업트럭 한 대와 가드레일 한 줄이 있고, 주변에는 돌과 흙, 풀과 보라색 꽃이 펼쳐진다. 뒤의 해안 능선과 바다도 참조 장소에 부합한다. 다만 배경이 비교적 선명하고 넓게 보이며, 오른쪽 전경의 찰리 상체가 큰 면적을 차지해 요청된 앰버 얼굴 중심 클로즈업과 다르다.",
        "entities": "앰버의 금발, 큰 눈, 둥근 얼굴과 남색 상의는 참조에 대체로 부합한다. 눈에 눈물이 맺혀 있고 이마에 상처가 있다. 오른쪽에는 찰리로 읽히는 별도 인물의 얼굴 일부와 손상된 기계식 몸체, 열린 가슴이 추가로 보인다. 몸체의 'B-200' 표기는 읽을 수 있어 무문자 조건에 어긋난다. 눈물을 닦는 손은 화면에 보이지 않는다.",
        "hard_violations": [
         "가시 인물을 앰버로 한정한 조건과 달리, 찰리의 얼굴 일부와 상체를 전경에 추가했다.",
         "오른쪽 몸체에 읽을 수 있는 'B-200' 표기가 있어 판독 가능한 글자 금지 조건을 위반한다."
        ],
        "physics": "앰버의 머리와 어깨, 화면 아래로 이어지는 몸통은 자연스럽게 연결되어 있다. 찰리의 열린 가슴과 외장도 몸통에 부착되어 있으며 떠 있는 부품은 보이지 않는다. 두 인물의 발과 지지면은 프레임 밖이므로 지면 접촉은 확인할 수 없지만, 보이는 부분에 불가능한 부유나 해부학적 문제는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.375,
    "B": 1.571
   },
   "adjusted": {
    "A": 1.125,
    "B": 1.571
   },
   "violations": {
    "A": [
     "[gpt-high] 가시 인물을 앰버로 한정한 조건과 달리, 찰리의 얼굴 일부와 상체를 전경에 추가했다.",
     "[gpt-high] 오른쪽 몸체에 읽을 수 있는 'B-200' 표기가 있어 판독 가능한 글자 금지 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1125,
   "B": 1571
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1125,
    "verdict_ko": "레퍼런스의 배경(트럭 포함), 앰버의 머리 상처, 찰리의 손상된 가슴 파츠를 매우 사실적으로 구현했으나, 눈물을 닦는 동작이 없고 로봇에 읽을 수 있는 글자(B-200)가 포함된 점이 감점 요인입니다.  ★위반: [gpt-high] 가시 인물을 앰버로 한정한 조건과 달리, 찰리의 얼굴 일부와 상체를 전경에 추가했다. / [gpt-high] 오른쪽 몸체에 읽을 수 있는 'B-200' 표기가 있어 판독 가능한 글자 금지 조건을 위반한다."
   },
   {
    "label": "B",
    "score": 1571,
    "verdict_ko": "눈물이 맺힌 얼굴과 뺨에 손을 올린 모습은 표현되었으나, 필수 유지 조건인 앰버의 머리 상처가 누락되었고 찰리와 배경의 트럭 등 핵심 상황적 디테일이 생략되었습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_coastal_search_hill_7092b9.png",
    "asset_id": "ab08fa93-ee50-4572-91f7-b41dbd747f21",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-cead-7bf8-ac3b-59c479c0f30f",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S71sh7__bgfirst_bg.png",
   "bg_asset_id": "0cca7227-2e69-4db8-ab19-8acea477fee5",
   "bg_record_key": "S71sh7::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "coastal_search_hill",
   "groupbg_asset_id": "ab08fa93-ee50-4572-91f7-b41dbd747f21"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S71sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:23:04.562468+00:00",
  "fingerprint": "43caaa1970c01fa5d67ce5a5057c78012b55ae8f01d8469d7580c5ab8c9c9121",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S71sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S71sh7_sel.png",
  "source_sha256": "c85c15a4defbe009b6c7f845f185b5df077794087d5f8257beed52201d7554a1",
  "file": "S71sh7_cine.png",
  "staged_sha256": "cc3d8a0c3a772e469a837c8b58db22953042b1f18b4c152f9348cea2697a921f",
  "latency_ms": 10409
 },
 "S71sh17::confined_fp_apt": {
  "applies": true,
  "reason_ko": "트럭 조종석(운전석 및 조수석) 내부에서 인물들이 어느 자리에 앉아 앞 유리창 밖을 바라보고 있는지 정확한 위치 배치가 중요하기 때문입니다.",
  "input_fingerprint": "824462859e254c68"
 },
 "S71sh17::signage": {
  "fp": "4210a21e1515e6aa",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "confinedfp::3f20774e773a": {
  "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/confinedfp_base_3f20774e773a.png",
  "place_text": "Inside the old truck's front cab, at the driver and passenger seats facing a vast violet field through the daylight windshield.",
  "input_fingerprint": "a815d3660a9511db"
 },
 "S71sh17::confined_fp": {
  "reads": {
   "controls": "The steering wheel is attached to the driver's station on the left.",
   "mirrors": "No mirrors or reflective surfaces are depicted in the diagram.",
   "camera": "The camera is positioned in the bottom center behind the seats, pointing straight forward toward the windshield.",
   "occupants": "현우 occupies the driver seat on the left, and 앰버 occupies the passenger seat on the right."
  },
  "mismatches": [],
  "scene_description_en": "The camera is positioned in the lower center of the truck cab, facing straight forward. In the near left foreground, 현우 occupies the driver seat, facing away from the camera toward the front. The steering wheel is located at his station on the left side of the dashboard. In the near right foreground, 앰버 occupies the passenger seat, also facing away from the camera. The center of the frame remains an unobstructed viewing gap between the two front seats. In the far center, the wide windshield spans the cab, revealing the outside violet field in the upper background. No mirrors or reflective surfaces are present in this view.",
  "fixed": false,
  "input_fingerprint": "d5debdc1a7c0e8d2"
 },
 "S71sh17": {
  "input_fingerprint": "227b5e0074dc353e",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 트럭 앞 유리창 너머로 끝없이 펼쳐진 보라색 제비꽃 밭을 멍하니 바라보는 현우와 앰버의 뒷모습.\n\nLOCATION (lock): Inside the old truck's front cab, at the driver and passenger seats facing a vast violet field through the daylight windshield. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Unobstructed windshield view of the violet field in the upper-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Truck windshield (The violet field is visible through it) — Seen from inside the cab, with the field beyond the glass rather than reflected on it; used as Creates an unobstructed central view between the occupants; Violet field (Extensively covered in blooming purple violets); used as Provides the shared distant object of attention and the visual payoff beyond the cab; Truck seats (Occupied by 현우 and 앰버) — Their rear portions border the central viewing gap; used as Establishes the interior depth and supports the two separated foreground figures.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight balances the cab and exterior sufficiently to retain both backs while allowing the purple violets to provide the scene's richer color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): An extensive field of purple violets stretches ahead, with Ji Dong-hyun's research institute visible in the distance. Charlie retains his damaged bodywork, chest opening and B-200's chest component. 현우: He is at the wheel after being startled awake from drowsy driving, with his earlier injuries still present. 앰버: She is riding in the truck after wiping away her tears, pointing forward. Her earlier head injury has no stated treatment.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera is positioned in the lower center of the truck cab, facing straight forward. In the near left foreground, 현우 occupies the driver seat, facing away from the camera toward the front. The steering wheel is located at his station on the left side of the dashboard. In the near right foreground, 앰버 occupies the passenger seat, also facing away from the camera. The center of the frame remains an unobstructed viewing gap between the two front seats. In the far center, the wide windshield spans the cab, revealing the outside violet field in the upper background. No mirrors or reflective surfaces are present in this view.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 트럭 앞 유리창 너머로 끝없이 펼쳐진 보라색 제비꽃 밭을 멍하니 바라보는 현우와 앰버의 뒷모습.\n\nLOCATION (lock): Inside the old truck's front cab, at the driver and passenger seats facing a vast violet field through the daylight windshield. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight balances the cab and exterior sufficiently to retain both backs while allowing the purple violets to provide the scene's richer color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): An extensive field of purple violets stretches ahead, with Ji Dong-hyun's research institute visible in the distance. Charlie retains his damaged bodywork, chest opening and B-200's chest component. 현우: He is at the wheel after being startled awake from drowsy driving, with his earlier injuries still present. 앰버: She is riding in the truck after wiping away her tears, pointing forward. Her earlier head injury has no stated treatment.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera is positioned in the lower center of the truck cab, facing straight forward. In the near left foreground, 현우 occupies the driver seat, facing away from the camera toward the front. The steering wheel is located at his station on the left side of the dashboard. In the near right foreground, 앰버 occupies the passenger seat, also facing away from the camera. The center of the frame remains an unobstructed viewing gap between the two front seats. In the far center, the wide windshield spans the cab, revealing the outside violet field in the upper background. No mirrors or reflective surfaces are present in this view.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dawn.\n\nSHOT TEXT (authoritative, Korean): 트럭 앞 유리창 너머로 끝없이 펼쳐진 보라색 제비꽃 밭을 멍하니 바라보는 현우와 앰버의 뒷모습.\n\nLOCATION (lock): Inside the old truck's front cab, at the driver and passenger seats facing a vast violet field through the daylight windshield. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight balances the cab and exterior sufficiently to retain both backs while allowing the purple violets to provide the scene's richer color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): An extensive field of purple violets stretches ahead, with Ji Dong-hyun's research institute visible in the distance. Charlie retains his damaged bodywork, chest opening and B-200's chest component. 현우: He is at the wheel after being startled awake from drowsy driving, with his earlier injuries still present. 앰버: She is riding in the truck after wiping away her tears, pointing forward. Her earlier head injury has no stated treatment.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S71sh17_confinedfp.png",
     "asset_id": null,
     "role": null
    },
    {
     "label": "현우",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "앰버",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S71sh17_confinedfp.png",
     "asset_id": null,
     "role": null
    },
    {
     "label": "현우",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "앰버",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우는 앞을 향하고 있으나, 앰버는 제비꽃 밭이 아닌 현우 쪽으로 고개를 돌려 시선을 향하고 있음.",
    "built_space": "트럭 내부, 운전석과 조수석 및 대시보드 배치가 구조에 맞게 구현됨.",
    "entities": "현우(검은 머리), 앰버(금발, 옆모습 노출), 보라색 제비꽃 밭이 모두 식별됨.",
    "hard_violations": [],
    "physics": "두 인물 모두 시트에 자연스럽게 앉아있으며, 현우의 손은 운전대를 잘 지지하고 있음."
   },
   {
    "label": "B",
    "direction": "현우와 앰버 모두 전방의 앞유리창 너머 제비꽃 밭으로 시선이 일치함.",
    "built_space": "트럭 캡 내부의 좌석과 창문 레이아웃이 기준 스케치 및 위치 설정과 일치함.",
    "entities": "현우와 앰버의 뒷모습, 끝없이 펼쳐진 제비꽃 밭, 멀리 보이는 연구소 건물이 모두 잘 나타남.",
    "hard_violations": [],
    "physics": "시트에 기댄 자세와 운전대를 잡은 손 등 무게 중심과 지지 상태가 자연스러움."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "앰버가 제비꽃 밭을 바라보는 뒷모습이 아니라 현우를 향한 옆모습으로 연출되어 샷 텍스트의 핵심 지시를 위반함."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "두 인물이 앞유리창 너머의 제비꽃 밭을 주시하는 뒷모습을 명시된 프레이밍과 위치에 맞게 정확히 구현함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 앞을 향하고 있으나, 앰버는 제비꽃 밭이 아닌 현우 쪽으로 고개를 돌려 시선을 향하고 있음.",
        "built_space": "트럭 내부, 운전석과 조수석 및 대시보드 배치가 구조에 맞게 구현됨.",
        "entities": "현우(검은 머리), 앰버(금발, 옆모습 노출), 보라색 제비꽃 밭이 모두 식별됨.",
        "hard_violations": [],
        "physics": "두 인물 모두 시트에 자연스럽게 앉아있으며, 현우의 손은 운전대를 잘 지지하고 있음."
       },
       {
        "label": "B",
        "direction": "현우와 앰버 모두 전방의 앞유리창 너머 제비꽃 밭으로 시선이 일치함.",
        "built_space": "트럭 캡 내부의 좌석과 창문 레이아웃이 기준 스케치 및 위치 설정과 일치함.",
        "entities": "현우와 앰버의 뒷모습, 끝없이 펼쳐진 제비꽃 밭, 멀리 보이는 연구소 건물이 모두 잘 나타남.",
        "hard_violations": [],
        "physics": "시트에 기댄 자세와 운전대를 잡은 손 등 무게 중심과 지지 상태가 자연스러움."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "앰버가 제비꽃 밭을 바라보는 뒷모습이 아니라 현우를 향한 옆모습으로 연출되어 샷 텍스트의 핵심 지시를 위반함."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "두 인물이 앞유리창 너머의 제비꽃 밭을 주시하는 뒷모습을 명시된 프레이밍과 위치에 맞게 정확히 구현함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우는 앞을 향하고 있으나, 앰버는 제비꽃 밭이 아닌 현우 쪽으로 고개를 돌려 시선을 향하고 있음.",
        "built_space": "트럭 내부, 운전석과 조수석 및 대시보드 배치가 구조에 맞게 구현됨.",
        "entities": "현우(검은 머리), 앰버(금발, 옆모습 노출), 보라색 제비꽃 밭이 모두 식별됨.",
        "hard_violations": [],
        "physics": "두 인물 모두 시트에 자연스럽게 앉아있으며, 현우의 손은 운전대를 잘 지지하고 있음."
       },
       {
        "label": "B",
        "direction": "현우와 앰버 모두 전방의 앞유리창 너머 제비꽃 밭으로 시선이 일치함.",
        "built_space": "트럭 캡 내부의 좌석과 창문 레이아웃이 기준 스케치 및 위치 설정과 일치함.",
        "entities": "현우와 앰버의 뒷모습, 끝없이 펼쳐진 제비꽃 밭, 멀리 보이는 연구소 건물이 모두 잘 나타남.",
        "hard_violations": [],
        "physics": "시트에 기댄 자세와 운전대를 잡은 손 등 무게 중심과 지지 상태가 자연스러움."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "두 사람의 뒷모습과 꽃밭을 향한 공동 주시, 좌석 사이로 열린 전면창의 와이드 구도가 충실하나 앰버의 전방 지목 동작은 확인되지 않는다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "운전실 배치와 와이드 구도는 맞지만 앰버가 꽃밭 대신 현우 쪽을 바라보며 옆얼굴을 드러내어 핵심 행동과 뒷모습 지시에서 벗어난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우와 앰버 모두 뒤통수가 보이고 머리가 전면창 밖 보라색 꽃밭을 향한다. 눈 자체는 보이지 않지만 두 사람의 공동 주시 방향은 요청과 맞는다. 현우의 오른손은 운전대를 잡고 있으며, 앰버가 전방을 가리키는 손은 보이지 않는다.",
        "built_space": "차량 내부 두 좌석 뒤 중앙에서 앞을 보는 구도다. 왼쪽 운전석과 오른쪽 조수석에 각각 한 명씩 앉아 있고, 좌석 등받이 두 개가 아래 양옆을 둘러싼다. 전면창 하나, 운전대 하나, 대시보드 하나, 선바이저 두 개, 전면 와이퍼 두 개와 양쪽 측면창이 보인다. 좌석 사이와 전면창 상단 중앙으로 꽃밭이 열려 있어 배치도와 일치한다. 꽃밭은 유리 반사가 아니라 차량 밖 풍경으로 보이며, 먼 건물도 작게 배치되어 있다.",
        "entities": "등장인물은 두 명뿐이다. 현우의 검은 헝클어진 머리, 젊은 남성 체격, 남색 상의가 참조와 부합한다. 앰버는 작은 여자아이 체격에 어깨까지 내려오는 금발과 남색 상의를 갖춰 참조에 가깝다. 얼굴이 가려져 정확한 나이와 한국계·혼혈 정체성, 얼굴 특징 및 기존 부상은 확인할 수 없다. 낡은 트럭과 광대한 보라색 꽃밭, 연구소로 읽힐 수 있는 먼 건물이 보인다. 꽃의 정확한 종은 식별하기 어렵다. 차가운 외광은 이른 아침에 어울리지만 새벽임이 뚜렷하지는 않다. 읽을 수 있는 글자나 불필요한 인물은 없다.",
        "hard_violations": [],
        "physics": "두 사람 모두 각자의 좌석에 앉아 있으며 좌석 쿠션과 등받이가 몸을 지지한다. 현우의 오른손은 운전대 테두리에 접촉하고 팔꿈치도 자연스럽게 굽혀져 있다. 앰버의 몸은 조수석과 정상적인 착석 관계를 이룬다. 하체 일부는 가려져 있지만 떠 있는 신체나 지지 없는 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "현우는 전면창과 꽃밭 쪽을 향하지만, 앰버는 고개를 왼쪽으로 돌려 현우 쪽을 바라본다. 따라서 두 사람이 함께 꽃밭을 멍하니 보는 순간이 아니다. 현우의 오른손은 운전대를 잡고 있고, 앰버의 전방 지목 동작은 확인되지 않는다.",
        "built_space": "카메라는 두 좌석 뒤 중앙에 있고 왼쪽 현우, 오른쪽 앰버의 배치는 배치도에 맞는다. 좌석 등받이 두 개, 운전대 하나, 대시보드 하나, 전면창 하나, 선바이저 두 개, 전면 와이퍼 두 개와 천장등 하나가 보인다. 두 사람 사이로 전면창 밖 꽃밭이 보이며 먼 건물도 배경 크기로 유지된다. 꽃밭은 반사가 아닌 창밖 풍경이다. A보다 인물과 등받이가 화면을 조금 더 크게 차지하지만 운전실 와이드 구도는 유지한다.",
        "entities": "현우와 앰버로 읽히는 두 사람만 있다. 현우의 검은 머리와 남색 상의, 젊은 남성 체격은 참조에 부합한다. 앰버의 금발과 어린 여자아이의 옆얼굴은 대체로 부합하지만, 보이는 상의는 참조의 남색보다 회색에 가깝다. 옆얼굴만으로 혼혈 정체성이나 정확한 얼굴 일치를 확정하기 어렵고 기존 머리 부상도 뚜렷하지 않다. 낡은 트럭, 넓은 보라색 꽃밭, 먼 연구소 형태의 건물이 보인다. 꽃의 정확한 종은 확정하기 어렵다. 안개 낀 따뜻한 외광은 새벽 분위기에 어울린다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 사람은 각각 정상 방향으로 좌석에 앉아 있고 쿠션과 등받이의 지지를 받는다. 앰버가 앉은 채 목을 왼쪽으로 돌리는 자세는 물리적으로 가능하다. 현우의 오른손이 운전대를 잡고 있어 사용 방향과 접촉도 자연스럽다. 지지 없이 떠 있는 몸이나 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "두 사람의 뒷모습과 꽃밭을 향한 공동 주시, 좌석 사이로 열린 전면창의 와이드 구도가 충실하나 앰버의 전방 지목 동작은 확인되지 않는다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "운전실 배치와 와이드 구도는 맞지만 앰버가 꽃밭 대신 현우 쪽을 바라보며 옆얼굴을 드러내어 핵심 행동과 뒷모습 지시에서 벗어난다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우와 앰버 모두 뒤통수가 보이고 머리가 전면창 밖 보라색 꽃밭을 향한다. 눈 자체는 보이지 않지만 두 사람의 공동 주시 방향은 요청과 맞는다. 현우의 오른손은 운전대를 잡고 있으며, 앰버가 전방을 가리키는 손은 보이지 않는다.",
        "built_space": "차량 내부 두 좌석 뒤 중앙에서 앞을 보는 구도다. 왼쪽 운전석과 오른쪽 조수석에 각각 한 명씩 앉아 있고, 좌석 등받이 두 개가 아래 양옆을 둘러싼다. 전면창 하나, 운전대 하나, 대시보드 하나, 선바이저 두 개, 전면 와이퍼 두 개와 양쪽 측면창이 보인다. 좌석 사이와 전면창 상단 중앙으로 꽃밭이 열려 있어 배치도와 일치한다. 꽃밭은 유리 반사가 아니라 차량 밖 풍경으로 보이며, 먼 건물도 작게 배치되어 있다.",
        "entities": "등장인물은 두 명뿐이다. 현우의 검은 헝클어진 머리, 젊은 남성 체격, 남색 상의가 참조와 부합한다. 앰버는 작은 여자아이 체격에 어깨까지 내려오는 금발과 남색 상의를 갖춰 참조에 가깝다. 얼굴이 가려져 정확한 나이와 한국계·혼혈 정체성, 얼굴 특징 및 기존 부상은 확인할 수 없다. 낡은 트럭과 광대한 보라색 꽃밭, 연구소로 읽힐 수 있는 먼 건물이 보인다. 꽃의 정확한 종은 식별하기 어렵다. 차가운 외광은 이른 아침에 어울리지만 새벽임이 뚜렷하지는 않다. 읽을 수 있는 글자나 불필요한 인물은 없다.",
        "hard_violations": [],
        "physics": "두 사람 모두 각자의 좌석에 앉아 있으며 좌석 쿠션과 등받이가 몸을 지지한다. 현우의 오른손은 운전대 테두리에 접촉하고 팔꿈치도 자연스럽게 굽혀져 있다. 앰버의 몸은 조수석과 정상적인 착석 관계를 이룬다. 하체 일부는 가려져 있지만 떠 있는 신체나 지지 없는 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "현우는 전면창과 꽃밭 쪽을 향하지만, 앰버는 고개를 왼쪽으로 돌려 현우 쪽을 바라본다. 따라서 두 사람이 함께 꽃밭을 멍하니 보는 순간이 아니다. 현우의 오른손은 운전대를 잡고 있고, 앰버의 전방 지목 동작은 확인되지 않는다.",
        "built_space": "카메라는 두 좌석 뒤 중앙에 있고 왼쪽 현우, 오른쪽 앰버의 배치는 배치도에 맞는다. 좌석 등받이 두 개, 운전대 하나, 대시보드 하나, 전면창 하나, 선바이저 두 개, 전면 와이퍼 두 개와 천장등 하나가 보인다. 두 사람 사이로 전면창 밖 꽃밭이 보이며 먼 건물도 배경 크기로 유지된다. 꽃밭은 반사가 아닌 창밖 풍경이다. A보다 인물과 등받이가 화면을 조금 더 크게 차지하지만 운전실 와이드 구도는 유지한다.",
        "entities": "현우와 앰버로 읽히는 두 사람만 있다. 현우의 검은 머리와 남색 상의, 젊은 남성 체격은 참조에 부합한다. 앰버의 금발과 어린 여자아이의 옆얼굴은 대체로 부합하지만, 보이는 상의는 참조의 남색보다 회색에 가깝다. 옆얼굴만으로 혼혈 정체성이나 정확한 얼굴 일치를 확정하기 어렵고 기존 머리 부상도 뚜렷하지 않다. 낡은 트럭, 넓은 보라색 꽃밭, 먼 연구소 형태의 건물이 보인다. 꽃의 정확한 종은 확정하기 어렵다. 안개 낀 따뜻한 외광은 새벽 분위기에 어울린다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 사람은 각각 정상 방향으로 좌석에 앉아 있고 쿠션과 등받이의 지지를 받는다. 앰버가 앉은 채 목을 왼쪽으로 돌리는 자세는 물리적으로 가능하다. 현우의 오른손이 운전대를 잡고 있어 사용 방향과 접촉도 자연스럽다. 지지 없이 떠 있는 몸이나 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.321,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.321,
    "B": 2.0
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "A": 1321,
   "B": 2000
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1321,
    "verdict_ko": "앰버가 제비꽃 밭을 바라보는 뒷모습이 아니라 현우를 향한 옆모습으로 연출되어 샷 텍스트의 핵심 지시를 위반함."
   },
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "두 인물이 앞유리창 너머의 제비꽃 밭을 주시하는 뒷모습을 명시된 프레이밍과 위치에 맞게 정확히 구현함."
   }
  ],
  "refs": [
   {
    "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S71sh17_confinedfp.png",
    "asset_id": null,
    "role": null
   },
   {
    "label": "현우",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "앰버",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-d38e-766c-ac5f-5ea361861cda",
  "confined_fp": {
   "base_key": "confinedfp::3f20774e773a",
   "apt_reason": "트럭 조종석(운전석 및 조수석) 내부에서 인물들이 어느 자리에 앉아 앞 유리창 밖을 바라보고 있는지 정확한 위치 배치가 중요하기 때문입니다.",
   "fixed": false,
   "mismatches": []
  },
  "ref_mode": "confined_fp: 도면+장면설명+엔티티",
  "share_plan": {
   "ref_plan": "background"
  },
  "lane_policy": "ab_select_ready"
 },
 "S71sh17::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:25:12.256598+00:00",
  "fingerprint": "576d9424473749445365fa54d12c0c6e250b9eecba88189d5a7a939494c6ce0b",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S71sh17_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S71sh17_sel.png",
  "source_sha256": "1973da4292cf32aad6147d685991e53d90249f426a71a8e7841810d6870af906",
  "file": "S71sh17_cine.png",
  "staged_sha256": "e1d49a9533b04bccf809d8b9a6bd9cabdb92dd25b28df4828935a0e460cac68f",
  "latency_ms": 9822
 },
 "S72sh38::signage": {
  "fp": "3342b1e6949800e8",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S72sh38": {
  "input_fingerprint": "280ca32f5165de8e",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 찰리의 가슴에서 뿜어진 푸른빛에 휩싸인 채 공중으로 번쩍 치켜 올려진 윤성찬의 자동차 공중 찰나.\n\nLOCATION (lock): Above the outdoor vehicle pursuit route beside the coastal research facility's violet fields, where the pursuing car is lifted into the air. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Entire airborne car with visible space beneath in the upper-left of the frame, midground; Truck carrying 찰리 farther ahead in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: 윤성찬's car (Airborne at the peak of its lift, before the crash) — Its left flank, rear quarter, and part of the underside are visible from the low camera; used as The complete vehicle and the clear space beneath it establish the physical reversal of the pursuit; 현우's truck (Ahead of the pursuing car with 찰리 aboard) — Seen obliquely from behind along the established pursuit direction; used as Provides the distant source of the action and preserves the vehicle-to-vehicle geography; Violet field (The pursuit is taking place among the violets); used as A lower strip of terrain supplies a stable reference for the car's elevation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The established sunset ambience is interrupted by the blue light emitted from 찰리's chest and enveloping the airborne car, with enough tonal separation to read its outline.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The pursuit continues through the violet-field area at dusk, with the pursuing car lifted into the air before its rollover impact. Charlie's chest opening exposes the interior, and his chest ring is active despite the accumulated body damage.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 찰리의 가슴에서 뿜어진 푸른빛에 휩싸인 채 공중으로 번쩍 치켜 올려진 윤성찬의 자동차 공중 찰나.\n\nLOCATION (lock): Above the outdoor vehicle pursuit route beside the coastal research facility's violet fields, where the pursuing car is lifted into the air. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Entire airborne car with visible space beneath in the upper-left of the frame, midground; Truck carrying 찰리 farther ahead in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: 윤성찬's car (Airborne at the peak of its lift, before the crash) — Its left flank, rear quarter, and part of the underside are visible from the low camera; used as The complete vehicle and the clear space beneath it establish the physical reversal of the pursuit; 현우's truck (Ahead of the pursuing car with 찰리 aboard) — Seen obliquely from behind along the established pursuit direction; used as Provides the distant source of the action and preserves the vehicle-to-vehicle geography; Violet field (The pursuit is taking place among the violets); used as A lower strip of terrain supplies a stable reference for the car's elevation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The established sunset ambience is interrupted by the blue light emitted from 찰리's chest and enveloping the airborne car, with enough tonal separation to read its outline.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The pursuit continues through the violet-field area at dusk, with the pursuing car lifted into the air before its rollover impact. Charlie's chest opening exposes the interior, and his chest ring is active despite the accumulated body damage.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 찰리의 가슴에서 뿜어진 푸른빛에 휩싸인 채 공중으로 번쩍 치켜 올려진 윤성찬의 자동차 공중 찰나.\n\nLOCATION (lock): Above the outdoor vehicle pursuit route beside the coastal research facility's violet fields, where the pursuing car is lifted into the air. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Entire airborne car with visible space beneath in the upper-left of the frame, midground; Truck carrying 찰리 farther ahead in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: 윤성찬's car (Airborne at the peak of its lift, before the crash) — Its left flank, rear quarter, and part of the underside are visible from the low camera; used as The complete vehicle and the clear space beneath it establish the physical reversal of the pursuit; 현우's truck (Ahead of the pursuing car with 찰리 aboard) — Seen obliquely from behind along the established pursuit direction; used as Provides the distant source of the action and preserves the vehicle-to-vehicle geography; Violet field (The pursuit is taking place among the violets); used as A lower strip of terrain supplies a stable reference for the car's elevation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The established sunset ambience is interrupted by the blue light emitted from 찰리's chest and enveloping the airborne car, with enough tonal separation to read its outline.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The pursuit continues through the violet-field area at dusk, with the pursuing car lifted into the air before its rollover impact. Charlie's chest opening exposes the interior, and his chest ring is active despite the accumulated body damage.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 찰리 right now, so 찰리's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 찰리: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름); 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "푸른빛의 광선이 찰리의 가슴이 아닌 오른손에서 뿜어져 나와 자동차를 향함.",
    "built_space": "레퍼런스와 동일한 도로. 자동차는 화면 좌측 상단, 트럭은 우측 중경에 위치함.",
    "entities": "공중에 띄워진 자동차 1대, 박스 형태의 트럭, 트럭 짐칸에 앉아있는 찰리.",
    "hard_violations": [],
    "physics": "자동차가 지지체 없이 공중에 떠 있으나 프롬프트의 지시된 상황에 부합함. 찰리는 트럭 짐칸에 앉아 지탱됨."
   },
   {
    "label": "B",
    "direction": "푸른빛의 광선이 찰리의 가슴 중앙에서 뿜어져 나와 자동차 전면부를 향함.",
    "built_space": "레퍼런스와 동일한 해질녘 도로와 우측 들판. 자동차는 화면 좌측 상단, 트럭은 우측 중경에 위치함.",
    "entities": "공중에 띄워진 자동차 1대, 소형 평판 트럭, 짐칸에 앉은 찰리(가슴 부위 링 활성화).",
    "hard_violations": [],
    "physics": "자동차가 지지체 없이 공중에 떠 있으나 프롬프트의 지시된 상황에 부합함. 찰리는 트럭 화물칸에 앉아 지탱됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "찰리의 가슴에서 빛이 뿜어지는 설정을 정확히 연출하고, 요구된 넓은 샷의 구도와 피사체 배치를 충실히 따름."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "푸른빛이 가슴이 아닌 손에서 발사되어 '가슴에서 뿜어진 푸른빛'이라는 핵심 묘사를 위반함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "푸른빛의 광선이 찰리의 가슴이 아닌 오른손에서 뿜어져 나와 자동차를 향함.",
        "built_space": "레퍼런스와 동일한 도로. 자동차는 화면 좌측 상단, 트럭은 우측 중경에 위치함.",
        "entities": "공중에 띄워진 자동차 1대, 박스 형태의 트럭, 트럭 짐칸에 앉아있는 찰리.",
        "hard_violations": [],
        "physics": "자동차가 지지체 없이 공중에 떠 있으나 프롬프트의 지시된 상황에 부합함. 찰리는 트럭 짐칸에 앉아 지탱됨."
       },
       {
        "label": "B",
        "direction": "푸른빛의 광선이 찰리의 가슴 중앙에서 뿜어져 나와 자동차 전면부를 향함.",
        "built_space": "레퍼런스와 동일한 해질녘 도로와 우측 들판. 자동차는 화면 좌측 상단, 트럭은 우측 중경에 위치함.",
        "entities": "공중에 띄워진 자동차 1대, 소형 평판 트럭, 짐칸에 앉은 찰리(가슴 부위 링 활성화).",
        "hard_violations": [],
        "physics": "자동차가 지지체 없이 공중에 떠 있으나 프롬프트의 지시된 상황에 부합함. 찰리는 트럭 화물칸에 앉아 지탱됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "찰리의 가슴에서 빛이 뿜어지는 설정을 정확히 연출하고, 요구된 넓은 샷의 구도와 피사체 배치를 충실히 따름."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "푸른빛이 가슴이 아닌 손에서 발사되어 '가슴에서 뿜어진 푸른빛'이라는 핵심 묘사를 위반함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "푸른빛의 광선이 찰리의 가슴이 아닌 오른손에서 뿜어져 나와 자동차를 향함.",
        "built_space": "레퍼런스와 동일한 도로. 자동차는 화면 좌측 상단, 트럭은 우측 중경에 위치함.",
        "entities": "공중에 띄워진 자동차 1대, 박스 형태의 트럭, 트럭 짐칸에 앉아있는 찰리.",
        "hard_violations": [],
        "physics": "자동차가 지지체 없이 공중에 떠 있으나 프롬프트의 지시된 상황에 부합함. 찰리는 트럭 짐칸에 앉아 지탱됨."
       },
       {
        "label": "B",
        "direction": "푸른빛의 광선이 찰리의 가슴 중앙에서 뿜어져 나와 자동차 전면부를 향함.",
        "built_space": "레퍼런스와 동일한 해질녘 도로와 우측 들판. 자동차는 화면 좌측 상단, 트럭은 우측 중경에 위치함.",
        "entities": "공중에 띄워진 자동차 1대, 소형 평판 트럭, 짐칸에 앉은 찰리(가슴 부위 링 활성화).",
        "hard_violations": [],
        "physics": "자동차가 지지체 없이 공중에 떠 있으나 프롬프트의 지시된 상황에 부합함. 찰리는 트럭 화물칸에 앉아 지탱됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "좌상단 공중 차량과 우측 후방 트럭의 와이드 배치, 가슴에서 차량으로 이어지는 청색광이 정확하지만 차량은 요청한 왼쪽 대신 오른쪽 측면을 드러낸다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "차량이 중경보다 전경에서 지나치게 크게 보이고, 청색광의 출발점도 가슴보다 뻗은 손으로 읽혀 핵심 구도와 작용 관계가 약하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 차량 모두 화면 오른쪽 안쪽으로 굽는 도로를 따라 전진하는 방향이며 트럭이 앞선다. 찰리는 트럭 뒤쪽에서 추격 차량을 향해 몸과 얼굴을 돌리고 있다. 청색 광선은 찰리의 가슴 고리에서 승용차 앞부분으로 정확히 연결된다. 차량의 뒤와 오른쪽 측면, 하부가 보여 요청된 왼쪽 측면과는 다르다.",
        "built_space": "굽은 포장도로 하나, 오른쪽의 열린 콘크리트 배수로 하나, 도로 양옆 꽃밭, 우측 전주열과 굽이의 가드레일이 참조 장소와 대응한다. 승용차 전체는 좌상단에 있고 아래로 도로와 빈 공간이 보인다. 소형 적재함 트럭 한 대는 우측의 더 먼 곳에 있으며, 찰리는 적재함 안에 자리한다. 배경을 과도하게 확대하지 않았다.",
        "entities": "손상된 검은 승용차 한 대, 소형 트럭 한 대, 찰리로 읽히는 성인 남성형 인물 한 명이 보인다. 찰리의 손상된 몸통과 청색 가슴 고리는 식별되지만 얼굴의 연령·민족적 특징 및 가슴 내부의 세부는 거리상 확정하기 어렵다. 윤성찬과 현우는 창 안에서 식별되지 않으며 다른 인물은 추가되지 않았다. 노을과 보라·분홍 꽃밭은 유지된다. 차체의 작은 배지는 있으나 글자는 명확히 판독되지 않는다.",
        "hard_violations": [],
        "physics": "승용차의 바퀴는 모두 도로에서 떨어져 있고 하부와 지면 사이가 분명하다. 가슴에서 차량까지 연결된 청색광이 요청된 들어 올림의 작용원으로 제시되므로 원인 없는 부유는 아니다. 차량은 아직 충돌하지 않은 상태다. 트럭은 바퀴로 도로에 지지되고 찰리의 하체는 적재함 안에 있어 적재함의 지지를 받는 배치다. 다만 청색광은 차 전체를 감싸기보다 앞부분에 집중된다."
       },
       {
        "label": "B",
        "direction": "트럭은 굽은 도로의 오른쪽 안쪽으로 앞서가고 승용차도 같은 방향을 향한다. 찰리는 뒤의 승용차를 바라보며 한 팔을 뻗는다. 청색광은 승용차에 닿지만 트럭 쪽 끝이 가슴 고리보다 뻗은 손에 연결되어 보인다. 승용차는 뒤와 오른쪽 측면, 넓은 하부를 드러내므로 요청된 왼쪽 측면은 아니다.",
        "built_space": "포장도로 하나와 우측 콘크리트 배수로 하나, 양옆 꽃밭, 전주열과 먼 가드레일이 참조 장소를 유지한다. 찰리는 뒤가 열린 상자형 트럭의 적재함 문턱에 앉아 있다. 승용차 전체와 그 아래 빈 공간은 확보되지만 차가 화면 왼쪽 대부분을 차지하며 중앙 높이까지 내려와, 지정된 좌상단 중경보다 큰 전경 피사체로 읽힌다.",
        "entities": "낡고 손상된 회색 승용차 한 대, 상자형 화물 트럭 한 대, 성인 남성 한 명이 보인다. 남성은 찰리 역할로 읽히며 손상된 상의와 몸통 중앙의 청색 고리가 있지만 노출된 가슴 내부는 분명하지 않다. 얼굴의 정확한 연령과 민족적 특징은 확정하기 어렵다. 윤성찬과 현우는 식별되지 않고 불필요한 추가 인물도 없다. 일몰과 꽃밭은 유지되며 판독 가능한 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "승용차는 네 바퀴가 지면에서 떨어져 있고 아래에 그림자와 빈 공간이 있다. 차량을 감싸며 트럭 쪽으로 연결되는 청색광이 들어 올림의 원인으로 제시되어 단순한 무근거 부유는 아니다. 다만 그 힘이 가슴에서 나온다는 연결은 약하다. 찰리의 엉덩이는 적재함 문턱에 지지되고 다리는 아래로 늘어져 있으며, 트럭은 바퀴로 도로에 지지된다. 충돌 전 공중 순간으로는 읽힌다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "좌상단 공중 차량과 우측 후방 트럭의 와이드 배치, 가슴에서 차량으로 이어지는 청색광이 정확하지만 차량은 요청한 왼쪽 대신 오른쪽 측면을 드러낸다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "차량이 중경보다 전경에서 지나치게 크게 보이고, 청색광의 출발점도 가슴보다 뻗은 손으로 읽혀 핵심 구도와 작용 관계가 약하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "두 차량 모두 화면 오른쪽 안쪽으로 굽는 도로를 따라 전진하는 방향이며 트럭이 앞선다. 찰리는 트럭 뒤쪽에서 추격 차량을 향해 몸과 얼굴을 돌리고 있다. 청색 광선은 찰리의 가슴 고리에서 승용차 앞부분으로 정확히 연결된다. 차량의 뒤와 오른쪽 측면, 하부가 보여 요청된 왼쪽 측면과는 다르다.",
        "built_space": "굽은 포장도로 하나, 오른쪽의 열린 콘크리트 배수로 하나, 도로 양옆 꽃밭, 우측 전주열과 굽이의 가드레일이 참조 장소와 대응한다. 승용차 전체는 좌상단에 있고 아래로 도로와 빈 공간이 보인다. 소형 적재함 트럭 한 대는 우측의 더 먼 곳에 있으며, 찰리는 적재함 안에 자리한다. 배경을 과도하게 확대하지 않았다.",
        "entities": "손상된 검은 승용차 한 대, 소형 트럭 한 대, 찰리로 읽히는 성인 남성형 인물 한 명이 보인다. 찰리의 손상된 몸통과 청색 가슴 고리는 식별되지만 얼굴의 연령·민족적 특징 및 가슴 내부의 세부는 거리상 확정하기 어렵다. 윤성찬과 현우는 창 안에서 식별되지 않으며 다른 인물은 추가되지 않았다. 노을과 보라·분홍 꽃밭은 유지된다. 차체의 작은 배지는 있으나 글자는 명확히 판독되지 않는다.",
        "hard_violations": [],
        "physics": "승용차의 바퀴는 모두 도로에서 떨어져 있고 하부와 지면 사이가 분명하다. 가슴에서 차량까지 연결된 청색광이 요청된 들어 올림의 작용원으로 제시되므로 원인 없는 부유는 아니다. 차량은 아직 충돌하지 않은 상태다. 트럭은 바퀴로 도로에 지지되고 찰리의 하체는 적재함 안에 있어 적재함의 지지를 받는 배치다. 다만 청색광은 차 전체를 감싸기보다 앞부분에 집중된다."
       },
       {
        "label": "A",
        "direction": "트럭은 굽은 도로의 오른쪽 안쪽으로 앞서가고 승용차도 같은 방향을 향한다. 찰리는 뒤의 승용차를 바라보며 한 팔을 뻗는다. 청색광은 승용차에 닿지만 트럭 쪽 끝이 가슴 고리보다 뻗은 손에 연결되어 보인다. 승용차는 뒤와 오른쪽 측면, 넓은 하부를 드러내므로 요청된 왼쪽 측면은 아니다.",
        "built_space": "포장도로 하나와 우측 콘크리트 배수로 하나, 양옆 꽃밭, 전주열과 먼 가드레일이 참조 장소를 유지한다. 찰리는 뒤가 열린 상자형 트럭의 적재함 문턱에 앉아 있다. 승용차 전체와 그 아래 빈 공간은 확보되지만 차가 화면 왼쪽 대부분을 차지하며 중앙 높이까지 내려와, 지정된 좌상단 중경보다 큰 전경 피사체로 읽힌다.",
        "entities": "낡고 손상된 회색 승용차 한 대, 상자형 화물 트럭 한 대, 성인 남성 한 명이 보인다. 남성은 찰리 역할로 읽히며 손상된 상의와 몸통 중앙의 청색 고리가 있지만 노출된 가슴 내부는 분명하지 않다. 얼굴의 정확한 연령과 민족적 특징은 확정하기 어렵다. 윤성찬과 현우는 식별되지 않고 불필요한 추가 인물도 없다. 일몰과 꽃밭은 유지되며 판독 가능한 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "승용차는 네 바퀴가 지면에서 떨어져 있고 아래에 그림자와 빈 공간이 있다. 차량을 감싸며 트럭 쪽으로 연결되는 청색광이 들어 올림의 원인으로 제시되어 단순한 무근거 부유는 아니다. 다만 그 힘이 가슴에서 나온다는 연결은 약하다. 찰리의 엉덩이는 적재함 문턱에 지지되고 다리는 아래로 늘어져 있으며, 트럭은 바퀴로 도로에 지지된다. 충돌 전 공중 순간으로는 읽힌다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.196,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.196,
    "B": 2.0
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1196
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "찰리의 가슴에서 빛이 뿜어지는 설정을 정확히 연출하고, 요구된 넓은 샷의 구도와 피사체 배치를 충실히 따름."
   },
   {
    "label": "A",
    "score": 1196,
    "verdict_ko": "푸른빛이 가슴이 아닌 손에서 발사되어 '가슴에서 뿜어진 푸른빛'이라는 핵심 묘사를 위반함."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L248B01.png",
    "asset_id": "53477491-7f2b-45f1-b840-dd1a1a2068e0",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-d6e7-79fe-9a7d-092628ca6f93",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S72sh38::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:26:07.329029+00:00",
  "fingerprint": "73a5bdbdda2b6c847ba545cab0825b19cb8d0804d5f828843011a321b67bc164",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S72sh38_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S72sh38_sel.png",
  "source_sha256": "991565214fa5ad6197b6af5060bd7bbba568028ed58cae5c8770aefbc9ba4df7",
  "file": "S72sh38_cine.png",
  "staged_sha256": "ae1dfa8b9e4bf8e2e34041e0c8cf154246a159313400e7b5250bdcbe4e6a4b32",
  "latency_ms": 11542
 },
 "S72sh43::signage": {
  "fp": "3e6305d7e277d74c",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::d2a0a165e0f87beb": {
  "subjects": [],
  "subject_text": "해남 지동현의 연구소 1층\n낡은 철문 안으로 펼쳐지는 비어 있는 연구소 1층. 먼지가 쌓인 바닥과 위층으로 이어지는 나선형 계단이 눈에 띈다.",
  "identity": "canonical",
  "scope_id": "L248",
  "scope_role": "location_interior",
  "scope_sha": "7488a540c22abf73"
 },
 "S72sh43::bgfirst_bg": {
  "input_fingerprint": "4cf06b55bf705dc2",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 피 묻은 손으로 자신의 옆구리를 강하게 움켜쥔 채 고통스럽게 눈을 감은 앰버의 웅크린 상체.\n\nLOCATION (lock): In the old truck's passenger area during the nighttime escape, with the injured child sheltered within the vehicle.\n\nTIME OF DAY (lock): sunset.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Truck interior (Enclosing 앰버 while the vehicle is moving) — Only an oblique strip of the interior remains behind her upper body; used as Maintains the confined setting while leaving the injury and protective posture unobstructed.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Low nighttime ambient illumination keeps her closed eyes, bloodied hand, and injured side readable without adding a specific light source or importing the later flashlight beams.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 피 묻은 손으로 자신의 옆구리를 강하게 움켜쥔 채 고통스럽게 눈을 감은 앰버의 웅크린 상체.\n\nLOCATION (lock): In the old truck's passenger area during the nighttime escape, with the injured child sheltered within the vehicle.\n\nTIME OF DAY (lock): sunset.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Truck interior (Enclosing 앰버 while the vehicle is moving) — Only an oblique strip of the interior remains behind her upper body; used as Maintains the confined setting while leaving the injury and protective posture unobstructed.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Low nighttime ambient illumination keeps her closed eyes, bloodied hand, and injured side readable without adding a specific light source or importing the later flashlight beams.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S72sh43__bgfirst_bg.png",
  "asset_id": "d805328b-fd8e-4d2d-b2d9-3aebab763862",
  "input_asset_ids": [
   "5cff0d94-8701-405a-b9fc-ab23b5fb7cba",
   "20b9995f-79f3-4c42-9a56-159746fe0f8d"
  ]
 },
 "S72sh43": {
  "input_fingerprint": "a97ba76f2ef7ac24",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 피 묻은 손으로 자신의 옆구리를 강하게 움켜쥔 채 고통스럽게 눈을 감은 앰버의 웅크린 상체.\n\nLOCATION (lock): In the old truck's passenger area during the nighttime escape, with the injured child sheltered within the vehicle. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Truck interior (Enclosing 앰버 while the vehicle is moving) — Only an oblique strip of the interior remains behind her upper body; used as Maintains the confined setting while leaving the injury and protective posture unobstructed.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Low nighttime ambient illumination keeps her closed eyes, bloodied hand, and injured side readable without adding a specific light source or importing the later flashlight beams.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The escape truck is speeding along the road at night. Charlie's damaged chest remains open, and his diagnostic view displays a 'CAUTION' warning when checking the injury. 앰버: She has a severe, actively bleeding wound in her side and is groaning in pain. The earlier injury to the back of her head also remains unresolved.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 피 묻은 손으로 자신의 옆구리를 강하게 움켜쥔 채 고통스럽게 눈을 감은 앰버의 웅크린 상체.\n\nLOCATION (lock): In the old truck's passenger area during the nighttime escape, with the injured child sheltered within the vehicle. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Truck interior (Enclosing 앰버 while the vehicle is moving) — Only an oblique strip of the interior remains behind her upper body; used as Maintains the confined setting while leaving the injury and protective posture unobstructed.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Low nighttime ambient illumination keeps her closed eyes, bloodied hand, and injured side readable without adding a specific light source or importing the later flashlight beams.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The escape truck is speeding along the road at night. Charlie's damaged chest remains open, and his diagnostic view displays a 'CAUTION' warning when checking the injury. 앰버: She has a severe, actively bleeding wound in her side and is groaning in pain. The earlier injury to the back of her head also remains unresolved.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 피 묻은 손으로 자신의 옆구리를 강하게 움켜쥔 채 고통스럽게 눈을 감은 앰버의 웅크린 상체.\n\nLOCATION (lock): In the old truck's passenger area during the nighttime escape, with the injured child sheltered within the vehicle. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Truck interior (Enclosing 앰버 while the vehicle is moving) — Only an oblique strip of the interior remains behind her upper body; used as Maintains the confined setting while leaving the injury and protective posture unobstructed.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Low nighttime ambient illumination keeps her closed eyes, bloodied hand, and injured side readable without adding a specific light source or importing the later flashlight beams.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The escape truck is speeding along the road at night. Charlie's damaged chest remains open, and his diagnostic view displays a 'CAUTION' warning when checking the injury. 앰버: She has a severe, actively bleeding wound in her side and is groaning in pain. The earlier injury to the back of her head also remains unresolved.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S72sh43__bgfirst_bg.png",
     "asset_id": "d805328b-fd8e-4d2d-b2d9-3aebab763862",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S72sh43.png",
     "asset_id": "5cff0d94-8701-405a-b9fc-ab23b5fb7cba",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L248B03.png",
     "asset_id": "20b9995f-79f3-4c42-9a56-159746fe0f8d",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 2,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "두 눈을 감고 고통스러운 표정으로 상체를 웅크린 채, 피 묻은 양손으로 복부 및 옆구리를 움켜쥐고 있음.",
    "built_space": "레퍼런스와 일치하는 트럭 내부의 회색 직물 시트에 위치하며, 배경으로 창문과 벽면이 비스듬하게 배치됨.",
    "entities": "금발의 10세 혼혈 여아(앰버)가 등장하며 손과 옷에 혈흔이 있으나 머리 쪽의 부상 흔적은 보이지 않음.",
    "hard_violations": [],
    "physics": "시트에 안정적으로 앉아 체중을 지탱하고 있으며, 웅크린 상체와 상처를 감싼 손의 위치가 물리적으로 타당함."
   },
   {
    "label": "B",
    "direction": "고통으로 눈을 질끈 감고 상체를 깊게 웅크린 상태로, 피 묻은 두 손을 이용해 옆구리를 강하게 압박하고 있음.",
    "built_space": "레퍼런스의 트럭 내부 시트에 앉아 있으며, 창밖의 가로등 불빛에 적용된 모션 블러를 통해 차량이 고속 주행 중임이 잘 나타남.",
    "entities": "금발의 10세 여아(앰버)와 일치하며, 손과 상의의 혈흔뿐만 아니라 프롬프트가 지시한 머리 쪽의 출혈(이전 부상)도 묘사됨.",
    "hard_violations": [],
    "physics": "시트에 엉덩이를 대고 앉아 있으며, 웅크린 상체와 상처를 강하게 쥐고 있는 손의 압박이 중력과 동작에 맞게 지지됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "옆구리를 움켜쥔 웅크린 자세뿐만 아니라, 창밖의 속도감과 명시된 머리 부상 등 지속 상태(Carried State)를 완벽하게 반영하여 프롬프트 지시를 가장 충실히 이행함."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "샷 텍스트가 요구한 고통스러운 표정과 자세는 잘 구현했으나, 달리는 트럭의 속도감과 앰버의 머리 부상에 대한 묘사가 누락됨."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 눈을 감고 고통스러운 표정으로 상체를 웅크린 채, 피 묻은 양손으로 복부 및 옆구리를 움켜쥐고 있음.",
        "built_space": "레퍼런스와 일치하는 트럭 내부의 회색 직물 시트에 위치하며, 배경으로 창문과 벽면이 비스듬하게 배치됨.",
        "entities": "금발의 10세 혼혈 여아(앰버)가 등장하며 손과 옷에 혈흔이 있으나 머리 쪽의 부상 흔적은 보이지 않음.",
        "hard_violations": [],
        "physics": "시트에 안정적으로 앉아 체중을 지탱하고 있으며, 웅크린 상체와 상처를 감싼 손의 위치가 물리적으로 타당함."
       },
       {
        "label": "B",
        "direction": "고통으로 눈을 질끈 감고 상체를 깊게 웅크린 상태로, 피 묻은 두 손을 이용해 옆구리를 강하게 압박하고 있음.",
        "built_space": "레퍼런스의 트럭 내부 시트에 앉아 있으며, 창밖의 가로등 불빛에 적용된 모션 블러를 통해 차량이 고속 주행 중임이 잘 나타남.",
        "entities": "금발의 10세 여아(앰버)와 일치하며, 손과 상의의 혈흔뿐만 아니라 프롬프트가 지시한 머리 쪽의 출혈(이전 부상)도 묘사됨.",
        "hard_violations": [],
        "physics": "시트에 엉덩이를 대고 앉아 있으며, 웅크린 상체와 상처를 강하게 쥐고 있는 손의 압박이 중력과 동작에 맞게 지지됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "옆구리를 움켜쥔 웅크린 자세뿐만 아니라, 창밖의 속도감과 명시된 머리 부상 등 지속 상태(Carried State)를 완벽하게 반영하여 프롬프트 지시를 가장 충실히 이행함."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "샷 텍스트가 요구한 고통스러운 표정과 자세는 잘 구현했으나, 달리는 트럭의 속도감과 앰버의 머리 부상에 대한 묘사가 누락됨."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "두 눈을 감고 고통스러운 표정으로 상체를 웅크린 채, 피 묻은 양손으로 복부 및 옆구리를 움켜쥐고 있음.",
        "built_space": "레퍼런스와 일치하는 트럭 내부의 회색 직물 시트에 위치하며, 배경으로 창문과 벽면이 비스듬하게 배치됨.",
        "entities": "금발의 10세 혼혈 여아(앰버)가 등장하며 손과 옷에 혈흔이 있으나 머리 쪽의 부상 흔적은 보이지 않음.",
        "hard_violations": [],
        "physics": "시트에 안정적으로 앉아 체중을 지탱하고 있으며, 웅크린 상체와 상처를 감싼 손의 위치가 물리적으로 타당함."
       },
       {
        "label": "B",
        "direction": "고통으로 눈을 질끈 감고 상체를 깊게 웅크린 상태로, 피 묻은 두 손을 이용해 옆구리를 강하게 압박하고 있음.",
        "built_space": "레퍼런스의 트럭 내부 시트에 앉아 있으며, 창밖의 가로등 불빛에 적용된 모션 블러를 통해 차량이 고속 주행 중임이 잘 나타남.",
        "entities": "금발의 10세 여아(앰버)와 일치하며, 손과 상의의 혈흔뿐만 아니라 프롬프트가 지시한 머리 쪽의 출혈(이전 부상)도 묘사됨.",
        "hard_violations": [],
        "physics": "시트에 엉덩이를 대고 앉아 있으며, 웅크린 상체와 상처를 강하게 쥐고 있는 손의 압박이 중력과 동작에 맞게 지지됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "배경 노출은 요구보다 넓지만, 눈을 질끈 감고 상체를 웅크린 채 피 묻은 양손으로 옆구리 쪽을 강하게 감싸는 핵심 행동이 더 정확하다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "더 밀착된 상체 구도와 일몰은 좋지만, 손이 옆구리보다 복부 정면을 누르고 다른 팔은 아래로 내려가 있어 핵심 움켜쥐기 동작이 A보다 약하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "눈은 단단히 감겨 있고 얼굴은 아래로 숙여져 있어 외부 대상을 바라보지 않는다. 양손이 자신의 오른쪽 옆구리와 앞배가 만나는 부위로 모여 출혈 부위를 감싼다. 손과 팔의 방향이 자신의 부상을 압박하는 행동으로 연결된다.",
        "built_space": "연속된 회색 천 벤치 한 개, 오른쪽 측면 창문 한 개, 위쪽에 일부 잘린 작은 개구부와 낡은 금속 벽이 보인다. 참고 장소의 재질과 배치에 부합하며, 아이는 벤치 좌판에 앉아 등받이를 뒤에 두고 앞으로 굽힌다. 다만 상체 뒤에 좁고 비스듬한 실내 띠만 남기라는 요구보다 등받이와 창문이 넓게 드러난다.",
        "entities": "금발의 약 10세 여자아이 한 명만 등장한다. 둥근 얼굴, 밝은 피부, 체격과 남색 반팔 상의가 인물 참고와 대체로 맞으며, 혼혈 배경 자체는 외관만으로 확정할 수 없다. 감긴 눈 때문에 큰 눈의 형태는 확인할 수 없다. 손과 상의에 젖은 피가 있고 머리카락 윗부분에도 혈흔이 있으나 후두부 상처 자체는 보이지 않는다. 창밖의 어두운 풍경과 수평으로 번진 불빛은 야간 주행을 나타내지만 일몰은 뚜렷하지 않다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "골반과 허벅지가 좌판에 받쳐져 있으며, 앉은 상태에서 허리를 굽히고 팔꿈치를 접어 양손으로 몸통을 압박한다. 손은 상의 및 다른 손과 접촉하고 있어 떠 있는 물체나 지지 없는 신체는 없다. 압박 부근의 천 주름과 피의 광택도 물리적인 접촉으로 읽힌다."
       },
       {
        "label": "B",
        "direction": "눈을 감고 얼굴을 아래로 숙이고 있다. 오른손은 몸을 가로질러 복부 정면의 혈흔을 누르지만 옆구리 바깥쪽을 감싸는 모습은 약하다. 왼팔은 아래쪽 허벅지 사이로 내려가며 손끝은 프레임 밖이다. 따라서 손의 목표는 자신의 부상이지만 압박 위치는 옆구리보다 앞배로 읽힌다.",
        "built_space": "회색 천 벤치 한 개, 오른쪽 창문 한 개, 위쪽의 부분적인 작은 개구부와 낡은 금속 벽이 보인다. 참고 장소의 고정 구조와 재질을 유지하고 아이도 좌판에 정상적으로 앉아 있다. A보다 상체가 크게 잡혔지만 넓은 등받이와 창문이 여전히 보여, 배경을 좁은 띠로 제한하라는 요구에는 덜 미친다.",
        "entities": "금발의 어린 여자아이 한 명이며 얼굴 윤곽, 체격, 남색 반팔 상의는 참고 인물과 대체로 일치한다. 정확한 혼혈 배경과 감긴 눈의 크기는 영상만으로 검증할 수 없다. 손, 팔, 상의와 바지에 혈흔이 있고 복부의 피는 젖은 광택을 보인다. 후두부는 보이지 않아 이전 상처의 지속 여부는 판단할 수 없다. 창밖의 주황빛 수평선은 일몰에 부합하지만 차량의 속도는 뚜렷하게 드러나지 않는다. 추가 인물이나 판독 가능한 글자는 없다.",
        "hard_violations": [],
        "physics": "엉덩이와 허벅지는 벤치 좌판에 지지되고 상체는 앞으로 굽어 있다. 오른손은 복부에 닿아 압박하며 왼팔은 어깨에서 자연스럽게 아래로 내려온다. 프레임 밖 손의 접촉점은 확인되지 않지만 팔 자체가 부유하는 것은 아니다. 불가능한 관절이나 지지 없는 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "배경 노출은 요구보다 넓지만, 눈을 질끈 감고 상체를 웅크린 채 피 묻은 양손으로 옆구리 쪽을 강하게 감싸는 핵심 행동이 더 정확하다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "더 밀착된 상체 구도와 일몰은 좋지만, 손이 옆구리보다 복부 정면을 누르고 다른 팔은 아래로 내려가 있어 핵심 움켜쥐기 동작이 A보다 약하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "눈은 단단히 감겨 있고 얼굴은 아래로 숙여져 있어 외부 대상을 바라보지 않는다. 양손이 자신의 오른쪽 옆구리와 앞배가 만나는 부위로 모여 출혈 부위를 감싼다. 손과 팔의 방향이 자신의 부상을 압박하는 행동으로 연결된다.",
        "built_space": "연속된 회색 천 벤치 한 개, 오른쪽 측면 창문 한 개, 위쪽에 일부 잘린 작은 개구부와 낡은 금속 벽이 보인다. 참고 장소의 재질과 배치에 부합하며, 아이는 벤치 좌판에 앉아 등받이를 뒤에 두고 앞으로 굽힌다. 다만 상체 뒤에 좁고 비스듬한 실내 띠만 남기라는 요구보다 등받이와 창문이 넓게 드러난다.",
        "entities": "금발의 약 10세 여자아이 한 명만 등장한다. 둥근 얼굴, 밝은 피부, 체격과 남색 반팔 상의가 인물 참고와 대체로 맞으며, 혼혈 배경 자체는 외관만으로 확정할 수 없다. 감긴 눈 때문에 큰 눈의 형태는 확인할 수 없다. 손과 상의에 젖은 피가 있고 머리카락 윗부분에도 혈흔이 있으나 후두부 상처 자체는 보이지 않는다. 창밖의 어두운 풍경과 수평으로 번진 불빛은 야간 주행을 나타내지만 일몰은 뚜렷하지 않다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "골반과 허벅지가 좌판에 받쳐져 있으며, 앉은 상태에서 허리를 굽히고 팔꿈치를 접어 양손으로 몸통을 압박한다. 손은 상의 및 다른 손과 접촉하고 있어 떠 있는 물체나 지지 없는 신체는 없다. 압박 부근의 천 주름과 피의 광택도 물리적인 접촉으로 읽힌다."
       },
       {
        "label": "A",
        "direction": "눈을 감고 얼굴을 아래로 숙이고 있다. 오른손은 몸을 가로질러 복부 정면의 혈흔을 누르지만 옆구리 바깥쪽을 감싸는 모습은 약하다. 왼팔은 아래쪽 허벅지 사이로 내려가며 손끝은 프레임 밖이다. 따라서 손의 목표는 자신의 부상이지만 압박 위치는 옆구리보다 앞배로 읽힌다.",
        "built_space": "회색 천 벤치 한 개, 오른쪽 창문 한 개, 위쪽의 부분적인 작은 개구부와 낡은 금속 벽이 보인다. 참고 장소의 고정 구조와 재질을 유지하고 아이도 좌판에 정상적으로 앉아 있다. A보다 상체가 크게 잡혔지만 넓은 등받이와 창문이 여전히 보여, 배경을 좁은 띠로 제한하라는 요구에는 덜 미친다.",
        "entities": "금발의 어린 여자아이 한 명이며 얼굴 윤곽, 체격, 남색 반팔 상의는 참고 인물과 대체로 일치한다. 정확한 혼혈 배경과 감긴 눈의 크기는 영상만으로 검증할 수 없다. 손, 팔, 상의와 바지에 혈흔이 있고 복부의 피는 젖은 광택을 보인다. 후두부는 보이지 않아 이전 상처의 지속 여부는 판단할 수 없다. 창밖의 주황빛 수평선은 일몰에 부합하지만 차량의 속도는 뚜렷하게 드러나지 않는다. 추가 인물이나 판독 가능한 글자는 없다.",
        "hard_violations": [],
        "physics": "엉덩이와 허벅지는 벤치 좌판에 지지되고 상체는 앞으로 굽어 있다. 오른손은 복부에 닿아 압박하며 왼팔은 어깨에서 자연스럽게 아래로 내려온다. 프레임 밖 손의 접촉점은 확인되지 않지만 팔 자체가 부유하는 것은 아니다. 불가능한 관절이나 지지 없는 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.446,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.446,
    "B": 2.0
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1446
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "옆구리를 움켜쥔 웅크린 자세뿐만 아니라, 창밖의 속도감과 명시된 머리 부상 등 지속 상태(Carried State)를 완벽하게 반영하여 프롬프트 지시를 가장 충실히 이행함."
   },
   {
    "label": "A",
    "score": 1446,
    "verdict_ko": "샷 텍스트가 요구한 고통스러운 표정과 자세는 잘 구현했으나, 달리는 트럭의 속도감과 앰버의 머리 부상에 대한 묘사가 누락됨."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L248B03.png",
    "asset_id": "20b9995f-79f3-4c42-9a56-159746fe0f8d",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-d898-7644-abac-2e5eb9c7ee9e",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S72sh43__bgfirst_bg.png",
   "bg_asset_id": "d805328b-fd8e-4d2d-b2d9-3aebab763862",
   "bg_record_key": "S72sh43::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S72sh58::signage": {
  "fp": "8220f4478eefb777",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S72sh58": {
  "input_fingerprint": "b42e39c42ad06ca8",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 열린 차 문 너머 눈부신 불빛 사이로 다정하게 미소를 띤 신부의 환한 얼굴.\n\nLOCATION (lock): Outside the stalled truck's open doorway on a rain-soaked road at night, amid the rescuers' bright flashlight beams. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Truck doorway (Open after the truck has stopped) — Viewed diagonally from the interior toward 신부 and 라울 outside; used as A narrow edge frames the recognition and preserves the inside-to-outside relationship.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Intense flashlight glare surrounds the doorway while controlled exposure preserves 신부's bright, gently smiling face and the darker foreground shoulder.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The truck has stopped on the road with its fuel exhausted, and its door is now open in the rain. Charlie remains inside with his accumulated body damage and exposed chest opening. 신부: He stands outside the open truck door, still wearing his clerical collar. Powerful flashlight beams shine into the truck through the open doorway.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 열린 차 문 너머 눈부신 불빛 사이로 다정하게 미소를 띤 신부의 환한 얼굴.\n\nLOCATION (lock): Outside the stalled truck's open doorway on a rain-soaked road at night, amid the rescuers' bright flashlight beams. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Truck doorway (Open after the truck has stopped) — Viewed diagonally from the interior toward 신부 and 라울 outside; used as A narrow edge frames the recognition and preserves the inside-to-outside relationship.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Intense flashlight glare surrounds the doorway while controlled exposure preserves 신부's bright, gently smiling face and the darker foreground shoulder.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The truck has stopped on the road with its fuel exhausted, and its door is now open in the rain. Charlie remains inside with his accumulated body damage and exposed chest opening. 신부: He stands outside the open truck door, still wearing his clerical collar. Powerful flashlight beams shine into the truck through the open doorway.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): sunset.\n\nSHOT TEXT (authoritative, Korean): 열린 차 문 너머 눈부신 불빛 사이로 다정하게 미소를 띤 신부의 환한 얼굴.\n\nLOCATION (lock): Outside the stalled truck's open doorway on a rain-soaked road at night, amid the rescuers' bright flashlight beams. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Truck doorway (Open after the truck has stopped) — Viewed diagonally from the interior toward 신부 and 라울 outside; used as A narrow edge frames the recognition and preserves the inside-to-outside relationship.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Intense flashlight glare surrounds the doorway while controlled exposure preserves 신부's bright, gently smiling face and the darker foreground shoulder.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The truck has stopped on the road with its fuel exhausted, and its door is now open in the rain. Charlie remains inside with his accumulated body damage and exposed chest opening. 신부: He stands outside the open truck door, still wearing his clerical collar. Powerful flashlight beams shine into the truck through the open doorway.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "신부가 전경을 향해 미소 지음. 배경에서 강한 플래시 빛이 비춤.",
    "built_space": "트럭 운전석 내부 구조(대시보드, 일반 차문). 이전 샷의 화물칸과 불일치.",
    "entities": "신부의 외모는 레퍼런스와 일치. 전경 인물 옷에 명확한 알파벳 텍스트 포함.",
    "hard_violations": [
     "[gemini-pro] 화면에 읽을 수 있는 텍스트('RE') 포함",
     "[gemini-pro] 이전 샷과 단절된 잘못된 실내 공간(운전석) 생성",
     "[gpt-high] 신부만 허용한 인물 제한과 달리, 배경에 반사띠 복장을 입은 구조대원들이 추가되어 있다."
    ],
    "physics": "신부는 비 내리는 외부 바닥에 서 있음."
   },
   {
    "label": "B",
    "direction": "신부가 전경을 바라보며 미소 지음. 뒤편에서 밝은 조명이 들어옴.",
    "built_space": "화물칸 내부에서 밖을 보는 구도. 두꺼운 문과 왼쪽 쿠션 좌석이 이전 샷과 일치.",
    "entities": "신부의 외모와 복장이 정확함. 전경에 어두운 옷을 입은 어깨가 배치됨.",
    "hard_violations": [
     "[gpt-high] 신부만 허용한 인물 제한과 달리, 배경 도로에 여러 구조대원의 몸이 추가되어 있다."
    ],
    "physics": "신부는 외부 바닥에 안정적으로 서 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "이전 샷의 화물칸 내부 공간을 일관되게 유지하며 신부의 표정과 지정된 프레이밍을 정확히 구현함."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "금지된 텍스트('RE')가 포함되었으며, 이전 샷과 일치하지 않는 운전석 내부를 생성하여 감점됨."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "신부가 전경을 향해 미소 지음. 배경에서 강한 플래시 빛이 비춤.",
        "built_space": "트럭 운전석 내부 구조(대시보드, 일반 차문). 이전 샷의 화물칸과 불일치.",
        "entities": "신부의 외모는 레퍼런스와 일치. 전경 인물 옷에 명확한 알파벳 텍스트 포함.",
        "hard_violations": [
         "화면에 읽을 수 있는 텍스트('RE') 포함",
         "이전 샷과 단절된 잘못된 실내 공간(운전석) 생성"
        ],
        "physics": "신부는 비 내리는 외부 바닥에 서 있음."
       },
       {
        "label": "B",
        "direction": "신부가 전경을 바라보며 미소 지음. 뒤편에서 밝은 조명이 들어옴.",
        "built_space": "화물칸 내부에서 밖을 보는 구도. 두꺼운 문과 왼쪽 쿠션 좌석이 이전 샷과 일치.",
        "entities": "신부의 외모와 복장이 정확함. 전경에 어두운 옷을 입은 어깨가 배치됨.",
        "hard_violations": [],
        "physics": "신부는 외부 바닥에 안정적으로 서 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "이전 샷의 화물칸 내부 공간을 일관되게 유지하며 신부의 표정과 지정된 프레이밍을 정확히 구현함."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "금지된 텍스트('RE')가 포함되었으며, 이전 샷과 일치하지 않는 운전석 내부를 생성하여 감점됨."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "신부가 전경을 향해 미소 지음. 배경에서 강한 플래시 빛이 비춤.",
        "built_space": "트럭 운전석 내부 구조(대시보드, 일반 차문). 이전 샷의 화물칸과 불일치.",
        "entities": "신부의 외모는 레퍼런스와 일치. 전경 인물 옷에 명확한 알파벳 텍스트 포함.",
        "hard_violations": [
         "화면에 읽을 수 있는 텍스트('RE') 포함",
         "이전 샷과 단절된 잘못된 실내 공간(운전석) 생성"
        ],
        "physics": "신부는 비 내리는 외부 바닥에 서 있음."
       },
       {
        "label": "B",
        "direction": "신부가 전경을 바라보며 미소 지음. 뒤편에서 밝은 조명이 들어옴.",
        "built_space": "화물칸 내부에서 밖을 보는 구도. 두꺼운 문과 왼쪽 쿠션 좌석이 이전 샷과 일치.",
        "entities": "신부의 외모와 복장이 정확함. 전경에 어두운 옷을 입은 어깨가 배치됨.",
        "hard_violations": [],
        "physics": "신부는 외부 바닥에 안정적으로 서 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "기존 차량의 금속 내장과 젖은 문은 비교적 잘 이어지지만, 신부의 얼굴이 너무 작고 전경 인물이 화면을 지배하며 허용되지 않은 구조대원들도 보인다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "신부의 환한 미소와 실내를 향한 시선을 더 가까이 담아 우세하지만, 추가 인물과 운전석 중심으로 바뀐 공간 때문에 그대로 사용할 수는 없다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "신부는 고개를 약간 들고 차량 안쪽, 화면 오른쪽 전경 인물의 얼굴 쪽을 바라보며 웃는다. 뒤쪽의 강한 빛은 열린 출입구를 통해 카메라와 실내 방향으로 들어온다. 개별 손전등이나 그것을 쥔 손은 눈부심 때문에 확인되지 않는다.",
        "built_space": "왼쪽에 회색 직물 좌석 일부, 낡은 금속 내벽, 경첩이 달린 열린 금속 문 한 짝이 보인다. 신부는 문 밖 젖은 도로 쪽에 있고 카메라는 실내에 있어 안팎 관계는 성립한다. 다만 문과 내벽이 넓은 면적을 차지해 좁은 가장자리라는 지정과 다르며, 신부는 얼굴 클로즈업보다 상반신 구도에 가깝다. 불가능한 거울 반사나 중복 문은 보이지 않는다.",
        "entities": "신부는 주름과 희끗한 빗어 넘긴 머리를 지닌 고령의 한국인 남성으로 보이며 참조의 얼굴 특징을 상당 부분 유지한다. 검은 성직복과 흰 성직자 칼라가 확인된다. 오른쪽에는 젖은 짙은 구조복과 반사띠처럼 보이는 부분을 가진 별도 인물이 크게 들어오며, 뒤쪽 도로에도 여러 사람의 몸과 다리가 보인다. 전경의 어두운 어깨 자체는 조명 지시에 있으나 이 복장과 인물 정체는 지정된 신부와 맞지 않는다. 젖은 노면, 구조 차량으로 보이는 배경 차량, 강한 불빛과 석양 하늘이 보인다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "신부만 허용한 인물 제한과 달리, 배경 도로에 여러 구조대원의 몸이 추가되어 있다."
        ],
        "physics": "신부는 몸을 차량 안쪽으로 조금 기울이고 있으며 하체는 화면 밖으로 이어진다. 발 접촉은 보이지 않지만 도로에 서 있는 자세와 모순되지 않는다. 전경 인물도 몸통이 화면 밖으로 이어져 공중에 떠 있다는 증거가 없다. 열린 문은 보이는 경첩으로 차체에 지지되며, 좌석은 실내에 고정되어 있다. 빛의 광원은 눈부심에 가려져 있어 손전등의 파지 여부까지 판단할 수 없다."
       },
       {
        "label": "B",
        "direction": "신부는 화면 오른쪽의 차량 내부 인물을 바라보며 다정하게 웃는다. 시선은 카메라 정면보다 전경 상대 쪽으로 향해 알아보는 순간으로 읽힌다. 배경의 강한 불빛은 열린 문을 통과해 실내와 카메라 쪽으로 들어오지만, 손전등 자체와 조사 각도는 분리해서 확인하기 어렵다.",
        "built_space": "열린 창문 달린 문 한 짝, 문 안쪽 손잡이 하나, 앞기둥의 보조 손잡이 하나, 왼쪽 아래 대시보드와 송풍구 하나가 보인다. 카메라는 운전실 안에서 문 밖 신부를 비스듬히 본다. 문과 신부의 위치는 물리적으로 성립하지만, 참조의 회색 직물 벤치와 노출 금속 내장 대신 운전석 설비가 강조되어 같은 장소의 연속성이 약하다. 신부의 얼굴은 A보다 크지만 문 전체와 전경 어깨가 여전히 넓게 들어온다. 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "신부는 60대 정도의 한국인 남성으로 보이고, 참조와 유사한 머리선, 희끗한 머리, 이마와 눈가 주름을 갖는다. 검은 성직복과 흰 칼라가 있으며 환하게 웃는 얼굴이 명확하다. 오른쪽 전경에는 머리 일부와 구조복 차림의 어깨를 가진 별도 인물이 있고, 배경 좌우에도 반사띠 복장의 사람들이 보인다. 지정된 어두운 전경 어깨는 구현했으나 그 사람의 정체와 복장은 인물 제한에 부합하지 않는다. 비, 젖은 차체, 도로, 석양과 구조 차량의 불빛이 보이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "신부만 허용한 인물 제한과 달리, 배경에 반사띠 복장을 입은 구조대원들이 추가되어 있다."
        ],
        "physics": "신부의 상체는 문 밖에서 자연스럽게 앞으로 기울어 있고 하체는 프레임 밖이다. 보이지 않는 발 때문에 부유로 판단할 근거는 없다. 전경 인물의 어깨와 머리는 연결되어 있으며 몸통이 화면 아래로 이어진다. 열린 문은 앞쪽 차체에 연결된 정상적인 차량 문 배치이고, 유리와 손잡이도 문에 지지되어 있다. 빗방울과 젖은 표면의 반사는 자연스럽다. 손전등을 쥔 손은 확인되지 않지만 떠 있는 소품도 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "기존 차량의 금속 내장과 젖은 문은 비교적 잘 이어지지만, 신부의 얼굴이 너무 작고 전경 인물이 화면을 지배하며 허용되지 않은 구조대원들도 보인다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "신부의 환한 미소와 실내를 향한 시선을 더 가까이 담아 우세하지만, 추가 인물과 운전석 중심으로 바뀐 공간 때문에 그대로 사용할 수는 없다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "신부는 고개를 약간 들고 차량 안쪽, 화면 오른쪽 전경 인물의 얼굴 쪽을 바라보며 웃는다. 뒤쪽의 강한 빛은 열린 출입구를 통해 카메라와 실내 방향으로 들어온다. 개별 손전등이나 그것을 쥔 손은 눈부심 때문에 확인되지 않는다.",
        "built_space": "왼쪽에 회색 직물 좌석 일부, 낡은 금속 내벽, 경첩이 달린 열린 금속 문 한 짝이 보인다. 신부는 문 밖 젖은 도로 쪽에 있고 카메라는 실내에 있어 안팎 관계는 성립한다. 다만 문과 내벽이 넓은 면적을 차지해 좁은 가장자리라는 지정과 다르며, 신부는 얼굴 클로즈업보다 상반신 구도에 가깝다. 불가능한 거울 반사나 중복 문은 보이지 않는다.",
        "entities": "신부는 주름과 희끗한 빗어 넘긴 머리를 지닌 고령의 한국인 남성으로 보이며 참조의 얼굴 특징을 상당 부분 유지한다. 검은 성직복과 흰 성직자 칼라가 확인된다. 오른쪽에는 젖은 짙은 구조복과 반사띠처럼 보이는 부분을 가진 별도 인물이 크게 들어오며, 뒤쪽 도로에도 여러 사람의 몸과 다리가 보인다. 전경의 어두운 어깨 자체는 조명 지시에 있으나 이 복장과 인물 정체는 지정된 신부와 맞지 않는다. 젖은 노면, 구조 차량으로 보이는 배경 차량, 강한 불빛과 석양 하늘이 보인다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "신부만 허용한 인물 제한과 달리, 배경 도로에 여러 구조대원의 몸이 추가되어 있다."
        ],
        "physics": "신부는 몸을 차량 안쪽으로 조금 기울이고 있으며 하체는 화면 밖으로 이어진다. 발 접촉은 보이지 않지만 도로에 서 있는 자세와 모순되지 않는다. 전경 인물도 몸통이 화면 밖으로 이어져 공중에 떠 있다는 증거가 없다. 열린 문은 보이는 경첩으로 차체에 지지되며, 좌석은 실내에 고정되어 있다. 빛의 광원은 눈부심에 가려져 있어 손전등의 파지 여부까지 판단할 수 없다."
       },
       {
        "label": "A",
        "direction": "신부는 화면 오른쪽의 차량 내부 인물을 바라보며 다정하게 웃는다. 시선은 카메라 정면보다 전경 상대 쪽으로 향해 알아보는 순간으로 읽힌다. 배경의 강한 불빛은 열린 문을 통과해 실내와 카메라 쪽으로 들어오지만, 손전등 자체와 조사 각도는 분리해서 확인하기 어렵다.",
        "built_space": "열린 창문 달린 문 한 짝, 문 안쪽 손잡이 하나, 앞기둥의 보조 손잡이 하나, 왼쪽 아래 대시보드와 송풍구 하나가 보인다. 카메라는 운전실 안에서 문 밖 신부를 비스듬히 본다. 문과 신부의 위치는 물리적으로 성립하지만, 참조의 회색 직물 벤치와 노출 금속 내장 대신 운전석 설비가 강조되어 같은 장소의 연속성이 약하다. 신부의 얼굴은 A보다 크지만 문 전체와 전경 어깨가 여전히 넓게 들어온다. 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "신부는 60대 정도의 한국인 남성으로 보이고, 참조와 유사한 머리선, 희끗한 머리, 이마와 눈가 주름을 갖는다. 검은 성직복과 흰 칼라가 있으며 환하게 웃는 얼굴이 명확하다. 오른쪽 전경에는 머리 일부와 구조복 차림의 어깨를 가진 별도 인물이 있고, 배경 좌우에도 반사띠 복장의 사람들이 보인다. 지정된 어두운 전경 어깨는 구현했으나 그 사람의 정체와 복장은 인물 제한에 부합하지 않는다. 비, 젖은 차체, 도로, 석양과 구조 차량의 불빛이 보이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "신부만 허용한 인물 제한과 달리, 배경에 반사띠 복장을 입은 구조대원들이 추가되어 있다."
        ],
        "physics": "신부의 상체는 문 밖에서 자연스럽게 앞으로 기울어 있고 하체는 프레임 밖이다. 보이지 않는 발 때문에 부유로 판단할 근거는 없다. 전경 인물의 어깨와 머리는 연결되어 있으며 몸통이 화면 아래로 이어진다. 열린 문은 앞쪽 차체에 연결된 정상적인 차량 문 배치이고, 유리와 손잡이도 문에 지지되어 있다. 빗방울과 젖은 표면의 반사는 자연스럽다. 손전등을 쥔 손은 확인되지 않지만 떠 있는 소품도 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.429,
    "B": 1.75
   },
   "adjusted": {
    "A": 1.179,
    "B": 1.5
   },
   "violations": {
    "A": [
     "[gemini-pro] 화면에 읽을 수 있는 텍스트('RE') 포함",
     "[gemini-pro] 이전 샷과 단절된 잘못된 실내 공간(운전석) 생성",
     "[gpt-high] 신부만 허용한 인물 제한과 달리, 배경에 반사띠 복장을 입은 구조대원들이 추가되어 있다."
    ],
    "B": [
     "[gpt-high] 신부만 허용한 인물 제한과 달리, 배경 도로에 여러 구조대원의 몸이 추가되어 있다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1500,
   "A": 1179
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1500,
    "verdict_ko": "이전 샷의 화물칸 내부 공간을 일관되게 유지하며 신부의 표정과 지정된 프레이밍을 정확히 구현함.  ★위반: [gpt-high] 신부만 허용한 인물 제한과 달리, 배경 도로에 여러 구조대원의 몸이 추가되어 있다."
   },
   {
    "label": "A",
    "score": 1179,
    "verdict_ko": "금지된 텍스트('RE')가 포함되었으며, 이전 샷과 일치하지 않는 운전석 내부를 생성하여 감점됨.  ★위반: [gemini-pro] 화면에 읽을 수 있는 텍스트('RE') 포함 / [gemini-pro] 이전 샷과 단절된 잘못된 실내 공간(운전석) 생성 / [gpt-high] 신부만 허용한 인물 제한과 달리, 배경에 반사띠 복장을 입은 구조대원들이 추가되어 있다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S72sh43_sel.png",
    "asset_id": "01973771-a955-4aae-87a5-759f37e4c40c",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 신부: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1402213>",
    "asset_id": "8696070d-ac09-4a5c-95f2-3daf1c015c32",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-dbe5-7aee-ab7f-01b9dfc96c63",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S72sh43"
  }
 },
 "S72sh58::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T09:47:51.361244+00:00",
  "fingerprint": "252419827a526356b2562d74098a510d05eeb658fcf5e0c89f2d5a728a632bb8",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S72sh58_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S72sh58_sel.png",
  "source_sha256": "72424f4612ea46dd5012dc1589b7a909805a281dcb4829a6b35fc351e7dccc03",
  "file": "S72sh58_cine.png",
  "staged_sha256": "c1dd62de693b20d76404c8a9aabdebf37e526e9bb6c83898716ca3155ab56c86",
  "latency_ms": 10230
 },
 "S73sh2::signage": {
  "fp": "60df091bc68a947f",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::9c2ab35875802e03": {
  "subjects": [],
  "subject_text": "목포 항구도시 보건소 병실\n응급치료용 침대가 놓인 소규모 보건소 병실. 단출한 실내에 침대와 주변 통로가 마련되어 있고 출입문이 가까이 있다.",
  "identity": "canonical",
  "scope_id": "L253",
  "scope_role": "location_interior",
  "scope_sha": "bd4d2d740e8fc0c7"
 },
 "S73sh2::bgfirst_bg": {
  "input_fingerprint": "5df15547258838c8",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 침대 곁에 서서 걱정스러운 표정으로 앰버를 내려다보는 현우의 상체.\n\nLOCATION (lock): Beside the child's bed inside a harbor-town public clinic, under ordinary nighttime treatment-room lighting.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Clinic bed (Occupied by 앰버, sleeping after emergency treatment) — The bedside edge runs diagonally across the lower-left of the composition; used as Connects her sleeping foreground presence with 현우's downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained nighttime clinic ambience maintains gentle facial modeling and readable shadow detail without specifying unsupported fixtures or colored light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 침대 곁에 서서 걱정스러운 표정으로 앰버를 내려다보는 현우의 상체.\n\nLOCATION (lock): Beside the child's bed inside a harbor-town public clinic, under ordinary nighttime treatment-room lighting.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Clinic bed (Occupied by 앰버, sleeping after emergency treatment) — The bedside edge runs diagonally across the lower-left of the composition; used as Connects her sleeping foreground presence with 현우's downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained nighttime clinic ambience maintains gentle facial modeling and readable shadow detail without specifying unsupported fixtures or colored light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S73sh2__bgfirst_bg.png",
  "asset_id": "a58436f5-7eb4-4977-a36e-faddd6803726",
  "input_asset_ids": [
   "480435f8-c957-407f-9883-a0ac016f25eb",
   "469b5a0c-e506-448f-a857-aa6ec1553e01"
  ]
 },
 "S73sh2": {
  "input_fingerprint": "e753ba3427ba8ff6",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 침대 곁에 서서 걱정스러운 표정으로 앰버를 내려다보는 현우의 상체.\n\nLOCATION (lock): Beside the child's bed inside a harbor-town public clinic, under ordinary nighttime treatment-room lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Clinic bed (Occupied by 앰버, sleeping after emergency treatment) — The bedside edge runs diagonally across the lower-left of the composition; used as Connects her sleeping foreground presence with 현우's downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained nighttime clinic ambience maintains gentle facial modeling and readable shadow detail without specifying unsupported fixtures or colored light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Amber is reclined on the clinic bed, supported by the mattress as she sleeps after emergency treatment and later develops a high fever. The source does not specify whether she rests on her back or side, her head's direction, or the positions of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The setting is the health center at night, with a bed in use. Charlie's previously damaged bodywork and exposed chest opening remain unrepaired. 현우: He remains beside the bed, worried, with dirty skin, shabby clothing and accumulated injuries.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 침대 곁에 서서 걱정스러운 표정으로 앰버를 내려다보는 현우의 상체.\n\nLOCATION (lock): Beside the child's bed inside a harbor-town public clinic, under ordinary nighttime treatment-room lighting. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Clinic bed (Occupied by 앰버, sleeping after emergency treatment) — The bedside edge runs diagonally across the lower-left of the composition; used as Connects her sleeping foreground presence with 현우's downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained nighttime clinic ambience maintains gentle facial modeling and readable shadow detail without specifying unsupported fixtures or colored light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Amber is reclined on the clinic bed, supported by the mattress as she sleeps after emergency treatment and later develops a high fever. The source does not specify whether she rests on her back or side, her head's direction, or the positions of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The setting is the health center at night, with a bed in use. Charlie's previously damaged bodywork and exposed chest opening remain unrepaired. 현우: He remains beside the bed, worried, with dirty skin, shabby clothing and accumulated injuries.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 침대 곁에 서서 걱정스러운 표정으로 앰버를 내려다보는 현우의 상체.\n\nLOCATION (lock): Beside the child's bed inside a harbor-town public clinic, under ordinary nighttime treatment-room lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Clinic bed (Occupied by 앰버, sleeping after emergency treatment) — The bedside edge runs diagonally across the lower-left of the composition; used as Connects her sleeping foreground presence with 현우's downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained nighttime clinic ambience maintains gentle facial modeling and readable shadow detail without specifying unsupported fixtures or colored light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Amber is reclined on the clinic bed, supported by the mattress as she sleeps after emergency treatment and later develops a high fever. The source does not specify whether she rests on her back or side, her head's direction, or the positions of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The setting is the health center at night, with a bed in use. Charlie's previously damaged bodywork and exposed chest opening remain unrepaired. 현우: He remains beside the bed, worried, with dirty skin, shabby clothing and accumulated injuries.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S73sh2__bgfirst_bg.png",
     "asset_id": "a58436f5-7eb4-4977-a36e-faddd6803726",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S73sh2.png",
     "asset_id": "480435f8-c957-407f-9883-a0ac016f25eb",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L253B02.png",
     "asset_id": "469b5a0c-e506-448f-a857-aa6ec1553e01",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우의 고개와 시선이 침대에 누워 잠든 앰버를 향해 명확하게 아래로 내려다보고 있음.",
    "built_space": "로케이션 사진과 완벽히 일치하는 병실 내부. 침대 머리맡의 의료용 벽면 패널, 배경의 창문, 창문 아래의 수납장 등이 정확한 위치와 비례로 배치되어 있음.",
    "entities": "현우는 참조 이미지의 앳된 얼굴, 헝클어진 검은 머리를 잘 반영하고 있으며 지저분한 옷차림과 상처의 디테일이 훌륭함. 앰버는 아이의 모습으로 침대에 눕혀져 있음.",
    "hard_violations": [],
    "physics": "현우는 바닥에 선 채로 상체를 침대 쪽으로 약간 숙이고 있어 무게 중심과 자세가 자연스러움. 앰버는 매트리스의 지지를 받으며 중력에 맞게 편안히 누워 있음."
   },
   {
    "label": "B",
    "direction": "현우가 고개를 숙여 침대에 있는 앰버에게 시선을 향하고 있음.",
    "built_space": "병실의 전반적인 분위기는 비슷하나, 로케이션 사진에 없는 두꺼운 금속 침대 난간이 화면 전경을 수평으로 크게 가로막고 있어 공간의 구조적 특징을 왜곡함.",
    "entities": "현우의 얼굴이 참조 이미지와 어느 정도 닮았으나 A에 비해 인상이 약간 다르게 표현됨. 앰버는 다문화 배경을 지닌 아이의 모습으로 침대에 누워 있음.",
    "hard_violations": [
     "[gemini-pro] 샷 텍스트의 '서서(standing)'라는 명확한 지시를 위반하고 현우의 어깨 높이가 침대와 거의 맞닿을 정도로 앉거나 웅크린 자세로 묘사됨",
     "[gemini-pro] 프레임 지시사항인 '침대 가장자리가 화면 좌측 하단을 대각선으로 가로지름'을 어기고 두꺼운 난간이 수평으로 배치됨"
    ],
    "physics": "앰버의 몸과 팔은 침대 위에 자연스럽게 놓여 있음. 현우의 신체 하단은 잘려 있으나, 상체와 머리의 위치로 보아 서 있지 않고 앉아 있음을 강하게 암시함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 10,
        "verdict_ko": "프롬프트가 지시한 현우가 서서 내려다보는 자세와 침대 가장자리가 좌측 하단을 대각선으로 가로지르는 프레이밍 구도를 완벽하게 구현하였으며, 캐릭터의 참조 이미지와 로케이션의 특징도 매우 훌륭하게 반영했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "샷 텍스트에서 명시한 '서서'라는 행동 지시를 어기고 앉아 있는 높이로 렌더링되었으며, 요구된 대각선 구도를 무시하고 원본에 없는 두꺼운 가드레일이 화면 하단을 수평으로 막고 있어 지침을 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 고개와 시선이 침대에 누워 잠든 앰버를 향해 명확하게 아래로 내려다보고 있음.",
        "built_space": "로케이션 사진과 완벽히 일치하는 병실 내부. 침대 머리맡의 의료용 벽면 패널, 배경의 창문, 창문 아래의 수납장 등이 정확한 위치와 비례로 배치되어 있음.",
        "entities": "현우는 참조 이미지의 앳된 얼굴, 헝클어진 검은 머리를 잘 반영하고 있으며 지저분한 옷차림과 상처의 디테일이 훌륭함. 앰버는 아이의 모습으로 침대에 눕혀져 있음.",
        "hard_violations": [],
        "physics": "현우는 바닥에 선 채로 상체를 침대 쪽으로 약간 숙이고 있어 무게 중심과 자세가 자연스러움. 앰버는 매트리스의 지지를 받으며 중력에 맞게 편안히 누워 있음."
       },
       {
        "label": "B",
        "direction": "현우가 고개를 숙여 침대에 있는 앰버에게 시선을 향하고 있음.",
        "built_space": "병실의 전반적인 분위기는 비슷하나, 로케이션 사진에 없는 두꺼운 금속 침대 난간이 화면 전경을 수평으로 크게 가로막고 있어 공간의 구조적 특징을 왜곡함.",
        "entities": "현우의 얼굴이 참조 이미지와 어느 정도 닮았으나 A에 비해 인상이 약간 다르게 표현됨. 앰버는 다문화 배경을 지닌 아이의 모습으로 침대에 누워 있음.",
        "hard_violations": [
         "샷 텍스트의 '서서(standing)'라는 명확한 지시를 위반하고 현우의 어깨 높이가 침대와 거의 맞닿을 정도로 앉거나 웅크린 자세로 묘사됨",
         "프레임 지시사항인 '침대 가장자리가 화면 좌측 하단을 대각선으로 가로지름'을 어기고 두꺼운 난간이 수평으로 배치됨"
        ],
        "physics": "앰버의 몸과 팔은 침대 위에 자연스럽게 놓여 있음. 현우의 신체 하단은 잘려 있으나, 상체와 머리의 위치로 보아 서 있지 않고 앉아 있음을 강하게 암시함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 10,
        "verdict_ko": "프롬프트가 지시한 현우가 서서 내려다보는 자세와 침대 가장자리가 좌측 하단을 대각선으로 가로지르는 프레이밍 구도를 완벽하게 구현하였으며, 캐릭터의 참조 이미지와 로케이션의 특징도 매우 훌륭하게 반영했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "샷 텍스트에서 명시한 '서서'라는 행동 지시를 어기고 앉아 있는 높이로 렌더링되었으며, 요구된 대각선 구도를 무시하고 원본에 없는 두꺼운 가드레일이 화면 하단을 수평으로 막고 있어 지침을 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 고개와 시선이 침대에 누워 잠든 앰버를 향해 명확하게 아래로 내려다보고 있음.",
        "built_space": "로케이션 사진과 완벽히 일치하는 병실 내부. 침대 머리맡의 의료용 벽면 패널, 배경의 창문, 창문 아래의 수납장 등이 정확한 위치와 비례로 배치되어 있음.",
        "entities": "현우는 참조 이미지의 앳된 얼굴, 헝클어진 검은 머리를 잘 반영하고 있으며 지저분한 옷차림과 상처의 디테일이 훌륭함. 앰버는 아이의 모습으로 침대에 눕혀져 있음.",
        "hard_violations": [],
        "physics": "현우는 바닥에 선 채로 상체를 침대 쪽으로 약간 숙이고 있어 무게 중심과 자세가 자연스러움. 앰버는 매트리스의 지지를 받으며 중력에 맞게 편안히 누워 있음."
       },
       {
        "label": "B",
        "direction": "현우가 고개를 숙여 침대에 있는 앰버에게 시선을 향하고 있음.",
        "built_space": "병실의 전반적인 분위기는 비슷하나, 로케이션 사진에 없는 두꺼운 금속 침대 난간이 화면 전경을 수평으로 크게 가로막고 있어 공간의 구조적 특징을 왜곡함.",
        "entities": "현우의 얼굴이 참조 이미지와 어느 정도 닮았으나 A에 비해 인상이 약간 다르게 표현됨. 앰버는 다문화 배경을 지닌 아이의 모습으로 침대에 누워 있음.",
        "hard_violations": [
         "샷 텍스트의 '서서(standing)'라는 명확한 지시를 위반하고 현우의 어깨 높이가 침대와 거의 맞닿을 정도로 앉거나 웅크린 자세로 묘사됨",
         "프레임 지시사항인 '침대 가장자리가 화면 좌측 하단을 대각선으로 가로지름'을 어기고 두꺼운 난간이 수평으로 배치됨"
        ],
        "physics": "앰버의 몸과 팔은 침대 위에 자연스럽게 놓여 있음. 현우의 신체 하단은 잘려 있으나, 상체와 머리의 위치로 보아 서 있지 않고 앉아 있음을 강하게 암시함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "걱정하며 앰버를 내려다보는 관계는 맞지만, 침대와 앰버를 넓게 보여주느라 현우의 상체 중심 미디엄 숏이라는 구도 지시가 약해졌다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "현우의 상체를 중심으로 왼쪽 아래 잠든 앰버와 하향 시선을 연결하며, 참조 진료실의 배치와 야간 분위기도 더 충실하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 고개와 눈을 왼쪽 아래로 향해 침대에 누운 앰버를 바라본다. 앰버는 눈을 감고 있으며 얼굴은 카메라 쪽으로 비스듬히 돌아 있다. 겨누거나 조작하는 물건은 없다.",
        "built_space": "침대 한 개가 화면 하단 대부분을 차지하고 앞쪽 난간이 대각선으로 지나간다. 현우는 침대 건너편에 있으며 머리부터 허리 부근까지 보인다. 벽면 의료 패널과 조명 각 한 개, 수액대와 펌프 각 한 개, 창문 한 조와 커튼, 침상 옆 수납장 한 개가 보인다. 오른쪽에는 참조에서 확인되지 않는 이동식 침상 테이블이 있다. 주요 재료와 설비는 유사하지만 침대의 비중이 커서 현우 중심의 구도가 약하다. 반사상은 없다.",
        "entities": "현우는 앳된 동아시아계 남성으로 보이며 헝클어진 검은 머리와 얼굴 윤곽이 참조에 대체로 가깝다. 피부의 오염과 상처, 낡고 해진 회색 반팔이 보이지만 참조의 남색 티셔츠와는 다르다. 앰버는 곱슬 짙은 머리의 어린 여자아이로, 눈을 감고 분홍색 상의를 입은 채 누워 있다. 앰버의 외모를 고정하는 참조는 없다. 두 사람 외 인물이나 찰리는 보이지 않는다. 판독 가능한 문구는 뚜렷하지 않다.",
        "hard_violations": [],
        "physics": "앰버는 옆으로 누워 머리를 베개에, 몸통을 매트리스에 맡기고 있다. 팔과 손도 침구 위에 놓여 있어 잠든 몸의 지지가 자연스럽다. 담요는 몸과 침대 위에 늘어져 있다. 현우의 하체는 침대에 가려져 발의 접지는 확인할 수 없지만 상체가 떠 있거나 불가능하게 배치된 증거는 없다. 수액 용기와 펌프는 수액대에 고정되어 있다."
       },
       {
        "label": "B",
        "direction": "현우는 고개를 숙이고 눈을 왼쪽 아래의 앰버 얼굴 쪽으로 향한다. 앰버는 눈을 감고 얼굴을 위로 향한 채 잠들어 있다. 현우의 시선과 잠든 아이의 전경 위치가 직접 연결된다. 방향을 확인해야 할 무기나 손에 든 기기는 없다.",
        "built_space": "현우의 상체가 화면 오른쪽을 크게 차지하고, 침대 한 개와 앰버가 왼쪽 아래 전경에 배치된다. 침대 가장자리와 침구가 하단을 대각선으로 가른다. 왼쪽 붙박이장 한 개, 벽면 의료 패널과 조명 각 한 개, 수액대와 펌프 각 한 개, 침상 옆 수납장 한 개, 뒤쪽 두 칸 창문 한 조와 커튼이 보여 참조 공간의 관계를 잘 유지한다. 오른쪽 벽의 낮은 콘센트들도 보인다. 현우가 침대 옆에 서 있을 공간이 확보되어 있고 불가능한 반사는 없다.",
        "entities": "현우는 앳된 동아시아계 남성으로 보이며 검은 헝클어진 머리와 얼굴이 참조에 대체로 부합한다. 볼과 목, 팔에 상처와 오염이 있고 낡은 회색 티셔츠 위에 해진 후드 겉옷을 입었다. 이는 누적된 부상과 초라한 차림에는 맞지만 참조의 남색 티셔츠 차림과는 다르다. 앰버는 긴 짙은 머리의 어린 여자아이로, 눈을 감고 담요를 덮고 있다. 외모를 고정하는 앰버 참조는 없다. 추가 인물이나 찰리는 없으며 명확하게 읽히는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "앰버의 머리와 목은 베개에, 몸통은 매트리스에 지지된다. 담요 밖으로 들린 팔다리는 없고 담요도 몸과 침구 위에 자연스럽게 놓여 있다. 현우는 침대 옆에서 상체를 조금 앞으로 기울인 서 있는 자세로 읽힌다. 발은 프레임 밖이지만 부유나 비정상적인 지지는 보이지 않는다. 수액 용기는 고리에 걸려 있고 펌프는 기둥에 부착되어 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "걱정하며 앰버를 내려다보는 관계는 맞지만, 침대와 앰버를 넓게 보여주느라 현우의 상체 중심 미디엄 숏이라는 구도 지시가 약해졌다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "현우의 상체를 중심으로 왼쪽 아래 잠든 앰버와 하향 시선을 연결하며, 참조 진료실의 배치와 야간 분위기도 더 충실하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 고개와 눈을 왼쪽 아래로 향해 침대에 누운 앰버를 바라본다. 앰버는 눈을 감고 있으며 얼굴은 카메라 쪽으로 비스듬히 돌아 있다. 겨누거나 조작하는 물건은 없다.",
        "built_space": "침대 한 개가 화면 하단 대부분을 차지하고 앞쪽 난간이 대각선으로 지나간다. 현우는 침대 건너편에 있으며 머리부터 허리 부근까지 보인다. 벽면 의료 패널과 조명 각 한 개, 수액대와 펌프 각 한 개, 창문 한 조와 커튼, 침상 옆 수납장 한 개가 보인다. 오른쪽에는 참조에서 확인되지 않는 이동식 침상 테이블이 있다. 주요 재료와 설비는 유사하지만 침대의 비중이 커서 현우 중심의 구도가 약하다. 반사상은 없다.",
        "entities": "현우는 앳된 동아시아계 남성으로 보이며 헝클어진 검은 머리와 얼굴 윤곽이 참조에 대체로 가깝다. 피부의 오염과 상처, 낡고 해진 회색 반팔이 보이지만 참조의 남색 티셔츠와는 다르다. 앰버는 곱슬 짙은 머리의 어린 여자아이로, 눈을 감고 분홍색 상의를 입은 채 누워 있다. 앰버의 외모를 고정하는 참조는 없다. 두 사람 외 인물이나 찰리는 보이지 않는다. 판독 가능한 문구는 뚜렷하지 않다.",
        "hard_violations": [],
        "physics": "앰버는 옆으로 누워 머리를 베개에, 몸통을 매트리스에 맡기고 있다. 팔과 손도 침구 위에 놓여 있어 잠든 몸의 지지가 자연스럽다. 담요는 몸과 침대 위에 늘어져 있다. 현우의 하체는 침대에 가려져 발의 접지는 확인할 수 없지만 상체가 떠 있거나 불가능하게 배치된 증거는 없다. 수액 용기와 펌프는 수액대에 고정되어 있다."
       },
       {
        "label": "A",
        "direction": "현우는 고개를 숙이고 눈을 왼쪽 아래의 앰버 얼굴 쪽으로 향한다. 앰버는 눈을 감고 얼굴을 위로 향한 채 잠들어 있다. 현우의 시선과 잠든 아이의 전경 위치가 직접 연결된다. 방향을 확인해야 할 무기나 손에 든 기기는 없다.",
        "built_space": "현우의 상체가 화면 오른쪽을 크게 차지하고, 침대 한 개와 앰버가 왼쪽 아래 전경에 배치된다. 침대 가장자리와 침구가 하단을 대각선으로 가른다. 왼쪽 붙박이장 한 개, 벽면 의료 패널과 조명 각 한 개, 수액대와 펌프 각 한 개, 침상 옆 수납장 한 개, 뒤쪽 두 칸 창문 한 조와 커튼이 보여 참조 공간의 관계를 잘 유지한다. 오른쪽 벽의 낮은 콘센트들도 보인다. 현우가 침대 옆에 서 있을 공간이 확보되어 있고 불가능한 반사는 없다.",
        "entities": "현우는 앳된 동아시아계 남성으로 보이며 검은 헝클어진 머리와 얼굴이 참조에 대체로 부합한다. 볼과 목, 팔에 상처와 오염이 있고 낡은 회색 티셔츠 위에 해진 후드 겉옷을 입었다. 이는 누적된 부상과 초라한 차림에는 맞지만 참조의 남색 티셔츠 차림과는 다르다. 앰버는 긴 짙은 머리의 어린 여자아이로, 눈을 감고 담요를 덮고 있다. 외모를 고정하는 앰버 참조는 없다. 추가 인물이나 찰리는 없으며 명확하게 읽히는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "앰버의 머리와 목은 베개에, 몸통은 매트리스에 지지된다. 담요 밖으로 들린 팔다리는 없고 담요도 몸과 침구 위에 자연스럽게 놓여 있다. 현우는 침대 옆에서 상체를 조금 앞으로 기울인 서 있는 자세로 읽힌다. 발은 프레임 밖이지만 부유나 비정상적인 지지는 보이지 않는다. 수액 용기는 고리에 걸려 있고 펌프는 기둥에 부착되어 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.178
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.928
   },
   "violations": {
    "B": [
     "[gemini-pro] 샷 텍스트의 '서서(standing)'라는 명확한 지시를 위반하고 현우의 어깨 높이가 침대와 거의 맞닿을 정도로 앉거나 웅크린 자세로 묘사됨",
     "[gemini-pro] 프레임 지시사항인 '침대 가장자리가 화면 좌측 하단을 대각선으로 가로지름'을 어기고 두꺼운 난간이 수평으로 배치됨"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 928
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "프롬프트가 지시한 현우가 서서 내려다보는 자세와 침대 가장자리가 좌측 하단을 대각선으로 가로지르는 프레이밍 구도를 완벽하게 구현하였으며, 캐릭터의 참조 이미지와 로케이션의 특징도 매우 훌륭하게 반영했습니다."
   },
   {
    "label": "B",
    "score": 928,
    "verdict_ko": "샷 텍스트에서 명시한 '서서'라는 행동 지시를 어기고 앉아 있는 높이로 렌더링되었으며, 요구된 대각선 구도를 무시하고 원본에 없는 두꺼운 가드레일이 화면 하단을 수평으로 막고 있어 지침을 위반했습니다.  ★위반: [gemini-pro] 샷 텍스트의 '서서(standing)'라는 명확한 지시를 위반하고 현우의 어깨 높이가 침대와 거의 맞닿을 정도로 앉거나 웅크린 자세로 묘사됨 / [gemini-pro] 프레임 지시사항인 '침대 가장자리가 화면 좌측 하단을 대각선으로 가로지름'을 어기고 두꺼운 난간이 수평으로 배치됨"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L253B02.png",
    "asset_id": "469b5a0c-e506-448f-a857-aa6ec1553e01",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-dd97-71e2-a718-fb5051650a8d",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S73sh2__bgfirst_bg.png",
   "bg_asset_id": "a58436f5-7eb4-4977-a36e-faddd6803726",
   "bg_record_key": "S73sh2::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S73sh2::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:32:53.985127+00:00",
  "fingerprint": "ba102d113b4350b6f31f83a4a7514f3d53559a822c33610b064c263f53b8faa0",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S73sh2_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S73sh2_sel.png",
  "source_sha256": "93f3c494c3f84036a2a3414f5f0862cf62af231e7c0210e23f480c8bb96755e1",
  "file": "S73sh2_cine.png",
  "staged_sha256": "6aeb9544f3693780354621bf3b4ce2789f80c506f4d65b32a274e2abc66a49b2",
  "latency_ms": 8765
 },
 "S73sh5::signage": {
  "fp": "ae0027474e2b8acb",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S73sh5": {
  "input_fingerprint": "1758af9170e89449",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 꼬질꼬질한 현우의 어깨를 양손으로 꽉 감싸 쥔 신부의 밀착된 자세.\n\nLOCATION (lock): At the bedside inside the harbor-town clinic's patient room, under nighttime clinic lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Clinic bedside edge (Beside the two men as 앰버 rests outside the crop) — A short oblique section appears along the lower edge; used as Retains continuity with the bedside while keeping both supporting hands unobstructed.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the clinic's restrained ambient illumination continuous, preserving 현우's visibly grubby appearance and the gentle modeling of the supporting hands.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the patient's bed, surrounding clinic surfaces, and nighttime interior lighting. Exclude the roadside vehicle, rain, and rescuers' flashlight beams from the preceding rescue.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed remains occupied, and Charlie's damaged bodywork remains unrepaired. 현우: He stays beside the bed, visibly dirty and battered, with shabby clothing. 신부: He stands close beside the bed with his hands extended in a consoling gesture, wearing his clerical collar.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 꼬질꼬질한 현우의 어깨를 양손으로 꽉 감싸 쥔 신부의 밀착된 자세.\n\nLOCATION (lock): At the bedside inside the harbor-town clinic's patient room, under nighttime clinic lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Clinic bedside edge (Beside the two men as 앰버 rests outside the crop) — A short oblique section appears along the lower edge; used as Retains continuity with the bedside while keeping both supporting hands unobstructed.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the clinic's restrained ambient illumination continuous, preserving 현우's visibly grubby appearance and the gentle modeling of the supporting hands.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the patient's bed, surrounding clinic surfaces, and nighttime interior lighting. Exclude the roadside vehicle, rain, and rescuers' flashlight beams from the preceding rescue.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed remains occupied, and Charlie's damaged bodywork remains unrepaired. 현우: He stays beside the bed, visibly dirty and battered, with shabby clothing. 신부: He stands close beside the bed with his hands extended in a consoling gesture, wearing his clerical collar.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 꼬질꼬질한 현우의 어깨를 양손으로 꽉 감싸 쥔 신부의 밀착된 자세.\n\nLOCATION (lock): At the bedside inside the harbor-town clinic's patient room, under nighttime clinic lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Clinic bedside edge (Beside the two men as 앰버 rests outside the crop) — A short oblique section appears along the lower edge; used as Retains continuity with the bedside while keeping both supporting hands unobstructed.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the clinic's restrained ambient illumination continuous, preserving 현우's visibly grubby appearance and the gentle modeling of the supporting hands.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the patient's bed, surrounding clinic surfaces, and nighttime interior lighting. Exclude the roadside vehicle, rain, and rescuers' flashlight beams from the preceding rescue.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed remains occupied, and Charlie's damaged bodywork remains unrepaired. 현우: He stays beside the bed, visibly dirty and battered, with shabby clothing. 신부: He stands close beside the bed with his hands extended in a consoling gesture, wearing his clerical collar.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "신부는 걱정스러운 시선으로 현우의 얼굴을 향해 있고, 현우는 고개를 숙인 채 아래를 내려다보고 있습니다.",
    "built_space": "야간의 병실 내부입니다. 뒤쪽으로 밤거리의 불빛이 보이는 창문이 있으며, 화면 왼쪽 하단 가장자리에 침대 난간의 일부가 사선으로 걸쳐 있어 공간의 연속성을 줍니다.",
    "entities": "현우(18세 남성, 헝클어진 머리, 상처와 오염이 있는 얼굴, 회색 후드티)와 신부(60대 남성, 로만 칼라가 있는 검은색 셔츠)가 참조 이미지와 일치하게 등장합니다. 앰버는 프레임에서 완전히 제외되었습니다.",
    "hard_violations": [],
    "physics": "신부의 두 손이 현우의 양어깨를 단단히 쥐고 몸을 지탱해주고 있으며, 인물들의 무게 중심과 밀착된 자세가 자연스럽습니다."
   },
   {
    "label": "B",
    "direction": "신부는 현우를 향해 시선을 두고 있으며 현우는 아래를 봅니다. 침대에 누운 앰버는 눈을 감고 있습니다.",
    "built_space": "수액 걸이와 병상, 뒤편의 창문이 넓게 보이는 야간의 병실입니다. 인물들의 위치가 침대 옆에 배치되어 있습니다.",
    "entities": "현우와 신부가 등장하지만, 프레임에 포함되지 않아야 할 앰버가 침대에 누운 채로 화면의 상당 부분을 차지하고 있습니다.",
    "hard_violations": [
     "[gemini-pro] 지문에서 등장하지 않아야 한다고 명시된 인물(앰버)이 프레임 안에 크게 포함됨 (\"never add a person the shot text does not show\" 및 \"앰버 rests outside the crop\" 지시 위반).",
     "[gemini-pro] 신부가 현우의 어깨를 감싸는 팔의 구조와 손의 위치가 해부학적으로 불가능하고 기형적으로 교차되어 있음.",
     "[gpt-high] 화면 밖에 있어야 하며 이번 숏의 허용 인물이 아닌 앰버의 얼굴과 상체를 노출했습니다."
    ],
    "physics": "신부의 왼쪽 팔이 현우의 가슴 앞을 가로질러 반대편 어깨를 잡고 있는데, 동시에 나타난 다른 손의 각도와 위치가 신체 구조상 불가능한 형태를 띠고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 10,
        "verdict_ko": "지시된 미디엄 샷 프레이밍을 정확히 구현하여 앰버를 프레임 밖으로 제외하고, 두 인물의 밀착된 자세와 양손으로 어깨를 쥔 모습을 완벽하게 묘사했습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "프레임에서 명시적으로 제외해야 할 인물(앰버)이 그대로 등장하였으며, 지정된 미디엄 샷 스케일을 따르지 않은 치명적인 오류가 있습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "신부는 걱정스러운 시선으로 현우의 얼굴을 향해 있고, 현우는 고개를 숙인 채 아래를 내려다보고 있습니다.",
        "built_space": "야간의 병실 내부입니다. 뒤쪽으로 밤거리의 불빛이 보이는 창문이 있으며, 화면 왼쪽 하단 가장자리에 침대 난간의 일부가 사선으로 걸쳐 있어 공간의 연속성을 줍니다.",
        "entities": "현우(18세 남성, 헝클어진 머리, 상처와 오염이 있는 얼굴, 회색 후드티)와 신부(60대 남성, 로만 칼라가 있는 검은색 셔츠)가 참조 이미지와 일치하게 등장합니다. 앰버는 프레임에서 완전히 제외되었습니다.",
        "hard_violations": [],
        "physics": "신부의 두 손이 현우의 양어깨를 단단히 쥐고 몸을 지탱해주고 있으며, 인물들의 무게 중심과 밀착된 자세가 자연스럽습니다."
       },
       {
        "label": "B",
        "direction": "신부는 현우를 향해 시선을 두고 있으며 현우는 아래를 봅니다. 침대에 누운 앰버는 눈을 감고 있습니다.",
        "built_space": "수액 걸이와 병상, 뒤편의 창문이 넓게 보이는 야간의 병실입니다. 인물들의 위치가 침대 옆에 배치되어 있습니다.",
        "entities": "현우와 신부가 등장하지만, 프레임에 포함되지 않아야 할 앰버가 침대에 누운 채로 화면의 상당 부분을 차지하고 있습니다.",
        "hard_violations": [
         "지문에서 등장하지 않아야 한다고 명시된 인물(앰버)이 프레임 안에 크게 포함됨 (\"never add a person the shot text does not show\" 및 \"앰버 rests outside the crop\" 지시 위반).",
         "신부가 현우의 어깨를 감싸는 팔의 구조와 손의 위치가 해부학적으로 불가능하고 기형적으로 교차되어 있음."
        ],
        "physics": "신부의 왼쪽 팔이 현우의 가슴 앞을 가로질러 반대편 어깨를 잡고 있는데, 동시에 나타난 다른 손의 각도와 위치가 신체 구조상 불가능한 형태를 띠고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 10,
        "verdict_ko": "지시된 미디엄 샷 프레이밍을 정확히 구현하여 앰버를 프레임 밖으로 제외하고, 두 인물의 밀착된 자세와 양손으로 어깨를 쥔 모습을 완벽하게 묘사했습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "프레임에서 명시적으로 제외해야 할 인물(앰버)이 그대로 등장하였으며, 지정된 미디엄 샷 스케일을 따르지 않은 치명적인 오류가 있습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "신부는 걱정스러운 시선으로 현우의 얼굴을 향해 있고, 현우는 고개를 숙인 채 아래를 내려다보고 있습니다.",
        "built_space": "야간의 병실 내부입니다. 뒤쪽으로 밤거리의 불빛이 보이는 창문이 있으며, 화면 왼쪽 하단 가장자리에 침대 난간의 일부가 사선으로 걸쳐 있어 공간의 연속성을 줍니다.",
        "entities": "현우(18세 남성, 헝클어진 머리, 상처와 오염이 있는 얼굴, 회색 후드티)와 신부(60대 남성, 로만 칼라가 있는 검은색 셔츠)가 참조 이미지와 일치하게 등장합니다. 앰버는 프레임에서 완전히 제외되었습니다.",
        "hard_violations": [],
        "physics": "신부의 두 손이 현우의 양어깨를 단단히 쥐고 몸을 지탱해주고 있으며, 인물들의 무게 중심과 밀착된 자세가 자연스럽습니다."
       },
       {
        "label": "B",
        "direction": "신부는 현우를 향해 시선을 두고 있으며 현우는 아래를 봅니다. 침대에 누운 앰버는 눈을 감고 있습니다.",
        "built_space": "수액 걸이와 병상, 뒤편의 창문이 넓게 보이는 야간의 병실입니다. 인물들의 위치가 침대 옆에 배치되어 있습니다.",
        "entities": "현우와 신부가 등장하지만, 프레임에 포함되지 않아야 할 앰버가 침대에 누운 채로 화면의 상당 부분을 차지하고 있습니다.",
        "hard_violations": [
         "지문에서 등장하지 않아야 한다고 명시된 인물(앰버)이 프레임 안에 크게 포함됨 (\"never add a person the shot text does not show\" 및 \"앰버 rests outside the crop\" 지시 위반).",
         "신부가 현우의 어깨를 감싸는 팔의 구조와 손의 위치가 해부학적으로 불가능하고 기형적으로 교차되어 있음."
        ],
        "physics": "신부의 왼쪽 팔이 현우의 가슴 앞을 가로질러 반대편 어깨를 잡고 있는데, 동시에 나타난 다른 손의 각도와 위치가 신체 구조상 불가능한 형태를 띠고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "병실과 현우의 외형은 잘 이어지지만, 제외해야 할 앰버를 노출하고 침대를 크게 담았으며 신부의 두 손도 명확히 보이지 않습니다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "두 남성에 집중한 미디엄 구도에서 신부가 현우의 양어깨를 두 손으로 감싸 쥔 밀착 동작을 충실히 구현했으며, 하단 침대 난간만 다소 크게 보입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "신부는 현우의 얼굴을 바라보고, 현우는 침대 쪽 아래로 시선을 내립니다. 신부의 앞쪽 손은 현우의 화면 오른쪽 어깨를 감싸지만 반대 손은 목 뒤쪽에 일부만 보여 양손의 파지 상태를 확인하기 어렵습니다.",
        "built_space": "왼쪽 수납장 하나, 벽 조명 하나와 의료 설비 패널 하나, 수액대 하나와 부착 장치 하나, 협탁 하나, 뒤쪽 창문이 이전 장면과 같은 배치로 보입니다. 두 남성은 침대 옆에 서 있습니다. 다만 침대가 화면 하단 대부분을 차지하고 환자의 얼굴과 상체까지 드러나므로, 짧은 비스듬한 침대 가장자리만 남기라는 구도와 다릅니다.",
        "entities": "현우는 앳된 동아시아계 남성으로, 헝클어진 검은 머리와 얼굴 상처, 더러운 회색 후드 및 티셔츠가 참고 이미지와 잘 이어집니다. 신부는 주름과 희끗한 머리가 있는 고령의 동아시아계 남성이며 성직자 칼라를 착용했습니다. 그러나 이번 화면에서 제외하도록 지정된 앰버의 얼굴과 몸도 보입니다. 읽을 수 있는 문구는 보이지 않습니다.",
        "hard_violations": [
         "화면 밖에 있어야 하며 이번 숏의 허용 인물이 아닌 앰버의 얼굴과 상체를 노출했습니다."
        ],
        "physics": "신부의 앞쪽 팔과 손은 자연스럽게 연결되어 현우의 어깨에 접촉합니다. 반대 손의 대부분은 가려져 있습니다. 두 남성은 선 자세로 하체가 프레임 밖에 있으며 공중에 뜬 정황은 없습니다. 앰버의 머리와 몸은 베개와 매트리스에 지지되어 있습니다. 수액 용기는 수액대에 걸려 있고 소품들은 협탁에 놓여 있습니다."
       },
       {
        "label": "B",
        "direction": "신부는 가까이 숙인 현우의 얼굴 쪽을 내려다보고, 현우는 고개와 시선을 아래로 떨어뜨립니다. 신부의 두 손이 각각 현우의 양어깨를 향해 뻗어 실제로 감싸 쥐고 있어 위로하는 동작의 대상과 접촉이 명확합니다.",
        "built_space": "왼쪽에 야간 창문과 커튼 한 구역, 왼쪽 아래에 협탁 일부, 뒤쪽에 밝은 병실 벽, 하단에 침대 난간 한 구간이 보입니다. 두 남성은 난간 뒤 침대 옆에 밀착해 있으며, 카메라 각도 변경으로 이전 장면의 나머지 설비가 제외된 것으로 읽힙니다. 환자는 보이지 않습니다. 난간은 두 손을 가리지 않지만 요청한 짧은 가장자리보다 조금 길고 두드러집니다.",
        "entities": "허용된 두 남성만 보입니다. 현우의 앳된 얼굴, 검은 헝클어진 머리, 볼의 상처와 때 묻은 회색 후드가 참고 이미지에 부합합니다. 신부의 나이 든 얼굴, 이마와 눈가 주름, 희끗한 머리 및 얼굴형도 참고 인물과 유사하며 성직자 칼라가 명확합니다. 앰버와 찰리는 프레임 밖이므로 그 상태는 확인할 수 없습니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "신부의 앞쪽 손은 현우의 가까운 어깨와 상완을 감싸고, 반대 손은 먼 쪽 어깨에 닿아 있습니다. 손가락의 굽힘과 팔의 연결은 이 동작으로 가능한 형태입니다. 현우는 상체를 신부 쪽으로 기울였으며 두 사람 모두 서 있는 자세로 읽힙니다. 발은 잘려 있지만 부유하거나 지지 없이 매달린 모습은 없습니다. 침대 난간도 아래로 이어지는 지지 부재가 보입니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "병실과 현우의 외형은 잘 이어지지만, 제외해야 할 앰버를 노출하고 침대를 크게 담았으며 신부의 두 손도 명확히 보이지 않습니다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "두 남성에 집중한 미디엄 구도에서 신부가 현우의 양어깨를 두 손으로 감싸 쥔 밀착 동작을 충실히 구현했으며, 하단 침대 난간만 다소 크게 보입니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "신부는 현우의 얼굴을 바라보고, 현우는 침대 쪽 아래로 시선을 내립니다. 신부의 앞쪽 손은 현우의 화면 오른쪽 어깨를 감싸지만 반대 손은 목 뒤쪽에 일부만 보여 양손의 파지 상태를 확인하기 어렵습니다.",
        "built_space": "왼쪽 수납장 하나, 벽 조명 하나와 의료 설비 패널 하나, 수액대 하나와 부착 장치 하나, 협탁 하나, 뒤쪽 창문이 이전 장면과 같은 배치로 보입니다. 두 남성은 침대 옆에 서 있습니다. 다만 침대가 화면 하단 대부분을 차지하고 환자의 얼굴과 상체까지 드러나므로, 짧은 비스듬한 침대 가장자리만 남기라는 구도와 다릅니다.",
        "entities": "현우는 앳된 동아시아계 남성으로, 헝클어진 검은 머리와 얼굴 상처, 더러운 회색 후드 및 티셔츠가 참고 이미지와 잘 이어집니다. 신부는 주름과 희끗한 머리가 있는 고령의 동아시아계 남성이며 성직자 칼라를 착용했습니다. 그러나 이번 화면에서 제외하도록 지정된 앰버의 얼굴과 몸도 보입니다. 읽을 수 있는 문구는 보이지 않습니다.",
        "hard_violations": [
         "화면 밖에 있어야 하며 이번 숏의 허용 인물이 아닌 앰버의 얼굴과 상체를 노출했습니다."
        ],
        "physics": "신부의 앞쪽 팔과 손은 자연스럽게 연결되어 현우의 어깨에 접촉합니다. 반대 손의 대부분은 가려져 있습니다. 두 남성은 선 자세로 하체가 프레임 밖에 있으며 공중에 뜬 정황은 없습니다. 앰버의 머리와 몸은 베개와 매트리스에 지지되어 있습니다. 수액 용기는 수액대에 걸려 있고 소품들은 협탁에 놓여 있습니다."
       },
       {
        "label": "A",
        "direction": "신부는 가까이 숙인 현우의 얼굴 쪽을 내려다보고, 현우는 고개와 시선을 아래로 떨어뜨립니다. 신부의 두 손이 각각 현우의 양어깨를 향해 뻗어 실제로 감싸 쥐고 있어 위로하는 동작의 대상과 접촉이 명확합니다.",
        "built_space": "왼쪽에 야간 창문과 커튼 한 구역, 왼쪽 아래에 협탁 일부, 뒤쪽에 밝은 병실 벽, 하단에 침대 난간 한 구간이 보입니다. 두 남성은 난간 뒤 침대 옆에 밀착해 있으며, 카메라 각도 변경으로 이전 장면의 나머지 설비가 제외된 것으로 읽힙니다. 환자는 보이지 않습니다. 난간은 두 손을 가리지 않지만 요청한 짧은 가장자리보다 조금 길고 두드러집니다.",
        "entities": "허용된 두 남성만 보입니다. 현우의 앳된 얼굴, 검은 헝클어진 머리, 볼의 상처와 때 묻은 회색 후드가 참고 이미지에 부합합니다. 신부의 나이 든 얼굴, 이마와 눈가 주름, 희끗한 머리 및 얼굴형도 참고 인물과 유사하며 성직자 칼라가 명확합니다. 앰버와 찰리는 프레임 밖이므로 그 상태는 확인할 수 없습니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [],
        "physics": "신부의 앞쪽 손은 현우의 가까운 어깨와 상완을 감싸고, 반대 손은 먼 쪽 어깨에 닿아 있습니다. 손가락의 굽힘과 팔의 연결은 이 동작으로 가능한 형태입니다. 현우는 상체를 신부 쪽으로 기울였으며 두 사람 모두 서 있는 자세로 읽힙니다. 발은 잘려 있지만 부유하거나 지지 없이 매달린 모습은 없습니다. 침대 난간도 아래로 이어지는 지지 부재가 보입니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.533
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.283
   },
   "violations": {
    "B": [
     "[gemini-pro] 지문에서 등장하지 않아야 한다고 명시된 인물(앰버)이 프레임 안에 크게 포함됨 (\"never add a person the shot text does not show\" 및 \"앰버 rests outside the crop\" 지시 위반).",
     "[gemini-pro] 신부가 현우의 어깨를 감싸는 팔의 구조와 손의 위치가 해부학적으로 불가능하고 기형적으로 교차되어 있음.",
     "[gpt-high] 화면 밖에 있어야 하며 이번 숏의 허용 인물이 아닌 앰버의 얼굴과 상체를 노출했습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 283
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지시된 미디엄 샷 프레이밍을 정확히 구현하여 앰버를 프레임 밖으로 제외하고, 두 인물의 밀착된 자세와 양손으로 어깨를 쥔 모습을 완벽하게 묘사했습니다."
   },
   {
    "label": "B",
    "score": 283,
    "verdict_ko": "프레임에서 명시적으로 제외해야 할 인물(앰버)이 그대로 등장하였으며, 지정된 미디엄 샷 스케일을 따르지 않은 치명적인 오류가 있습니다.  ★위반: [gemini-pro] 지문에서 등장하지 않아야 한다고 명시된 인물(앰버)이 프레임 안에 크게 포함됨 (\"never add a person the shot text does not show\" 및 \"앰버 rests outside the crop\" 지시 위반). / [gemini-pro] 신부가 현우의 어깨를 감싸는 팔의 구조와 손의 위치가 해부학적으로 불가능하고 기형적으로 교차되어 있음. / [gpt-high] 화면 밖에 있어야 하며 이번 숏의 허용 인물이 아닌 앰버의 얼굴과 상체를 노출했습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S73sh2_sel.png",
    "asset_id": "74f0f898-0527-4c05-856b-9bb6d8711d36",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 신부: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1402213>",
    "asset_id": "8696070d-ac09-4a5c-95f2-3daf1c015c32",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-e0ea-725d-b27c-d2dc55e12dce",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S73sh2"
  }
 },
 "S73sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:34:03.334936+00:00",
  "fingerprint": "2f31434f7ff68ed610c157ab18f70e384224f81562bc0c377950b7cb1778377f",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S73sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S73sh5_sel.png",
  "source_sha256": "bd0325b4b9f7dbed7741cf33994768de113816016f1f8605ce9cd1ccb13021c6",
  "file": "S73sh5_cine.png",
  "staged_sha256": "4fc8ece948c7d3368cf71396c583ec20e38a1d745de65599ad5cef53ab625830",
  "latency_ms": 10994
 },
 "S73sh11::signage": {
  "fp": "77462a96f9fb24b2",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S73sh11": {
  "input_fingerprint": "128479f3ed3d82f6",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 서로를 양팔로 빈틈없이 끌어안은 현우와 쿠마의 밀착된 전신.\n\nLOCATION (lock): In the open floor area of the clinic patient room near the child's bed, under nighttime interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Clinic entrance (The entrance through which 쿠마 has just arrived) — Seen obliquely behind the embracing pair without specifying the door's position; used as Preserves the direction of his arrival in the wider composition; Clinic floor (Visible beneath both men's feet); used as Makes their distinct weight distribution and complete head-to-foot embrace readable.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the same restrained clinic ambience and controlled contrast, allowing the closeness of the embrace rather than a lighting shift to convey relief.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same clinic bed, surrounding room surfaces, and nighttime lighting. Exclude the stranded truck and outdoor rescue lights.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed remains occupied at night. Charlie retains his accumulated body damage and exposed chest opening. 현우: He stands with both arms raised in an embrace, still dirty and visibly battered. 쿠마: He has entered the health center and stands with both arms raised in an embrace.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 서로를 양팔로 빈틈없이 끌어안은 현우와 쿠마의 밀착된 전신.\n\nLOCATION (lock): In the open floor area of the clinic patient room near the child's bed, under nighttime interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Clinic entrance (The entrance through which 쿠마 has just arrived) — Seen obliquely behind the embracing pair without specifying the door's position; used as Preserves the direction of his arrival in the wider composition; Clinic floor (Visible beneath both men's feet); used as Makes their distinct weight distribution and complete head-to-foot embrace readable.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the same restrained clinic ambience and controlled contrast, allowing the closeness of the embrace rather than a lighting shift to convey relief.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same clinic bed, surrounding room surfaces, and nighttime lighting. Exclude the stranded truck and outdoor rescue lights.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed remains occupied at night. Charlie retains his accumulated body damage and exposed chest opening. 현우: He stands with both arms raised in an embrace, still dirty and visibly battered. 쿠마: He has entered the health center and stands with both arms raised in an embrace.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 서로를 양팔로 빈틈없이 끌어안은 현우와 쿠마의 밀착된 전신.\n\nLOCATION (lock): In the open floor area of the clinic patient room near the child's bed, under nighttime interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Clinic entrance (The entrance through which 쿠마 has just arrived) — Seen obliquely behind the embracing pair without specifying the door's position; used as Preserves the direction of his arrival in the wider composition; Clinic floor (Visible beneath both men's feet); used as Makes their distinct weight distribution and complete head-to-foot embrace readable.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the same restrained clinic ambience and controlled contrast, allowing the closeness of the embrace rather than a lighting shift to convey relief.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same clinic bed, surrounding room surfaces, and nighttime lighting. Exclude the stranded truck and outdoor rescue lights.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed remains occupied at night. Charlie retains his accumulated body damage and exposed chest opening. 현우: He stands with both arms raised in an embrace, still dirty and visibly battered. 쿠마: He has entered the health center and stands with both arms raised in an embrace.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우와 쿠마가 서로를 바라보며 끌어안고 있음. 현우의 시선은 아래를 향하고 쿠마의 시선은 현우의 어깨 너머를 향함.",
    "built_space": "병실 내부. 왼쪽 뒤에 열린 문과 복도가 보이고, 오른쪽에 침대가 있음. 하지만 인물들의 무릎 위까지만 프레이밍되어 프롬프트가 요구한 두 사람의 발 아래 바닥을 확인할 수 없음.",
    "entities": "현우(회색 후드, 상처 입은 얼굴)와 쿠마(파란색 재킷)가 등장함. 레퍼런스와 일치하는 외모. 침대에는 가슴에 심각한 상처가 있는 아이(찰리)가 누워있음.",
    "hard_violations": [
     "[gpt-high] 현우와 쿠마만 허용된 숏에 병상 환자의 얼굴과 몸을 추가로 노출했다."
    ],
    "physics": "두 사람이 끌어안고 있으나 화면 하단이 잘려 발이 바닥을 지탱하는 모습이 보이지 않음. 아이는 침대 위에 올바르게 누워 있음."
   },
   {
    "label": "B",
    "direction": "현우와 쿠마가 병실 한가운데 서서 양팔로 서로를 꽉 끌어안고 있음. 쿠마의 시선은 현우의 어깨 너머 허공을 향함.",
    "built_space": "병실 내부. 프롬프트가 요구한 대로 와이드 샷으로 설정되어 두 사람의 발 아래 바닥 전체가 보임. 왼쪽에 열린 문, 중앙에 창문, 오른쪽에 병상이 알맞은 비율로 배치됨.",
    "entities": "쿠마(얼굴과 체형 일치)와 현우(뒷모습, 회색 후드 일치)의 밀착된 전신이 보임. 침대에 찰리의 뚜렷한 모습은 보이지 않음.",
    "hard_violations": [],
    "physics": "두 사람이 바닥에 안정적으로 두 발을 딛고 서서 체중을 나누며 끌어안고 있는 자세가 자연스럽고 물리적으로 타당함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "요구된 와이드 샷 프레이밍을 정확히 구현하여 두 사람의 밀착된 전신과 발 아래 바닥을 모두 포착했으며, 병실의 배경 요소와 카메라 구도 지침을 가장 충실하게 따랐습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "캐릭터의 외형과 침대의 세부 묘사는 잘 반영되었으나, 가장 중요한 와이드 샷 및 전신(발 아래 바닥 포함) 프레이밍 지시를 완전히 무시하고 무릎 위에서 잘린 샷을 생성하여 우선순위 2번을 크게 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우와 쿠마가 서로를 바라보며 끌어안고 있음. 현우의 시선은 아래를 향하고 쿠마의 시선은 현우의 어깨 너머를 향함.",
        "built_space": "병실 내부. 왼쪽 뒤에 열린 문과 복도가 보이고, 오른쪽에 침대가 있음. 하지만 인물들의 무릎 위까지만 프레이밍되어 프롬프트가 요구한 두 사람의 발 아래 바닥을 확인할 수 없음.",
        "entities": "현우(회색 후드, 상처 입은 얼굴)와 쿠마(파란색 재킷)가 등장함. 레퍼런스와 일치하는 외모. 침대에는 가슴에 심각한 상처가 있는 아이(찰리)가 누워있음.",
        "hard_violations": [],
        "physics": "두 사람이 끌어안고 있으나 화면 하단이 잘려 발이 바닥을 지탱하는 모습이 보이지 않음. 아이는 침대 위에 올바르게 누워 있음."
       },
       {
        "label": "B",
        "direction": "현우와 쿠마가 병실 한가운데 서서 양팔로 서로를 꽉 끌어안고 있음. 쿠마의 시선은 현우의 어깨 너머 허공을 향함.",
        "built_space": "병실 내부. 프롬프트가 요구한 대로 와이드 샷으로 설정되어 두 사람의 발 아래 바닥 전체가 보임. 왼쪽에 열린 문, 중앙에 창문, 오른쪽에 병상이 알맞은 비율로 배치됨.",
        "entities": "쿠마(얼굴과 체형 일치)와 현우(뒷모습, 회색 후드 일치)의 밀착된 전신이 보임. 침대에 찰리의 뚜렷한 모습은 보이지 않음.",
        "hard_violations": [],
        "physics": "두 사람이 바닥에 안정적으로 두 발을 딛고 서서 체중을 나누며 끌어안고 있는 자세가 자연스럽고 물리적으로 타당함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "요구된 와이드 샷 프레이밍을 정확히 구현하여 두 사람의 밀착된 전신과 발 아래 바닥을 모두 포착했으며, 병실의 배경 요소와 카메라 구도 지침을 가장 충실하게 따랐습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "캐릭터의 외형과 침대의 세부 묘사는 잘 반영되었으나, 가장 중요한 와이드 샷 및 전신(발 아래 바닥 포함) 프레이밍 지시를 완전히 무시하고 무릎 위에서 잘린 샷을 생성하여 우선순위 2번을 크게 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우와 쿠마가 서로를 바라보며 끌어안고 있음. 현우의 시선은 아래를 향하고 쿠마의 시선은 현우의 어깨 너머를 향함.",
        "built_space": "병실 내부. 왼쪽 뒤에 열린 문과 복도가 보이고, 오른쪽에 침대가 있음. 하지만 인물들의 무릎 위까지만 프레이밍되어 프롬프트가 요구한 두 사람의 발 아래 바닥을 확인할 수 없음.",
        "entities": "현우(회색 후드, 상처 입은 얼굴)와 쿠마(파란색 재킷)가 등장함. 레퍼런스와 일치하는 외모. 침대에는 가슴에 심각한 상처가 있는 아이(찰리)가 누워있음.",
        "hard_violations": [],
        "physics": "두 사람이 끌어안고 있으나 화면 하단이 잘려 발이 바닥을 지탱하는 모습이 보이지 않음. 아이는 침대 위에 올바르게 누워 있음."
       },
       {
        "label": "B",
        "direction": "현우와 쿠마가 병실 한가운데 서서 양팔로 서로를 꽉 끌어안고 있음. 쿠마의 시선은 현우의 어깨 너머 허공을 향함.",
        "built_space": "병실 내부. 프롬프트가 요구한 대로 와이드 샷으로 설정되어 두 사람의 발 아래 바닥 전체가 보임. 왼쪽에 열린 문, 중앙에 창문, 오른쪽에 병상이 알맞은 비율로 배치됨.",
        "entities": "쿠마(얼굴과 체형 일치)와 현우(뒷모습, 회색 후드 일치)의 밀착된 전신이 보임. 침대에 찰리의 뚜렷한 모습은 보이지 않음.",
        "hard_violations": [],
        "physics": "두 사람이 바닥에 안정적으로 두 발을 딛고 서서 체중을 나누며 끌어안고 있는 자세가 자연스럽고 물리적으로 타당함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "두 사람의 밀착 포옹을 머리부터 신발까지 담고 출입구와 바닥도 보여 주지만, 비어 있는 병상은 환자가 계속 누워 있어야 한다는 상태 조건과 어긋난다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "현우의 상처와 포옹은 잘 보이나, 허용되지 않은 세 번째 인물을 노출하고 두 사람의 하체와 발을 잘라 핵심 전신 와이드 구도를 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 등을 카메라 쪽으로 둔 채 얼굴을 쿠마의 어깨 쪽에 묻고 있다. 쿠마의 눈은 현우의 어깨 너머 왼쪽 아래를 향한다. 두 사람의 팔은 상대의 등과 몸통을 감싸며, 카메라를 응시하지 않는다. 무기나 방향을 확인할 휴대 소품은 없다.",
        "built_space": "두 사람은 병상 왼쪽의 열린 바닥에 서 있다. 뒤 왼쪽에 열린 출입문 하나, 오른쪽에 병상 하나와 의료용 벽면 패널 하나, 수액대 하나, 의자 하나가 보인다. 왼쪽에는 싱크대와 수납장이 있고, 뒤쪽에는 커튼이 달린 야간 창이 있으며 문 너머에도 창 하나가 보인다. 출입구는 인물 뒤에 비스듬히 보이고 두 사람의 발 아래 바닥도 충분히 드러난다. 다만 병상 전체가 사실상 비어 있어 병상 점유 상태를 유지하지 못했다. 참조의 밝은 벽, 금속 창틀, 커튼과 크림색 병상 난간 계열은 이어진다.",
        "entities": "보이는 사람은 젊은 남성 두 명뿐이다. 현우의 헝클어진 검은 머리와 심하게 더러워진 회색 후드는 이전 장면과 맞지만 얼굴은 가려져 정확한 얼굴 일치와 얼굴 상처는 확인하기 어렵다. 쿠마는 짧은 검은 머리의 젊은 성인 남성으로 참조와 대체로 부합하지만 남색 티셔츠 대신 카키색 겉옷이 두드러진다. 두 사람의 구체적인 혈통은 외관만으로 확인할 수 없다. 찰리나 흉부 손상은 보이지 않으며, 빈 침대 때문에 단순히 환자가 화면 밖에 있다고 보기도 어렵다. 왼쪽 게시물에는 글자 같은 흔적이 있으나 확실히 읽히는 문구는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "두 사람 모두 신발이 바닥에 닿아 있고, 발을 조금 어긋나게 놓아 서로 기대는 무게를 지탱한다. 보이는 손들은 상대의 등과 옆구리에 접촉하며 포옹하는 팔의 굽힘도 가능하다. 일부 팔은 몸 뒤에 가려져 있으나 추가 팔다리나 불가능한 연결은 보이지 않는다. 침대는 바퀴로, 수납장과 의자는 바닥으로 지지된다. 지지 없이 떠 있는 신체나 물체는 없다."
       },
       {
        "label": "B",
        "direction": "현우는 눈을 감거나 아래로 내린 채 쿠마의 어깨에 얼굴을 붙이고 있다. 쿠마는 현우의 머리 너머 화면 오른쪽을 향한다. 두 사람의 팔과 손은 상대의 등과 옆구리를 감싼다. 병상의 세 번째 인물은 얼굴을 위로 향한 채 누워 있으며 포옹에 참여하지 않는다.",
        "built_space": "포옹하는 두 사람은 병상 앞 열린 공간에 있지만 화면 아래에서 다리와 발이 잘린다. 뒤 왼쪽에 열린 병실 출입문 하나와 그 너머 복도 문 하나, 왼쪽 벽과 뒤쪽 벽에 각각 창 구역 하나가 보인다. 오른쪽에는 병상 하나, 의료용 벽면 패널 하나, 의자 하나와 일부 협탁이 있다. 야간 창과 밝은 병실 벽은 참조의 분위기를 잇지만, 구도는 요청한 전신 와이드보다 훨씬 가깝다. 병상은 점유되어 있으나 그 환자를 화면에 드러낸 것은 인물 제한과 충돌한다.",
        "entities": "현우는 앳된 얼굴, 헝클어진 검은 머리, 뺨의 상처와 때 묻은 회색 후드를 보여 참조와 잘 연결된다. 쿠마는 짧은 검은 머리의 젊은 남성이지만 얼굴 상당 부분이 가려져 정확한 일치는 제한적으로만 확인된다. 남색 겉옷은 참조의 남색 티셔츠와 다르다. 두 사람의 구체적인 혈통은 외관만으로 확정할 수 없다. 오른쪽 병상에는 노출된 흉부 상처를 가진 세 번째 인물이 명확히 보인다. 이는 찰리의 상태 설명에는 대응하지만, 이 숏에는 현우와 쿠마만 보여야 한다는 명시적 제한을 어긴다. 복도 게시물의 문구는 식별되지 않는다.",
        "hard_violations": [
         "현우와 쿠마만 허용된 숏에 병상 환자의 얼굴과 몸을 추가로 노출했다."
        ],
        "physics": "포옹하는 팔과 손은 상대의 몸에 닿아 있으며 어깨와 팔꿈치의 연결도 자연스럽다. 두 사람의 발은 프레임 밖이므로 접지와 체중 분배는 확인할 수 없지만, 공중에 떠 있다고 판단할 시각적 근거는 없다. 병상 환자는 머리를 베개에, 몸통과 팔을 매트리스에 기대어 지지받는다. 침대와 주변 가구도 바닥 위에 놓여 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "두 사람의 밀착 포옹을 머리부터 신발까지 담고 출입구와 바닥도 보여 주지만, 비어 있는 병상은 환자가 계속 누워 있어야 한다는 상태 조건과 어긋난다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "현우의 상처와 포옹은 잘 보이나, 허용되지 않은 세 번째 인물을 노출하고 두 사람의 하체와 발을 잘라 핵심 전신 와이드 구도를 위반한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 등을 카메라 쪽으로 둔 채 얼굴을 쿠마의 어깨 쪽에 묻고 있다. 쿠마의 눈은 현우의 어깨 너머 왼쪽 아래를 향한다. 두 사람의 팔은 상대의 등과 몸통을 감싸며, 카메라를 응시하지 않는다. 무기나 방향을 확인할 휴대 소품은 없다.",
        "built_space": "두 사람은 병상 왼쪽의 열린 바닥에 서 있다. 뒤 왼쪽에 열린 출입문 하나, 오른쪽에 병상 하나와 의료용 벽면 패널 하나, 수액대 하나, 의자 하나가 보인다. 왼쪽에는 싱크대와 수납장이 있고, 뒤쪽에는 커튼이 달린 야간 창이 있으며 문 너머에도 창 하나가 보인다. 출입구는 인물 뒤에 비스듬히 보이고 두 사람의 발 아래 바닥도 충분히 드러난다. 다만 병상 전체가 사실상 비어 있어 병상 점유 상태를 유지하지 못했다. 참조의 밝은 벽, 금속 창틀, 커튼과 크림색 병상 난간 계열은 이어진다.",
        "entities": "보이는 사람은 젊은 남성 두 명뿐이다. 현우의 헝클어진 검은 머리와 심하게 더러워진 회색 후드는 이전 장면과 맞지만 얼굴은 가려져 정확한 얼굴 일치와 얼굴 상처는 확인하기 어렵다. 쿠마는 짧은 검은 머리의 젊은 성인 남성으로 참조와 대체로 부합하지만 남색 티셔츠 대신 카키색 겉옷이 두드러진다. 두 사람의 구체적인 혈통은 외관만으로 확인할 수 없다. 찰리나 흉부 손상은 보이지 않으며, 빈 침대 때문에 단순히 환자가 화면 밖에 있다고 보기도 어렵다. 왼쪽 게시물에는 글자 같은 흔적이 있으나 확실히 읽히는 문구는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "두 사람 모두 신발이 바닥에 닿아 있고, 발을 조금 어긋나게 놓아 서로 기대는 무게를 지탱한다. 보이는 손들은 상대의 등과 옆구리에 접촉하며 포옹하는 팔의 굽힘도 가능하다. 일부 팔은 몸 뒤에 가려져 있으나 추가 팔다리나 불가능한 연결은 보이지 않는다. 침대는 바퀴로, 수납장과 의자는 바닥으로 지지된다. 지지 없이 떠 있는 신체나 물체는 없다."
       },
       {
        "label": "A",
        "direction": "현우는 눈을 감거나 아래로 내린 채 쿠마의 어깨에 얼굴을 붙이고 있다. 쿠마는 현우의 머리 너머 화면 오른쪽을 향한다. 두 사람의 팔과 손은 상대의 등과 옆구리를 감싼다. 병상의 세 번째 인물은 얼굴을 위로 향한 채 누워 있으며 포옹에 참여하지 않는다.",
        "built_space": "포옹하는 두 사람은 병상 앞 열린 공간에 있지만 화면 아래에서 다리와 발이 잘린다. 뒤 왼쪽에 열린 병실 출입문 하나와 그 너머 복도 문 하나, 왼쪽 벽과 뒤쪽 벽에 각각 창 구역 하나가 보인다. 오른쪽에는 병상 하나, 의료용 벽면 패널 하나, 의자 하나와 일부 협탁이 있다. 야간 창과 밝은 병실 벽은 참조의 분위기를 잇지만, 구도는 요청한 전신 와이드보다 훨씬 가깝다. 병상은 점유되어 있으나 그 환자를 화면에 드러낸 것은 인물 제한과 충돌한다.",
        "entities": "현우는 앳된 얼굴, 헝클어진 검은 머리, 뺨의 상처와 때 묻은 회색 후드를 보여 참조와 잘 연결된다. 쿠마는 짧은 검은 머리의 젊은 남성이지만 얼굴 상당 부분이 가려져 정확한 일치는 제한적으로만 확인된다. 남색 겉옷은 참조의 남색 티셔츠와 다르다. 두 사람의 구체적인 혈통은 외관만으로 확정할 수 없다. 오른쪽 병상에는 노출된 흉부 상처를 가진 세 번째 인물이 명확히 보인다. 이는 찰리의 상태 설명에는 대응하지만, 이 숏에는 현우와 쿠마만 보여야 한다는 명시적 제한을 어긴다. 복도 게시물의 문구는 식별되지 않는다.",
        "hard_violations": [
         "현우와 쿠마만 허용된 숏에 병상 환자의 얼굴과 몸을 추가로 노출했다."
        ],
        "physics": "포옹하는 팔과 손은 상대의 몸에 닿아 있으며 어깨와 팔꿈치의 연결도 자연스럽다. 두 사람의 발은 프레임 밖이므로 접지와 체중 분배는 확인할 수 없지만, 공중에 떠 있다고 판단할 시각적 근거는 없다. 병상 환자는 머리를 베개에, 몸통과 팔을 매트리스에 기대어 지지받는다. 침대와 주변 가구도 바닥 위에 놓여 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.73,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.48,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gpt-high] 현우와 쿠마만 허용된 숏에 병상 환자의 얼굴과 몸을 추가로 노출했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 480
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "요구된 와이드 샷 프레이밍을 정확히 구현하여 두 사람의 밀착된 전신과 발 아래 바닥을 모두 포착했으며, 병실의 배경 요소와 카메라 구도 지침을 가장 충실하게 따랐습니다."
   },
   {
    "label": "A",
    "score": 480,
    "verdict_ko": "캐릭터의 외형과 침대의 세부 묘사는 잘 반영되었으나, 가장 중요한 와이드 샷 및 전신(발 아래 바닥 포함) 프레이밍 지시를 완전히 무시하고 무릎 위에서 잘린 샷을 생성하여 우선순위 2번을 크게 위반했습니다.  ★위반: [gpt-high] 현우와 쿠마만 허용된 숏에 병상 환자의 얼굴과 몸을 추가로 노출했다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S73sh5_sel.png",
    "asset_id": "680d1747-0dce-4f56-af2e-87592d281da0",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 쿠마: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1218434>",
    "asset_id": "aed54006-63aa-42f8-8d50-5ff0f4385074",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-e2a7-7cd0-8d34-27b593237407",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S73sh5"
  }
 },
 "S73sh11::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:35:22.341642+00:00",
  "fingerprint": "4a26624e2017a28e821f57b3fb52efcd01d2449a4881497a34644973681a27ec",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S73sh11_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S73sh11_sel.png",
  "source_sha256": "86b3ccb2a9dd9bf32279bc3e3e1787e448b4a7f6263d681bf129ee6c87b9d691",
  "file": "S73sh11_cine.png",
  "staged_sha256": "3c3a08d3ed1211827f40772d2511f1d5aa7d4bd97fc6824c9910fcb43352fc5b",
  "latency_ms": 10161
 },
 "S74sh8::signage": {
  "fp": "ee8fec592733164f",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S74sh8": {
  "input_fingerprint": "ff6ad2b0ba579305",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 접시 위에 담긴 피부가 벗겨지고 흉측한 돌연변이 생선회 덩어리 클로즈업.\n\nLOCATION (lock): On a dining table at a street-side seafood eatery in the quiet harbor village at night. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Plate of mutant fish (Holding skinned, misshapen pieces of raw mutant fish) — The serving side is visible obliquely from above; used as Central focal detail, kept small enough to retain tabletop context; Restaurant tabletop (Supporting the served plate) — The top surface surrounds the plate; used as Scale reference and uncluttered framing margin.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained nighttime ambient illumination and controlled contrast reveal the fish's malformed contours without theatrical color or added atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A plate of sliced, radiation-mutated fish has been served at the harbor street restaurant at night.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 접시 위에 담긴 피부가 벗겨지고 흉측한 돌연변이 생선회 덩어리 클로즈업.\n\nLOCATION (lock): On a dining table at a street-side seafood eatery in the quiet harbor village at night. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Plate of mutant fish (Holding skinned, misshapen pieces of raw mutant fish) — The serving side is visible obliquely from above; used as Central focal detail, kept small enough to retain tabletop context; Restaurant tabletop (Supporting the served plate) — The top surface surrounds the plate; used as Scale reference and uncluttered framing margin.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained nighttime ambient illumination and controlled contrast reveal the fish's malformed contours without theatrical color or added atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A plate of sliced, radiation-mutated fish has been served at the harbor street restaurant at night.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 접시 위에 담긴 피부가 벗겨지고 흉측한 돌연변이 생선회 덩어리 클로즈업.\n\nLOCATION (lock): On a dining table at a street-side seafood eatery in the quiet harbor village at night. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Plate of mutant fish (Holding skinned, misshapen pieces of raw mutant fish) — The serving side is visible obliquely from above; used as Central focal detail, kept small enough to retain tabletop context; Restaurant tabletop (Supporting the served plate) — The top surface surrounds the plate; used as Scale reference and uncluttered framing margin.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained nighttime ambient illumination and controlled contrast reveal the fish's malformed contours without theatrical color or added atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A plate of sliced, radiation-mutated fish has been served at the harbor street restaurant at night.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "특정 대상을 향한 시선이나 방향성 없음.",
    "built_space": "야간 항구 식당. 금속 테이블 위에 접시가 위치하며, 배경에 포장마차와 항구의 흐릿한 야간 조명이 보임.",
    "entities": "생선회 접시(껍질과 비늘이 제거되지 않고 평범한 생선 조각처럼 보임), 젊은 남성의 손과 팔(샷 텍스트에 명시되지 않은 인물).",
    "hard_violations": [
     "[gemini-pro] 샷 텍스트에 언급되지 않은 인물(손과 팔)이 화면에 임의로 추가됨",
     "[gpt-high] 사람을 보여 주지 않는 샷 지시와 명시적 인물 추가 금지에도 접시 뒤에 사람의 상반신과 팔, 손을 넣었다."
    ],
    "physics": "손과 접시 모두 테이블 표면에 물리적으로 자연스럽게 지지되어 있음."
   },
   {
    "label": "B",
    "direction": "특정 대상을 향한 시선이나 방향성 없음.",
    "built_space": "야간 항구 식당. 나무 테이블 위에 접시와 간장 종지가 놓여 있고, 배경에 어선과 야외 식당 구조물이 배치됨.",
    "entities": "흉측한 돌연변이 생선회 덩어리(두껍고 기괴한 형태), 굵고 주름진 노인의 손(샷 텍스트에 없으며 20대 설정과 어긋남).",
    "hard_violations": [
     "[gemini-pro] 샷 텍스트에 언급되지 않은 인물(손)이 화면에 임의로 추가됨",
     "[gemini-pro] 화면에 등장한 신체 일부(손)가 유일한 인물 설정인 20대 남성의 나이와 명백히 모순되는 노인의 손임",
     "[gpt-high] 사람을 보여 주지 않는 샷 지시와 명시적 인물 추가 금지에도 왼쪽에 사람의 팔과 손을 넣었다.",
     "[gpt-high] 접시와 식탁만 지정되고 추가 발명이 금지된 화면에 작은 그릇 두 개와 금속 식기를 추가했다."
    ],
    "physics": "손, 접시, 젓가락, 종지 등이 테이블 표면에 안정적으로 지지되어 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "샷 텍스트에 없는 인물의 손이 등장하여 지침을 위반했으며, 생선회에 비늘이 그대로 있어 '피부가 벗겨진' 돌연변이라는 묘사를 제대로 살리지 못했습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "돌연변이 생선의 형태는 텍스트에 부합하나, 샷 텍스트에 없는 인물의 손이 추가되었고 20대 설정과 전혀 맞지 않는 노인의 손이 그려져 심각한 오류가 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "특정 대상을 향한 시선이나 방향성 없음.",
        "built_space": "야간 항구 식당. 금속 테이블 위에 접시가 위치하며, 배경에 포장마차와 항구의 흐릿한 야간 조명이 보임.",
        "entities": "생선회 접시(껍질과 비늘이 제거되지 않고 평범한 생선 조각처럼 보임), 젊은 남성의 손과 팔(샷 텍스트에 명시되지 않은 인물).",
        "hard_violations": [
         "샷 텍스트에 언급되지 않은 인물(손과 팔)이 화면에 임의로 추가됨"
        ],
        "physics": "손과 접시 모두 테이블 표면에 물리적으로 자연스럽게 지지되어 있음."
       },
       {
        "label": "B",
        "direction": "특정 대상을 향한 시선이나 방향성 없음.",
        "built_space": "야간 항구 식당. 나무 테이블 위에 접시와 간장 종지가 놓여 있고, 배경에 어선과 야외 식당 구조물이 배치됨.",
        "entities": "흉측한 돌연변이 생선회 덩어리(두껍고 기괴한 형태), 굵고 주름진 노인의 손(샷 텍스트에 없으며 20대 설정과 어긋남).",
        "hard_violations": [
         "샷 텍스트에 언급되지 않은 인물(손)이 화면에 임의로 추가됨",
         "화면에 등장한 신체 일부(손)가 유일한 인물 설정인 20대 남성의 나이와 명백히 모순되는 노인의 손임"
        ],
        "physics": "손, 접시, 젓가락, 종지 등이 테이블 표면에 안정적으로 지지되어 있음."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "샷 텍스트에 없는 인물의 손이 등장하여 지침을 위반했으며, 생선회에 비늘이 그대로 있어 '피부가 벗겨진' 돌연변이라는 묘사를 제대로 살리지 못했습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "돌연변이 생선의 형태는 텍스트에 부합하나, 샷 텍스트에 없는 인물의 손이 추가되었고 20대 설정과 전혀 맞지 않는 노인의 손이 그려져 심각한 오류가 발생했습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "특정 대상을 향한 시선이나 방향성 없음.",
        "built_space": "야간 항구 식당. 금속 테이블 위에 접시가 위치하며, 배경에 포장마차와 항구의 흐릿한 야간 조명이 보임.",
        "entities": "생선회 접시(껍질과 비늘이 제거되지 않고 평범한 생선 조각처럼 보임), 젊은 남성의 손과 팔(샷 텍스트에 명시되지 않은 인물).",
        "hard_violations": [
         "샷 텍스트에 언급되지 않은 인물(손과 팔)이 화면에 임의로 추가됨"
        ],
        "physics": "손과 접시 모두 테이블 표면에 물리적으로 자연스럽게 지지되어 있음."
       },
       {
        "label": "B",
        "direction": "특정 대상을 향한 시선이나 방향성 없음.",
        "built_space": "야간 항구 식당. 나무 테이블 위에 접시와 간장 종지가 놓여 있고, 배경에 어선과 야외 식당 구조물이 배치됨.",
        "entities": "흉측한 돌연변이 생선회 덩어리(두껍고 기괴한 형태), 굵고 주름진 노인의 손(샷 텍스트에 없으며 20대 설정과 어긋남).",
        "hard_violations": [
         "샷 텍스트에 언급되지 않은 인물(손)이 화면에 임의로 추가됨",
         "화면에 등장한 신체 일부(손)가 유일한 인물 설정인 20대 남성의 나이와 명백히 모순되는 노인의 손임"
        ],
        "physics": "손, 접시, 젓가락, 종지 등이 테이블 표면에 안정적으로 지지되어 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "울퉁불퉁하고 흉측한 돌연변이 생선 덩어리는 더 충실하지만, 지시되지 않은 사람과 식기를 추가했고 껍질도 상당 부분 남아 있다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "접시와 식탁의 근접 구도는 맞지만, 불필요한 사람의 상반신이 크게 들어오며 생선도 피부를 벗긴 돌연변이보다 껍질 붙은 일반 토막에 가깝다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "접시의 담는 면이 카메라 쪽으로 비스듬히 위에서 보인다. 생선 덩어리는 접시 중앙에 모여 있다. 왼쪽 사람은 얼굴이 잘려 시선 방향을 확인할 수 없으며, 손은 생선을 집지 않고 식탁에 놓여 있다.",
        "built_space": "앞쪽 나무 식탁 위에 큰 접시 하나, 뒤쪽 작은 그릇 하나, 오른쪽 가장자리에 잘린 작은 그릇 하나와 금속 식기 두 가닥이 보인다. 배경에는 별도 식탁과 차양 기둥, 조명, 판매대 및 정박한 배들이 보인다. 접시 주변 식탁 여백은 있지만, 요구된 단순한 식탁 중심 화면보다 배경과 부가물이 많다. 장소 참조 사진은 없어 특정 구조의 일치 여부는 판단할 수 없다.",
        "entities": "접시에는 젖은 생선 덩어리 여러 개가 있고, 혹처럼 부푼 형태가 돌연변이의 흉측함을 드러낸다. 다만 회갈색 외피가 넓게 남아 있어 피부가 벗겨진 생선회라는 조건에는 미달한다. 왼쪽에는 검은 소매와 손이 추가되어 있다. 얼굴이 없어 쿠마의 나이·혈통·머리카락은 확인할 수 없으며, 이 샷에는 애초에 사람이 요구되지 않는다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "사람을 보여 주지 않는 샷 지시와 명시적 인물 추가 금지에도 왼쪽에 사람의 팔과 손을 넣었다.",
         "접시와 식탁만 지정되고 추가 발명이 금지된 화면에 작은 그릇 두 개와 금속 식기를 추가했다."
        ],
        "physics": "생선 덩어리는 접시 바닥이나 다른 덩어리에 받쳐져 있고, 접시는 식탁에 놓여 있다. 손과 작은 그릇, 금속 식기에도 식탁의 지지가 보인다. 공중에 떠 있거나 지지 없이 매달린 물체는 없다."
       },
       {
        "label": "B",
        "direction": "접시의 담는 면을 비스듬히 위에서 내려다보며, 절단된 생선 살과 껍질 면이 함께 카메라에 드러난다. 뒤쪽 사람의 얼굴은 화면 밖이므로 시선은 확인할 수 없다. 손은 접시 뒤 가장자리 가까이에 놓여 있지만 생선을 집거나 가리키지 않는다.",
        "built_space": "금속 식탁 하나 위에 큰 접시 하나가 있으며 다른 전경 식기는 없다. 접시 주변으로 식탁 표면이 충분히 보인다. 그러나 왼쪽 위를 사람의 상반신과 팔이 크게 차지하고, 상단에는 항구의 포장면과 판매대, 배들이 넓게 들어와 접시만을 다루는 클로즈업의 집중을 흐린다.",
        "entities": "접시에는 두껍고 불규칙하게 썬 생선 토막들이 있다. 밝은 생살은 보이지만 여러 조각의 가장자리에 비늘 무늬가 있는 껍질이 뚜렷하게 남아 있고, 돌연변이 특유의 기형성은 약하다. 검은 겉옷을 입은 사람의 상반신과 한 손이 추가되었다. 얼굴과 머리카락이 보이지 않아 지정 인물의 신원 특성은 확인할 수 없다. 배경 표지는 흐려 읽을 수 없다.",
        "hard_violations": [
         "사람을 보여 주지 않는 샷 지시와 명시적 인물 추가 금지에도 접시 뒤에 사람의 상반신과 팔, 손을 넣었다."
        ],
        "physics": "생선 토막들은 접시와 서로에게 받쳐져 있고, 접시는 금속 식탁에 안정적으로 놓여 있다. 사람의 팔은 식탁에 기대어 있으며 손도 식탁 높이에 놓여 있다. 지지 없는 부유나 불가능한 동작은 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "울퉁불퉁하고 흉측한 돌연변이 생선 덩어리는 더 충실하지만, 지시되지 않은 사람과 식기를 추가했고 껍질도 상당 부분 남아 있다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "접시와 식탁의 근접 구도는 맞지만, 불필요한 사람의 상반신이 크게 들어오며 생선도 피부를 벗긴 돌연변이보다 껍질 붙은 일반 토막에 가깝다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "접시의 담는 면이 카메라 쪽으로 비스듬히 위에서 보인다. 생선 덩어리는 접시 중앙에 모여 있다. 왼쪽 사람은 얼굴이 잘려 시선 방향을 확인할 수 없으며, 손은 생선을 집지 않고 식탁에 놓여 있다.",
        "built_space": "앞쪽 나무 식탁 위에 큰 접시 하나, 뒤쪽 작은 그릇 하나, 오른쪽 가장자리에 잘린 작은 그릇 하나와 금속 식기 두 가닥이 보인다. 배경에는 별도 식탁과 차양 기둥, 조명, 판매대 및 정박한 배들이 보인다. 접시 주변 식탁 여백은 있지만, 요구된 단순한 식탁 중심 화면보다 배경과 부가물이 많다. 장소 참조 사진은 없어 특정 구조의 일치 여부는 판단할 수 없다.",
        "entities": "접시에는 젖은 생선 덩어리 여러 개가 있고, 혹처럼 부푼 형태가 돌연변이의 흉측함을 드러낸다. 다만 회갈색 외피가 넓게 남아 있어 피부가 벗겨진 생선회라는 조건에는 미달한다. 왼쪽에는 검은 소매와 손이 추가되어 있다. 얼굴이 없어 쿠마의 나이·혈통·머리카락은 확인할 수 없으며, 이 샷에는 애초에 사람이 요구되지 않는다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "사람을 보여 주지 않는 샷 지시와 명시적 인물 추가 금지에도 왼쪽에 사람의 팔과 손을 넣었다.",
         "접시와 식탁만 지정되고 추가 발명이 금지된 화면에 작은 그릇 두 개와 금속 식기를 추가했다."
        ],
        "physics": "생선 덩어리는 접시 바닥이나 다른 덩어리에 받쳐져 있고, 접시는 식탁에 놓여 있다. 손과 작은 그릇, 금속 식기에도 식탁의 지지가 보인다. 공중에 떠 있거나 지지 없이 매달린 물체는 없다."
       },
       {
        "label": "A",
        "direction": "접시의 담는 면을 비스듬히 위에서 내려다보며, 절단된 생선 살과 껍질 면이 함께 카메라에 드러난다. 뒤쪽 사람의 얼굴은 화면 밖이므로 시선은 확인할 수 없다. 손은 접시 뒤 가장자리 가까이에 놓여 있지만 생선을 집거나 가리키지 않는다.",
        "built_space": "금속 식탁 하나 위에 큰 접시 하나가 있으며 다른 전경 식기는 없다. 접시 주변으로 식탁 표면이 충분히 보인다. 그러나 왼쪽 위를 사람의 상반신과 팔이 크게 차지하고, 상단에는 항구의 포장면과 판매대, 배들이 넓게 들어와 접시만을 다루는 클로즈업의 집중을 흐린다.",
        "entities": "접시에는 두껍고 불규칙하게 썬 생선 토막들이 있다. 밝은 생살은 보이지만 여러 조각의 가장자리에 비늘 무늬가 있는 껍질이 뚜렷하게 남아 있고, 돌연변이 특유의 기형성은 약하다. 검은 겉옷을 입은 사람의 상반신과 한 손이 추가되었다. 얼굴과 머리카락이 보이지 않아 지정 인물의 신원 특성은 확인할 수 없다. 배경 표지는 흐려 읽을 수 없다.",
        "hard_violations": [
         "사람을 보여 주지 않는 샷 지시와 명시적 인물 추가 금지에도 접시 뒤에 사람의 상반신과 팔, 손을 넣었다."
        ],
        "physics": "생선 토막들은 접시와 서로에게 받쳐져 있고, 접시는 금속 식탁에 안정적으로 놓여 있다. 사람의 팔은 식탁에 기대어 있으며 손도 식탁 높이에 놓여 있다. 지지 없는 부유나 불가능한 동작은 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.667,
    "B": 1.667
   },
   "adjusted": {
    "A": 1.417,
    "B": 1.417
   },
   "violations": {
    "A": [
     "[gemini-pro] 샷 텍스트에 언급되지 않은 인물(손과 팔)이 화면에 임의로 추가됨",
     "[gpt-high] 사람을 보여 주지 않는 샷 지시와 명시적 인물 추가 금지에도 접시 뒤에 사람의 상반신과 팔, 손을 넣었다."
    ],
    "B": [
     "[gemini-pro] 샷 텍스트에 언급되지 않은 인물(손)이 화면에 임의로 추가됨",
     "[gemini-pro] 화면에 등장한 신체 일부(손)가 유일한 인물 설정인 20대 남성의 나이와 명백히 모순되는 노인의 손임",
     "[gpt-high] 사람을 보여 주지 않는 샷 지시와 명시적 인물 추가 금지에도 왼쪽에 사람의 팔과 손을 넣었다.",
     "[gpt-high] 접시와 식탁만 지정되고 추가 발명이 금지된 화면에 작은 그릇 두 개와 금속 식기를 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1417,
   "B": 1417
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1417,
    "verdict_ko": "샷 텍스트에 없는 인물의 손이 등장하여 지침을 위반했으며, 생선회에 비늘이 그대로 있어 '피부가 벗겨진' 돌연변이라는 묘사를 제대로 살리지 못했습니다.  ★위반: [gemini-pro] 샷 텍스트에 언급되지 않은 인물(손과 팔)이 화면에 임의로 추가됨 / [gpt-high] 사람을 보여 주지 않는 샷 지시와 명시적 인물 추가 금지에도 접시 뒤에 사람의 상반신과 팔, 손을 넣었다."
   },
   {
    "label": "B",
    "score": 1417,
    "verdict_ko": "돌연변이 생선의 형태는 텍스트에 부합하나, 샷 텍스트에 없는 인물의 손이 추가되었고 20대 설정과 전혀 맞지 않는 노인의 손이 그려져 심각한 오류가 발생했습니다.  ★위반: [gemini-pro] 샷 텍스트에 언급되지 않은 인물(손)이 화면에 임의로 추가됨 / [gemini-pro] 화면에 등장한 신체 일부(손)가 유일한 인물 설정인 20대 남성의 나이와 명백히 모순되는 노인의 손임 / [gpt-high] 사람을 보여 주지 않는 샷 지시와 명시적 인물 추가 금지에도 왼쪽에 사람의 팔과 손을 넣었다. / [gpt-high] 접시와 식탁만 지정되고 추가 발명이 금지된 화면에 작은 그릇 두 개와 금속 식기를 추가했다."
   }
  ],
  "refs": [],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-e45e-7c3d-9e5f-eaa297ea5de1",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S74sh8::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:36:35.786166+00:00",
  "fingerprint": "bfefad1326ed0bcdaa2aabafb75b5fc9cf5203d5936d904b9c7f08a1d0fdd3fb",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S74sh8_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S74sh8_sel.png",
  "source_sha256": "8f5948a3673b92b0b26c031177f9b8115aeb21647bdce98fbf6d5cbe853c9915",
  "file": "S74sh8_cine.png",
  "staged_sha256": "b3da1480d272cc65bc76b0bd9b419938b37c6d25f511dc5bad6c521b01ae42f3",
  "latency_ms": 9650
 },
 "S74sh9::signage": {
  "fp": "8bd30630902575c7",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S74sh9": {
  "input_fingerprint": "4f4863cc5cf6b9ce",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 젓가락으로 생선회를 허공에 치켜든 채 태연한 표정으로 앞을 주시하는 쿠마의 상체.\n\nLOCATION (lock): At the outdoor table of the harbor village's street-side seafood eatery, near the nighttime fish trade. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Raised chopsticks and fish (A piece of raw fish is held aloft) — The chopsticks angle upward beside 쿠마's face rather than toward the lens; used as Small foreground-to-midground link between the plate and 쿠마's calm expression; Restaurant table (The served plate remains on the table) — Only the near tabletop and part of the plate are visible along the lower edge; used as Continuity anchor for the completed rise.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding nighttime ambient illumination and restrained contrast, giving equal readability to 쿠마's expression and the lifted food.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The plate of sliced, radiation-mutated fish remains served at the harbor street restaurant. 쿠마: He remains at the harbor street restaurant.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 쿠마 right now, so 쿠마's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 쿠마: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 젓가락으로 생선회를 허공에 치켜든 채 태연한 표정으로 앞을 주시하는 쿠마의 상체.\n\nLOCATION (lock): At the outdoor table of the harbor village's street-side seafood eatery, near the nighttime fish trade. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Raised chopsticks and fish (A piece of raw fish is held aloft) — The chopsticks angle upward beside 쿠마's face rather than toward the lens; used as Small foreground-to-midground link between the plate and 쿠마's calm expression; Restaurant table (The served plate remains on the table) — Only the near tabletop and part of the plate are visible along the lower edge; used as Continuity anchor for the completed rise.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding nighttime ambient illumination and restrained contrast, giving equal readability to 쿠마's expression and the lifted food.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The plate of sliced, radiation-mutated fish remains served at the harbor street restaurant. 쿠마: He remains at the harbor street restaurant.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 쿠마 right now, so 쿠마's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 쿠마: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 젓가락으로 생선회를 허공에 치켜든 채 태연한 표정으로 앞을 주시하는 쿠마의 상체.\n\nLOCATION (lock): At the outdoor table of the harbor village's street-side seafood eatery, near the nighttime fish trade. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Raised chopsticks and fish (A piece of raw fish is held aloft) — The chopsticks angle upward beside 쿠마's face rather than toward the lens; used as Small foreground-to-midground link between the plate and 쿠마's calm expression; Restaurant table (The served plate remains on the table) — Only the near tabletop and part of the plate are visible along the lower edge; used as Continuity anchor for the completed rise.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the preceding nighttime ambient illumination and restrained contrast, giving equal readability to 쿠마's expression and the lifted food.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The plate of sliced, radiation-mutated fish remains served at the harbor street restaurant. 쿠마: He remains at the harbor street restaurant.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 쿠마 right now, so 쿠마's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 쿠마: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 쿠마 (20대 남성, 중국계 혼혈, 젊은 성인의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "정면을 응시하나, 젓가락이 얼굴 옆이 아닌 가슴 높이에서 카메라 쪽으로 가로로 뻗어 있음.",
    "built_space": "야간 항구 배경, 금속 테이블과 접시가 하단에 적절히 위치함.",
    "entities": "쿠마의 인상착의와 생선회 소품이 레퍼런스와 일치함.",
    "hard_violations": [
     "[gpt-high] 쿠마 오른쪽 뒤의 판매대에 추가 인물이 보여, 화면에 쿠마만 허용한 조건을 위반한다."
    ],
    "physics": "오른손으로 젓가락을 쥐고 생선회를 들고 있음."
   },
   {
    "label": "B",
    "direction": "시선은 정면을 향하며, 젓가락이 지시대로 얼굴 옆에서 위를 향해 치켜들려 있음.",
    "built_space": "야간 항구 배경에 금속 테이블과 회접시가 화면 하단에 정확히 배치됨.",
    "entities": "쿠마의 외모, 의상, 생선회 소품이 레퍼런스와 정확히 일치함.",
    "hard_violations": [
     "[gpt-high] 왼쪽 점포 앞에 검은 상의와 노란 앞치마처럼 보이는 옷을 입은 추가 인물이 있어, 쿠마 이외의 사람을 금지한 조건을 위반한다."
    ],
    "physics": "손이 젓가락을 자연스럽게 쥐고 생선회를 안정적으로 허공에 들고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "얼굴 옆으로 젓가락을 위를 향해 치켜든 정확한 구도와 차분한 표정 등 프롬프트의 핵심 지시를 훌륭하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "인물과 배경은 일치하나, 젓가락을 얼굴 옆으로 위를 향하게 들라는 구도 지시를 따르지 않고 가슴 높이에서 가로로 들고 있습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "시선은 정면을 향하며, 젓가락이 지시대로 얼굴 옆에서 위를 향해 치켜들려 있음.",
        "built_space": "야간 항구 배경에 금속 테이블과 회접시가 화면 하단에 정확히 배치됨.",
        "entities": "쿠마의 외모, 의상, 생선회 소품이 레퍼런스와 정확히 일치함.",
        "hard_violations": [],
        "physics": "손이 젓가락을 자연스럽게 쥐고 생선회를 안정적으로 허공에 들고 있음."
       },
       {
        "label": "A",
        "direction": "정면을 응시하나, 젓가락이 얼굴 옆이 아닌 가슴 높이에서 카메라 쪽으로 가로로 뻗어 있음.",
        "built_space": "야간 항구 배경, 금속 테이블과 접시가 하단에 적절히 위치함.",
        "entities": "쿠마의 인상착의와 생선회 소품이 레퍼런스와 일치함.",
        "hard_violations": [],
        "physics": "오른손으로 젓가락을 쥐고 생선회를 들고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "얼굴 옆으로 젓가락을 위를 향해 치켜든 정확한 구도와 차분한 표정 등 프롬프트의 핵심 지시를 훌륭하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "인물과 배경은 일치하나, 젓가락을 얼굴 옆으로 위를 향하게 들라는 구도 지시를 따르지 않고 가슴 높이에서 가로로 들고 있습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "시선은 정면을 향하며, 젓가락이 지시대로 얼굴 옆에서 위를 향해 치켜들려 있음.",
        "built_space": "야간 항구 배경에 금속 테이블과 회접시가 화면 하단에 정확히 배치됨.",
        "entities": "쿠마의 외모, 의상, 생선회 소품이 레퍼런스와 정확히 일치함.",
        "hard_violations": [],
        "physics": "손이 젓가락을 자연스럽게 쥐고 생선회를 안정적으로 허공에 들고 있음."
       },
       {
        "label": "A",
        "direction": "정면을 응시하나, 젓가락이 얼굴 옆이 아닌 가슴 높이에서 카메라 쪽으로 가로로 뻗어 있음.",
        "built_space": "야간 항구 배경, 금속 테이블과 접시가 하단에 적절히 위치함.",
        "entities": "쿠마의 인상착의와 생선회 소품이 레퍼런스와 일치함.",
        "hard_violations": [],
        "physics": "오른손으로 젓가락을 쥐고 생선회를 들고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "얼굴 옆으로 젓가락을 치켜든 방향과 하단에만 걸친 식탁·접시는 정확하지만, 왼쪽 배경에 금지된 추가 인물이 있어 실격이다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "항구의 재질과 쿠마의 외모는 이어지지만, 추가 인물이 있으며 젓가락이 얼굴 옆 위쪽이 아니라 가슴 앞으로 내려가고 식탁도 과도하게 드러난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "쿠마는 생선이 아닌 카메라 쪽 정면을 태연하게 주시한다. 오른손의 젓가락 두 개가 화면 왼쪽 아래에서 오른쪽 위로 뻗어 얼굴 왼편의 생선회를 집고 있다. 끝이 렌즈를 겨냥하지 않으며, 얼굴 옆으로 치켜든다는 지시와 맞는다.",
        "built_space": "앞쪽 금속 식탁 하나와 접시 하나의 일부만 하단에 보인다. 쿠마는 식탁 뒤 붉은 좌석에 앉아 있다. 왼쪽에는 점포 전면과 긴 탁자 하나, 붉은 의자가 있고, 오른쪽에는 지붕 달린 판매대 하나와 통·의자들이 있다. 뒤쪽 배와 젖은 포장도로는 항구 환경에 맞으며, 식탁과 노면의 조명 반사도 가능하다. 다만 참조보다 왼쪽 점포 전면이 크게 추가로 드러나 정확히 같은 장소인지는 확정하기 어렵다.",
        "entities": "검은 짧은 머리의 젊은 동아시아계 남성 한 명이 중심에 있으며 얼굴, 체격, 짙은 티셔츠와 어두운 재킷은 쿠마 참조와 대체로 일치한다. 중국계 혼혈 여부는 외관만으로 검증할 수 없다. 금속 젓가락 한 쌍, 껍질 붙은 생선회 한 점, 남은 회가 담긴 밝은 접시와 긁힌 금속 식탁이 보인다. 생선의 두꺼운 살과 무늬 있는 껍질은 이전 장면과 이어진다. 왼쪽 점포 앞에는 쿠마가 아닌 사람이 명확히 보인다. 읽을 수 있는 글자는 식별되지 않는다.",
        "hard_violations": [
         "왼쪽 점포 앞에 검은 상의와 노란 앞치마처럼 보이는 옷을 입은 추가 인물이 있어, 쿠마 이외의 사람을 금지한 조건을 위반한다."
        ],
        "physics": "들어 올린 오른팔은 팔꿈치 부근이 식탁에 닿고, 손가락이 젓가락을 쥔다. 생선은 젓가락 끝에 집혀 아래로 늘어져 있어 지지와 중력 방향이 자연스럽다. 접시는 식탁에 놓여 있고 몸통은 뒤쪽 좌석에 앉은 자세로 이어진다. 지지 없이 떠 있는 물체나 불가능한 관절은 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "쿠마는 카메라 쪽 앞을 차분하게 바라본다. 젓가락은 화면 왼쪽 손에서 오른쪽 아래로 거의 수평으로 뻗어 가슴 앞의 생선회를 집는다. 렌즈를 직접 겨냥하지는 않지만, 얼굴 옆에서 위로 기울어야 한다는 핵심 방향과 다르다.",
        "built_space": "금속 식탁 하나가 하단의 상당 부분을 차지하고 접시 하나의 넓은 부분이 보인다. 쿠마는 식탁 뒤 붉은 의자에 앉아 있다. 왼쪽에는 금속 탁자 두 개와 차양 지지기둥이, 뒤에는 지붕 달린 판매대 하나와 통들이 있으며 오른쪽에 정박한 배들이 보인다. 젖은 도로와 판매대·선박의 배치는 이전 장면과 잘 연결되지만, 가까운 식탁과 접시를 하단에 조금만 보이라는 구도보다 노출 면적이 크다. 표면 반사는 물리적으로 가능하다.",
        "entities": "쿠마의 짧은 검은 머리, 젊은 남성 얼굴, 짙은 티셔츠와 어두운 재킷은 참조와 대체로 맞는다. 중국계 혼혈이라는 배경 자체는 이미지로 확인할 수 없다. 오른손에 금속 젓가락 한 쌍이 있고, 껍질 붙은 회 한 점과 접시에 남은 두꺼운 회가 보인다. 금속 식탁의 흠집과 물기는 이전 장면에 부합한다. 중앙 뒤 판매대에는 머리와 상체가 보이는 작은 배경 인물이 있다. 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [
         "쿠마 오른쪽 뒤의 판매대에 추가 인물이 보여, 화면에 쿠마만 허용한 조건을 위반한다."
        ],
        "physics": "오른손이 젓가락을 쥐고 그 끝이 생선 윗부분을 집어 지지한다. 생선은 아래로 늘어져 있고, 들어 올린 팔은 어깨와 자연스럽게 연결된다. 쿠마는 붉은 좌석에 앉아 있으며 접시는 식탁에 놓여 있다. 떠 있는 무지지 물체나 명백히 불가능한 자세는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "얼굴 옆으로 젓가락을 치켜든 방향과 하단에만 걸친 식탁·접시는 정확하지만, 왼쪽 배경에 금지된 추가 인물이 있어 실격이다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "항구의 재질과 쿠마의 외모는 이어지지만, 추가 인물이 있으며 젓가락이 얼굴 옆 위쪽이 아니라 가슴 앞으로 내려가고 식탁도 과도하게 드러난다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "쿠마는 생선이 아닌 카메라 쪽 정면을 태연하게 주시한다. 오른손의 젓가락 두 개가 화면 왼쪽 아래에서 오른쪽 위로 뻗어 얼굴 왼편의 생선회를 집고 있다. 끝이 렌즈를 겨냥하지 않으며, 얼굴 옆으로 치켜든다는 지시와 맞는다.",
        "built_space": "앞쪽 금속 식탁 하나와 접시 하나의 일부만 하단에 보인다. 쿠마는 식탁 뒤 붉은 좌석에 앉아 있다. 왼쪽에는 점포 전면과 긴 탁자 하나, 붉은 의자가 있고, 오른쪽에는 지붕 달린 판매대 하나와 통·의자들이 있다. 뒤쪽 배와 젖은 포장도로는 항구 환경에 맞으며, 식탁과 노면의 조명 반사도 가능하다. 다만 참조보다 왼쪽 점포 전면이 크게 추가로 드러나 정확히 같은 장소인지는 확정하기 어렵다.",
        "entities": "검은 짧은 머리의 젊은 동아시아계 남성 한 명이 중심에 있으며 얼굴, 체격, 짙은 티셔츠와 어두운 재킷은 쿠마 참조와 대체로 일치한다. 중국계 혼혈 여부는 외관만으로 검증할 수 없다. 금속 젓가락 한 쌍, 껍질 붙은 생선회 한 점, 남은 회가 담긴 밝은 접시와 긁힌 금속 식탁이 보인다. 생선의 두꺼운 살과 무늬 있는 껍질은 이전 장면과 이어진다. 왼쪽 점포 앞에는 쿠마가 아닌 사람이 명확히 보인다. 읽을 수 있는 글자는 식별되지 않는다.",
        "hard_violations": [
         "왼쪽 점포 앞에 검은 상의와 노란 앞치마처럼 보이는 옷을 입은 추가 인물이 있어, 쿠마 이외의 사람을 금지한 조건을 위반한다."
        ],
        "physics": "들어 올린 오른팔은 팔꿈치 부근이 식탁에 닿고, 손가락이 젓가락을 쥔다. 생선은 젓가락 끝에 집혀 아래로 늘어져 있어 지지와 중력 방향이 자연스럽다. 접시는 식탁에 놓여 있고 몸통은 뒤쪽 좌석에 앉은 자세로 이어진다. 지지 없이 떠 있는 물체나 불가능한 관절은 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "쿠마는 카메라 쪽 앞을 차분하게 바라본다. 젓가락은 화면 왼쪽 손에서 오른쪽 아래로 거의 수평으로 뻗어 가슴 앞의 생선회를 집는다. 렌즈를 직접 겨냥하지는 않지만, 얼굴 옆에서 위로 기울어야 한다는 핵심 방향과 다르다.",
        "built_space": "금속 식탁 하나가 하단의 상당 부분을 차지하고 접시 하나의 넓은 부분이 보인다. 쿠마는 식탁 뒤 붉은 의자에 앉아 있다. 왼쪽에는 금속 탁자 두 개와 차양 지지기둥이, 뒤에는 지붕 달린 판매대 하나와 통들이 있으며 오른쪽에 정박한 배들이 보인다. 젖은 도로와 판매대·선박의 배치는 이전 장면과 잘 연결되지만, 가까운 식탁과 접시를 하단에 조금만 보이라는 구도보다 노출 면적이 크다. 표면 반사는 물리적으로 가능하다.",
        "entities": "쿠마의 짧은 검은 머리, 젊은 남성 얼굴, 짙은 티셔츠와 어두운 재킷은 참조와 대체로 맞는다. 중국계 혼혈이라는 배경 자체는 이미지로 확인할 수 없다. 오른손에 금속 젓가락 한 쌍이 있고, 껍질 붙은 회 한 점과 접시에 남은 두꺼운 회가 보인다. 금속 식탁의 흠집과 물기는 이전 장면에 부합한다. 중앙 뒤 판매대에는 머리와 상체가 보이는 작은 배경 인물이 있다. 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [
         "쿠마 오른쪽 뒤의 판매대에 추가 인물이 보여, 화면에 쿠마만 허용한 조건을 위반한다."
        ],
        "physics": "오른손이 젓가락을 쥐고 그 끝이 생선 윗부분을 집어 지지한다. 생선은 아래로 늘어져 있고, 들어 올린 팔은 어깨와 자연스럽게 연결된다. 쿠마는 붉은 좌석에 앉아 있으며 접시는 식탁에 놓여 있다. 떠 있는 무지지 물체나 명백히 불가능한 자세는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.464,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.214,
    "B": 1.75
   },
   "violations": {
    "B": [
     "[gpt-high] 왼쪽 점포 앞에 검은 상의와 노란 앞치마처럼 보이는 옷을 입은 추가 인물이 있어, 쿠마 이외의 사람을 금지한 조건을 위반한다."
    ],
    "A": [
     "[gpt-high] 쿠마 오른쪽 뒤의 판매대에 추가 인물이 보여, 화면에 쿠마만 허용한 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 1750,
   "A": 1214
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "얼굴 옆으로 젓가락을 위를 향해 치켜든 정확한 구도와 차분한 표정 등 프롬프트의 핵심 지시를 훌륭하게 구현했습니다.  ★위반: [gpt-high] 왼쪽 점포 앞에 검은 상의와 노란 앞치마처럼 보이는 옷을 입은 추가 인물이 있어, 쿠마 이외의 사람을 금지한 조건을 위반한다."
   },
   {
    "label": "A",
    "score": 1214,
    "verdict_ko": "인물과 배경은 일치하나, 젓가락을 얼굴 옆으로 위를 향하게 들라는 구도 지시를 따르지 않고 가슴 높이에서 가로로 들고 있습니다.  ★위반: [gpt-high] 쿠마 오른쪽 뒤의 판매대에 추가 인물이 보여, 화면에 쿠마만 허용한 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features, lighting mood and each person's clothing are LOCKED to this photo; never copy its camera framing. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S74sh8_sel.png",
    "asset_id": "4dddf26c-90fb-4c6c-810a-c95bb5bff52e",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 쿠마: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1218434>",
    "asset_id": "aed54006-63aa-42f8-8d50-5ff0f4385074",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-e5fa-7eff-84b9-91d395442805",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S74sh8"
  }
 },
 "S74sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:37:35.783683+00:00",
  "fingerprint": "b213a8584ef249a894c530614774a1bfff7a67bf2cb87842f19f5ee9bb0db7fe",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S74sh9_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S74sh9_sel.png",
  "source_sha256": "51fc0738cd79cda2670de82debf29ca265b8b0fa37b70001d7fe1c3b6f7e2982",
  "file": "S74sh9_cine.png",
  "staged_sha256": "eb9c9e12f7d19fe0db8491f7c383fa46c44f2d4692414c0dc45ad13aa81e6c29",
  "latency_ms": 10863
 },
 "S75sh5::signage": {
  "fp": "c7aec2cf7688f135",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::ae001ce363d68bb2": {
  "subjects": [
   {
    "subject_native": "목포공항 입구 및 간판",
    "search_terms_native": [
     "목포공항 입구",
     "목포공항 터미널",
     "폐쇄 목포공항 전경"
    ],
    "language_lock_native": "반드시 한국어로만 검색을 진행하고 다른 언어로 번역하거나 외래 검색어를 추가하지 마십시오.",
    "reason_ko": "과거 실재했던 목포공항 특유의 소형 지방공항 터미널 외형 및 진입로 간판 양식을 정확히 반영하기 위함입니다."
   }
  ],
  "subject_text": "목포공항 입구와 공터\n오래 방치된 비행장 입구와 넓은 공터. 갈라진 도로에 이끼가 번지고 공항 간판은 낡아 일부가 떨어져 있다.",
  "identity": "canonical",
  "scope_id": "L256",
  "scope_role": "location_exterior",
  "scope_sha": "90bb07f185654788"
 },
 "era_fail::4398bc8971d3c7a3": {
  "stage": "research",
  "subject": "목포공항 입구 및 간판",
  "terms": [
   "목포공항 입구",
   "목포공항 터미널",
   "폐쇄 목포공항 전경"
  ],
  "status": "no_usable",
  "queries": [
   [
    "목포공항 입구 간판 터미널",
    "폐쇄 목포공항 전경"
   ],
   [
    "목포공항 터미널 2007년 10월 8일",
    "목포공항 폐쇄 2007년 11월 8일 해군"
   ]
  ],
  "candidate_urls": [
   "https://upload.wikimedia.org/wikipedia/commons/thumb/c/c0/Mokpo_airport_South_Korea_20071008.jpg/3840px-Mokpo_airport_South_Korea_20071008.jpg",
   "https://t1.daumcdn.net/news/202510/28/551729-qOzPZGl/20251028200227896ghdt.jpg",
   "https://wimg.munhwa.com/news/cms/2026/04/02/news-p.v1.20260402.d76e6eb1727a4b79852635612ac3a90c_P1.jpg",
   "https://mblogthumb-phinf.pstatic.net/MjAyNTAxMTZfMTAx/MDAxNzM2OTk1ODk3MDI5.-6sUDK3RU8s2QrJOanEcDnjQmjTBtyLROJP1zfmrRK8g.bcQjHNSENj-2AFgUzyq_dNRB9GpInSbB6z1q8ZxZIKsg.JPEG/KakaoTalk_20250116_115102911_19.jpg?type=w800"
  ],
  "coarse": {
   "eligible": [],
   "chosen_index": 0,
   "reason": "종류·보임·기준을 다 만족하는 후보가 없다 — 이 라운드에선 안 고른다",
   "single_judge": true,
   "rejected_judges": {}
  },
  "verdicts": [
   {
    "index": 1,
    "object_type_match": "yes",
    "visible": true,
    "criteria_match": "unsure",
    "similarity": 100
   },
   {
    "index": 2,
    "object_type_match": "no",
    "visible": true,
    "criteria_match": "unsure",
    "similarity": 10
   },
   {
    "index": 3,
    "object_type_match": "no",
    "visible": true,
    "criteria_match": "unsure",
    "similarity": 10
   },
   {
    "index": 4,
    "object_type_match": "no",
    "visible": true,
    "criteria_match": "unsure",
    "similarity": 50
   }
  ],
  "chosen_reason_ko": "",
  "attempts": 8
 },
 "S75sh5::bgfirst_bg": {
  "input_fingerprint": "5d27d9bcc99f7e15",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 창고 안, 먼지가 뽀얗게 쌓인 낡은 경비행기 무더기들을 올려다보는 현우의 멍한 얼굴.\n\nLOCATION (lock): Inside the abandoned airfield's large aircraft-storage hangar, among dusty light aircraft in the dim nighttime interior.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Stored light aircraft (Old, abandoned, and covered with accumulated dust) — Partial aircraft contours remain oblique and out of focus behind 현우; the aircraft receiving his gaze lie beyond the frame; used as Soft contextual fragments rather than an oversized aircraft foreground.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued ambient illumination appropriate to the nighttime warehouse preserves facial detail with controlled contrast and no added colored source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 창고 안, 먼지가 뽀얗게 쌓인 낡은 경비행기 무더기들을 올려다보는 현우의 멍한 얼굴.\n\nLOCATION (lock): Inside the abandoned airfield's large aircraft-storage hangar, among dusty light aircraft in the dim nighttime interior.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Stored light aircraft (Old, abandoned, and covered with accumulated dust) — Partial aircraft contours remain oblique and out of focus behind 현우; the aircraft receiving his gaze lie beyond the frame; used as Soft contextual fragments rather than an oversized aircraft foreground.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued ambient illumination appropriate to the nighttime warehouse preserves facial detail with controlled contrast and no added colored source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S75sh5__bgfirst_bg.png",
  "asset_id": "6dae19f5-dd83-4234-9035-688927798687",
  "input_asset_ids": [
   "78ed0e38-f357-47d7-96e6-bcaf44bbe506",
   "16795a55-0399-4e31-9bed-0e880d7fe7e2"
  ]
 },
 "S75sh5": {
  "input_fingerprint": "021e897d6f6b3d1b",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 창고 안, 먼지가 뽀얗게 쌓인 낡은 경비행기 무더기들을 올려다보는 현우의 멍한 얼굴.\n\nLOCATION (lock): Inside the abandoned airfield's large aircraft-storage hangar, among dusty light aircraft in the dim nighttime interior. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Stored light aircraft (Old, abandoned, and covered with accumulated dust) — Partial aircraft contours remain oblique and out of focus behind 현우; the aircraft receiving his gaze lie beyond the frame; used as Soft contextual fragments rather than an oversized aircraft foreground.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued ambient illumination appropriate to the nighttime warehouse preserves facial detail with controlled contrast and no added colored source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The large hangar door is open, revealing numerous old, nonfunctional light aircraft. The abandoned airport has moss, broken paving and a deteriorated, partly fallen sign. 현우: He is inside the hangar, still dirty and battered from the journey.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 창고 안, 먼지가 뽀얗게 쌓인 낡은 경비행기 무더기들을 올려다보는 현우의 멍한 얼굴.\n\nLOCATION (lock): Inside the abandoned airfield's large aircraft-storage hangar, among dusty light aircraft in the dim nighttime interior. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Stored light aircraft (Old, abandoned, and covered with accumulated dust) — Partial aircraft contours remain oblique and out of focus behind 현우; the aircraft receiving his gaze lie beyond the frame; used as Soft contextual fragments rather than an oversized aircraft foreground.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued ambient illumination appropriate to the nighttime warehouse preserves facial detail with controlled contrast and no added colored source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The large hangar door is open, revealing numerous old, nonfunctional light aircraft. The abandoned airport has moss, broken paving and a deteriorated, partly fallen sign. 현우: He is inside the hangar, still dirty and battered from the journey.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 창고 안, 먼지가 뽀얗게 쌓인 낡은 경비행기 무더기들을 올려다보는 현우의 멍한 얼굴.\n\nLOCATION (lock): Inside the abandoned airfield's large aircraft-storage hangar, among dusty light aircraft in the dim nighttime interior. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Stored light aircraft (Old, abandoned, and covered with accumulated dust) — Partial aircraft contours remain oblique and out of focus behind 현우; the aircraft receiving his gaze lie beyond the frame; used as Soft contextual fragments rather than an oversized aircraft foreground.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued ambient illumination appropriate to the nighttime warehouse preserves facial detail with controlled contrast and no added colored source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The large hangar door is open, revealing numerous old, nonfunctional light aircraft. The abandoned airport has moss, broken paving and a deteriorated, partly fallen sign. 현우: He is inside the hangar, still dirty and battered from the journey.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S75sh5__bgfirst_bg.png",
     "asset_id": "6dae19f5-dd83-4234-9035-688927798687",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S75sh5.png",
     "asset_id": "78ed0e38-f357-47d7-96e6-bcaf44bbe506",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L256B02.png",
     "asset_id": "16795a55-0399-4e31-9bed-0e880d7fe7e2",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "시선은 화면 좌측 상단의 프레임 밖을 향하고 있음.",
    "built_space": "격납고 내부로, 배경의 비행기들이 아웃포커스 없이 선명하게 렌더링되어 참조 사진과 거의 동일하게 보임.",
    "entities": "현우(참조 이미지 일치)의 얼굴과 옷에 먼지가 묻어 있음.",
    "hard_violations": [],
    "physics": "인물에 가해진 조명이 배경과 다소 겉돌지만 명백한 물리적 오류는 없음."
   },
   {
    "label": "B",
    "direction": "시선은 화면 밖 위쪽에 있는 가상의 비행기 무더기를 향하고 있음.",
    "built_space": "어두운 격납고 내부이며, 지시된 대로 뒤편의 경비행기들이 초점이 나간 상태로 부드럽게 배치됨.",
    "entities": "현우(참조 이미지 일치, 헝클어진 검은 머리, 앳된 얼굴)가 흙먼지를 뒤집어쓴 채 등장함.",
    "hard_violations": [],
    "physics": "인물의 자세와 의상 착용 상태가 자연스러우며 떠 있는 물체 없음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "배경의 경비행기를 아웃포커스로 처리하라는 프레이밍 지시를 정확히 따랐으며, 멍하게 위를 올려다보는 현우의 표정과 더러워진 상태를 자연스럽게 구현했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "배경의 비행기들이 아웃포커스되지 않고 지나치게 선명하게 렌더링되어 카메라 및 프레임 지시를 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "시선은 화면 밖 위쪽에 있는 가상의 비행기 무더기를 향하고 있음.",
        "built_space": "어두운 격납고 내부이며, 지시된 대로 뒤편의 경비행기들이 초점이 나간 상태로 부드럽게 배치됨.",
        "entities": "현우(참조 이미지 일치, 헝클어진 검은 머리, 앳된 얼굴)가 흙먼지를 뒤집어쓴 채 등장함.",
        "hard_violations": [],
        "physics": "인물의 자세와 의상 착용 상태가 자연스러우며 떠 있는 물체 없음."
       },
       {
        "label": "A",
        "direction": "시선은 화면 좌측 상단의 프레임 밖을 향하고 있음.",
        "built_space": "격납고 내부로, 배경의 비행기들이 아웃포커스 없이 선명하게 렌더링되어 참조 사진과 거의 동일하게 보임.",
        "entities": "현우(참조 이미지 일치)의 얼굴과 옷에 먼지가 묻어 있음.",
        "hard_violations": [],
        "physics": "인물에 가해진 조명이 배경과 다소 겉돌지만 명백한 물리적 오류는 없음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "배경의 경비행기를 아웃포커스로 처리하라는 프레이밍 지시를 정확히 따랐으며, 멍하게 위를 올려다보는 현우의 표정과 더러워진 상태를 자연스럽게 구현했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "배경의 비행기들이 아웃포커스되지 않고 지나치게 선명하게 렌더링되어 카메라 및 프레임 지시를 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "시선은 화면 밖 위쪽에 있는 가상의 비행기 무더기를 향하고 있음.",
        "built_space": "어두운 격납고 내부이며, 지시된 대로 뒤편의 경비행기들이 초점이 나간 상태로 부드럽게 배치됨.",
        "entities": "현우(참조 이미지 일치, 헝클어진 검은 머리, 앳된 얼굴)가 흙먼지를 뒤집어쓴 채 등장함.",
        "hard_violations": [],
        "physics": "인물의 자세와 의상 착용 상태가 자연스러우며 떠 있는 물체 없음."
       },
       {
        "label": "A",
        "direction": "시선은 화면 좌측 상단의 프레임 밖을 향하고 있음.",
        "built_space": "격납고 내부로, 배경의 비행기들이 아웃포커스 없이 선명하게 렌더링되어 참조 사진과 거의 동일하게 보임.",
        "entities": "현우(참조 이미지 일치)의 얼굴과 옷에 먼지가 묻어 있음.",
        "hard_violations": [],
        "physics": "인물에 가해진 조명이 배경과 다소 겉돌지만 명백한 물리적 오류는 없음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "멍하게 올려다보는 얼굴 중심의 클로즈업과 흐린 배경은 더 충실하지만, 왼쪽 기체가 지나치게 크게 드러나고 참조에 없는 겉셔츠가 추가됐다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "현우의 의상과 격납고 구조는 잘 맞지만, 항공기와 건축 공간을 선명하고 넓게 보여 주어 얼굴 뒤의 흐릿한 기체 파편만 남기라는 구도 지시에서 더 멀어진다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 턱을 들고 화면 위쪽의 프레임 밖을 바라본다. 렌즈나 뒤쪽 항공기를 응시하지 않으며, 바라보는 실제 항공기는 보이지 않아 시선 대상의 정체는 확인할 수 없다. 프레임 밖 기체를 올려다본다는 지시와 방향상 부합한다.",
        "built_space": "금속 벽체와 상부 창열이 있는 격납고 내부다. 왼쪽에는 객실 창과 붉은 띠가 있는 기체 한 대가 크게 잘려 있고, 오른쪽에는 프로펠러와 파란 띠가 있는 다른 기체가 보인다. 현우는 그 앞쪽에 서 있다. 창과 천장 설비는 흐림과 크롭 때문에 정확한 개수를 셀 수 없다. 장소의 재료와 분위기는 맞지만 왼쪽 동체와 가로지르는 날개가 작은 배경 조각보다 훨씬 큰 비중을 차지한다.",
        "entities": "보이는 사람은 현우 한 명이며, 앳된 동아시아계 남성의 얼굴과 헝클어진 검은 머리가 참조와 대체로 일치한다. 국적과 정확한 나이는 외관만으로 확인할 수 없다. 정상적인 눈과 힘 빠진 입으로 멍한 표정을 보이며 얼굴에 때가 묻어 있다. 남색 티셔츠 위에 참조에는 없는 회갈색 겉셔츠를 입었다. 낡고 먼지 낀 경비행기들은 확인되며 읽을 수 있는 문자는 없다.",
        "hard_violations": [],
        "physics": "머리는 목과 어깨에 자연스럽게 연결되고 고개를 드는 자세도 가능하다. 발은 클로즈업 밖이므로 지면 접촉을 확인할 수 없지만 부유를 시사하는 모습은 없다. 날개는 동체에 연결되어 있고 항공기의 착륙장치는 대부분 가려져 있다. 손에 든 물건이나 공중에 떠 있는 독립 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "현우는 고개와 눈을 화면 왼쪽 위로 돌려 프레임 밖을 올려다본다. 시선은 왼쪽에 보이는 낮은 기수보다 위쪽으로 향하며 렌즈를 보지 않는다. 실제 응시 대상은 보이지 않아 특정 기체와의 연결은 확인할 수 없지만 요구된 상향 시선에는 부합한다.",
        "built_space": "박공지붕 철골 트러스, 양쪽의 높은 창열, 금속 뒷벽과 중앙의 작은 문이 보이며 천장 등기구는 세 개 식별된다. 좌우 전경 기체 두 대와 그 뒤의 기체 두 대가 비교적 분명하게 보이고, 현우는 중앙 통로의 앞쪽 오른편에 있다. 참조 격납고의 배열과 재료는 잘 유지되지만 공간 전체와 항공기들이 너무 선명하게 설명되어 부드러운 배경 조각이라는 지시를 충족하지 못한다.",
        "entities": "현우 한 명만 보인다. 앳된 동아시아계 남성의 얼굴, 검은 헝클어진 머리와 남색 둥근목 티셔츠가 참조에 대체로 맞는다. 정확한 나이와 국적은 외관만으로 판별할 수 없다. 얼굴과 옷에 먼지와 얼룩이 있고, 자연스러운 눈과 살짝 벌어진 입으로 멍한 반응을 표현한다. 먼지가 쌓인 오래된 프로펠러 경비행기들이 있으며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우의 목과 어깨가 고개를 자연스럽게 지탱하며 상향 회전 자세에 해부학적 문제는 없다. 하체는 화면 밖이다. 왼쪽 전경 기체와 뒤쪽 기체들의 바퀴가 바닥에 닿아 있고 날개와 프로펠러는 각각 동체와 허브에 연결되어 있다. 지지 없이 떠 있는 몸이나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "멍하게 올려다보는 얼굴 중심의 클로즈업과 흐린 배경은 더 충실하지만, 왼쪽 기체가 지나치게 크게 드러나고 참조에 없는 겉셔츠가 추가됐다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "현우의 의상과 격납고 구조는 잘 맞지만, 항공기와 건축 공간을 선명하고 넓게 보여 주어 얼굴 뒤의 흐릿한 기체 파편만 남기라는 구도 지시에서 더 멀어진다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 턱을 들고 화면 위쪽의 프레임 밖을 바라본다. 렌즈나 뒤쪽 항공기를 응시하지 않으며, 바라보는 실제 항공기는 보이지 않아 시선 대상의 정체는 확인할 수 없다. 프레임 밖 기체를 올려다본다는 지시와 방향상 부합한다.",
        "built_space": "금속 벽체와 상부 창열이 있는 격납고 내부다. 왼쪽에는 객실 창과 붉은 띠가 있는 기체 한 대가 크게 잘려 있고, 오른쪽에는 프로펠러와 파란 띠가 있는 다른 기체가 보인다. 현우는 그 앞쪽에 서 있다. 창과 천장 설비는 흐림과 크롭 때문에 정확한 개수를 셀 수 없다. 장소의 재료와 분위기는 맞지만 왼쪽 동체와 가로지르는 날개가 작은 배경 조각보다 훨씬 큰 비중을 차지한다.",
        "entities": "보이는 사람은 현우 한 명이며, 앳된 동아시아계 남성의 얼굴과 헝클어진 검은 머리가 참조와 대체로 일치한다. 국적과 정확한 나이는 외관만으로 확인할 수 없다. 정상적인 눈과 힘 빠진 입으로 멍한 표정을 보이며 얼굴에 때가 묻어 있다. 남색 티셔츠 위에 참조에는 없는 회갈색 겉셔츠를 입었다. 낡고 먼지 낀 경비행기들은 확인되며 읽을 수 있는 문자는 없다.",
        "hard_violations": [],
        "physics": "머리는 목과 어깨에 자연스럽게 연결되고 고개를 드는 자세도 가능하다. 발은 클로즈업 밖이므로 지면 접촉을 확인할 수 없지만 부유를 시사하는 모습은 없다. 날개는 동체에 연결되어 있고 항공기의 착륙장치는 대부분 가려져 있다. 손에 든 물건이나 공중에 떠 있는 독립 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "현우는 고개와 눈을 화면 왼쪽 위로 돌려 프레임 밖을 올려다본다. 시선은 왼쪽에 보이는 낮은 기수보다 위쪽으로 향하며 렌즈를 보지 않는다. 실제 응시 대상은 보이지 않아 특정 기체와의 연결은 확인할 수 없지만 요구된 상향 시선에는 부합한다.",
        "built_space": "박공지붕 철골 트러스, 양쪽의 높은 창열, 금속 뒷벽과 중앙의 작은 문이 보이며 천장 등기구는 세 개 식별된다. 좌우 전경 기체 두 대와 그 뒤의 기체 두 대가 비교적 분명하게 보이고, 현우는 중앙 통로의 앞쪽 오른편에 있다. 참조 격납고의 배열과 재료는 잘 유지되지만 공간 전체와 항공기들이 너무 선명하게 설명되어 부드러운 배경 조각이라는 지시를 충족하지 못한다.",
        "entities": "현우 한 명만 보인다. 앳된 동아시아계 남성의 얼굴, 검은 헝클어진 머리와 남색 둥근목 티셔츠가 참조에 대체로 맞는다. 정확한 나이와 국적은 외관만으로 판별할 수 없다. 얼굴과 옷에 먼지와 얼룩이 있고, 자연스러운 눈과 살짝 벌어진 입으로 멍한 반응을 표현한다. 먼지가 쌓인 오래된 프로펠러 경비행기들이 있으며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우의 목과 어깨가 고개를 자연스럽게 지탱하며 상향 회전 자세에 해부학적 문제는 없다. 하체는 화면 밖이다. 왼쪽 전경 기체와 뒤쪽 기체들의 바퀴가 바닥에 닿아 있고 날개와 프로펠러는 각각 동체와 허브에 연결되어 있다. 지지 없이 떠 있는 몸이나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.405,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.405,
    "B": 2.0
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1405
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "배경의 경비행기를 아웃포커스로 처리하라는 프레이밍 지시를 정확히 따랐으며, 멍하게 위를 올려다보는 현우의 표정과 더러워진 상태를 자연스럽게 구현했습니다."
   },
   {
    "label": "A",
    "score": 1405,
    "verdict_ko": "배경의 비행기들이 아웃포커스되지 않고 지나치게 선명하게 렌더링되어 카메라 및 프레임 지시를 위반했습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L256B02.png",
    "asset_id": "16795a55-0399-4e31-9bed-0e880d7fe7e2",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-e7af-7017-a740-2c5a5e240b16",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S75sh5__bgfirst_bg.png",
   "bg_asset_id": "6dae19f5-dd83-4234-9035-688927798687",
   "bg_record_key": "S75sh5::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S75sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:39:36.399489+00:00",
  "fingerprint": "4df64434ef78e283afe58b6eeb3b083bd712a4f23bf311932eb5266be61ece48",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S75sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S75sh5_sel.png",
  "source_sha256": "d661cdb01a5338ad876dfa20afeb8d77985af061e3dd2998cf564827732d2f24",
  "file": "S75sh5_cine.png",
  "staged_sha256": "9a1be7ece1386af88df8d52a1f4cf716fd73f9e6689ec3b4667f51a0c190f60e",
  "latency_ms": 9333
 },
 "S75sh10::signage": {
  "fp": "90b29363b8ba304d",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S75sh10": {
  "input_fingerprint": "c6224e858bde9fb9",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 현우와 쿠마를 향해 뛰어오는 mid-stride 자세로, 한 발이 지면에서 떨어져 있고 겉옷 자락이 뒤로 휘날리는 신부의 전신.\n\nLOCATION (lock): At the open entrance of the abandoned airfield's aircraft-storage hangar at night, where the arriving priest reaches the others. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Open warehouse entrance behind the approaching priest in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Warehouse entrance (Open after 현우 and 쿠마 entered) — Seen diagonally from inside, with the exterior approach visible through the opening; used as Shared spatial anchor connecting the foreground men and the approaching 신부; Warehouse floor (Visible beneath the running figure); used as Clear ground reference beneath the lifted foot.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained nighttime ambient illumination with enough tonal separation to read the running figure against the entrance, without importing the outdoor firelight into the warehouse.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The hangar remains open with the old light aircraft still nonfunctional, and the airport campfire remains lit. 신부: He has arrived at the airport in alarm, wearing his clerical collar.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 현우와 쿠마를 향해 뛰어오는 mid-stride 자세로, 한 발이 지면에서 떨어져 있고 겉옷 자락이 뒤로 휘날리는 신부의 전신.\n\nLOCATION (lock): At the open entrance of the abandoned airfield's aircraft-storage hangar at night, where the arriving priest reaches the others. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Open warehouse entrance behind the approaching priest in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Warehouse entrance (Open after 현우 and 쿠마 entered) — Seen diagonally from inside, with the exterior approach visible through the opening; used as Shared spatial anchor connecting the foreground men and the approaching 신부; Warehouse floor (Visible beneath the running figure); used as Clear ground reference beneath the lifted foot.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained nighttime ambient illumination with enough tonal separation to read the running figure against the entrance, without importing the outdoor firelight into the warehouse.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The hangar remains open with the old light aircraft still nonfunctional, and the airport campfire remains lit. 신부: He has arrived at the airport in alarm, wearing his clerical collar.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 현우와 쿠마를 향해 뛰어오는 mid-stride 자세로, 한 발이 지면에서 떨어져 있고 겉옷 자락이 뒤로 휘날리는 신부의 전신.\n\nLOCATION (lock): At the open entrance of the abandoned airfield's aircraft-storage hangar at night, where the arriving priest reaches the others. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Open warehouse entrance behind the approaching priest in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Warehouse entrance (Open after 현우 and 쿠마 entered) — Seen diagonally from inside, with the exterior approach visible through the opening; used as Shared spatial anchor connecting the foreground men and the approaching 신부; Warehouse floor (Visible beneath the running figure); used as Clear ground reference beneath the lifted foot.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained nighttime ambient illumination with enough tonal separation to read the running figure against the entrance, without importing the outdoor firelight into the warehouse.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The hangar remains open with the old light aircraft still nonfunctional, and the airport campfire remains lit. 신부: He has arrived at the airport in alarm, wearing his clerical collar.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 신부 (한국인 남성, 60대, 나이 든 얼굴, 이마와 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "신부는 카메라 너머 앞쪽(현우와 쿠마의 위치)을 향해 달려오며 정면을 응시함.",
    "built_space": "격납고 내부에서 밖을 사선(대각선)으로 바라보는 앵글. 좌측 뒤로 열린 입구와 외부 모닥불이 보이고 우측에는 비행기가 위치함.",
    "entities": "레퍼런스와 일치하는 60대 신부 얼굴. 로만 칼라와 검은색 긴 겉옷 착용.",
    "hard_violations": [
     "[gpt-high] 야외 모닥불 오른쪽에 앉은 인물이 추가되어, 신부 외의 인물을 등장시키지 말라는 제한을 위반한다."
    ],
    "physics": "왼발이 지면을 딛고 오른발이 공중에 뜬 달리는 자세이며, 겉옷 자락이 움직임에 맞춰 뒤로 휘날림."
   },
   {
    "label": "B",
    "direction": "신부가 화면 정면을 향해 앞으로 달려오고 있음.",
    "built_space": "격납고 내부에서 외부를 정면으로 바라보는 평면적 구도. 좌측에 이전 레퍼런스와 동일한 앵글의 비행기가 위치함.",
    "entities": "신부 캐릭터의 인상착의 및 복장(로만 칼라, 겉옷)이 레퍼런스와 일치함.",
    "hard_violations": [],
    "physics": "달리는 포즈로 겉옷이 펄럭이나, 왼발 끝이 바닥에 닿은 상태인지 공중에 뜬 상태인지 약간 모호하게 렌더링됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "프롬프트가 요구한 '창고 내부에서 사선으로 바라보는 구도'와 '중우측 인물 배치'를 완벽히 구현했으며, 빛과 질감 표현이 자연스럽습니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "사선 구도(diagonally from inside) 지시를 어기고 정면 구도를 취했으며, 이전 샷의 비행기를 공간적 맥락 없이 그대로 복사해 붙인 듯한 배치가 어색합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "신부는 카메라 너머 앞쪽(현우와 쿠마의 위치)을 향해 달려오며 정면을 응시함.",
        "built_space": "격납고 내부에서 밖을 사선(대각선)으로 바라보는 앵글. 좌측 뒤로 열린 입구와 외부 모닥불이 보이고 우측에는 비행기가 위치함.",
        "entities": "레퍼런스와 일치하는 60대 신부 얼굴. 로만 칼라와 검은색 긴 겉옷 착용.",
        "hard_violations": [],
        "physics": "왼발이 지면을 딛고 오른발이 공중에 뜬 달리는 자세이며, 겉옷 자락이 움직임에 맞춰 뒤로 휘날림."
       },
       {
        "label": "B",
        "direction": "신부가 화면 정면을 향해 앞으로 달려오고 있음.",
        "built_space": "격납고 내부에서 외부를 정면으로 바라보는 평면적 구도. 좌측에 이전 레퍼런스와 동일한 앵글의 비행기가 위치함.",
        "entities": "신부 캐릭터의 인상착의 및 복장(로만 칼라, 겉옷)이 레퍼런스와 일치함.",
        "hard_violations": [],
        "physics": "달리는 포즈로 겉옷이 펄럭이나, 왼발 끝이 바닥에 닿은 상태인지 공중에 뜬 상태인지 약간 모호하게 렌더링됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "프롬프트가 요구한 '창고 내부에서 사선으로 바라보는 구도'와 '중우측 인물 배치'를 완벽히 구현했으며, 빛과 질감 표현이 자연스럽습니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "사선 구도(diagonally from inside) 지시를 어기고 정면 구도를 취했으며, 이전 샷의 비행기를 공간적 맥락 없이 그대로 복사해 붙인 듯한 배치가 어색합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "신부는 카메라 너머 앞쪽(현우와 쿠마의 위치)을 향해 달려오며 정면을 응시함.",
        "built_space": "격납고 내부에서 밖을 사선(대각선)으로 바라보는 앵글. 좌측 뒤로 열린 입구와 외부 모닥불이 보이고 우측에는 비행기가 위치함.",
        "entities": "레퍼런스와 일치하는 60대 신부 얼굴. 로만 칼라와 검은색 긴 겉옷 착용.",
        "hard_violations": [],
        "physics": "왼발이 지면을 딛고 오른발이 공중에 뜬 달리는 자세이며, 겉옷 자락이 움직임에 맞춰 뒤로 휘날림."
       },
       {
        "label": "B",
        "direction": "신부가 화면 정면을 향해 앞으로 달려오고 있음.",
        "built_space": "격납고 내부에서 외부를 정면으로 바라보는 평면적 구도. 좌측에 이전 레퍼런스와 동일한 앵글의 비행기가 위치함.",
        "entities": "신부 캐릭터의 인상착의 및 복장(로만 칼라, 겉옷)이 레퍼런스와 일치함.",
        "hard_violations": [],
        "physics": "달리는 포즈로 겉옷이 펄럭이나, 왼발 끝이 바닥에 닿은 상태인지 공중에 뜬 상태인지 약간 모호하게 렌더링됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "중앙 오른쪽 신부의 전신과 그 뒤 열린 출입구를 정확히 배치하고, 달리는 동작·날리는 옷자락·낡은 항공기와 야간 공간의 연속성을 잘 구현했다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "달리는 신부의 전신은 구현했지만 출입구가 왼쪽으로 치우쳐 지정 배치와 어긋나고, 야외 모닥불 옆에 허용되지 않은 인물이 추가됐다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "신부의 얼굴과 시선은 카메라보다 약간 왼쪽의 실내 전방을 향하며, 몸도 출입구에서 그쪽으로 접근한다. 화면 밖 현우와 쿠마를 향해 달려오는 동선으로 읽힌다. 겉옷 자락은 진행 방향의 뒤쪽인 화면 오른쪽으로 날린다. 무기나 겨누는 소품은 없다.",
        "built_space": "카메라는 격납고 안에서 열린 대형 출입구 하나를 비스듬히 바라본다. 출입구는 중앙부터 오른쪽 배경에 있고 신부가 그 앞 중앙 오른쪽에 있다. 왼쪽에는 잘린 적갈색 띠의 항공기 한 대와 그 뒤 청색 띠의 쌍발 경항공기 한 대가 보인다. 금속 벽판, 철골, 높은 창과 닳은 콘크리트 바닥이 참조 장소의 재질과 연결된다. 바깥 모닥불 하나는 작게 보이며 실내에 강한 주황색 불빛을 퍼뜨리지 않는다.",
        "entities": "뚜렷하게 보이는 인물은 신부 한 명이다. 주름진 얼굴과 뒤로 빗은 회흑색 머리의 한국인 노년 남성으로, 참조의 연령·성별·얼굴 인상과 대체로 맞는다. 흰 성직자 칼라와 검은 긴 성직복이 보인다. 낡고 정지한 항공기들과 켜진 야외 모닥불이 있으며, 현우와 쿠마는 화면 밖이다. 읽을 수 있는 글자나 표시는 없다.",
        "hard_violations": [],
        "physics": "앞으로 내민 발은 바닥 바로 위에서 착지를 준비하고, 반대쪽 발은 무릎을 굽힌 채 뒤로 들려 있다. 양발이 잠깐 떠 있는 달리기의 비행 국면으로, 다리의 전후 분리와 팔의 교차 운동이 도약 직후부터 착지까지의 동작을 설명한다. 발 아래에 연속된 바닥과 가까운 그림자가 있어 근거 없는 부유로 보이지 않는다. 옷은 어깨와 몸에 걸려 뒤로 흐르고, 항공기는 착륙장치로 바닥에 지지된다."
       },
       {
        "label": "B",
        "direction": "신부는 화면 왼쪽 전방을 바라보며 카메라 쪽으로 달려온다. 화면 밖 현우와 쿠마를 향하는 동작으로 해석할 수 있지만, 출입구가 몸의 뒤가 아니라 왼쪽에 있어 바깥에서 곧장 접근하는 동선은 A보다 덜 분명하다. 코트 자락은 뒤쪽인 화면 오른쪽으로 날린다.",
        "built_space": "열린 출입구 하나가 화면 왼쪽 대부분을 차지하고, 신부는 중앙 오른쪽의 금속 벽과 문짝 앞에 있다. 따라서 중앙 오른쪽 신부 뒤에 출입구를 두라는 배치와 다르다. 오른쪽 내부에는 겹친 항공기 날개와 동체 일부가 보인다. 철골 지붕, 높은 창, 낡은 금속 문과 콘크리트 바닥은 참조의 격납고 재질에 부합한다. 야외 모닥불 하나가 왼쪽 개구부 너머에 보이고 실내 조명은 절제되어 있다.",
        "entities": "주인공은 회흑색 머리와 주름진 얼굴의 한국인 노년 남성으로, 참조 신부와 대체로 부합한다. 검은 셔츠의 흰 성직자 칼라, 검은 바지와 긴 코트를 착용했다. 그러나 야외 모닥불 오른쪽에 작은 착석 인물이 추가로 보인다. 낡은 항공기와 켜진 모닥불은 유지되며 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "야외 모닥불 오른쪽에 앉은 인물이 추가되어, 신부 외의 인물을 등장시키지 말라는 제한을 위반한다."
        ],
        "physics": "앞발은 바닥 가까이 내려오고 뒷발은 무릎을 굽혀 들려 있어 달리기의 짧은 공중 국면으로 읽힌다. 앞으로 이동하는 체중, 굽힌 팔, 착지할 바닥과 그림자가 함께 보여 설명 없는 공중 부유는 아니다. 코트는 어깨에 지지된 채 뒤로 펼쳐지고 항공기는 바닥의 착륙장치에 지지된다. 야외의 작은 인물은 낮은 좌석에 앉은 모습으로 보인다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "중앙 오른쪽 신부의 전신과 그 뒤 열린 출입구를 정확히 배치하고, 달리는 동작·날리는 옷자락·낡은 항공기와 야간 공간의 연속성을 잘 구현했다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "달리는 신부의 전신은 구현했지만 출입구가 왼쪽으로 치우쳐 지정 배치와 어긋나고, 야외 모닥불 옆에 허용되지 않은 인물이 추가됐다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "신부의 얼굴과 시선은 카메라보다 약간 왼쪽의 실내 전방을 향하며, 몸도 출입구에서 그쪽으로 접근한다. 화면 밖 현우와 쿠마를 향해 달려오는 동선으로 읽힌다. 겉옷 자락은 진행 방향의 뒤쪽인 화면 오른쪽으로 날린다. 무기나 겨누는 소품은 없다.",
        "built_space": "카메라는 격납고 안에서 열린 대형 출입구 하나를 비스듬히 바라본다. 출입구는 중앙부터 오른쪽 배경에 있고 신부가 그 앞 중앙 오른쪽에 있다. 왼쪽에는 잘린 적갈색 띠의 항공기 한 대와 그 뒤 청색 띠의 쌍발 경항공기 한 대가 보인다. 금속 벽판, 철골, 높은 창과 닳은 콘크리트 바닥이 참조 장소의 재질과 연결된다. 바깥 모닥불 하나는 작게 보이며 실내에 강한 주황색 불빛을 퍼뜨리지 않는다.",
        "entities": "뚜렷하게 보이는 인물은 신부 한 명이다. 주름진 얼굴과 뒤로 빗은 회흑색 머리의 한국인 노년 남성으로, 참조의 연령·성별·얼굴 인상과 대체로 맞는다. 흰 성직자 칼라와 검은 긴 성직복이 보인다. 낡고 정지한 항공기들과 켜진 야외 모닥불이 있으며, 현우와 쿠마는 화면 밖이다. 읽을 수 있는 글자나 표시는 없다.",
        "hard_violations": [],
        "physics": "앞으로 내민 발은 바닥 바로 위에서 착지를 준비하고, 반대쪽 발은 무릎을 굽힌 채 뒤로 들려 있다. 양발이 잠깐 떠 있는 달리기의 비행 국면으로, 다리의 전후 분리와 팔의 교차 운동이 도약 직후부터 착지까지의 동작을 설명한다. 발 아래에 연속된 바닥과 가까운 그림자가 있어 근거 없는 부유로 보이지 않는다. 옷은 어깨와 몸에 걸려 뒤로 흐르고, 항공기는 착륙장치로 바닥에 지지된다."
       },
       {
        "label": "A",
        "direction": "신부는 화면 왼쪽 전방을 바라보며 카메라 쪽으로 달려온다. 화면 밖 현우와 쿠마를 향하는 동작으로 해석할 수 있지만, 출입구가 몸의 뒤가 아니라 왼쪽에 있어 바깥에서 곧장 접근하는 동선은 A보다 덜 분명하다. 코트 자락은 뒤쪽인 화면 오른쪽으로 날린다.",
        "built_space": "열린 출입구 하나가 화면 왼쪽 대부분을 차지하고, 신부는 중앙 오른쪽의 금속 벽과 문짝 앞에 있다. 따라서 중앙 오른쪽 신부 뒤에 출입구를 두라는 배치와 다르다. 오른쪽 내부에는 겹친 항공기 날개와 동체 일부가 보인다. 철골 지붕, 높은 창, 낡은 금속 문과 콘크리트 바닥은 참조의 격납고 재질에 부합한다. 야외 모닥불 하나가 왼쪽 개구부 너머에 보이고 실내 조명은 절제되어 있다.",
        "entities": "주인공은 회흑색 머리와 주름진 얼굴의 한국인 노년 남성으로, 참조 신부와 대체로 부합한다. 검은 셔츠의 흰 성직자 칼라, 검은 바지와 긴 코트를 착용했다. 그러나 야외 모닥불 오른쪽에 작은 착석 인물이 추가로 보인다. 낡은 항공기와 켜진 모닥불은 유지되며 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "야외 모닥불 오른쪽에 앉은 인물이 추가되어, 신부 외의 인물을 등장시키지 말라는 제한을 위반한다."
        ],
        "physics": "앞발은 바닥 가까이 내려오고 뒷발은 무릎을 굽혀 들려 있어 달리기의 짧은 공중 국면으로 읽힌다. 앞으로 이동하는 체중, 굽힌 팔, 착지할 바닥과 그림자가 함께 보여 설명 없는 공중 부유는 아니다. 코트는 어깨에 지지된 채 뒤로 펼쳐지고 항공기는 바닥의 착륙장치에 지지된다. 야외의 작은 인물은 낮은 좌석에 앉은 모습으로 보인다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.444,
    "B": 1.625
   },
   "adjusted": {
    "A": 1.194,
    "B": 1.625
   },
   "violations": {
    "A": [
     "[gpt-high] 야외 모닥불 오른쪽에 앉은 인물이 추가되어, 신부 외의 인물을 등장시키지 말라는 제한을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1194,
   "B": 1625
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1194,
    "verdict_ko": "프롬프트가 요구한 '창고 내부에서 사선으로 바라보는 구도'와 '중우측 인물 배치'를 완벽히 구현했으며, 빛과 질감 표현이 자연스럽습니다.  ★위반: [gpt-high] 야외 모닥불 오른쪽에 앉은 인물이 추가되어, 신부 외의 인물을 등장시키지 말라는 제한을 위반한다."
   },
   {
    "label": "B",
    "score": 1625,
    "verdict_ko": "사선 구도(diagonally from inside) 지시를 어기고 정면 구도를 취했으며, 이전 샷의 비행기를 공간적 맥락 없이 그대로 복사해 붙인 듯한 배치가 어색합니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S75sh5_sel.png",
    "asset_id": "80915e03-5b2d-49db-a50d-37a66b64d032",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 신부: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1402213>",
    "asset_id": "8696070d-ac09-4a5c-95f2-3daf1c015c32",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-b955-75d5-b760-3a4200cda9ee",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S75sh5"
  },
  "lane_policy": "ab_select_bypass:prev"
 },
 "S75sh10::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:40:39.436150+00:00",
  "fingerprint": "6ea8a838c3ef6cbd229fe19105f2160afc510d71dc275e0f177e30c493ffb1d1",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S75sh10_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S75sh10_sel.png",
  "source_sha256": "ac213408deb004ded3e323af4228be49252f769242daab54b29e1515186b3178",
  "file": "S75sh10_cine.png",
  "staged_sha256": "1d4158b275af20dfb87e190907e39fa10864ede72ab038892a1cbe51fe6521c6",
  "latency_ms": 9698
 },
 "S76sh1::signage": {
  "fp": "8e9414617eb6cc5c",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S76sh1": {
  "input_fingerprint": "3d107689bfbbc3ce",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 보건소 침대에 누워 식은땀을 뻘뻘 흘리며 찡그린 앰버의 고통스러운 얼굴 클로즈업.\n\nLOCATION (lock): On the child's bed inside the harbor-town clinic, under nighttime treatment-room lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Clinic bed (Occupied by 앰버) — A narrow part of the supporting surface is visible beneath her head and shoulder; used as Minimal physical context for the isolated face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Quiet interior ambient illumination gently reveals 앰버's cold sweat and strained expression without stylized fever colors or perceptual distortion.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the clinic bed, surrounding room surfaces, and the established nighttime lighting. Exclude aircraft and hangar equipment from the intervening airport scene.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Amber is reclined on the clinic bed, supported by the mattress as she sleeps after emergency treatment and later develops a high fever. The source does not specify whether she rests on her back or side, her head's direction, or the positions of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed is still in use at night. Charlie retains his unrepaired dents, punctures and exposed chest opening. 앰버: She remains in bed after emergency treatment of her side wound, now visibly sweating with a high fever.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 보건소 침대에 누워 식은땀을 뻘뻘 흘리며 찡그린 앰버의 고통스러운 얼굴 클로즈업.\n\nLOCATION (lock): On the child's bed inside the harbor-town clinic, under nighttime treatment-room lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Clinic bed (Occupied by 앰버) — A narrow part of the supporting surface is visible beneath her head and shoulder; used as Minimal physical context for the isolated face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Quiet interior ambient illumination gently reveals 앰버's cold sweat and strained expression without stylized fever colors or perceptual distortion.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the clinic bed, surrounding room surfaces, and the established nighttime lighting. Exclude aircraft and hangar equipment from the intervening airport scene.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Amber is reclined on the clinic bed, supported by the mattress as she sleeps after emergency treatment and later develops a high fever. The source does not specify whether she rests on her back or side, her head's direction, or the positions of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed is still in use at night. Charlie retains his unrepaired dents, punctures and exposed chest opening. 앰버: She remains in bed after emergency treatment of her side wound, now visibly sweating with a high fever.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 보건소 침대에 누워 식은땀을 뻘뻘 흘리며 찡그린 앰버의 고통스러운 얼굴 클로즈업.\n\nLOCATION (lock): On the child's bed inside the harbor-town clinic, under nighttime treatment-room lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Clinic bed (Occupied by 앰버) — A narrow part of the supporting surface is visible beneath her head and shoulder; used as Minimal physical context for the isolated face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Quiet interior ambient illumination gently reveals 앰버's cold sweat and strained expression without stylized fever colors or perceptual distortion.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the clinic bed, surrounding room surfaces, and the established nighttime lighting. Exclude aircraft and hangar equipment from the intervening airport scene.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Amber is reclined on the clinic bed, supported by the mattress as she sleeps after emergency treatment and later develops a high fever. The source does not specify whether she rests on her back or side, her head's direction, or the positions of her arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed is still in use at night. Charlie retains his unrepaired dents, punctures and exposed chest opening. 앰버: She remains in bed after emergency treatment of her side wound, now visibly sweating with a high fever.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "시선은 초점 없이 약간 아래쪽 허공을 향하고 있습니다.",
    "built_space": "병원 침대 위 베개에 누워 있으며, 프레임 배경에 병실 창문과 침대 난간 등 공간의 구조물들이 꽤 넓게 보입니다.",
    "entities": "앰버(금발, 10세 여아 추정)가 얼굴에 땀을 흘리며 약간 찡그린 표정으로 누워 있습니다.",
    "hard_violations": [],
    "physics": "머리와 목, 어깨가 베개와 침대 매트리스에 자연스럽게 닿아 무게가 지지되고 있습니다."
   },
   {
    "label": "B",
    "direction": "시선은 고통스러운 듯 살짝 위쪽 허공을 멍하게 응시하고 있습니다.",
    "built_space": "머리가 닿은 베개와 이불 끝부분만 살짝 보이며, 주변 배경 구조물이 철저히 배제되어 얼굴에 집중할 수 있는 좁은 공간감을 형성합니다.",
    "entities": "앰버(금발, 10세 여아 추정)가 얼굴 전체에 굵은 식은땀을 뻘뻘 흘리고 있으며, 미간을 강하게 찡그린 고통스러운 얼굴이 명확히 묘사되었습니다.",
    "hard_violations": [],
    "physics": "머리 전체가 베개에 완전히 밀착되어 중력에 맞게 지지되어 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 10,
        "verdict_ko": "지시된 클로즈업 프레이밍과 '최소한의 물리적 맥락' 조건을 완벽히 충족하며, 식은땀과 고통스러운 표정 묘사가 매우 사실적이고 탁월합니다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "인물의 표정과 상태 묘사는 훌륭하지만, 배경에 창문과 침대 난간이 많이 노출되어 프롬프트가 지시한 '얼굴을 고립시키는 최소한의 배경' 조건에 비해 다소 넓은 프레이밍입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 초점 없이 약간 아래쪽 허공을 향하고 있습니다.",
        "built_space": "병원 침대 위 베개에 누워 있으며, 프레임 배경에 병실 창문과 침대 난간 등 공간의 구조물들이 꽤 넓게 보입니다.",
        "entities": "앰버(금발, 10세 여아 추정)가 얼굴에 땀을 흘리며 약간 찡그린 표정으로 누워 있습니다.",
        "hard_violations": [],
        "physics": "머리와 목, 어깨가 베개와 침대 매트리스에 자연스럽게 닿아 무게가 지지되고 있습니다."
       },
       {
        "label": "B",
        "direction": "시선은 고통스러운 듯 살짝 위쪽 허공을 멍하게 응시하고 있습니다.",
        "built_space": "머리가 닿은 베개와 이불 끝부분만 살짝 보이며, 주변 배경 구조물이 철저히 배제되어 얼굴에 집중할 수 있는 좁은 공간감을 형성합니다.",
        "entities": "앰버(금발, 10세 여아 추정)가 얼굴 전체에 굵은 식은땀을 뻘뻘 흘리고 있으며, 미간을 강하게 찡그린 고통스러운 얼굴이 명확히 묘사되었습니다.",
        "hard_violations": [],
        "physics": "머리 전체가 베개에 완전히 밀착되어 중력에 맞게 지지되어 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 10,
        "verdict_ko": "지시된 클로즈업 프레이밍과 '최소한의 물리적 맥락' 조건을 완벽히 충족하며, 식은땀과 고통스러운 표정 묘사가 매우 사실적이고 탁월합니다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "인물의 표정과 상태 묘사는 훌륭하지만, 배경에 창문과 침대 난간이 많이 노출되어 프롬프트가 지시한 '얼굴을 고립시키는 최소한의 배경' 조건에 비해 다소 넓은 프레이밍입니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 초점 없이 약간 아래쪽 허공을 향하고 있습니다.",
        "built_space": "병원 침대 위 베개에 누워 있으며, 프레임 배경에 병실 창문과 침대 난간 등 공간의 구조물들이 꽤 넓게 보입니다.",
        "entities": "앰버(금발, 10세 여아 추정)가 얼굴에 땀을 흘리며 약간 찡그린 표정으로 누워 있습니다.",
        "hard_violations": [],
        "physics": "머리와 목, 어깨가 베개와 침대 매트리스에 자연스럽게 닿아 무게가 지지되고 있습니다."
       },
       {
        "label": "B",
        "direction": "시선은 고통스러운 듯 살짝 위쪽 허공을 멍하게 응시하고 있습니다.",
        "built_space": "머리가 닿은 베개와 이불 끝부분만 살짝 보이며, 주변 배경 구조물이 철저히 배제되어 얼굴에 집중할 수 있는 좁은 공간감을 형성합니다.",
        "entities": "앰버(금발, 10세 여아 추정)가 얼굴 전체에 굵은 식은땀을 뻘뻘 흘리고 있으며, 미간을 강하게 찡그린 고통스러운 얼굴이 명확히 묘사되었습니다.",
        "hard_violations": [],
        "physics": "머리 전체가 베개에 완전히 밀착되어 중력에 맞게 지지되어 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "얼굴을 더 밀착해 담고 깊게 찡그린 미간과 흐르는 식은땀을 선명하게 보여 권위 있는 샷 지시에 가장 가깝지만, 눈을 뜬 모습은 수면 상태 지시와 다르다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "야간 병실의 조명과 침구 연속성은 좋지만, 얼굴 밖 배경과 이불의 비중이 더 크고 고통의 찡그림이 약하며 수면 상태도 구현하지 않았다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴은 위쪽을 향해 약간 돌아가 있고, 뜬 눈은 화면 오른쪽 위의 프레임 밖을 바라본다. 시선의 대상은 보이지 않으며 샷이 요구한 특정 대상도 없다. 다만 잠든 상태로 보이지는 않는다.",
        "built_space": "머리 아래 흰 베개 하나와 하단의 연한 파란 침구 일부가 보인다. 침대 난간이나 창문 등 고정 설비는 거의 모두 프레임 밖이다. 얼굴 중심의 밀착 클로즈업으로 침상의 지지면만 남기는 구도에 가깝다. 설비 중복이나 불가능한 반사는 없다.",
        "entities": "금발의 약 10세 여자아이 한 명이며 둥근 얼굴과 자연스러운 눈이 보인다. 한국계 백인 혼혈이라는 구체적 배경은 외모만으로 확정할 수 없다. 이마·볼·목에 땀방울이 맺히고 미간을 강하게 찡그려 고통을 드러낸다. 흰 베개와 파란 침구는 참고 장소와 부합하지만 조명은 참고보다 중성적이다. 다른 사람, 항공 장비, 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "뒤통수와 머리카락이 베개에 눌려 있고 목과 어깨는 침상에 누운 몸으로 이어진다. 떠 있거나 힘을 주어 들어 올린 신체 부위는 보이지 않는다. 침구는 몸 위에 얹혀 있으며 땀방울은 피부에 붙거나 아래로 흐르는 형태여서 물리적으로 자연스럽다."
       },
       {
        "label": "B",
        "direction": "머리가 카메라 쪽으로 조금 돌아가 있고 반쯤 뜬 눈은 대체로 카메라 방향을 향한다. 별도의 시선 대상은 없다. 눈꺼풀이 무겁지만 눈을 감고 자는 모습은 아니다.",
        "built_space": "흰 베개 하나, 파란 담요, 뒤쪽 침대 머리판 하나의 일부와 창문 영역 하나가 보인다. 침대에 누운 위치와 머리판의 관계는 자연스럽고 참고의 따뜻한 병실 조명과 어두운 창밖도 이어진다. 다만 얼굴 외에 벽·창문·담요까지 상당 부분 포함해 최소한의 지지면만 남기는 요구에는 A보다 덜 밀착되어 있다. 중복 설비나 반사 문제는 없다.",
        "entities": "금발의 약 10세 여자아이 한 명으로 둥근 얼굴과 정상적인 눈의 형태가 보인다. 구체적인 혼혈 배경은 시각적으로 확정할 수 없다. 회색 상의 일부와 흰 침구, 파란 담요가 보이며 참고 병상의 재질과 색에 가깝다. 이마와 얼굴에 식은땀이 있지만 미간과 입 주변의 긴장은 A보다 약하다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "뒤통수와 옆머리가 베개에 충분히 받쳐지고 어깨와 몸통은 침상 위에 놓여 있다. 담요가 가슴 위에 자연스럽게 걸쳐지며 지지 없이 떠 있는 신체나 물체는 없다. 머리를 약간 돌린 자세도 베개의 지지로 유지 가능하다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "얼굴을 더 밀착해 담고 깊게 찡그린 미간과 흐르는 식은땀을 선명하게 보여 권위 있는 샷 지시에 가장 가깝지만, 눈을 뜬 모습은 수면 상태 지시와 다르다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "야간 병실의 조명과 침구 연속성은 좋지만, 얼굴 밖 배경과 이불의 비중이 더 크고 고통의 찡그림이 약하며 수면 상태도 구현하지 않았다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴은 위쪽을 향해 약간 돌아가 있고, 뜬 눈은 화면 오른쪽 위의 프레임 밖을 바라본다. 시선의 대상은 보이지 않으며 샷이 요구한 특정 대상도 없다. 다만 잠든 상태로 보이지는 않는다.",
        "built_space": "머리 아래 흰 베개 하나와 하단의 연한 파란 침구 일부가 보인다. 침대 난간이나 창문 등 고정 설비는 거의 모두 프레임 밖이다. 얼굴 중심의 밀착 클로즈업으로 침상의 지지면만 남기는 구도에 가깝다. 설비 중복이나 불가능한 반사는 없다.",
        "entities": "금발의 약 10세 여자아이 한 명이며 둥근 얼굴과 자연스러운 눈이 보인다. 한국계 백인 혼혈이라는 구체적 배경은 외모만으로 확정할 수 없다. 이마·볼·목에 땀방울이 맺히고 미간을 강하게 찡그려 고통을 드러낸다. 흰 베개와 파란 침구는 참고 장소와 부합하지만 조명은 참고보다 중성적이다. 다른 사람, 항공 장비, 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "뒤통수와 머리카락이 베개에 눌려 있고 목과 어깨는 침상에 누운 몸으로 이어진다. 떠 있거나 힘을 주어 들어 올린 신체 부위는 보이지 않는다. 침구는 몸 위에 얹혀 있으며 땀방울은 피부에 붙거나 아래로 흐르는 형태여서 물리적으로 자연스럽다."
       },
       {
        "label": "A",
        "direction": "머리가 카메라 쪽으로 조금 돌아가 있고 반쯤 뜬 눈은 대체로 카메라 방향을 향한다. 별도의 시선 대상은 없다. 눈꺼풀이 무겁지만 눈을 감고 자는 모습은 아니다.",
        "built_space": "흰 베개 하나, 파란 담요, 뒤쪽 침대 머리판 하나의 일부와 창문 영역 하나가 보인다. 침대에 누운 위치와 머리판의 관계는 자연스럽고 참고의 따뜻한 병실 조명과 어두운 창밖도 이어진다. 다만 얼굴 외에 벽·창문·담요까지 상당 부분 포함해 최소한의 지지면만 남기는 요구에는 A보다 덜 밀착되어 있다. 중복 설비나 반사 문제는 없다.",
        "entities": "금발의 약 10세 여자아이 한 명으로 둥근 얼굴과 정상적인 눈의 형태가 보인다. 구체적인 혼혈 배경은 시각적으로 확정할 수 없다. 회색 상의 일부와 흰 침구, 파란 담요가 보이며 참고 병상의 재질과 색에 가깝다. 이마와 얼굴에 식은땀이 있지만 미간과 입 주변의 긴장은 A보다 약하다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "뒤통수와 옆머리가 베개에 충분히 받쳐지고 어깨와 몸통은 침상 위에 놓여 있다. 담요가 가슴 위에 자연스럽게 걸쳐지며 지지 없이 떠 있는 신체나 물체는 없다. 머리를 약간 돌린 자세도 베개의 지지로 유지 가능하다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.675,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.675,
    "B": 2.0
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1675
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "지시된 클로즈업 프레이밍과 '최소한의 물리적 맥락' 조건을 완벽히 충족하며, 식은땀과 고통스러운 표정 묘사가 매우 사실적이고 탁월합니다."
   },
   {
    "label": "A",
    "score": 1675,
    "verdict_ko": "인물의 표정과 상태 묘사는 훌륭하지만, 배경에 창문과 침대 난간이 많이 노출되어 프롬프트가 지시한 '얼굴을 고립시키는 최소한의 배경' 조건에 비해 다소 넓은 프레이밍입니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S73sh2_sel.png",
    "asset_id": "74f0f898-0527-4c05-856b-9bb6d8711d36",
    "role": "prev_still"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-bb00-7f88-b735-8d3367474a4f",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S73sh2"
  },
  "locked_char_refs_excluded": [
   "앰버(C03)"
  ]
 },
 "S76sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:41:32.126417+00:00",
  "fingerprint": "aa60fc8936561ee15f369009a0f58d0119a24f9fcddb25c7e6721a00ea9e605e",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S76sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S76sh1_sel.png",
  "source_sha256": "007f47be447b64d5f9c97fc1a952592e92f1e7832b8258df53a39459419902f2",
  "file": "S76sh1_cine.png",
  "staged_sha256": "54154ffdbffc974c0cb1cd6a7b602ce8a50e5414d09adb97b0318d6bcefd3dd5",
  "latency_ms": 10307
 },
 "S76sh6::signage": {
  "fp": "77311fab223af5ca",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S76sh6": {
  "input_fingerprint": "311032661ad68456",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 중년 의사의 옷소매를 양손으로 꽉 틀어쥔 현우의 절박한 상체.\n\nLOCATION (lock): Beside the sick child's bed inside the clinic patient room, under ordinary nighttime clinical lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Doctor's sleeves (Clenched in both of 현우's hands) — The near sleeve runs along the right edge while both grasped areas remain visible below 현우's face; used as Physical evidence of the plea, linked directly to the facial performance; Clinic bed edge (Beside the conversation) — A small partial edge remains in the lower background without revealing the patient; used as Continuity anchor for the bedside camera path.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the clinic's neutral ambient illumination and controlled contrast, keeping the pleading face and tightly gripping hands equally legible.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same clinic furnishings, bed, and nighttime interior lighting. Exclude the airport's aircraft, tools, and warehouse fixtures.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed remains occupied. Charlie's previously damaged bodywork remains unrepaired. 현우: He remains in the health center, visibly dirty and battered.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 중년 의사의 옷소매를 양손으로 꽉 틀어쥔 현우의 절박한 상체.\n\nLOCATION (lock): Beside the sick child's bed inside the clinic patient room, under ordinary nighttime clinical lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Doctor's sleeves (Clenched in both of 현우's hands) — The near sleeve runs along the right edge while both grasped areas remain visible below 현우's face; used as Physical evidence of the plea, linked directly to the facial performance; Clinic bed edge (Beside the conversation) — A small partial edge remains in the lower background without revealing the patient; used as Continuity anchor for the bedside camera path.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the clinic's neutral ambient illumination and controlled contrast, keeping the pleading face and tightly gripping hands equally legible.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same clinic furnishings, bed, and nighttime interior lighting. Exclude the airport's aircraft, tools, and warehouse fixtures.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed remains occupied. Charlie's previously damaged bodywork remains unrepaired. 현우: He remains in the health center, visibly dirty and battered.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 중년 의사의 옷소매를 양손으로 꽉 틀어쥔 현우의 절박한 상체.\n\nLOCATION (lock): Beside the sick child's bed inside the clinic patient room, under ordinary nighttime clinical lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Doctor's sleeves (Clenched in both of 현우's hands) — The near sleeve runs along the right edge while both grasped areas remain visible below 현우's face; used as Physical evidence of the plea, linked directly to the facial performance; Clinic bed edge (Beside the conversation) — A small partial edge remains in the lower background without revealing the patient; used as Continuity anchor for the bedside camera path.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the clinic's neutral ambient illumination and controlled contrast, keeping the pleading face and tightly gripping hands equally legible.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same clinic furnishings, bed, and nighttime interior lighting. Exclude the airport's aircraft, tools, and warehouse fixtures.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed remains occupied. Charlie's previously damaged bodywork remains unrepaired. 현우: He remains in the health center, visibly dirty and battered.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선이 화면 우측 상단에 위치한 의사의 얼굴을 향해 위로 향하고 있음.",
    "built_space": "병실 내부. 왼쪽 뒤편에 흰 베개가 놓인 환자 침대가 있고, 그 옆에 의료용 서랍장과 야간 창문이 보임.",
    "entities": "현우(얼굴의 얼룩, 헝클어진 머리, 남색 셔츠)가 회색 상의를 입은 의사의 소매를 잡고 있으나, 의사의 턱과 몸통이 과도하게 노출됨.",
    "hard_violations": [
     "[gemini-pro] physically impossible anatomy (화면 우측 하단에서 소매를 잡고 있는 손이 현우의 몸 위치상 물리적으로 뻗을 수 없는 각도에서 튀어나옴)"
    ],
    "physics": "위쪽 손은 소매 앞섶을 쥐고 있으나, 아래쪽 손은 의사의 등 뒤쪽이나 화면 밖의 불가능한 위치에서 뻗어 나와 신체 연결이 성립하지 않음."
   },
   {
    "label": "B",
    "direction": "현우가 화면 우측 가장자리에 위치한 보이지 않는 의사를 향해 간절한 시선을 보내고 있음.",
    "built_space": "병실 내부. 좌측 하단 배경에 환자 침대의 베개 모서리가 보이며, 우측 뒤편에 의료 기기와 수납장이 배치된 야간 실내 환경.",
    "entities": "현우(눈물 맺힌 눈, 얼룩진 얼굴, 남색 셔츠)가 의사의 흰색 가운 소매를 양손으로 단단히 틀어쥐고 있음.",
    "hard_violations": [],
    "physics": "현우의 양손이 가운 소매를 물리적으로 자연스럽게 꽉 쥐고 있으며, 손의 장력에 따라 옷감의 주름과 무게감이 사실적으로 표현됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "프롬프트가 요구한 절박한 표정과 얼굴 바로 아래에서 소매를 움켜쥔 양손의 구도를 해부학적 오류 없이 완벽하게 재현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "화면 우측 하단에서 소매를 쥔 손이 현우의 신체 구조상 불가능한 방향에서 뻗어 나와 하드 위반에 해당하며, 의사의 얼굴이 노출되어 지시를 어겼습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "현우가 화면 우측 가장자리에 위치한 보이지 않는 의사를 향해 간절한 시선을 보내고 있음.",
        "built_space": "병실 내부. 좌측 하단 배경에 환자 침대의 베개 모서리가 보이며, 우측 뒤편에 의료 기기와 수납장이 배치된 야간 실내 환경.",
        "entities": "현우(눈물 맺힌 눈, 얼룩진 얼굴, 남색 셔츠)가 의사의 흰색 가운 소매를 양손으로 단단히 틀어쥐고 있음.",
        "hard_violations": [],
        "physics": "현우의 양손이 가운 소매를 물리적으로 자연스럽게 꽉 쥐고 있으며, 손의 장력에 따라 옷감의 주름과 무게감이 사실적으로 표현됨."
       },
       {
        "label": "A",
        "direction": "현우의 시선이 화면 우측 상단에 위치한 의사의 얼굴을 향해 위로 향하고 있음.",
        "built_space": "병실 내부. 왼쪽 뒤편에 흰 베개가 놓인 환자 침대가 있고, 그 옆에 의료용 서랍장과 야간 창문이 보임.",
        "entities": "현우(얼굴의 얼룩, 헝클어진 머리, 남색 셔츠)가 회색 상의를 입은 의사의 소매를 잡고 있으나, 의사의 턱과 몸통이 과도하게 노출됨.",
        "hard_violations": [
         "physically impossible anatomy (화면 우측 하단에서 소매를 잡고 있는 손이 현우의 몸 위치상 물리적으로 뻗을 수 없는 각도에서 튀어나옴)"
        ],
        "physics": "위쪽 손은 소매 앞섶을 쥐고 있으나, 아래쪽 손은 의사의 등 뒤쪽이나 화면 밖의 불가능한 위치에서 뻗어 나와 신체 연결이 성립하지 않음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "프롬프트가 요구한 절박한 표정과 얼굴 바로 아래에서 소매를 움켜쥔 양손의 구도를 해부학적 오류 없이 완벽하게 재현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "화면 우측 하단에서 소매를 쥔 손이 현우의 신체 구조상 불가능한 방향에서 뻗어 나와 하드 위반에 해당하며, 의사의 얼굴이 노출되어 지시를 어겼습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우가 화면 우측 가장자리에 위치한 보이지 않는 의사를 향해 간절한 시선을 보내고 있음.",
        "built_space": "병실 내부. 좌측 하단 배경에 환자 침대의 베개 모서리가 보이며, 우측 뒤편에 의료 기기와 수납장이 배치된 야간 실내 환경.",
        "entities": "현우(눈물 맺힌 눈, 얼룩진 얼굴, 남색 셔츠)가 의사의 흰색 가운 소매를 양손으로 단단히 틀어쥐고 있음.",
        "hard_violations": [],
        "physics": "현우의 양손이 가운 소매를 물리적으로 자연스럽게 꽉 쥐고 있으며, 손의 장력에 따라 옷감의 주름과 무게감이 사실적으로 표현됨."
       },
       {
        "label": "A",
        "direction": "현우의 시선이 화면 우측 상단에 위치한 의사의 얼굴을 향해 위로 향하고 있음.",
        "built_space": "병실 내부. 왼쪽 뒤편에 흰 베개가 놓인 환자 침대가 있고, 그 옆에 의료용 서랍장과 야간 창문이 보임.",
        "entities": "현우(얼굴의 얼룩, 헝클어진 머리, 남색 셔츠)가 회색 상의를 입은 의사의 소매를 잡고 있으나, 의사의 턱과 몸통이 과도하게 노출됨.",
        "hard_violations": [
         "physically impossible anatomy (화면 우측 하단에서 소매를 잡고 있는 손이 현우의 몸 위치상 물리적으로 뻗을 수 없는 각도에서 튀어나옴)"
        ],
        "physics": "위쪽 손은 소매 앞섶을 쥐고 있으나, 아래쪽 손은 의사의 등 뒤쪽이나 화면 밖의 불가능한 위치에서 뻗어 나와 신체 연결이 성립하지 않음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "현우의 절박한 시선과 두 손의 움켜쥠은 선명하지만, 상체가 더 타이트하게 잘리고 의사의 소매가 화면 중앙을 크게 가로질러 지정된 오른쪽 가장자리 배치에서 벗어난다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "현우의 상체를 담은 미디엄 구도와 의사를 향한 시선, 얼굴 아래 두 움켜쥔 지점 및 오른쪽 소매 배치가 더 정확하나 침대 노출은 요구보다 많다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 고개와 눈을 오른쪽 위로 들어 프레임 밖 의사의 얼굴 위치를 바라본다. 두 손은 의사의 흰 소매를 서로 떨어진 두 지점에서 붙잡고 있으며, 모두 현우의 얼굴 아래에 있다. 의사의 얼굴과 시선은 보이지 않는다.",
        "built_space": "뒤쪽 왼편에 침대 한 개의 머리판, 베개 한 개와 매트리스 가장자리가 보이고 오른쪽 아래에는 푸른 침구가 보인다. 현우는 침대 옆 낮은 위치, 의사는 오른쪽 전경에 있다. 침대는 작은 하단 가장자리만 남기라는 지시보다 넓게 드러난다. 가까운 소매는 오른쪽 위에서 화면 중앙 아래로 대각선으로 뻗어 가장자리에 머물지 않는다. 반사면이나 중복된 침대는 보이지 않는다.",
        "entities": "현우는 앳된 동아시아계 남성으로, 헝클어진 검은 머리와 남색 둥근목 티셔츠가 인물 참고와 대체로 일치한다. 얼굴과 손의 때 및 찰과상이 보인다. 한국계 미국인이라는 국적은 외관만으로 확인할 수 없다. 의사는 흰 가운을 입은 몸통과 팔만 보여 중년 여부는 확인할 수 없다. 환자의 얼굴이나 식별 가능한 몸은 드러나지 않아 침대 점유 상태는 확정하기 어렵다. 읽을 수 있는 글자나 찰리는 보이지 않는다.",
        "hard_violations": [],
        "physics": "현우의 두 손가락이 소매 천을 감싸고 움켜쥔 부분에 주름을 만들어 실제 접촉과 당김이 읽힌다. 의사의 팔은 오른쪽 몸통에 이어져 있고 소매는 그 팔에 걸쳐 있다. 두 사람의 하체와 바닥 접촉은 프레임 밖이지만, 공중에 떠 있거나 지지 없이 놓인 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "현우의 눈과 올라간 턱은 오른쪽 위에 일부 보이는 의사의 얼굴을 향한다. 의사는 고개를 현우 쪽으로 숙였으나 눈은 잘려 시선 자체는 확인할 수 없다. 두 손은 오른쪽 소매의 두 지점을 얼굴 아래에서 강하게 움켜쥔다.",
        "built_space": "왼쪽 하단 배경에 침대 한 개와 베개 한 개, 뒤쪽 중앙에 서랍장 한 개, 오른쪽 뒤에 어두운 창 한 개가 보인다. 왼쪽 벽에는 안내물 한 장과 설비 패널 한 개가 있고 왼쪽 아래에 수납장 일부가 걸린다. 현우와 의사는 침대 옆에서 마주하며 가까운 소매는 화면 오른쪽을 따라 내려온다. 상체와 두 손을 함께 담은 미디엄 구도에 더 가깝지만 침대와 베개는 요구한 작은 가장자리보다 많이 보인다. 참고의 좁은 구도에 없던 설비들의 정확한 연속성은 확인할 수 없다.",
        "entities": "현우는 참고와 유사한 앳된 동아시아계 남성의 얼굴, 검은 헝클어진 머리, 남색 티셔츠를 지녔으며 얼굴과 팔에 때와 작은 상처가 있다. 의사는 회색 계열 가운을 입고 주름진 코와 입 주변이 일부 보여 중년 이상의 남성으로 읽힌다. 환자는 보이지 않으며 침대가 계속 점유되어 있는지는 가려진 영역 때문에 판단할 수 없다. 벽 안내물의 글자는 흐려 읽히지 않고 추가 인물이나 찰리는 보이지 않는다.",
        "hard_violations": [],
        "physics": "두 손이 의사의 소매를 직접 감싸 쥐고 있으며 손 주변 천이 모이고 눌려 있어 붙잡는 힘이 자연스럽다. 의사의 팔과 소매는 어깨 및 몸통에 연결된다. 현우의 팔도 상체에서 자연스럽게 뻗는다. 하체 지지점은 구도 밖이지만 부유를 나타내는 자세는 없으며, 베개와 침구는 침대에 놓여 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "현우의 절박한 시선과 두 손의 움켜쥠은 선명하지만, 상체가 더 타이트하게 잘리고 의사의 소매가 화면 중앙을 크게 가로질러 지정된 오른쪽 가장자리 배치에서 벗어난다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "현우의 상체를 담은 미디엄 구도와 의사를 향한 시선, 얼굴 아래 두 움켜쥔 지점 및 오른쪽 소매 배치가 더 정확하나 침대 노출은 요구보다 많다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 고개와 눈을 오른쪽 위로 들어 프레임 밖 의사의 얼굴 위치를 바라본다. 두 손은 의사의 흰 소매를 서로 떨어진 두 지점에서 붙잡고 있으며, 모두 현우의 얼굴 아래에 있다. 의사의 얼굴과 시선은 보이지 않는다.",
        "built_space": "뒤쪽 왼편에 침대 한 개의 머리판, 베개 한 개와 매트리스 가장자리가 보이고 오른쪽 아래에는 푸른 침구가 보인다. 현우는 침대 옆 낮은 위치, 의사는 오른쪽 전경에 있다. 침대는 작은 하단 가장자리만 남기라는 지시보다 넓게 드러난다. 가까운 소매는 오른쪽 위에서 화면 중앙 아래로 대각선으로 뻗어 가장자리에 머물지 않는다. 반사면이나 중복된 침대는 보이지 않는다.",
        "entities": "현우는 앳된 동아시아계 남성으로, 헝클어진 검은 머리와 남색 둥근목 티셔츠가 인물 참고와 대체로 일치한다. 얼굴과 손의 때 및 찰과상이 보인다. 한국계 미국인이라는 국적은 외관만으로 확인할 수 없다. 의사는 흰 가운을 입은 몸통과 팔만 보여 중년 여부는 확인할 수 없다. 환자의 얼굴이나 식별 가능한 몸은 드러나지 않아 침대 점유 상태는 확정하기 어렵다. 읽을 수 있는 글자나 찰리는 보이지 않는다.",
        "hard_violations": [],
        "physics": "현우의 두 손가락이 소매 천을 감싸고 움켜쥔 부분에 주름을 만들어 실제 접촉과 당김이 읽힌다. 의사의 팔은 오른쪽 몸통에 이어져 있고 소매는 그 팔에 걸쳐 있다. 두 사람의 하체와 바닥 접촉은 프레임 밖이지만, 공중에 떠 있거나 지지 없이 놓인 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "현우의 눈과 올라간 턱은 오른쪽 위에 일부 보이는 의사의 얼굴을 향한다. 의사는 고개를 현우 쪽으로 숙였으나 눈은 잘려 시선 자체는 확인할 수 없다. 두 손은 오른쪽 소매의 두 지점을 얼굴 아래에서 강하게 움켜쥔다.",
        "built_space": "왼쪽 하단 배경에 침대 한 개와 베개 한 개, 뒤쪽 중앙에 서랍장 한 개, 오른쪽 뒤에 어두운 창 한 개가 보인다. 왼쪽 벽에는 안내물 한 장과 설비 패널 한 개가 있고 왼쪽 아래에 수납장 일부가 걸린다. 현우와 의사는 침대 옆에서 마주하며 가까운 소매는 화면 오른쪽을 따라 내려온다. 상체와 두 손을 함께 담은 미디엄 구도에 더 가깝지만 침대와 베개는 요구한 작은 가장자리보다 많이 보인다. 참고의 좁은 구도에 없던 설비들의 정확한 연속성은 확인할 수 없다.",
        "entities": "현우는 참고와 유사한 앳된 동아시아계 남성의 얼굴, 검은 헝클어진 머리, 남색 티셔츠를 지녔으며 얼굴과 팔에 때와 작은 상처가 있다. 의사는 회색 계열 가운을 입고 주름진 코와 입 주변이 일부 보여 중년 이상의 남성으로 읽힌다. 환자는 보이지 않으며 침대가 계속 점유되어 있는지는 가려진 영역 때문에 판단할 수 없다. 벽 안내물의 글자는 흐려 읽히지 않고 추가 인물이나 찰리는 보이지 않는다.",
        "hard_violations": [],
        "physics": "두 손이 의사의 소매를 직접 감싸 쥐고 있으며 손 주변 천이 모이고 눌려 있어 붙잡는 힘이 자연스럽다. 의사의 팔과 소매는 어깨 및 몸통에 연결된다. 현우의 팔도 상체에서 자연스럽게 뻗는다. 하체 지지점은 구도 밖이지만 부유를 나타내는 자세는 없으며, 베개와 침구는 침대에 놓여 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.333,
    "B": 1.875
   },
   "adjusted": {
    "A": 1.083,
    "B": 1.875
   },
   "violations": {
    "A": [
     "[gemini-pro] physically impossible anatomy (화면 우측 하단에서 소매를 잡고 있는 손이 현우의 몸 위치상 물리적으로 뻗을 수 없는 각도에서 튀어나옴)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1875,
   "A": 1083
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1875,
    "verdict_ko": "프롬프트가 요구한 절박한 표정과 얼굴 바로 아래에서 소매를 움켜쥔 양손의 구도를 해부학적 오류 없이 완벽하게 재현했습니다."
   },
   {
    "label": "A",
    "score": 1083,
    "verdict_ko": "화면 우측 하단에서 소매를 쥔 손이 현우의 신체 구조상 불가능한 방향에서 뻗어 나와 하드 위반에 해당하며, 의사의 얼굴이 노출되어 지시를 어겼습니다.  ★위반: [gemini-pro] physically impossible anatomy (화면 우측 하단에서 소매를 잡고 있는 손이 현우의 몸 위치상 물리적으로 뻗을 수 없는 각도에서 튀어나옴)"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S76sh1_sel.png",
    "asset_id": "62f972bb-2e34-4b86-931b-692bc64b4834",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-bca7-7fa1-baf9-238a1db94ea7",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S76sh1"
  }
 },
 "S76sh6::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:43:07.256516+00:00",
  "fingerprint": "2f6a93a649fcfae46fdeeee6a1a8077a28567b6c9a87752082adf37c1ded466e",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S76sh6_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S76sh6_sel.png",
  "source_sha256": "254ea39b3464790d5e496b25d0705e2921b0ae108bedae5e4ee1f07ad93c3192",
  "file": "S76sh6_cine.png",
  "staged_sha256": "eccdd4a66fb53fbecb80debcc3ad1a1426010c3401b92de08047cb7141f2cf6e",
  "latency_ms": 10075
 },
 "S76sh12::signage": {
  "fp": "8b8ede2d2c856f26",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S76sh12": {
  "input_fingerprint": "124357c4f8bd5939",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 허공을 향해 커다란 금속 손을 번쩍 치켜든 찰리의 역동적인 상체.\n\nLOCATION (lock): In the group gathered beside the clinic bed, inside the nighttime-lit patient room. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Clinic bed (Beside 찰리's position) — Only a small edge is retained at the lower margin, with the patient outside the framing; used as Spatial continuity with the bedside scene without competing with the raised arm.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Continue the same restrained clinic illumination, using controlled highlights to articulate the metal hand without adding an artificial glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed remains occupied, and Charlie's damaged chest and bodywork have not yet been repaired.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 허공을 향해 커다란 금속 손을 번쩍 치켜든 찰리의 역동적인 상체.\n\nLOCATION (lock): In the group gathered beside the clinic bed, inside the nighttime-lit patient room. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Clinic bed (Beside 찰리's position) — Only a small edge is retained at the lower margin, with the patient outside the framing; used as Spatial continuity with the bedside scene without competing with the raised arm.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Continue the same restrained clinic illumination, using controlled highlights to articulate the metal hand without adding an artificial glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed remains occupied, and Charlie's damaged chest and bodywork have not yet been repaired.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 허공을 향해 커다란 금속 손을 번쩍 치켜든 찰리의 역동적인 상체.\n\nLOCATION (lock): In the group gathered beside the clinic bed, inside the nighttime-lit patient room. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Clinic bed (Beside 찰리's position) — Only a small edge is retained at the lower margin, with the patient outside the framing; used as Spatial continuity with the bedside scene without competing with the raised arm.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Continue the same restrained clinic illumination, using controlled highlights to articulate the metal hand without adding an artificial glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The health-center bed remains occupied, and Charlie's damaged chest and bodywork have not yet been repaired.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "찰리의 얼굴과 높이 치켜든 오른팔이 위쪽 허공을 뚜렷하게 향하고 있음.",
    "built_space": "야간 조명이 켜진 병실. 화면 하단 테두리에 병원 침대의 가장자리가 요구사항대로 작게 걸쳐 있음.",
    "entities": "모래색 장갑과 하얀 마스크 형태의 얼굴을 가진 찰리가 레퍼런스와 일치하며, 가슴 부위에 긁히고 파손된 흔적이 잘 묘사됨.",
    "hard_violations": [],
    "physics": "침대 곁에 단단히 발을 디디고 몸을 뻗어 올린 자연스러운 무게중심을 보여줌."
   },
   {
    "label": "B",
    "direction": "찰리의 시선은 정면을 향하며, 오른팔은 위로 치켜들기보다는 앞으로 가볍게 뻗은 정지 동작에 가까움.",
    "built_space": "병실 내부. 좌측 하단에 병원 침대의 등받이 부분이 다소 크게 노출됨.",
    "entities": "찰리의 외형은 캐릭터 레퍼런스와 일치하나, 지시문에 명시된 흉부 장갑의 파손 흔적이 보이지 않음.",
    "hard_violations": [],
    "physics": "지면에 안정적으로 서 있으나 포즈에 움직임이나 힘이 느껴지지 않음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "요청된 '역동적인 상체'와 '허공을 향해 번쩍 치켜든 손'의 포즈를 극적인 구도로 완벽하게 구현했으며, 명시된 흉부 장갑의 손상 흔적과 침대 가장자리 프레이밍 지시도 정확히 따랐습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "지시문이 요구한 '역동적인 상체'와 달리 자세가 몹시 정적이고 경직되어 있으며, 흉부의 손상된 묘사가 누락되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 얼굴과 높이 치켜든 오른팔이 위쪽 허공을 뚜렷하게 향하고 있음.",
        "built_space": "야간 조명이 켜진 병실. 화면 하단 테두리에 병원 침대의 가장자리가 요구사항대로 작게 걸쳐 있음.",
        "entities": "모래색 장갑과 하얀 마스크 형태의 얼굴을 가진 찰리가 레퍼런스와 일치하며, 가슴 부위에 긁히고 파손된 흔적이 잘 묘사됨.",
        "hard_violations": [],
        "physics": "침대 곁에 단단히 발을 디디고 몸을 뻗어 올린 자연스러운 무게중심을 보여줌."
       },
       {
        "label": "B",
        "direction": "찰리의 시선은 정면을 향하며, 오른팔은 위로 치켜들기보다는 앞으로 가볍게 뻗은 정지 동작에 가까움.",
        "built_space": "병실 내부. 좌측 하단에 병원 침대의 등받이 부분이 다소 크게 노출됨.",
        "entities": "찰리의 외형은 캐릭터 레퍼런스와 일치하나, 지시문에 명시된 흉부 장갑의 파손 흔적이 보이지 않음.",
        "hard_violations": [],
        "physics": "지면에 안정적으로 서 있으나 포즈에 움직임이나 힘이 느껴지지 않음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "요청된 '역동적인 상체'와 '허공을 향해 번쩍 치켜든 손'의 포즈를 극적인 구도로 완벽하게 구현했으며, 명시된 흉부 장갑의 손상 흔적과 침대 가장자리 프레이밍 지시도 정확히 따랐습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "지시문이 요구한 '역동적인 상체'와 달리 자세가 몹시 정적이고 경직되어 있으며, 흉부의 손상된 묘사가 누락되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 얼굴과 높이 치켜든 오른팔이 위쪽 허공을 뚜렷하게 향하고 있음.",
        "built_space": "야간 조명이 켜진 병실. 화면 하단 테두리에 병원 침대의 가장자리가 요구사항대로 작게 걸쳐 있음.",
        "entities": "모래색 장갑과 하얀 마스크 형태의 얼굴을 가진 찰리가 레퍼런스와 일치하며, 가슴 부위에 긁히고 파손된 흔적이 잘 묘사됨.",
        "hard_violations": [],
        "physics": "침대 곁에 단단히 발을 디디고 몸을 뻗어 올린 자연스러운 무게중심을 보여줌."
       },
       {
        "label": "B",
        "direction": "찰리의 시선은 정면을 향하며, 오른팔은 위로 치켜들기보다는 앞으로 가볍게 뻗은 정지 동작에 가까움.",
        "built_space": "병실 내부. 좌측 하단에 병원 침대의 등받이 부분이 다소 크게 노출됨.",
        "entities": "찰리의 외형은 캐릭터 레퍼런스와 일치하나, 지시문에 명시된 흉부 장갑의 파손 흔적이 보이지 않음.",
        "hard_violations": [],
        "physics": "지면에 안정적으로 서 있으나 포즈에 움직임이나 힘이 느껴지지 않음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "금속 손을 올린 상체와 찰리의 외형은 맞지만, 전경 침대판이 지나치게 크게 들어오며 팔을 번쩍 치켜드는 역동성과 미수리 가슴 손상이 B보다 약하다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "허공으로 높이 치켜든 거대한 손과 역동적인 상체, 선명한 가슴 손상을 충실히 구현하며 침대도 하단으로 제한했지만 침구와 난간의 노출은 지시보다 조금 많다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "화면 왼쪽 팔을 대각선 위로 뻗고 손바닥을 카메라 쪽으로 펼쳤다. 손끝은 왼쪽 위 빈 공간을 향하므로 허공을 향한다는 지시에는 맞는다. 얼굴도 대체로 들어 올린 손 쪽을 향하지만, 팔은 머리 위로 번쩍 들기보다 앞쪽으로 내민 인상이 강하다.",
        "built_space": "찰리는 침대 뒤쪽 옆에 서 있다. 침대 끝판 하나가 왼쪽 하단을 크게 차지하고 파란 침구 일부가 보이며 환자는 보이지 않는다. 왼쪽 벽 설비 패널 하나, 수액대 하나, 뒤쪽 수납장, 오른쪽 블라인드 창과 모니터 하나, 천장 조명과 커튼 레일이 보인다. 회색 계열 병실 재료는 연결되지만, 침대 끝판이 하단의 작은 가장자리만 남기라는 구도보다 훨씬 크다. 불가능한 반사나 명백한 설비 중복은 보이지 않는다.",
        "entities": "등장 인물은 찰리 하나뿐이다. 육중한 어깨와 긴 팔, 각진 샌드 베이지 장갑판, 흰 분절형 마스크와 점·선 형태의 얼굴은 캐릭터 참조와 대체로 일치한다. 다리는 대부분 프레임 밖이다. 장갑판에 긁힘과 마모는 있지만 가슴의 미수리 파손은 뚜렷하지 않다. 이전 장면의 사람은 나타나지 않고 읽을 수 있는 글자도 확인되지 않는다.",
        "hard_violations": [],
        "physics": "들어 올린 손은 손목·팔꿈치·어깨 관절을 통해 몸통에 연결되어 있고 반대쪽 팔은 아래로 내려가 있다. 팔의 자세는 관절 회전으로 가능한 범위이며 분리되거나 떠 있는 물체는 없다. 하체와 바닥 접촉은 프레임 밖이지만 몸통은 직립한 자세로 이어져 공중부양으로 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "커다란 금속 손을 화면 오른쪽 위 허공으로 치켜들었고 손등이 카메라 쪽을 향한다. 손가락은 머리보다 훨씬 높은 위치에서 펼쳐져 있다. 얼굴은 팔에 일부 가려져 있지만 고개가 올라간 팔과 같은 방향으로 들려 있어 동작의 목표가 위쪽 빈 공간으로 읽힌다.",
        "built_space": "낮은 시점에서 침대 옆 찰리의 상체를 바라본다. 침대 하나의 끝판 일부는 왼쪽 하단에, 옆 난간 하나와 파란 침구는 하단에 놓이며 환자는 프레임 밖이다. 왼쪽 벽 설비 패널 하나, 수액대 하나와 수액 용기 하나, 뒤쪽 수납 공간과 작은 의료 장비, 천장의 조명 기구들이 보인다. 침대가 손과 경쟁하지 않도록 아래로 제한되었으나 난간과 침구는 아주 작은 가장자리보다는 넓다. 참조의 병실 재료와는 대체로 연결되며 조명은 더 어둡고 차갑다.",
        "entities": "찰리만 보이며 추가 사람이나 환자의 신체는 없다. 큰 어깨, 길고 육중한 팔, 샌드 베이지 장갑판, 부분적으로 보이는 흰 마스크와 붉은 눈이 참조 정체성과 일치한다. 가슴 장갑판에는 깊게 찢기고 패인 손상이 있어 아직 수리되지 않은 상태가 분명하다. 짧은 다리는 구도 밖이므로 판단하지 않는다. 배경 표지는 흐려 읽을 수 없고 금속 표면에는 인위적 광채보다 제한된 반사가 보인다.",
        "hard_violations": [],
        "physics": "높이 든 손과 전완은 손목·팔꿈치·어깨의 기계 관절로 연속 연결되어 있다. 굽힌 팔꿈치와 위로 돌아간 상완, 약간 뒤로 기운 몸통이 무거운 팔을 치켜드는 동작을 구성한다. 몸통은 하단의 가려진 골반 쪽으로 이어지며 떠 있는 신체나 지지 없는 물체는 없다. 침구는 침대 위에 놓이고 난간은 수직 지지대로 고정되어 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "금속 손을 올린 상체와 찰리의 외형은 맞지만, 전경 침대판이 지나치게 크게 들어오며 팔을 번쩍 치켜드는 역동성과 미수리 가슴 손상이 B보다 약하다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "허공으로 높이 치켜든 거대한 손과 역동적인 상체, 선명한 가슴 손상을 충실히 구현하며 침대도 하단으로 제한했지만 침구와 난간의 노출은 지시보다 조금 많다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "화면 왼쪽 팔을 대각선 위로 뻗고 손바닥을 카메라 쪽으로 펼쳤다. 손끝은 왼쪽 위 빈 공간을 향하므로 허공을 향한다는 지시에는 맞는다. 얼굴도 대체로 들어 올린 손 쪽을 향하지만, 팔은 머리 위로 번쩍 들기보다 앞쪽으로 내민 인상이 강하다.",
        "built_space": "찰리는 침대 뒤쪽 옆에 서 있다. 침대 끝판 하나가 왼쪽 하단을 크게 차지하고 파란 침구 일부가 보이며 환자는 보이지 않는다. 왼쪽 벽 설비 패널 하나, 수액대 하나, 뒤쪽 수납장, 오른쪽 블라인드 창과 모니터 하나, 천장 조명과 커튼 레일이 보인다. 회색 계열 병실 재료는 연결되지만, 침대 끝판이 하단의 작은 가장자리만 남기라는 구도보다 훨씬 크다. 불가능한 반사나 명백한 설비 중복은 보이지 않는다.",
        "entities": "등장 인물은 찰리 하나뿐이다. 육중한 어깨와 긴 팔, 각진 샌드 베이지 장갑판, 흰 분절형 마스크와 점·선 형태의 얼굴은 캐릭터 참조와 대체로 일치한다. 다리는 대부분 프레임 밖이다. 장갑판에 긁힘과 마모는 있지만 가슴의 미수리 파손은 뚜렷하지 않다. 이전 장면의 사람은 나타나지 않고 읽을 수 있는 글자도 확인되지 않는다.",
        "hard_violations": [],
        "physics": "들어 올린 손은 손목·팔꿈치·어깨 관절을 통해 몸통에 연결되어 있고 반대쪽 팔은 아래로 내려가 있다. 팔의 자세는 관절 회전으로 가능한 범위이며 분리되거나 떠 있는 물체는 없다. 하체와 바닥 접촉은 프레임 밖이지만 몸통은 직립한 자세로 이어져 공중부양으로 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "커다란 금속 손을 화면 오른쪽 위 허공으로 치켜들었고 손등이 카메라 쪽을 향한다. 손가락은 머리보다 훨씬 높은 위치에서 펼쳐져 있다. 얼굴은 팔에 일부 가려져 있지만 고개가 올라간 팔과 같은 방향으로 들려 있어 동작의 목표가 위쪽 빈 공간으로 읽힌다.",
        "built_space": "낮은 시점에서 침대 옆 찰리의 상체를 바라본다. 침대 하나의 끝판 일부는 왼쪽 하단에, 옆 난간 하나와 파란 침구는 하단에 놓이며 환자는 프레임 밖이다. 왼쪽 벽 설비 패널 하나, 수액대 하나와 수액 용기 하나, 뒤쪽 수납 공간과 작은 의료 장비, 천장의 조명 기구들이 보인다. 침대가 손과 경쟁하지 않도록 아래로 제한되었으나 난간과 침구는 아주 작은 가장자리보다는 넓다. 참조의 병실 재료와는 대체로 연결되며 조명은 더 어둡고 차갑다.",
        "entities": "찰리만 보이며 추가 사람이나 환자의 신체는 없다. 큰 어깨, 길고 육중한 팔, 샌드 베이지 장갑판, 부분적으로 보이는 흰 마스크와 붉은 눈이 참조 정체성과 일치한다. 가슴 장갑판에는 깊게 찢기고 패인 손상이 있어 아직 수리되지 않은 상태가 분명하다. 짧은 다리는 구도 밖이므로 판단하지 않는다. 배경 표지는 흐려 읽을 수 없고 금속 표면에는 인위적 광채보다 제한된 반사가 보인다.",
        "hard_violations": [],
        "physics": "높이 든 손과 전완은 손목·팔꿈치·어깨의 기계 관절로 연속 연결되어 있다. 굽힌 팔꿈치와 위로 돌아간 상완, 약간 뒤로 기운 몸통이 무거운 팔을 치켜드는 동작을 구성한다. 몸통은 하단의 가려진 골반 쪽으로 이어지며 떠 있는 신체나 지지 없는 물체는 없다. 침구는 침대 위에 놓이고 난간은 수직 지지대로 고정되어 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.349
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.349
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1349
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "요청된 '역동적인 상체'와 '허공을 향해 번쩍 치켜든 손'의 포즈를 극적인 구도로 완벽하게 구현했으며, 명시된 흉부 장갑의 손상 흔적과 침대 가장자리 프레이밍 지시도 정확히 따랐습니다."
   },
   {
    "label": "B",
    "score": 1349,
    "verdict_ko": "지시문이 요구한 '역동적인 상체'와 달리 자세가 몹시 정적이고 경직되어 있으며, 흉부의 손상된 묘사가 누락되었습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S76sh6_sel.png",
    "asset_id": "0da16e26-f33a-46a6-a4bc-b03005046ef4",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-be58-7fd5-9a6c-60974d52dbbb",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S76sh6"
  }
 },
 "S76sh12::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T11:59:55.580974+00:00",
  "fingerprint": "9f3022b57881ca8f130ca2bc73eb232c2d57607c134ba5c200dce483b7b8b46d",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S76sh12_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S76sh12_sel.png",
  "source_sha256": "37d3b240c1947f112f7cb227c48680ea0232f04b836caeb9d2d8efd3bad4b8e0",
  "file": "S76sh12_cine.png",
  "staged_sha256": "8e6985f34a17c2e3e8a2ae897c12262bdebf18d75667a5b8df10b2f7cc9b36ca",
  "latency_ms": 10153
 },
 "S77sh22::signage": {
  "fp": "01b988ad1f408053",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S77sh22": {
  "input_fingerprint": "13de94207149f315",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 작동하는 비행기들 한가운데 서서 양팔을 허리에 얹은 채 당당한 포즈를 취한 찰리의 전신.\n\nLOCATION (lock): Inside the abandoned airfield's aircraft-storage hangar, among newly running aircraft and bulbs fluctuating brightly overhead. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Operating light aircraft (Several previously disabled aircraft have started and are emitting black exhaust) — Different oblique sides of the aircraft flank 찰리 and recede behind him; used as Visible proof of the repair, distributed across the composition with realistic scale; Hangar floor (Visible around 찰리 and the aircraft); used as Ground plane that anchors the full-body pose and aircraft spacing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Existing nighttime hangar illumination provides controlled separation through the aircrafts' black exhaust, keeping 찰리 readable without carrying the earlier electrical flicker forward as a new effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Several formerly nonfunctional light aircraft now have running engines and emit black exhaust; a bulb has burst during the power surge. Charlie remains physically battered, with his chest opening exposed despite successfully powering the aircraft.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 작동하는 비행기들 한가운데 서서 양팔을 허리에 얹은 채 당당한 포즈를 취한 찰리의 전신.\n\nLOCATION (lock): Inside the abandoned airfield's aircraft-storage hangar, among newly running aircraft and bulbs fluctuating brightly overhead. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Operating light aircraft (Several previously disabled aircraft have started and are emitting black exhaust) — Different oblique sides of the aircraft flank 찰리 and recede behind him; used as Visible proof of the repair, distributed across the composition with realistic scale; Hangar floor (Visible around 찰리 and the aircraft); used as Ground plane that anchors the full-body pose and aircraft spacing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Existing nighttime hangar illumination provides controlled separation through the aircrafts' black exhaust, keeping 찰리 readable without carrying the earlier electrical flicker forward as a new effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Several formerly nonfunctional light aircraft now have running engines and emit black exhaust; a bulb has burst during the power surge. Charlie remains physically battered, with his chest opening exposed despite successfully powering the aircraft.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 작동하는 비행기들 한가운데 서서 양팔을 허리에 얹은 채 당당한 포즈를 취한 찰리의 전신.\n\nLOCATION (lock): Inside the abandoned airfield's aircraft-storage hangar, among newly running aircraft and bulbs fluctuating brightly overhead. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Operating light aircraft (Several previously disabled aircraft have started and are emitting black exhaust) — Different oblique sides of the aircraft flank 찰리 and recede behind him; used as Visible proof of the repair, distributed across the composition with realistic scale; Hangar floor (Visible around 찰리 and the aircraft); used as Ground plane that anchors the full-body pose and aircraft spacing.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Existing nighttime hangar illumination provides controlled separation through the aircrafts' black exhaust, keeping 찰리 readable without carrying the earlier electrical flicker forward as a new effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Several formerly nonfunctional light aircraft now have running engines and emit black exhaust; a bulb has burst during the power surge. Charlie remains physically battered, with his chest opening exposed despite successfully powering the aircraft.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "찰리는 카메라를 정면으로 바라보며 양손을 허리에 얹고 있음. 양옆의 비행기들은 비스듬히 안쪽을 향함.",
    "built_space": "금속 뼈대가 드러난 격납고 내부. 천장에 5개의 조명이 켜져 있으며, 바닥에 여러 대의 경비행기가 배치됨.",
    "entities": "찰리는 샌드 베이지색 고릴라형 외형과 흰색 마스크를 정확히 재현했으며 가슴 장갑이 열려 있음. 주변 비행기들도 레퍼런스의 외형과 줄무늬를 일치시킴.",
    "hard_violations": [],
    "physics": "찰리는 두 발로 바닥을 견고하게 딛고 서 있음. 비행기들은 랜딩 기어로 바닥에 지지되어 있으며 프로펠러가 정상적으로 회전함."
   },
   {
    "label": "B",
    "direction": "찰리가 정면을 향해 양팔을 허리에 얹고 서 있음. 비행기들의 기수가 찰리를 향해 비스듬히 놓여 있음.",
    "built_space": "철골 구조의 격납고 천장 아래 5개의 조명 패널이 켜져 있음. 비행기들이 콘크리트 바닥 공간을 채움.",
    "entities": "찰리의 외형은 레퍼런스와 일치하나 가슴 장갑은 뜯겨나간 구멍 형태로 묘사됨. 비행기들은 검은 연기를 뿜고 있음.",
    "hard_violations": [],
    "physics": "찰리는 바닥에 안정적으로 서 있음. 비행기 바퀴들이 지면에 닿아 있고, 배기가스가 위로 솟아오름."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "작동 중인 비행기들 사이에서 양손을 허리에 얹은 당당한 포즈와 열린 가슴 장갑의 디테일을 프롬프트에 맞게 매우 훌륭하게 구현함."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지정된 샷 크기와 포즈를 잘 따랐으나, 가슴 장갑의 개방 상태와 배기가스의 퍼짐이 A에 비해 약간 덜 자연스러움."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 카메라를 정면으로 바라보며 양손을 허리에 얹고 있음. 양옆의 비행기들은 비스듬히 안쪽을 향함.",
        "built_space": "금속 뼈대가 드러난 격납고 내부. 천장에 5개의 조명이 켜져 있으며, 바닥에 여러 대의 경비행기가 배치됨.",
        "entities": "찰리는 샌드 베이지색 고릴라형 외형과 흰색 마스크를 정확히 재현했으며 가슴 장갑이 열려 있음. 주변 비행기들도 레퍼런스의 외형과 줄무늬를 일치시킴.",
        "hard_violations": [],
        "physics": "찰리는 두 발로 바닥을 견고하게 딛고 서 있음. 비행기들은 랜딩 기어로 바닥에 지지되어 있으며 프로펠러가 정상적으로 회전함."
       },
       {
        "label": "B",
        "direction": "찰리가 정면을 향해 양팔을 허리에 얹고 서 있음. 비행기들의 기수가 찰리를 향해 비스듬히 놓여 있음.",
        "built_space": "철골 구조의 격납고 천장 아래 5개의 조명 패널이 켜져 있음. 비행기들이 콘크리트 바닥 공간을 채움.",
        "entities": "찰리의 외형은 레퍼런스와 일치하나 가슴 장갑은 뜯겨나간 구멍 형태로 묘사됨. 비행기들은 검은 연기를 뿜고 있음.",
        "hard_violations": [],
        "physics": "찰리는 바닥에 안정적으로 서 있음. 비행기 바퀴들이 지면에 닿아 있고, 배기가스가 위로 솟아오름."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "작동 중인 비행기들 사이에서 양손을 허리에 얹은 당당한 포즈와 열린 가슴 장갑의 디테일을 프롬프트에 맞게 매우 훌륭하게 구현함."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지정된 샷 크기와 포즈를 잘 따랐으나, 가슴 장갑의 개방 상태와 배기가스의 퍼짐이 A에 비해 약간 덜 자연스러움."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 카메라를 정면으로 바라보며 양손을 허리에 얹고 있음. 양옆의 비행기들은 비스듬히 안쪽을 향함.",
        "built_space": "금속 뼈대가 드러난 격납고 내부. 천장에 5개의 조명이 켜져 있으며, 바닥에 여러 대의 경비행기가 배치됨.",
        "entities": "찰리는 샌드 베이지색 고릴라형 외형과 흰색 마스크를 정확히 재현했으며 가슴 장갑이 열려 있음. 주변 비행기들도 레퍼런스의 외형과 줄무늬를 일치시킴.",
        "hard_violations": [],
        "physics": "찰리는 두 발로 바닥을 견고하게 딛고 서 있음. 비행기들은 랜딩 기어로 바닥에 지지되어 있으며 프로펠러가 정상적으로 회전함."
       },
       {
        "label": "B",
        "direction": "찰리가 정면을 향해 양팔을 허리에 얹고 서 있음. 비행기들의 기수가 찰리를 향해 비스듬히 놓여 있음.",
        "built_space": "철골 구조의 격납고 천장 아래 5개의 조명 패널이 켜져 있음. 비행기들이 콘크리트 바닥 공간을 채움.",
        "entities": "찰리의 외형은 레퍼런스와 일치하나 가슴 장갑은 뜯겨나간 구멍 형태로 묘사됨. 비행기들은 검은 연기를 뿜고 있음.",
        "hard_violations": [],
        "physics": "찰리는 바닥에 안정적으로 서 있음. 비행기 바퀴들이 지면에 닿아 있고, 배기가스가 위로 솟아오름."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "비행기 사이에서 양손을 허리에 둔 찰리의 전신 와이드숏을 충족하며, 회전하는 프로펠러와 검은 배기, 크게 파손되어 열린 가슴이 작동 성공 직후의 상태를 더 명확히 전달한다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "전신 와이드숏과 당당한 허리손 자세는 충실하지만, 일부 프로펠러의 작동 표현과 가슴 개구부의 파손 상태가 B보다 약하게 드러난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 얼굴과 눈은 거의 카메라 정면을 향하며, 양쪽 팔꿈치는 바깥으로 벌어지고 손은 허리 양옆에 닿는다. 앞쪽 두 비행기의 기수는 각각 화면 중앙 쪽으로 비스듬히 향하고 뒤쪽 기체도 중앙 통로를 둘러싼다. 특정 응시 대상이나 이동 방향을 요구하지 않은 장면에 맞는다.",
        "built_space": "철골 지붕과 골강판 벽, 양측 상부 창열, 얼룩지고 갈라진 콘크리트 바닥이 보인다. 켜진 천장 조명은 여섯 곳이며, 경비행기는 전경 좌우 한 대씩과 후경 좌우 한 대씩 적어도 네 대가 식별된다. 찰리는 중앙 통로에 서 있고 전신과 주변 바닥이 모두 들어온다. 참고 장소의 낡은 격납고 재료와 적갈색·청색 줄무늬 기체를 유지한다. 파열된 전구는 식별되지 않으며, 참고 사진만으로 조명 총수의 일치 여부는 확인할 수 없다.",
        "entities": "찰리 한 개체만 있으며 이전 장면의 남성은 없다. 흰 각진 마스크형 얼굴, 주황색 눈, 안테나, 샌드 베이지 장갑판과 육중한 팔은 캐릭터 참고와 일치한다. 얼굴이 기계 마스크로 덮여 있어 인간의 나이·성별·민족적 외양은 확인할 수 없다. 가슴 중앙에 작은 개구부와 노출 배선이 있고 장갑에 마모가 있으나, 심하게 얻어맞은 상태의 표현은 비교적 약하다. 낡은 경비행기들과 여러 검은 연기 기둥이 보인다.",
        "hard_violations": [],
        "physics": "찰리의 두 발바닥이 바닥에 닿아 체중을 지지하며 양손도 허리 장갑에 접촉한다. 비행기들은 착륙장치 바퀴로 바닥에 지지되고 천장 조명은 지붕 구조에 매달려 있다. 일부 프로펠러는 흐리지만 오른쪽 전경 프로펠러는 날개가 선명해 여러 엔진의 가동 증거가 균일하지 않다. 검은 연기는 기체 뒤에서 상승하지만 배기구와의 연결은 가려져 있다. 지지 없이 떠 있는 몸체나 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "찰리는 얼굴을 카메라 정면으로 들고 양팔을 벌려 양손을 허리 양옆에 얹는다. 좌우 전경 비행기의 기수는 중앙 통로 쪽을 비스듬히 향하며 후경 기체도 서로 다른 사선 측면을 드러낸다. 찰리가 비행기들 한가운데 당당히 서 있다는 관계가 명확하다.",
        "built_space": "철골 트러스 지붕, 골강판 벽, 양측 창열과 마모된 콘크리트 바닥이 보인다. 천장에는 켜진 조명 여섯 곳이 식별되고, 경비행기는 전경 좌우 두 대와 후경 좌우 두 대가 확실히 보인다. 찰리는 날개 사이 중앙 통로에 있으며 머리부터 발끝까지 포함된 와이드숏이다. 기체 일부는 화면 가장자리에서 잘리지만 주변 바닥과 후경으로 이어지는 배치는 유지된다. 참고의 격납고 재료와 낡은 줄무늬 항공기 외관에 부합한다. 파열된 전구는 확인되지 않는다.",
        "entities": "등장 개체는 찰리 하나뿐이다. 흰 마스크형 얼굴과 주황색 눈, 안테나, 각진 베이지 장갑, 큰 어깨와 긴 팔이 참고 캐릭터와 일치한다. 인간의 얼굴이나 피부는 드러나지 않는다. 가슴 장갑이 크게 깨져 내부 장치와 배선이 노출되어 있으며 그을림과 긁힘이 보여, 손상된 채 가슴 개구부를 드러낸다는 상태 지시가 분명하다. 주변에는 낡은 경비행기들과 검은 배기가 있고 이전 장면 인물은 없다.",
        "hard_violations": [],
        "physics": "찰리는 벌린 두 발을 바닥에 붙여 서 있고 양손을 허리 장갑에 댄다. 기계 관절의 굽힘도 해당 자세를 지지한다. 비행기는 착륙장치로 바닥에 서 있으며 전경 프로펠러에 회전 흐림이 보여 엔진 가동을 뒷받침한다. 연기는 기체 주변에서 위로 퍼지고, 조명은 천장에 고정되어 있다. 배기구 자체는 대부분 가려져 있지만 지지 없는 부유 물체나 불가능한 신체 자세는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "비행기 사이에서 양손을 허리에 둔 찰리의 전신 와이드숏을 충족하며, 회전하는 프로펠러와 검은 배기, 크게 파손되어 열린 가슴이 작동 성공 직후의 상태를 더 명확히 전달한다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "전신 와이드숏과 당당한 허리손 자세는 충실하지만, 일부 프로펠러의 작동 표현과 가슴 개구부의 파손 상태가 B보다 약하게 드러난다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 얼굴과 눈은 거의 카메라 정면을 향하며, 양쪽 팔꿈치는 바깥으로 벌어지고 손은 허리 양옆에 닿는다. 앞쪽 두 비행기의 기수는 각각 화면 중앙 쪽으로 비스듬히 향하고 뒤쪽 기체도 중앙 통로를 둘러싼다. 특정 응시 대상이나 이동 방향을 요구하지 않은 장면에 맞는다.",
        "built_space": "철골 지붕과 골강판 벽, 양측 상부 창열, 얼룩지고 갈라진 콘크리트 바닥이 보인다. 켜진 천장 조명은 여섯 곳이며, 경비행기는 전경 좌우 한 대씩과 후경 좌우 한 대씩 적어도 네 대가 식별된다. 찰리는 중앙 통로에 서 있고 전신과 주변 바닥이 모두 들어온다. 참고 장소의 낡은 격납고 재료와 적갈색·청색 줄무늬 기체를 유지한다. 파열된 전구는 식별되지 않으며, 참고 사진만으로 조명 총수의 일치 여부는 확인할 수 없다.",
        "entities": "찰리 한 개체만 있으며 이전 장면의 남성은 없다. 흰 각진 마스크형 얼굴, 주황색 눈, 안테나, 샌드 베이지 장갑판과 육중한 팔은 캐릭터 참고와 일치한다. 얼굴이 기계 마스크로 덮여 있어 인간의 나이·성별·민족적 외양은 확인할 수 없다. 가슴 중앙에 작은 개구부와 노출 배선이 있고 장갑에 마모가 있으나, 심하게 얻어맞은 상태의 표현은 비교적 약하다. 낡은 경비행기들과 여러 검은 연기 기둥이 보인다.",
        "hard_violations": [],
        "physics": "찰리의 두 발바닥이 바닥에 닿아 체중을 지지하며 양손도 허리 장갑에 접촉한다. 비행기들은 착륙장치 바퀴로 바닥에 지지되고 천장 조명은 지붕 구조에 매달려 있다. 일부 프로펠러는 흐리지만 오른쪽 전경 프로펠러는 날개가 선명해 여러 엔진의 가동 증거가 균일하지 않다. 검은 연기는 기체 뒤에서 상승하지만 배기구와의 연결은 가려져 있다. 지지 없이 떠 있는 몸체나 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "찰리는 얼굴을 카메라 정면으로 들고 양팔을 벌려 양손을 허리 양옆에 얹는다. 좌우 전경 비행기의 기수는 중앙 통로 쪽을 비스듬히 향하며 후경 기체도 서로 다른 사선 측면을 드러낸다. 찰리가 비행기들 한가운데 당당히 서 있다는 관계가 명확하다.",
        "built_space": "철골 트러스 지붕, 골강판 벽, 양측 창열과 마모된 콘크리트 바닥이 보인다. 천장에는 켜진 조명 여섯 곳이 식별되고, 경비행기는 전경 좌우 두 대와 후경 좌우 두 대가 확실히 보인다. 찰리는 날개 사이 중앙 통로에 있으며 머리부터 발끝까지 포함된 와이드숏이다. 기체 일부는 화면 가장자리에서 잘리지만 주변 바닥과 후경으로 이어지는 배치는 유지된다. 참고의 격납고 재료와 낡은 줄무늬 항공기 외관에 부합한다. 파열된 전구는 확인되지 않는다.",
        "entities": "등장 개체는 찰리 하나뿐이다. 흰 마스크형 얼굴과 주황색 눈, 안테나, 각진 베이지 장갑, 큰 어깨와 긴 팔이 참고 캐릭터와 일치한다. 인간의 얼굴이나 피부는 드러나지 않는다. 가슴 장갑이 크게 깨져 내부 장치와 배선이 노출되어 있으며 그을림과 긁힘이 보여, 손상된 채 가슴 개구부를 드러낸다는 상태 지시가 분명하다. 주변에는 낡은 경비행기들과 검은 배기가 있고 이전 장면 인물은 없다.",
        "hard_violations": [],
        "physics": "찰리는 벌린 두 발을 바닥에 붙여 서 있고 양손을 허리 장갑에 댄다. 기계 관절의 굽힘도 해당 자세를 지지한다. 비행기는 착륙장치로 바닥에 서 있으며 전경 프로펠러에 회전 흐림이 보여 엔진 가동을 뒷받침한다. 연기는 기체 주변에서 위로 퍼지고, 조명은 천장에 고정되어 있다. 배기구 자체는 대부분 가려져 있지만 지지 없는 부유 물체나 불가능한 신체 자세는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.764
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.764
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1764
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "작동 중인 비행기들 사이에서 양손을 허리에 얹은 당당한 포즈와 열린 가슴 장갑의 디테일을 프롬프트에 맞게 매우 훌륭하게 구현함."
   },
   {
    "label": "B",
    "score": 1764,
    "verdict_ko": "지정된 샷 크기와 포즈를 잘 따랐으나, 가슴 장갑의 개방 상태와 배기가스의 퍼짐이 A에 비해 약간 덜 자연스러움."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S75sh5_sel.png",
    "asset_id": "80915e03-5b2d-49db-a50d-37a66b64d032",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-c008-7b22-a881-45064fe1a4eb",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S75sh5"
  }
 },
 "S77sh22::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T12:00:56.669627+00:00",
  "fingerprint": "63a3b69479a932af3ca2a8babffeb9e7e081dd98e70808f34842fcf03921aaa5",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S77sh22_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S77sh22_sel.png",
  "source_sha256": "a13b5dc87a8a0bc8dffe0d07036e8c9b349c69485b3d4a8fdf49c013a107b03f",
  "file": "S77sh22_cine.png",
  "staged_sha256": "df80f3d6f8af778e5493d6430e4d377ff566b1f664f5b88c540c19783a0990c7",
  "latency_ms": 9857
 },
 "S77sh43::signage": {
  "fp": "735e290f8cbaf775",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::fe03d1436e8ce9a0": {
  "subjects": [],
  "subject_text": "목포공항 격납고\n거대한 창고 문과 넓은 바닥을 갖춘 낡은 격납고. 높고 깊은 실내 위로 전구 조명이 설치되어 있다.",
  "identity": "canonical",
  "scope_id": "L257",
  "scope_role": "location_interior",
  "scope_sha": "a9d7f00b16b22d2b"
 },
 "S77sh43::bgfirst_bg": {
  "input_fingerprint": "0cf81d7fb6a2acd1",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 멀어지는 헬리콥터를 향해 커다란 금속 팔을 힘차게 흔드는 찰리와 그 옆에 우뚝 선 현우의 뒷모습.\n\nLOCATION (lock): On the open airfield apron in daylight, beside the helicopter's takeoff point and outside the hangar.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Distant airborne helicopter in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Departing helicopter (Airborne and receding from the two figures) — Seen obliquely from below and behind as it moves away; used as Small upper-right destination for both figures' attention and 찰리's farewell; Airport ground (The two figures remain on the ground after the helicopter's departure); used as Lower framing band establishing their separation from the airborne passengers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained daylight and controlled tonal contrast preserve the two rear figures against the open space around the receding helicopter.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 멀어지는 헬리콥터를 향해 커다란 금속 팔을 힘차게 흔드는 찰리와 그 옆에 우뚝 선 현우의 뒷모습.\n\nLOCATION (lock): On the open airfield apron in daylight, beside the helicopter's takeoff point and outside the hangar.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Distant airborne helicopter in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Departing helicopter (Airborne and receding from the two figures) — Seen obliquely from below and behind as it moves away; used as Small upper-right destination for both figures' attention and 찰리's farewell; Airport ground (The two figures remain on the ground after the helicopter's departure); used as Lower framing band establishing their separation from the airborne passengers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained daylight and controlled tonal contrast preserve the two rear figures against the open space around the receding helicopter.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S77sh43__bgfirst_bg.png",
  "asset_id": "aadf8628-2de5-41d1-ab0c-9466e0ee03d4",
  "input_asset_ids": [
   "5d426069-57f7-4d5a-b31e-f5848a2bcb37",
   "6a2d2684-afeb-4aa6-b80b-232ee22c7f48"
  ]
 },
 "S77sh43": {
  "input_fingerprint": "bccd7e606245e065",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 멀어지는 헬리콥터를 향해 커다란 금속 팔을 힘차게 흔드는 찰리와 그 옆에 우뚝 선 현우의 뒷모습.\n\nLOCATION (lock): On the open airfield apron in daylight, beside the helicopter's takeoff point and outside the hangar. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Distant airborne helicopter in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Departing helicopter (Airborne and receding from the two figures) — Seen obliquely from below and behind as it moves away; used as Small upper-right destination for both figures' attention and 찰리's farewell; Airport ground (The two figures remain on the ground after the helicopter's departure); used as Lower framing band establishing their separation from the airborne passengers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained daylight and controlled tonal contrast preserve the two rear figures against the open space around the receding helicopter.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): It is now daylight, and the helicopter is airborne and receding from the airport. Charlie remains on the ground with his unrepaired body damage and exposed chest opening. 현우: He remains on the ground at the airport, still dirty and battered.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 멀어지는 헬리콥터를 향해 커다란 금속 팔을 힘차게 흔드는 찰리와 그 옆에 우뚝 선 현우의 뒷모습.\n\nLOCATION (lock): On the open airfield apron in daylight, beside the helicopter's takeoff point and outside the hangar. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Distant airborne helicopter in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Departing helicopter (Airborne and receding from the two figures) — Seen obliquely from below and behind as it moves away; used as Small upper-right destination for both figures' attention and 찰리's farewell; Airport ground (The two figures remain on the ground after the helicopter's departure); used as Lower framing band establishing their separation from the airborne passengers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained daylight and controlled tonal contrast preserve the two rear figures against the open space around the receding helicopter.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): It is now daylight, and the helicopter is airborne and receding from the airport. Charlie remains on the ground with his unrepaired body damage and exposed chest opening. 현우: He remains on the ground at the airport, still dirty and battered.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 멀어지는 헬리콥터를 향해 커다란 금속 팔을 힘차게 흔드는 찰리와 그 옆에 우뚝 선 현우의 뒷모습.\n\nLOCATION (lock): On the open airfield apron in daylight, beside the helicopter's takeoff point and outside the hangar. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Distant airborne helicopter in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Departing helicopter (Airborne and receding from the two figures) — Seen obliquely from below and behind as it moves away; used as Small upper-right destination for both figures' attention and 찰리's farewell; Airport ground (The two figures remain on the ground after the helicopter's departure); used as Lower framing band establishing their separation from the airborne passengers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained daylight and controlled tonal contrast preserve the two rear figures against the open space around the receding helicopter.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): It is now daylight, and the helicopter is airborne and receding from the airport. Charlie remains on the ground with his unrepaired body damage and exposed chest opening. 현우: He remains on the ground at the airport, still dirty and battered.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S77sh43__bgfirst_bg.png",
     "asset_id": "aadf8628-2de5-41d1-ab0c-9466e0ee03d4",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S77sh43.png",
     "asset_id": "5d426069-57f7-4d5a-b31e-f5848a2bcb37",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L257B01.png",
     "asset_id": "6a2d2684-afeb-4aa6-b80b-232ee22c7f48",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "찰리와 현우 모두 우측 상단 하늘에 있는 헬리콥터를 향해 서 있으며, 찰리는 오른팔을 뻗어 그쪽을 겨냥하고 있음.",
    "built_space": "기준 사진과 동일한 왼쪽의 격납고 외벽, 활주로 바닥, 원경의 산맥이 야간 환경에 맞게 배치되어 있음.",
    "entities": "우측 상단의 헬리콥터, 뒷모습의 현우(검은 머리, 흰 티셔츠), 뒷모습의 찰리(베이지색 로봇)가 존재하나, 인물들이 2D 애니메이션 스타일로 그려짐. 찰리의 등에 일부 파손 흔적이 보임.",
    "hard_violations": [
     "[gemini-pro] 실사화 위반: 인물(찰리와 현우)이 실제 물리적 재질이 아닌 평면적인 2D 그래픽/일러스트레이션으로 렌더링됨."
    ],
    "physics": "두 인물 모두 지면에 발을 딛고 서 있으며, 찰리의 들린 팔은 몸체에 연결되어 지탱됨."
   },
   {
    "label": "B",
    "direction": "찰리와 현우가 우측 상단 하늘의 헬리콥터를 바라보고 있으며, 찰리가 오른팔을 헬리콥터 방향으로 들어 올림.",
    "built_space": "기준 사진의 장소인 왼쪽 격납고 셔터와 외벽, 콘크리트 활주로, 배경의 산맥이 정확한 위치에 야간 조명으로 구현됨.",
    "entities": "우측 상단에 위치한 헬리콥터, 더러워진 작업복과 헝클어진 머리를 한 현우의 뒷모습, 육중한 체구와 모래색 장갑을 지닌 찰리의 뒷모습이 모두 실사 재질로 렌더링됨. 찰리의 등에 커다란 기계부 노출 손상이 있음.",
    "hard_violations": [],
    "physics": "두 인물은 활주로 바닥에 안정적으로 서 있으며, 찰리의 들어 올린 팔은 어깨 관절에 의해 자연스럽게 지지됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "캐릭터가 2D 일러스트처럼 평면적으로 렌더링되어 실사 사진 지침을 심각하게 위반했습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "실사 렌더링 지침을 준수하였으며, 멀어지는 헬리콥터와 파손된 로봇이 손을 흔드는 뒷모습 구도를 정확히 구현했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리와 현우 모두 우측 상단 하늘에 있는 헬리콥터를 향해 서 있으며, 찰리는 오른팔을 뻗어 그쪽을 겨냥하고 있음.",
        "built_space": "기준 사진과 동일한 왼쪽의 격납고 외벽, 활주로 바닥, 원경의 산맥이 야간 환경에 맞게 배치되어 있음.",
        "entities": "우측 상단의 헬리콥터, 뒷모습의 현우(검은 머리, 흰 티셔츠), 뒷모습의 찰리(베이지색 로봇)가 존재하나, 인물들이 2D 애니메이션 스타일로 그려짐. 찰리의 등에 일부 파손 흔적이 보임.",
        "hard_violations": [
         "실사화 위반: 인물(찰리와 현우)이 실제 물리적 재질이 아닌 평면적인 2D 그래픽/일러스트레이션으로 렌더링됨."
        ],
        "physics": "두 인물 모두 지면에 발을 딛고 서 있으며, 찰리의 들린 팔은 몸체에 연결되어 지탱됨."
       },
       {
        "label": "B",
        "direction": "찰리와 현우가 우측 상단 하늘의 헬리콥터를 바라보고 있으며, 찰리가 오른팔을 헬리콥터 방향으로 들어 올림.",
        "built_space": "기준 사진의 장소인 왼쪽 격납고 셔터와 외벽, 콘크리트 활주로, 배경의 산맥이 정확한 위치에 야간 조명으로 구현됨.",
        "entities": "우측 상단에 위치한 헬리콥터, 더러워진 작업복과 헝클어진 머리를 한 현우의 뒷모습, 육중한 체구와 모래색 장갑을 지닌 찰리의 뒷모습이 모두 실사 재질로 렌더링됨. 찰리의 등에 커다란 기계부 노출 손상이 있음.",
        "hard_violations": [],
        "physics": "두 인물은 활주로 바닥에 안정적으로 서 있으며, 찰리의 들어 올린 팔은 어깨 관절에 의해 자연스럽게 지지됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "캐릭터가 2D 일러스트처럼 평면적으로 렌더링되어 실사 사진 지침을 심각하게 위반했습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "실사 렌더링 지침을 준수하였으며, 멀어지는 헬리콥터와 파손된 로봇이 손을 흔드는 뒷모습 구도를 정확히 구현했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리와 현우 모두 우측 상단 하늘에 있는 헬리콥터를 향해 서 있으며, 찰리는 오른팔을 뻗어 그쪽을 겨냥하고 있음.",
        "built_space": "기준 사진과 동일한 왼쪽의 격납고 외벽, 활주로 바닥, 원경의 산맥이 야간 환경에 맞게 배치되어 있음.",
        "entities": "우측 상단의 헬리콥터, 뒷모습의 현우(검은 머리, 흰 티셔츠), 뒷모습의 찰리(베이지색 로봇)가 존재하나, 인물들이 2D 애니메이션 스타일로 그려짐. 찰리의 등에 일부 파손 흔적이 보임.",
        "hard_violations": [
         "실사화 위반: 인물(찰리와 현우)이 실제 물리적 재질이 아닌 평면적인 2D 그래픽/일러스트레이션으로 렌더링됨."
        ],
        "physics": "두 인물 모두 지면에 발을 딛고 서 있으며, 찰리의 들린 팔은 몸체에 연결되어 지탱됨."
       },
       {
        "label": "B",
        "direction": "찰리와 현우가 우측 상단 하늘의 헬리콥터를 바라보고 있으며, 찰리가 오른팔을 헬리콥터 방향으로 들어 올림.",
        "built_space": "기준 사진의 장소인 왼쪽 격납고 셔터와 외벽, 콘크리트 활주로, 배경의 산맥이 정확한 위치에 야간 조명으로 구현됨.",
        "entities": "우측 상단에 위치한 헬리콥터, 더러워진 작업복과 헝클어진 머리를 한 현우의 뒷모습, 육중한 체구와 모래색 장갑을 지닌 찰리의 뒷모습이 모두 실사 재질로 렌더링됨. 찰리의 등에 커다란 기계부 노출 손상이 있음.",
        "hard_violations": [],
        "physics": "두 인물은 활주로 바닥에 안정적으로 서 있으며, 찰리의 들어 올린 팔은 어깨 관절에 의해 자연스럽게 지지됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "밤의 비행장과 작별하는 뒷모습, 찰리의 육중한 체형은 잘 맞지만, 찰리가 화면 높이 대부분을 차지해 지정된 와이드 구도의 열린 공간이 줄고 현우의 겉옷도 참조와 다릅니다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "찰리의 체형이 참조보다 날씬하지만, 나란히 선 두 뒷모습과 오른쪽 위로 멀어지는 작은 헬리콥터를 넓은 공간으로 분리한 와이드 구도가 핵심 연출에 더 충실합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 인물은 카메라에 등을 보이고 비행장 너머를 향합니다. 찰리는 오른팔을 헬리콥터가 있는 오른쪽 위로 들어 손을 펼쳤습니다. 현우의 머리도 그쪽 하늘을 향하지만 눈은 보이지 않아 정확한 시선은 확인할 수 없습니다. 헬리콥터는 꼬리가 왼쪽, 기수가 오른쪽으로 놓여 두 인물에게서 오른쪽 멀리 떠나는 모습이며, 아래쪽과 후측면이 보입니다.",
        "built_space": "왼쪽에 골강판 격납고 한 동과 큰 출입구 한 곳, 상부의 연속 창열, 출입구 옆 방호 기둥들이 보입니다. 두 인물은 격납고 밖 콘크리트 계류장에 서 있습니다. 균열과 물기, 노란 곡선 표식 및 먼 산은 장소 참조와 부합합니다. 전신을 포함하지만 찰리가 화면 높이 대부분을 차지하여 지면과 열린 공간의 비중이 상대적으로 작습니다. 물기 위 조명 반사는 가능한 배치입니다.",
        "entities": "찰리 한 명, 현우 한 명, 헬리콥터 한 대가 보이며 추가 인물이나 읽을 수 있는 글자는 없습니다. 찰리의 샌드 베이지 장갑판, 긴 금속 팔과 육중한 몸통은 참조에 가깝고 등에는 손상과 노출된 기계 부품이 있습니다. 흰 얼굴과 가슴 개구부는 뒷모습 때문에 확인할 수 없습니다. 현우는 헝클어진 검은 머리의 마른 젊은 남성으로 보이지만 얼굴과 정확한 연령·민족적 외양은 확인할 수 없습니다. 더러운 긴소매 겉옷과 카고 바지는 참조의 남색 반소매 상의와 다릅니다. 하늘과 인공조명은 명시된 밤 시간에 맞습니다.",
        "hard_violations": [],
        "physics": "찰리의 양발과 현우의 양발이 지면에 닿아 체중을 지지합니다. 찰리의 들어 올린 팔은 어깨·팔꿈치·손목 관절로 몸통에 연결되어 있으며 작별하며 흔드는 동작의 한 순간으로 가능합니다. 헬리콥터는 회전하는 주회전익으로 비행 중이며, 지지 없이 떠 있는 별도 물체는 보이지 않습니다."
       },
       {
        "label": "B",
        "direction": "두 인물은 등을 보인 채 오른쪽 위 하늘의 헬리콥터를 향해 서 있습니다. 찰리의 오른팔은 그 방향으로 비스듬히 올라가고 손바닥은 펼쳐져 작별 인사로 읽힙니다. 현우도 머리를 약간 오른쪽으로 향하지만 눈 자체는 보이지 않습니다. 헬리콥터는 왼쪽의 꼬리에서 오른쪽 기수로 이어지는 방향으로 날며, 아래에서 본 후측면이 드러나 멀어지는 배치와 부합합니다.",
        "built_space": "왼쪽에 골강판 격납고 한 동, 큰 열린 출입구 한 곳, 상부 창열과 출입구 주변 방호 기둥들이 보입니다. 두 인물은 그 밖의 계류장에 나란히 서 있고, 지면의 균열·물기·노란 곡선 표식과 먼 산이 참조 장소를 유지합니다. 인물들은 왼쪽과 중앙 아래에, 작은 헬리콥터는 오른쪽 위에 배치되어 넓은 하늘이 둘 사이를 분리합니다. 젖은 바닥의 격납고 조명 반사도 공간 배치상 가능합니다.",
        "entities": "찰리 한 명, 현우 한 명과 헬리콥터 한 대가 있으며 추가 인물이나 읽을 수 있는 글자는 없습니다. 찰리는 베이지 금속 장갑과 등 부분 손상을 갖췄지만, 참조보다 몸통과 팔이 가늘고 다리가 길어 고릴라형의 육중한 비율이 약합니다. 얼굴과 가슴 개구부는 뒤쪽 구도에서 보이지 않습니다. 현우는 헝클어진 검은 머리와 마른 청년 체형을 갖추고 더러운 반소매 상의를 입었으나, 상의는 참조의 남색보다 회색에 가깝습니다. 얼굴이 가려져 정확한 신원과 연령·민족적 외양은 검증할 수 없습니다. 달빛과 인공조명이 있는 밤입니다.",
        "hard_violations": [],
        "physics": "두 인물 모두 양발로 계류장 바닥을 딛고 있습니다. 찰리는 다리를 벌려 지지하고 어깨와 팔꿈치를 통해 오른팔을 들어 올려, 큰 팔을 흔드는 동작이 물리적으로 가능합니다. 현우도 지면에 체중이 실린 직립 자세입니다. 헬리콥터는 회전익의 회전 흔적이 보여 비행 지지가 설명되며, 지지 없는 신체나 물체는 보이지 않습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "밤의 비행장과 작별하는 뒷모습, 찰리의 육중한 체형은 잘 맞지만, 찰리가 화면 높이 대부분을 차지해 지정된 와이드 구도의 열린 공간이 줄고 현우의 겉옷도 참조와 다릅니다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "찰리의 체형이 참조보다 날씬하지만, 나란히 선 두 뒷모습과 오른쪽 위로 멀어지는 작은 헬리콥터를 넓은 공간으로 분리한 와이드 구도가 핵심 연출에 더 충실합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "두 인물은 카메라에 등을 보이고 비행장 너머를 향합니다. 찰리는 오른팔을 헬리콥터가 있는 오른쪽 위로 들어 손을 펼쳤습니다. 현우의 머리도 그쪽 하늘을 향하지만 눈은 보이지 않아 정확한 시선은 확인할 수 없습니다. 헬리콥터는 꼬리가 왼쪽, 기수가 오른쪽으로 놓여 두 인물에게서 오른쪽 멀리 떠나는 모습이며, 아래쪽과 후측면이 보입니다.",
        "built_space": "왼쪽에 골강판 격납고 한 동과 큰 출입구 한 곳, 상부의 연속 창열, 출입구 옆 방호 기둥들이 보입니다. 두 인물은 격납고 밖 콘크리트 계류장에 서 있습니다. 균열과 물기, 노란 곡선 표식 및 먼 산은 장소 참조와 부합합니다. 전신을 포함하지만 찰리가 화면 높이 대부분을 차지하여 지면과 열린 공간의 비중이 상대적으로 작습니다. 물기 위 조명 반사는 가능한 배치입니다.",
        "entities": "찰리 한 명, 현우 한 명, 헬리콥터 한 대가 보이며 추가 인물이나 읽을 수 있는 글자는 없습니다. 찰리의 샌드 베이지 장갑판, 긴 금속 팔과 육중한 몸통은 참조에 가깝고 등에는 손상과 노출된 기계 부품이 있습니다. 흰 얼굴과 가슴 개구부는 뒷모습 때문에 확인할 수 없습니다. 현우는 헝클어진 검은 머리의 마른 젊은 남성으로 보이지만 얼굴과 정확한 연령·민족적 외양은 확인할 수 없습니다. 더러운 긴소매 겉옷과 카고 바지는 참조의 남색 반소매 상의와 다릅니다. 하늘과 인공조명은 명시된 밤 시간에 맞습니다.",
        "hard_violations": [],
        "physics": "찰리의 양발과 현우의 양발이 지면에 닿아 체중을 지지합니다. 찰리의 들어 올린 팔은 어깨·팔꿈치·손목 관절로 몸통에 연결되어 있으며 작별하며 흔드는 동작의 한 순간으로 가능합니다. 헬리콥터는 회전하는 주회전익으로 비행 중이며, 지지 없이 떠 있는 별도 물체는 보이지 않습니다."
       },
       {
        "label": "A",
        "direction": "두 인물은 등을 보인 채 오른쪽 위 하늘의 헬리콥터를 향해 서 있습니다. 찰리의 오른팔은 그 방향으로 비스듬히 올라가고 손바닥은 펼쳐져 작별 인사로 읽힙니다. 현우도 머리를 약간 오른쪽으로 향하지만 눈 자체는 보이지 않습니다. 헬리콥터는 왼쪽의 꼬리에서 오른쪽 기수로 이어지는 방향으로 날며, 아래에서 본 후측면이 드러나 멀어지는 배치와 부합합니다.",
        "built_space": "왼쪽에 골강판 격납고 한 동, 큰 열린 출입구 한 곳, 상부 창열과 출입구 주변 방호 기둥들이 보입니다. 두 인물은 그 밖의 계류장에 나란히 서 있고, 지면의 균열·물기·노란 곡선 표식과 먼 산이 참조 장소를 유지합니다. 인물들은 왼쪽과 중앙 아래에, 작은 헬리콥터는 오른쪽 위에 배치되어 넓은 하늘이 둘 사이를 분리합니다. 젖은 바닥의 격납고 조명 반사도 공간 배치상 가능합니다.",
        "entities": "찰리 한 명, 현우 한 명과 헬리콥터 한 대가 있으며 추가 인물이나 읽을 수 있는 글자는 없습니다. 찰리는 베이지 금속 장갑과 등 부분 손상을 갖췄지만, 참조보다 몸통과 팔이 가늘고 다리가 길어 고릴라형의 육중한 비율이 약합니다. 얼굴과 가슴 개구부는 뒤쪽 구도에서 보이지 않습니다. 현우는 헝클어진 검은 머리와 마른 청년 체형을 갖추고 더러운 반소매 상의를 입었으나, 상의는 참조의 남색보다 회색에 가깝습니다. 얼굴이 가려져 정확한 신원과 연령·민족적 외양은 검증할 수 없습니다. 달빛과 인공조명이 있는 밤입니다.",
        "hard_violations": [],
        "physics": "두 인물 모두 양발로 계류장 바닥을 딛고 있습니다. 찰리는 다리를 벌려 지지하고 어깨와 팔꿈치를 통해 오른팔을 들어 올려, 큰 팔을 흔드는 동작이 물리적으로 가능합니다. 현우도 지면에 체중이 실린 직립 자세입니다. 헬리콥터는 회전익의 회전 흔적이 보여 비행 지지가 설명되며, 지지 없는 신체나 물체는 보이지 않습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.429,
    "B": 1.875
   },
   "adjusted": {
    "A": 1.179,
    "B": 1.875
   },
   "violations": {
    "A": [
     "[gemini-pro] 실사화 위반: 인물(찰리와 현우)이 실제 물리적 재질이 아닌 평면적인 2D 그래픽/일러스트레이션으로 렌더링됨."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "A": 1179,
   "B": 1875
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1179,
    "verdict_ko": "캐릭터가 2D 일러스트처럼 평면적으로 렌더링되어 실사 사진 지침을 심각하게 위반했습니다.  ★위반: [gemini-pro] 실사화 위반: 인물(찰리와 현우)이 실제 물리적 재질이 아닌 평면적인 2D 그래픽/일러스트레이션으로 렌더링됨."
   },
   {
    "label": "B",
    "score": 1875,
    "verdict_ko": "실사 렌더링 지침을 준수하였으며, 멀어지는 헬리콥터와 파손된 로봇이 손을 흔드는 뒷모습 구도를 정확히 구현했습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L257B01.png",
    "asset_id": "6a2d2684-afeb-4aa6-b80b-232ee22c7f48",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-c1ba-7040-9a4e-a4c4b0ebb91f",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S77sh43__bgfirst_bg.png",
   "bg_asset_id": "aadf8628-2de5-41d1-ab0c-9466e0ee03d4",
   "bg_record_key": "S77sh43::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S77sh43::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:46:37.618454+00:00",
  "fingerprint": "f837bdb7de034fe0fb3e4dbace9d315b573925e8d6470ec055fe60c31d404593",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S77sh43_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S77sh43_sel.png",
  "source_sha256": "1a40c87f1ee7173738399e52ed6007b4c093d1705694ce83a2648e99b7664ec3",
  "file": "S77sh43_cine.png",
  "staged_sha256": "7951cfe3739a41f6251f416d0475f3737932c8107acf16ceed768b4700ec7e27",
  "latency_ms": 10621
 },
 "S77sh53::signage": {
  "fp": "b019ba9531380923",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S77sh53": {
  "input_fingerprint": "c3ca90d56186f4a3",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 다시 앞을 향해 씩씩하게 걷는 도중 다리를 뻗어 딛은 mid-action 상태의 현우와 그 뒤를 따르며 한쪽 다리가 들린 mid-action 순간의 찰리가 멀어지는 뒷모습.\n\nLOCATION (lock): On the open ground leading away from the abandoned airfield's helicopter departure area in daylight. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Route continuing away from the camera in the upper-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Airport departure route (현우 and 찰리 are walking away along it) — The visible route recedes from the lower foreground toward the upper-center distance; used as Depth anchor for their increasing separation from the stationary camera; Backpacks (Worn by both departing figures) — Their outward-facing backs are visible against the figures' rear silhouettes; used as Small narrative details supporting the shared onward journey.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established restrained daylight and even tonal continuity, allowing the farewell's tenderness to persist without a new lighting cue.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The helicopter has disappeared into the distance. Charlie is now wearing a backpack over his still-damaged body, with no repair of the chest opening established. 현우: He walks away from the airport wearing a backpack, still dirty and battered.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 다시 앞을 향해 씩씩하게 걷는 도중 다리를 뻗어 딛은 mid-action 상태의 현우와 그 뒤를 따르며 한쪽 다리가 들린 mid-action 순간의 찰리가 멀어지는 뒷모습.\n\nLOCATION (lock): On the open ground leading away from the abandoned airfield's helicopter departure area in daylight. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Route continuing away from the camera in the upper-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Airport departure route (현우 and 찰리 are walking away along it) — The visible route recedes from the lower foreground toward the upper-center distance; used as Depth anchor for their increasing separation from the stationary camera; Backpacks (Worn by both departing figures) — Their outward-facing backs are visible against the figures' rear silhouettes; used as Small narrative details supporting the shared onward journey.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established restrained daylight and even tonal continuity, allowing the farewell's tenderness to persist without a new lighting cue.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The helicopter has disappeared into the distance. Charlie is now wearing a backpack over his still-damaged body, with no repair of the chest opening established. 현우: He walks away from the airport wearing a backpack, still dirty and battered.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 다시 앞을 향해 씩씩하게 걷는 도중 다리를 뻗어 딛은 mid-action 상태의 현우와 그 뒤를 따르며 한쪽 다리가 들린 mid-action 순간의 찰리가 멀어지는 뒷모습.\n\nLOCATION (lock): On the open ground leading away from the abandoned airfield's helicopter departure area in daylight. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Route continuing away from the camera in the upper-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Airport departure route (현우 and 찰리 are walking away along it) — The visible route recedes from the lower foreground toward the upper-center distance; used as Depth anchor for their increasing separation from the stationary camera; Backpacks (Worn by both departing figures) — Their outward-facing backs are visible against the figures' rear silhouettes; used as Small narrative details supporting the shared onward journey.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established restrained daylight and even tonal continuity, allowing the farewell's tenderness to persist without a new lighting cue.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The helicopter has disappeared into the distance. Charlie is now wearing a backpack over his still-damaged body, with no repair of the chest opening established. 현우: He walks away from the airport wearing a backpack, still dirty and battered.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "두 인물 모두 카메라를 등지고 멀어지는 방향을 향하고 있습니다.",
    "built_space": "이전 샷과 동일한 버려진 비행장 활주로와 좌측의 격납고 건물이 정확한 위치에 유지되어 있습니다.",
    "entities": "현우는 뒷모습으로 백팩을 메고 있으나, 찰리는 프롬프트의 지시와 달리 백팩을 메고 있지 않습니다. 캐릭터의 외형은 레퍼런스와 일치합니다.",
    "hard_violations": [
     "[gemini-pro] 찰리의 두 발이 모두 지면에서 떨어져 허공에 떠 있으며, 이를 지탱하거나 점프한 이유를 설명할 수 있는 물리적 근거가 전혀 없습니다.",
     "[gpt-high] 찰리의 양발 모두 포장면에서 떨어져 있고, 몸을 지지하는 접점이나 보행으로 설명할 수 있는 도약 동작이 없어 공중에 떠 있다."
    ],
    "physics": "현우는 한쪽 발로 지면을 딛고 걷고 있으나, 찰리는 어떠한 지지대나 추진력 없이 공중에 부양해 있습니다."
   },
   {
    "label": "B",
    "direction": "현우와 찰리 모두 카메라를 등지고 프레임 중앙부의 멀어지는 길을 향해 나아가고 있습니다.",
    "built_space": "이전 샷의 활주로, 좌측 격납고, 배경의 조명 및 산 실루엣 등이 지시된 위치와 원근감에 맞게 정확히 배치되어 있습니다.",
    "entities": "현우와 찰리 모두 명시된 대로 백팩을 메고 있으며, 뒷모습을 통해 두 캐릭터의 디테일과 착장이 레퍼런스와 일치함을 확인할 수 있습니다.",
    "hard_violations": [
     "[gpt-high] 현우 뒤를 따라야 하는 찰리가 현우보다 더 먼 지점에 배치되어 선후 관계가 반대로 보인다."
    ],
    "physics": "현우는 한쪽 발을 딛고 다른 발을 뻗는 자연스러운 걷기 동작을 취하고 있으며, 찰리 역시 한쪽 발을 땅에 굳건히 지탱하고 다른 다리를 들어 올린 걷기(mid-action) 상태를 물리적으로 올바르게 보여줍니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 10,
        "verdict_ko": "두 캐릭터 모두 백팩을 메고 멀어지는 역동적인 걷기 동작을 정확하게 구현하였으며 배경 및 설정과 완벽하게 일치합니다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "찰리가 백팩을 메고 있지 않으며, 아무런 지탱 없이 공중에 떠 있는 심각한 물리적 오류(Hard Violation)가 있습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 인물 모두 카메라를 등지고 멀어지는 방향을 향하고 있습니다.",
        "built_space": "이전 샷과 동일한 버려진 비행장 활주로와 좌측의 격납고 건물이 정확한 위치에 유지되어 있습니다.",
        "entities": "현우는 뒷모습으로 백팩을 메고 있으나, 찰리는 프롬프트의 지시와 달리 백팩을 메고 있지 않습니다. 캐릭터의 외형은 레퍼런스와 일치합니다.",
        "hard_violations": [
         "찰리의 두 발이 모두 지면에서 떨어져 허공에 떠 있으며, 이를 지탱하거나 점프한 이유를 설명할 수 있는 물리적 근거가 전혀 없습니다."
        ],
        "physics": "현우는 한쪽 발로 지면을 딛고 걷고 있으나, 찰리는 어떠한 지지대나 추진력 없이 공중에 부양해 있습니다."
       },
       {
        "label": "B",
        "direction": "현우와 찰리 모두 카메라를 등지고 프레임 중앙부의 멀어지는 길을 향해 나아가고 있습니다.",
        "built_space": "이전 샷의 활주로, 좌측 격납고, 배경의 조명 및 산 실루엣 등이 지시된 위치와 원근감에 맞게 정확히 배치되어 있습니다.",
        "entities": "현우와 찰리 모두 명시된 대로 백팩을 메고 있으며, 뒷모습을 통해 두 캐릭터의 디테일과 착장이 레퍼런스와 일치함을 확인할 수 있습니다.",
        "hard_violations": [],
        "physics": "현우는 한쪽 발을 딛고 다른 발을 뻗는 자연스러운 걷기 동작을 취하고 있으며, 찰리 역시 한쪽 발을 땅에 굳건히 지탱하고 다른 다리를 들어 올린 걷기(mid-action) 상태를 물리적으로 올바르게 보여줍니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 10,
        "verdict_ko": "두 캐릭터 모두 백팩을 메고 멀어지는 역동적인 걷기 동작을 정확하게 구현하였으며 배경 및 설정과 완벽하게 일치합니다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "찰리가 백팩을 메고 있지 않으며, 아무런 지탱 없이 공중에 떠 있는 심각한 물리적 오류(Hard Violation)가 있습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "두 인물 모두 카메라를 등지고 멀어지는 방향을 향하고 있습니다.",
        "built_space": "이전 샷과 동일한 버려진 비행장 활주로와 좌측의 격납고 건물이 정확한 위치에 유지되어 있습니다.",
        "entities": "현우는 뒷모습으로 백팩을 메고 있으나, 찰리는 프롬프트의 지시와 달리 백팩을 메고 있지 않습니다. 캐릭터의 외형은 레퍼런스와 일치합니다.",
        "hard_violations": [
         "찰리의 두 발이 모두 지면에서 떨어져 허공에 떠 있으며, 이를 지탱하거나 점프한 이유를 설명할 수 있는 물리적 근거가 전혀 없습니다."
        ],
        "physics": "현우는 한쪽 발로 지면을 딛고 걷고 있으나, 찰리는 어떠한 지지대나 추진력 없이 공중에 부양해 있습니다."
       },
       {
        "label": "B",
        "direction": "현우와 찰리 모두 카메라를 등지고 프레임 중앙부의 멀어지는 길을 향해 나아가고 있습니다.",
        "built_space": "이전 샷의 활주로, 좌측 격납고, 배경의 조명 및 산 실루엣 등이 지시된 위치와 원근감에 맞게 정확히 배치되어 있습니다.",
        "entities": "현우와 찰리 모두 명시된 대로 백팩을 메고 있으며, 뒷모습을 통해 두 캐릭터의 디테일과 착장이 레퍼런스와 일치함을 확인할 수 있습니다.",
        "hard_violations": [],
        "physics": "현우는 한쪽 발을 딛고 다른 발을 뻗는 자연스러운 걷기 동작을 취하고 있으며, 찰리 역시 한쪽 발을 땅에 굳건히 지탱하고 다른 다리를 들어 올린 걷기(mid-action) 상태를 물리적으로 올바르게 보여줍니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "야간 와이드 숏, 두 배낭과 지면에 지지된 보행은 충실하지만, 찰리가 현우보다 앞에 있어 뒤따르는 배치가 뒤집혔다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "찰리의 두 발이 지지 없이 공중에 떠 있고 배낭도 없으며, 인물들이 크게 잡혀 멀어지는 와이드 숏의 공간감도 약하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 인물 모두 뒤통수와 등을 카메라에 보이며 화면 위쪽 공항 공터로 걷는다. 얼굴과 눈은 보이지 않지만 머리와 몸의 방향은 진행 방향과 맞는다. 다만 왼쪽 찰리가 오른쪽 현우보다 화면 깊숙한 곳에 있어, 현우를 뒤따르기보다 앞서가는 배치다.",
        "built_space": "왼쪽 가장자리에 열린 격납고 일부가 있고 그 뒤로 별도의 격납고형 건물 한 동이 보인다. 전경의 갈라지고 젖은 콘크리트와 바닥 선이 상단 중앙의 먼 공터로 이어진다. 원경에는 조명 기둥 여러 개와 낮은 공항 건물, 산 능선이 있다. 기존 장소의 금속 외벽과 낡은 포장 재질은 이어지지만 뒤쪽 건물은 참고 사진에서 확인되지 않는다. 두 인물은 모두 장애물 없는 포장면 위에 있다.",
        "entities": "현우와 찰리 두 인물만 보이고 헬리콥터는 없다. 현우의 헝클어진 검은 머리, 마른 청년 체형, 더럽혀진 회갈색 상하의는 이전 장면과 부합한다. 후면이므로 얼굴과 정확한 나이·민족적 외모는 확인할 수 없다. 찰리는 베이지 장갑판, 큰 어깨와 긴 팔, 짧은 다리를 유지한다. 흰 얼굴과 흉부 손상은 후면 및 배낭에 가려 확인되지 않는다. 두 인물 모두 천 배낭을 메고 있으며 읽을 수 있는 글자는 없다. 야간 표현은 시간 잠금과 이전 장면에 맞는다.",
        "hard_violations": [
         "현우 뒤를 따라야 하는 찰리가 현우보다 더 먼 지점에 배치되어 선후 관계가 반대로 보인다."
        ],
        "physics": "찰리는 왼발을 포장면에 디디고 오른발을 들어 올려 지지점이 분명한 보행 자세다. 현우도 오른발이 지면을 지지하고 왼발이 뒤쪽으로 들려 있어 걸음 중간으로 읽힌다. 두 배낭은 어깨끈으로 몸에 고정되어 있다. 지지 없이 떠 있는 몸이나 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "왼쪽 현우와 오른쪽 찰리 모두 공터의 먼 방향으로 등을 돌리고 있다. 머리도 대체로 진행 방향을 향하며 카메라를 돌아보지 않는다. 둘이 거의 나란히 있어 찰리가 뒤따른다는 깊이 차이는 뚜렷하지 않다.",
        "built_space": "왼쪽에 격납고 외벽 한 면이 있고, 그 앞의 젖고 갈라진 포장면이 원경 공항 건물과 조명 기둥 여러 개로 이어진다. 배경 산 능선과 야간 조명은 이전 장면에 부합한다. 두 인물은 공터 중앙에 있으나 화면 높이 대부분을 차지해 A보다 가까운 전신 구도이며, 하단에서 상단 중앙으로 이어지는 출발 경로의 깊이가 덜 강조된다.",
        "entities": "현우와 찰리만 보이며 헬리콥터는 없다. 현우는 검은 헝클어진 머리와 더러운 회갈색 옷, 배낭을 유지한다. 후면이므로 얼굴의 정체성은 직접 확인할 수 없다. 찰리의 베이지 장갑판과 육중한 긴 팔은 참고와 유사하고 몸통의 파손 개구부도 보이지만, 반드시 메고 있어야 할 배낭이 전혀 없다. 흰 마스크 얼굴은 보이지 않는다. 읽을 수 있는 글자는 확인되지 않는다.",
        "hard_violations": [
         "찰리의 양발 모두 포장면에서 떨어져 있고, 몸을 지지하는 접점이나 보행으로 설명할 수 있는 도약 동작이 없어 공중에 떠 있다."
        ],
        "physics": "현우는 오른발로 지면을 지지하고 왼발을 들어 걷는 자세이며 배낭은 어깨끈으로 지지된다. 반면 찰리는 두 발바닥 아래로 지면과의 틈이 보이고 그림자도 떨어져 있다. 발, 손, 장치 중 어느 것도 몸을 지지하지 않으며, 도약의 추진이나 착지를 설명하는 자세도 없어 요구된 한쪽 다리를 든 보행으로 성립하지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "야간 와이드 숏, 두 배낭과 지면에 지지된 보행은 충실하지만, 찰리가 현우보다 앞에 있어 뒤따르는 배치가 뒤집혔다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "찰리의 두 발이 지지 없이 공중에 떠 있고 배낭도 없으며, 인물들이 크게 잡혀 멀어지는 와이드 숏의 공간감도 약하다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "두 인물 모두 뒤통수와 등을 카메라에 보이며 화면 위쪽 공항 공터로 걷는다. 얼굴과 눈은 보이지 않지만 머리와 몸의 방향은 진행 방향과 맞는다. 다만 왼쪽 찰리가 오른쪽 현우보다 화면 깊숙한 곳에 있어, 현우를 뒤따르기보다 앞서가는 배치다.",
        "built_space": "왼쪽 가장자리에 열린 격납고 일부가 있고 그 뒤로 별도의 격납고형 건물 한 동이 보인다. 전경의 갈라지고 젖은 콘크리트와 바닥 선이 상단 중앙의 먼 공터로 이어진다. 원경에는 조명 기둥 여러 개와 낮은 공항 건물, 산 능선이 있다. 기존 장소의 금속 외벽과 낡은 포장 재질은 이어지지만 뒤쪽 건물은 참고 사진에서 확인되지 않는다. 두 인물은 모두 장애물 없는 포장면 위에 있다.",
        "entities": "현우와 찰리 두 인물만 보이고 헬리콥터는 없다. 현우의 헝클어진 검은 머리, 마른 청년 체형, 더럽혀진 회갈색 상하의는 이전 장면과 부합한다. 후면이므로 얼굴과 정확한 나이·민족적 외모는 확인할 수 없다. 찰리는 베이지 장갑판, 큰 어깨와 긴 팔, 짧은 다리를 유지한다. 흰 얼굴과 흉부 손상은 후면 및 배낭에 가려 확인되지 않는다. 두 인물 모두 천 배낭을 메고 있으며 읽을 수 있는 글자는 없다. 야간 표현은 시간 잠금과 이전 장면에 맞는다.",
        "hard_violations": [
         "현우 뒤를 따라야 하는 찰리가 현우보다 더 먼 지점에 배치되어 선후 관계가 반대로 보인다."
        ],
        "physics": "찰리는 왼발을 포장면에 디디고 오른발을 들어 올려 지지점이 분명한 보행 자세다. 현우도 오른발이 지면을 지지하고 왼발이 뒤쪽으로 들려 있어 걸음 중간으로 읽힌다. 두 배낭은 어깨끈으로 몸에 고정되어 있다. 지지 없이 떠 있는 몸이나 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "왼쪽 현우와 오른쪽 찰리 모두 공터의 먼 방향으로 등을 돌리고 있다. 머리도 대체로 진행 방향을 향하며 카메라를 돌아보지 않는다. 둘이 거의 나란히 있어 찰리가 뒤따른다는 깊이 차이는 뚜렷하지 않다.",
        "built_space": "왼쪽에 격납고 외벽 한 면이 있고, 그 앞의 젖고 갈라진 포장면이 원경 공항 건물과 조명 기둥 여러 개로 이어진다. 배경 산 능선과 야간 조명은 이전 장면에 부합한다. 두 인물은 공터 중앙에 있으나 화면 높이 대부분을 차지해 A보다 가까운 전신 구도이며, 하단에서 상단 중앙으로 이어지는 출발 경로의 깊이가 덜 강조된다.",
        "entities": "현우와 찰리만 보이며 헬리콥터는 없다. 현우는 검은 헝클어진 머리와 더러운 회갈색 옷, 배낭을 유지한다. 후면이므로 얼굴의 정체성은 직접 확인할 수 없다. 찰리의 베이지 장갑판과 육중한 긴 팔은 참고와 유사하고 몸통의 파손 개구부도 보이지만, 반드시 메고 있어야 할 배낭이 전혀 없다. 흰 마스크 얼굴은 보이지 않는다. 읽을 수 있는 글자는 확인되지 않는다.",
        "hard_violations": [
         "찰리의 양발 모두 포장면에서 떨어져 있고, 몸을 지지하는 접점이나 보행으로 설명할 수 있는 도약 동작이 없어 공중에 떠 있다."
        ],
        "physics": "현우는 오른발로 지면을 지지하고 왼발을 들어 걷는 자세이며 배낭은 어깨끈으로 지지된다. 반면 찰리는 두 발바닥 아래로 지면과의 틈이 보이고 그림자도 떨어져 있다. 발, 손, 장치 중 어느 것도 몸을 지지하지 않으며, 도약의 추진이나 착지를 설명하는 자세도 없어 요구된 한쪽 다리를 든 보행으로 성립하지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.6,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.35,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 찰리의 두 발이 모두 지면에서 떨어져 허공에 떠 있으며, 이를 지탱하거나 점프한 이유를 설명할 수 있는 물리적 근거가 전혀 없습니다.",
     "[gpt-high] 찰리의 양발 모두 포장면에서 떨어져 있고, 몸을 지지하는 접점이나 보행으로 설명할 수 있는 도약 동작이 없어 공중에 떠 있다."
    ],
    "B": [
     "[gpt-high] 현우 뒤를 따라야 하는 찰리가 현우보다 더 먼 지점에 배치되어 선후 관계가 반대로 보인다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 1750,
   "A": 350
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "두 캐릭터 모두 백팩을 메고 멀어지는 역동적인 걷기 동작을 정확하게 구현하였으며 배경 및 설정과 완벽하게 일치합니다.  ★위반: [gpt-high] 현우 뒤를 따라야 하는 찰리가 현우보다 더 먼 지점에 배치되어 선후 관계가 반대로 보인다."
   },
   {
    "label": "A",
    "score": 350,
    "verdict_ko": "찰리가 백팩을 메고 있지 않으며, 아무런 지탱 없이 공중에 떠 있는 심각한 물리적 오류(Hard Violation)가 있습니다.  ★위반: [gemini-pro] 찰리의 두 발이 모두 지면에서 떨어져 허공에 떠 있으며, 이를 지탱하거나 점프한 이유를 설명할 수 있는 물리적 근거가 전혀 없습니다. / [gpt-high] 찰리의 양발 모두 포장면에서 떨어져 있고, 몸을 지지하는 접점이나 보행으로 설명할 수 있는 도약 동작이 없어 공중에 떠 있다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S77sh43_sel.png",
    "asset_id": "548daad7-0bc3-473b-b4c2-3e0b35a94322",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-c509-7bd8-897f-1cf125d61319",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S77sh43"
  }
 },
 "S77sh53::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:47:41.411259+00:00",
  "fingerprint": "c3c00150c1a8fd50aee8de0da2d0b9b29d6114e2d34c5dd870c77672f4738ab4",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S77sh53_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S77sh53_sel.png",
  "source_sha256": "397f7ee754908d5cb531b35d4d0b7de86897e3ae5c310db40d7897ba8177a9ee",
  "file": "S77sh53_cine.png",
  "staged_sha256": "a4939960296e02d5e22283bbf45c031e3c421def0658b41e8af71df8e0149272",
  "latency_ms": 10016
 },
 "S78sh2::signage": {
  "fp": "ae9740d88c7c9b67",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::e1bbaed3c4359ef2": {
  "subjects": [],
  "subject_text": "목포 항구 부둣가와 생선 가판대\n출렁이는 바다와 맞닿은 낡은 부두. 가장자리에 굵은 닻줄과 계류 시설이 있고 가까운 가판대에는 생선이 늘어서 있다.",
  "identity": "canonical",
  "scope_id": "L260",
  "scope_role": "location_exterior",
  "scope_sha": "9499f1d955760266"
 },
 "groupbg::harbor_waiting_quay": {
  "input_fingerprint": "177135af6ca670da",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "harbor_waiting_quay",
    "tags": [
     "S78sh2"
    ]
   },
   "context_sig": "1bb958259d7437ee"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On the harbor quay's exposed edge at dusk, beside fish vendors and moored boats.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n목포 항구 부둣가와 생선 가판대: 배가 정박해 있고 수산물을 거래하는 저녁 시간대의 방파제 주변. (특징: 출렁이는 검은 바닷물과 배를 묶어두는 시멘트 부두; 밧줄로 고정된 소형 어선들; 가판대 위에 놓인 어류들과 덮어놓은 비닐; 커다란 모포를 뒤집어쓰고 웅크린 인물)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 커다란 모포를 뒤집어쓴 찰리와 현우가 부둣가에 걸터앉아있다.\n\nTIME OF DAY (lock): dusk.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On the harbor quay's exposed edge at dusk, beside fish vendors and moored boats.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n목포 항구 부둣가와 생선 가판대: 배가 정박해 있고 수산물을 거래하는 저녁 시간대의 방파제 주변. (특징: 출렁이는 검은 바닷물과 배를 묶어두는 시멘트 부두; 밧줄로 고정된 소형 어선들; 가판대 위에 놓인 어류들과 덮어놓은 비닐; 커다란 모포를 뒤집어쓰고 웅크린 인물)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 커다란 모포를 뒤집어쓴 찰리와 현우가 부둣가에 걸터앉아있다.\n\nTIME OF DAY (lock): dusk.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_harbor_waiting_quay_bf7967.png",
  "asset_id": "0467ac8e-a464-46ad-b316-68bc8f56f1d4",
  "input_asset_ids": [
   "aed4a301-2447-4839-a8a5-7d2e86e4dc08"
  ],
  "origin_tag": "S78sh2",
  "place_text": "On the harbor quay's exposed edge at dusk, beside fish vendors and moored boats.",
  "origin_inputs": {
   "place_text": "On the harbor quay's exposed edge at dusk, beside fish vendors and moored boats.",
   "time_of_day_en": "dusk",
   "conti_asset_id": "aed4a301-2447-4839-a8a5-7d2e86e4dc08"
  }
 },
 "S78sh2::bgfirst_bg": {
  "input_fingerprint": "15e5c8425aab4c98",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 커다란 모포를 머리끝까지 푹 뒤집어쓴 찰리와 현우가 부둣가에 나란히 앉아있는 전신.\n\nLOCATION (lock): On the harbor quay's exposed edge at dusk, beside fish vendors and moored boats.\n\nTIME OF DAY (lock): dusk.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Dock edge (Occupied by the seated pair) — Runs diagonally behind and beneath the seated figures; used as Establishes their position at the harbor boundary; Moored boats (Stationary in the harbor) — Seen obliquely beyond the pair, with no single boat occupying a large portion of the frame; used as Provides the destination of their searching attention; Sea (Undulating); used as Separates the seated figures from the moored boats.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Dim evening ambient light and restrained contrast preserve detail in the covered figures without diminishing the harbor's dusk.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 커다란 모포를 머리끝까지 푹 뒤집어쓴 찰리와 현우가 부둣가에 나란히 앉아있는 전신.\n\nLOCATION (lock): On the harbor quay's exposed edge at dusk, beside fish vendors and moored boats.\n\nTIME OF DAY (lock): dusk.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Dock edge (Occupied by the seated pair) — Runs diagonally behind and beneath the seated figures; used as Establishes their position at the harbor boundary; Moored boats (Stationary in the harbor) — Seen obliquely beyond the pair, with no single boat occupying a large portion of the frame; used as Provides the destination of their searching attention; Sea (Undulating); used as Separates the seated figures from the moored boats.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Dim evening ambient light and restrained contrast preserve detail in the covered figures without diminishing the harbor's dusk.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S78sh2__bgfirst_bg.png",
  "asset_id": "f461492e-6de8-4f80-be43-4a1003286582",
  "input_asset_ids": [
   "aed4a301-2447-4839-a8a5-7d2e86e4dc08",
   "0467ac8e-a464-46ad-b316-68bc8f56f1d4"
  ]
 },
 "S78sh2": {
  "input_fingerprint": "11f56961efe3da71",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dusk.\n\nSHOT TEXT (authoritative, Korean): 커다란 모포를 머리끝까지 푹 뒤집어쓴 찰리와 현우가 부둣가에 나란히 앉아있는 전신.\n\nLOCATION (lock): On the harbor quay's exposed edge at dusk, beside fish vendors and moored boats. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Dock edge (Occupied by the seated pair) — Runs diagonally behind and beneath the seated figures; used as Establishes their position at the harbor boundary; Moored boats (Stationary in the harbor) — Seen obliquely beyond the pair, with no single boat occupying a large portion of the frame; used as Provides the destination of their searching attention; Sea (Undulating); used as Separates the seated figures from the moored boats.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Dim evening ambient light and restrained contrast preserve detail in the covered figures without diminishing the harbor's dusk.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Boats are moored along the harbor in the evening. Charlie sits on the quay under a large blanket, concealing his damaged body; the backpack brought from the airport has no stated removal. 현우: He sits on the quay, still dirty and battered, carrying the backpack from the airport.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dusk.\n\nSHOT TEXT (authoritative, Korean): 커다란 모포를 머리끝까지 푹 뒤집어쓴 찰리와 현우가 부둣가에 나란히 앉아있는 전신.\n\nLOCATION (lock): On the harbor quay's exposed edge at dusk, beside fish vendors and moored boats. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Dock edge (Occupied by the seated pair) — Runs diagonally behind and beneath the seated figures; used as Establishes their position at the harbor boundary; Moored boats (Stationary in the harbor) — Seen obliquely beyond the pair, with no single boat occupying a large portion of the frame; used as Provides the destination of their searching attention; Sea (Undulating); used as Separates the seated figures from the moored boats.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Dim evening ambient light and restrained contrast preserve detail in the covered figures without diminishing the harbor's dusk.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Boats are moored along the harbor in the evening. Charlie sits on the quay under a large blanket, concealing his damaged body; the backpack brought from the airport has no stated removal. 현우: He sits on the quay, still dirty and battered, carrying the backpack from the airport.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dusk.\n\nSHOT TEXT (authoritative, Korean): 커다란 모포를 머리끝까지 푹 뒤집어쓴 찰리와 현우가 부둣가에 나란히 앉아있는 전신.\n\nLOCATION (lock): On the harbor quay's exposed edge at dusk, beside fish vendors and moored boats. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Dock edge (Occupied by the seated pair) — Runs diagonally behind and beneath the seated figures; used as Establishes their position at the harbor boundary; Moored boats (Stationary in the harbor) — Seen obliquely beyond the pair, with no single boat occupying a large portion of the frame; used as Provides the destination of their searching attention; Sea (Undulating); used as Separates the seated figures from the moored boats.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Dim evening ambient light and restrained contrast preserve detail in the covered figures without diminishing the harbor's dusk.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Boats are moored along the harbor in the evening. Charlie sits on the quay under a large blanket, concealing his damaged body; the backpack brought from the airport has no stated removal. 현우: He sits on the quay, still dirty and battered, carrying the backpack from the airport.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S78sh2__bgfirst_bg.png",
     "asset_id": "f461492e-6de8-4f80-be43-4a1003286582",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S78sh2.png",
     "asset_id": "aed4a301-2447-4839-a8a5-7d2e86e4dc08",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_harbor_waiting_quay_bf7967.png",
     "asset_id": "0467ac8e-a464-46ad-b316-68bc8f56f1d4",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "두 인물이 부둣가 밖 바다와 정박된 배들을 향해 시선을 두고 앉아 있음.",
    "built_space": "부둣가 모서리에 두 인물이 나란히 앉아 있으며, 뒤편으로 어시장 천막과 정박된 배들이 레퍼런스와 일치하게 배치됨.",
    "entities": "두 인물이 커다란 모포를 덮어쓰고 있으나, 왼쪽 인물(찰리)의 다리가 로봇 장갑판이 아닌 회색 운동복과 운동화를 착용한 인간의 다리로 묘사됨. 우측 인물 옆에 배낭이 놓여 있음.",
    "hard_violations": [
     "[gemini-pro] 찰리의 외형(로봇) 대신 운동복과 운동화를 착용한 인간의 다리가 생성되어, 프롬프트가 허용하지 않은 인물이 등장함(invented people)."
    ],
    "physics": "두 인물 모두 부둣가 바닥에 엉덩이와 발을 대고 안정적으로 앉아 있음."
   },
   {
    "label": "B",
    "direction": "두 인물이 바다와 배를 향해 시선을 던지며 나란히 앉아 있음.",
    "built_space": "부둣가 끄트머리에 인물들이 앉아 있고, 좌측에 어시장 상인과 천막이 있으며 바다 쪽으로 배들이 정박해 있음.",
    "entities": "왼쪽의 거대한 인물(찰리)은 모포를 완전히 덮어쓰고 있으나, 오른쪽 인물(현우)은 모포를 어깨에만 걸치고 머리가 노출되어 있음. 현우의 손이 배낭을 짚고 있음.",
    "hard_violations": [
     "[gpt-high] 두 주인공 외에는 사람이 없어야 하지만, 왼쪽 생선 판매대에 상인 한 명이 추가되었습니다."
    ],
    "physics": "인물들이 바닥에 자연스럽게 앉아 체중을 지지하고 있으며, 현우의 손이 배낭 위에 얹혀 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "모포를 뒤집어쓴 형태는 지문에 부합하나, 로봇인 찰리의 하반신이 운동복을 입은 인간으로 묘사되어 캐릭터 설정에 치명적인 오류가 발생했습니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "현우의 머리가 노출되어 '머리끝까지 푹 뒤집어쓴'이라는 핵심 지문을 어겼으나, 찰리의 육중한 체형과 전체적인 배경 및 소품 구성은 안정적입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 인물이 부둣가 밖 바다와 정박된 배들을 향해 시선을 두고 앉아 있음.",
        "built_space": "부둣가 모서리에 두 인물이 나란히 앉아 있으며, 뒤편으로 어시장 천막과 정박된 배들이 레퍼런스와 일치하게 배치됨.",
        "entities": "두 인물이 커다란 모포를 덮어쓰고 있으나, 왼쪽 인물(찰리)의 다리가 로봇 장갑판이 아닌 회색 운동복과 운동화를 착용한 인간의 다리로 묘사됨. 우측 인물 옆에 배낭이 놓여 있음.",
        "hard_violations": [
         "찰리의 외형(로봇) 대신 운동복과 운동화를 착용한 인간의 다리가 생성되어, 프롬프트가 허용하지 않은 인물이 등장함(invented people)."
        ],
        "physics": "두 인물 모두 부둣가 바닥에 엉덩이와 발을 대고 안정적으로 앉아 있음."
       },
       {
        "label": "B",
        "direction": "두 인물이 바다와 배를 향해 시선을 던지며 나란히 앉아 있음.",
        "built_space": "부둣가 끄트머리에 인물들이 앉아 있고, 좌측에 어시장 상인과 천막이 있으며 바다 쪽으로 배들이 정박해 있음.",
        "entities": "왼쪽의 거대한 인물(찰리)은 모포를 완전히 덮어쓰고 있으나, 오른쪽 인물(현우)은 모포를 어깨에만 걸치고 머리가 노출되어 있음. 현우의 손이 배낭을 짚고 있음.",
        "hard_violations": [],
        "physics": "인물들이 바닥에 자연스럽게 앉아 체중을 지지하고 있으며, 현우의 손이 배낭 위에 얹혀 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "모포를 뒤집어쓴 형태는 지문에 부합하나, 로봇인 찰리의 하반신이 운동복을 입은 인간으로 묘사되어 캐릭터 설정에 치명적인 오류가 발생했습니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "현우의 머리가 노출되어 '머리끝까지 푹 뒤집어쓴'이라는 핵심 지문을 어겼으나, 찰리의 육중한 체형과 전체적인 배경 및 소품 구성은 안정적입니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "두 인물이 부둣가 밖 바다와 정박된 배들을 향해 시선을 두고 앉아 있음.",
        "built_space": "부둣가 모서리에 두 인물이 나란히 앉아 있으며, 뒤편으로 어시장 천막과 정박된 배들이 레퍼런스와 일치하게 배치됨.",
        "entities": "두 인물이 커다란 모포를 덮어쓰고 있으나, 왼쪽 인물(찰리)의 다리가 로봇 장갑판이 아닌 회색 운동복과 운동화를 착용한 인간의 다리로 묘사됨. 우측 인물 옆에 배낭이 놓여 있음.",
        "hard_violations": [
         "찰리의 외형(로봇) 대신 운동복과 운동화를 착용한 인간의 다리가 생성되어, 프롬프트가 허용하지 않은 인물이 등장함(invented people)."
        ],
        "physics": "두 인물 모두 부둣가 바닥에 엉덩이와 발을 대고 안정적으로 앉아 있음."
       },
       {
        "label": "B",
        "direction": "두 인물이 바다와 배를 향해 시선을 던지며 나란히 앉아 있음.",
        "built_space": "부둣가 끄트머리에 인물들이 앉아 있고, 좌측에 어시장 상인과 천막이 있으며 바다 쪽으로 배들이 정박해 있음.",
        "entities": "왼쪽의 거대한 인물(찰리)은 모포를 완전히 덮어쓰고 있으나, 오른쪽 인물(현우)은 모포를 어깨에만 걸치고 머리가 노출되어 있음. 현우의 손이 배낭을 짚고 있음.",
        "hard_violations": [],
        "physics": "인물들이 바닥에 자연스럽게 앉아 체중을 지지하고 있으며, 현우의 손이 배낭 위에 얹혀 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "허용되지 않은 상인이 추가되어 탈락하며, 현우는 머리를 모포로 덮지 않아 핵심 연출도 어긋납니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "두 인물만 나란히 앉힌 전신 와이드와 머리를 덮은 모포는 더 충실하지만, 현우의 머리·얼굴 일부가 드러나고 찰리의 노출된 다리와 운동화가 지정된 몸체와 다릅니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 고개를 오른쪽 아래로 숙여 가까운 수면을 보는 듯하며, 먼 계류선에 시선이 닿는지는 불확실합니다. 찰리는 머리 전체가 가려져 시선을 확인할 수 없습니다. 왼쪽 상인은 작업대 쪽을 내려다봅니다. 무기나 겨누는 물체는 없습니다.",
        "built_space": "왼쪽에 연속된 천막 판매대, 생선 상자와 통, 바퀴 달린 수레 한 대가 있고, 부두를 따라 가로등과 계선주가 반복됩니다. 오른쪽 전경의 큰 계선주 한 개, 중경의 작은 계선주들, 방파제 끝 등대 한 개와 오른쪽의 여러 계류선이 참조 장소의 구성을 따릅니다. 두 인물은 전경의 낮은 콘크리트 단에 앉고 부두 경계는 뒤쪽으로 비스듬히 이어집니다. 바다가 인물과 배 사이를 분리하며 특정 배가 화면을 지배하지 않습니다.",
        "entities": "찰리로 보이는 왼쪽 인물은 큰 모포에 머리와 몸이 가려져 장갑판과 얼굴을 확인할 수 없습니다. 오른쪽 현우는 검은 헝클어진 머리의 젊은 동아시아계 남성으로 보이지만, 옆뒤 모습이라 정확한 얼굴 일치는 확인하기 어렵습니다. 현우의 머리는 완전히 노출되어 있습니다. 손등에는 상처처럼 보이는 붉은 자국이 있고, 손 아래 배낭 한 개가 있습니다. 왼쪽 판매대에는 요청하지 않은 상인 한 명이 추가되어 있습니다. 황혼의 바다, 생선 판매대와 배는 보입니다.",
        "hard_violations": [
         "두 주인공 외에는 사람이 없어야 하지만, 왼쪽 생선 판매대에 상인 한 명이 추가되었습니다."
        ],
        "physics": "두 인물의 체중은 낮은 콘크리트 단이 받치며, 모포는 머리·어깨와 부두 바닥에 걸쳐 자연스럽게 늘어집니다. 현우의 손은 배낭 위에 놓이고 배낭 바닥은 포장면에 닿습니다. 다리 일부는 모포와 단에 가려지지만 몸이 공중에 뜬 증거는 없습니다. 배는 수면에 떠 있고 전경 밧줄은 계선주에 감겨 있습니다."
       },
       {
        "label": "B",
        "direction": "현우는 오른쪽 항구와 계류선이 있는 방향으로 얼굴을 돌리고 있습니다. 눈은 작게 가려져 특정 배를 보고 있다고 단정할 수 없습니다. 찰리는 모포로 얼굴을 완전히 가려 시선이 보이지 않으며, 무릎과 발은 카메라 쪽 육지를 향합니다. 무기나 겨누는 물체는 없습니다.",
        "built_space": "왼쪽 천막 판매대와 상자·통, 수레 한 대, 부두를 따라 이어지는 가로등과 계선주, 방파제의 등대 한 개, 오른쪽 계류선 무리가 참조 장소와 대응합니다. 전경 오른쪽에는 큰 계선주 한 개가 있습니다. 두 인물은 같은 낮은 콘크리트 단에 나란히 앉아 발을 앞쪽 포장면으로 내리고 있습니다. 단과 부두 경계가 인물 아래와 뒤로 비스듬히 이어지고, 배들은 바다 건너 작은 배경 요소로 남아 있습니다.",
        "entities": "인물은 두 명뿐입니다. 큰 모포가 두 사람의 정수리를 덮지만 현우의 검은 머리와 옆얼굴 일부가 밖으로 나와 완전한 은폐는 아닙니다. 현우는 젊은 동아시아계 남성으로 보이며 얼굴의 세부 일치와 부상은 확인하기 어렵습니다. 찰리로 배치된 왼쪽 인물의 얼굴과 상체는 가려졌으나, 드러난 부분은 사람처럼 가는 바지 다리와 운동화여서 참조의 육중한 로봇 하체·장갑 발과 다릅니다. 배낭 한 개가 현우 바로 옆 단 위에 놓여 있습니다. 황혼, 생선 판매대, 바다와 계류선은 구현되어 있습니다.",
        "hard_violations": [],
        "physics": "두 사람의 엉덩이는 콘크리트 단에 지지되고 무릎은 굽혀져 있으며, 발은 바로 아래 포장면에 닿거나 그 가까이 내려와 있습니다. 모포는 두 머리와 어깨·무릎이 받치고 가운데 부분은 중력에 따라 처집니다. 배낭은 단 위에 놓여 지지됩니다. 지지 없이 떠 있는 몸이나 물체는 보이지 않습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "허용되지 않은 상인이 추가되어 탈락하며, 현우는 머리를 모포로 덮지 않아 핵심 연출도 어긋납니다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "두 인물만 나란히 앉힌 전신 와이드와 머리를 덮은 모포는 더 충실하지만, 현우의 머리·얼굴 일부가 드러나고 찰리의 노출된 다리와 운동화가 지정된 몸체와 다릅니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 고개를 오른쪽 아래로 숙여 가까운 수면을 보는 듯하며, 먼 계류선에 시선이 닿는지는 불확실합니다. 찰리는 머리 전체가 가려져 시선을 확인할 수 없습니다. 왼쪽 상인은 작업대 쪽을 내려다봅니다. 무기나 겨누는 물체는 없습니다.",
        "built_space": "왼쪽에 연속된 천막 판매대, 생선 상자와 통, 바퀴 달린 수레 한 대가 있고, 부두를 따라 가로등과 계선주가 반복됩니다. 오른쪽 전경의 큰 계선주 한 개, 중경의 작은 계선주들, 방파제 끝 등대 한 개와 오른쪽의 여러 계류선이 참조 장소의 구성을 따릅니다. 두 인물은 전경의 낮은 콘크리트 단에 앉고 부두 경계는 뒤쪽으로 비스듬히 이어집니다. 바다가 인물과 배 사이를 분리하며 특정 배가 화면을 지배하지 않습니다.",
        "entities": "찰리로 보이는 왼쪽 인물은 큰 모포에 머리와 몸이 가려져 장갑판과 얼굴을 확인할 수 없습니다. 오른쪽 현우는 검은 헝클어진 머리의 젊은 동아시아계 남성으로 보이지만, 옆뒤 모습이라 정확한 얼굴 일치는 확인하기 어렵습니다. 현우의 머리는 완전히 노출되어 있습니다. 손등에는 상처처럼 보이는 붉은 자국이 있고, 손 아래 배낭 한 개가 있습니다. 왼쪽 판매대에는 요청하지 않은 상인 한 명이 추가되어 있습니다. 황혼의 바다, 생선 판매대와 배는 보입니다.",
        "hard_violations": [
         "두 주인공 외에는 사람이 없어야 하지만, 왼쪽 생선 판매대에 상인 한 명이 추가되었습니다."
        ],
        "physics": "두 인물의 체중은 낮은 콘크리트 단이 받치며, 모포는 머리·어깨와 부두 바닥에 걸쳐 자연스럽게 늘어집니다. 현우의 손은 배낭 위에 놓이고 배낭 바닥은 포장면에 닿습니다. 다리 일부는 모포와 단에 가려지지만 몸이 공중에 뜬 증거는 없습니다. 배는 수면에 떠 있고 전경 밧줄은 계선주에 감겨 있습니다."
       },
       {
        "label": "A",
        "direction": "현우는 오른쪽 항구와 계류선이 있는 방향으로 얼굴을 돌리고 있습니다. 눈은 작게 가려져 특정 배를 보고 있다고 단정할 수 없습니다. 찰리는 모포로 얼굴을 완전히 가려 시선이 보이지 않으며, 무릎과 발은 카메라 쪽 육지를 향합니다. 무기나 겨누는 물체는 없습니다.",
        "built_space": "왼쪽 천막 판매대와 상자·통, 수레 한 대, 부두를 따라 이어지는 가로등과 계선주, 방파제의 등대 한 개, 오른쪽 계류선 무리가 참조 장소와 대응합니다. 전경 오른쪽에는 큰 계선주 한 개가 있습니다. 두 인물은 같은 낮은 콘크리트 단에 나란히 앉아 발을 앞쪽 포장면으로 내리고 있습니다. 단과 부두 경계가 인물 아래와 뒤로 비스듬히 이어지고, 배들은 바다 건너 작은 배경 요소로 남아 있습니다.",
        "entities": "인물은 두 명뿐입니다. 큰 모포가 두 사람의 정수리를 덮지만 현우의 검은 머리와 옆얼굴 일부가 밖으로 나와 완전한 은폐는 아닙니다. 현우는 젊은 동아시아계 남성으로 보이며 얼굴의 세부 일치와 부상은 확인하기 어렵습니다. 찰리로 배치된 왼쪽 인물의 얼굴과 상체는 가려졌으나, 드러난 부분은 사람처럼 가는 바지 다리와 운동화여서 참조의 육중한 로봇 하체·장갑 발과 다릅니다. 배낭 한 개가 현우 바로 옆 단 위에 놓여 있습니다. 황혼, 생선 판매대, 바다와 계류선은 구현되어 있습니다.",
        "hard_violations": [],
        "physics": "두 사람의 엉덩이는 콘크리트 단에 지지되고 무릎은 굽혀져 있으며, 발은 바로 아래 포장면에 닿거나 그 가까이 내려와 있습니다. 모포는 두 머리와 어깨·무릎이 받치고 가운데 부분은 중력에 따라 처집니다. 배낭은 단 위에 놓여 지지됩니다. 지지 없이 떠 있는 몸이나 물체는 보이지 않습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.5,
    "B": 1.5
   },
   "adjusted": {
    "A": 1.25,
    "B": 1.25
   },
   "violations": {
    "A": [
     "[gemini-pro] 찰리의 외형(로봇) 대신 운동복과 운동화를 착용한 인간의 다리가 생성되어, 프롬프트가 허용하지 않은 인물이 등장함(invented people)."
    ],
    "B": [
     "[gpt-high] 두 주인공 외에는 사람이 없어야 하지만, 왼쪽 생선 판매대에 상인 한 명이 추가되었습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "A": 1250,
   "B": 1250
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1250,
    "verdict_ko": "모포를 뒤집어쓴 형태는 지문에 부합하나, 로봇인 찰리의 하반신이 운동복을 입은 인간으로 묘사되어 캐릭터 설정에 치명적인 오류가 발생했습니다.  ★위반: [gemini-pro] 찰리의 외형(로봇) 대신 운동복과 운동화를 착용한 인간의 다리가 생성되어, 프롬프트가 허용하지 않은 인물이 등장함(invented people)."
   },
   {
    "label": "B",
    "score": 1250,
    "verdict_ko": "현우의 머리가 노출되어 '머리끝까지 푹 뒤집어쓴'이라는 핵심 지문을 어겼으나, 찰리의 육중한 체형과 전체적인 배경 및 소품 구성은 안정적입니다.  ★위반: [gpt-high] 두 주인공 외에는 사람이 없어야 하지만, 왼쪽 생선 판매대에 상인 한 명이 추가되었습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_harbor_waiting_quay_bf7967.png",
    "asset_id": "0467ac8e-a464-46ad-b316-68bc8f56f1d4",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-c6c3-7c07-8365-31b405edac0b",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S78sh2__bgfirst_bg.png",
   "bg_asset_id": "f461492e-6de8-4f80-be43-4a1003286582",
   "bg_record_key": "S78sh2::bgfirst_bg",
   "chain_winner": true,
   "authority": "groupbg",
   "group_key": "harbor_waiting_quay",
   "groupbg_asset_id": "0467ac8e-a464-46ad-b316-68bc8f56f1d4"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S78sh2::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:49:55.194843+00:00",
  "fingerprint": "0b7600e72984523e5a037c8c954107a9f66807f2dbca7219ce089a7922a70d4a",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S78sh2_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S78sh2_sel.png",
  "source_sha256": "e235a58efb6563719b1fbdd04d33a7b79c053e8b5fdd83202c13df3aec0427c4",
  "file": "S78sh2_cine.png",
  "staged_sha256": "b467251d41e47240af4f5be651ffd26983937b84c3b608eee6439f53c1ce487c",
  "latency_ms": 10727
 },
 "S78sh13::signage": {
  "fp": "7eca9ae6b44448ef",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "groupbg::harbor_repair_berth": {
  "input_fingerprint": "b1bcb0ab845f2c42",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "harbor_repair_berth",
    "tags": [
     "S78sh13"
    ]
   },
   "context_sig": "b8355d48debd79dc"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At the outdoor repair spot beside an old boat moored in the harbor, where its captain pauses work at dusk.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n목포 항구 부둣가와 생선 가판대: 배가 정박해 있고 수산물을 거래하는 저녁 시간대의 방파제 주변. (특징: 출렁이는 검은 바닷물과 배를 묶어두는 시멘트 부두; 밧줄로 고정된 소형 어선들; 가판대 위에 놓인 어류들과 덮어놓은 비닐; 커다란 모포를 뒤집어쓰고 웅크린 인물)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 선원이 어디론가로 현우를 인도한다.\n- 낡은 배를 정비하는 40대 초반의 크리스와 몇몇 선원들 보인다.\n\nTIME OF DAY (lock): dusk.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At the outdoor repair spot beside an old boat moored in the harbor, where its captain pauses work at dusk.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n목포 항구 부둣가와 생선 가판대: 배가 정박해 있고 수산물을 거래하는 저녁 시간대의 방파제 주변. (특징: 출렁이는 검은 바닷물과 배를 묶어두는 시멘트 부두; 밧줄로 고정된 소형 어선들; 가판대 위에 놓인 어류들과 덮어놓은 비닐; 커다란 모포를 뒤집어쓰고 웅크린 인물)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 선원이 어디론가로 현우를 인도한다.\n- 낡은 배를 정비하는 40대 초반의 크리스와 몇몇 선원들 보인다.\n\nTIME OF DAY (lock): dusk.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_harbor_repair_berth_41df90.png",
  "asset_id": "7aa8b28b-26f7-41d5-b9ce-185137890bc3",
  "input_asset_ids": [
   "4effe083-7d46-4dd7-a28a-33b44efc4e7b"
  ],
  "origin_tag": "S78sh13",
  "place_text": "At the outdoor repair spot beside an old boat moored in the harbor, where its captain pauses work at dusk.",
  "origin_inputs": {
   "place_text": "At the outdoor repair spot beside an old boat moored in the harbor, where its captain pauses work at dusk.",
   "time_of_day_en": "dusk",
   "conti_asset_id": "4effe083-7d46-4dd7-a28a-33b44efc4e7b"
  }
 },
 "S78sh13::bgfirst_bg": {
  "input_fingerprint": "3ae8336ec28768d4",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 용접기를 내린 채 차가운 눈빛으로 현우와 모포 쓴 찰리를 빤히 주시하는 크리스의 얼굴 클로즈업.\n\nLOCATION (lock): At the outdoor repair spot beside an old boat moored in the harbor, where its captain pauses work at dusk.\n\nTIME OF DAY (lock): dusk.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old boat (Undergoing maintenance) — Only an oblique fragment of its dock-facing side is visible behind 크리스; used as Keeps the close scrutiny grounded in the boarding encounter.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established dim evening ambience with controlled facial contrast and no change of lighting at the closer distance.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 용접기를 내린 채 차가운 눈빛으로 현우와 모포 쓴 찰리를 빤히 주시하는 크리스의 얼굴 클로즈업.\n\nLOCATION (lock): At the outdoor repair spot beside an old boat moored in the harbor, where its captain pauses work at dusk.\n\nTIME OF DAY (lock): dusk.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old boat (Undergoing maintenance) — Only an oblique fragment of its dock-facing side is visible behind 크리스; used as Keeps the close scrutiny grounded in the boarding encounter.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established dim evening ambience with controlled facial contrast and no change of lighting at the closer distance.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S78sh13__bgfirst_bg.png",
  "asset_id": "640ad084-58c6-44a6-b7fc-1706ec522e95",
  "input_asset_ids": [
   "4effe083-7d46-4dd7-a28a-33b44efc4e7b",
   "7aa8b28b-26f7-41d5-b9ce-185137890bc3"
  ]
 },
 "S78sh13": {
  "input_fingerprint": "70eb343a2504cf48",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dusk.\n\nSHOT TEXT (authoritative, Korean): 용접기를 내린 채 차가운 눈빛으로 현우와 모포 쓴 찰리를 빤히 주시하는 크리스의 얼굴 클로즈업.\n\nLOCATION (lock): At the outdoor repair spot beside an old boat moored in the harbor, where its captain pauses work at dusk. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old boat (Undergoing maintenance) — Only an oblique fragment of its dock-facing side is visible behind 크리스; used as Keeps the close scrutiny grounded in the boarding encounter.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established dim evening ambience with controlled facial contrast and no change of lighting at the closer distance.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old boat remains moored for maintenance. Charlie's large blanket conceals his damaged body, and his travel backpack has not been discarded. 크리스: He is beside the old boat, where he has been carrying out maintenance.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 크리스 right now, so 크리스's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 크리스: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 크리스 (한국인 남성, 40대 초반, 중년 초입의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dusk.\n\nSHOT TEXT (authoritative, Korean): 용접기를 내린 채 차가운 눈빛으로 현우와 모포 쓴 찰리를 빤히 주시하는 크리스의 얼굴 클로즈업.\n\nLOCATION (lock): At the outdoor repair spot beside an old boat moored in the harbor, where its captain pauses work at dusk. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old boat (Undergoing maintenance) — Only an oblique fragment of its dock-facing side is visible behind 크리스; used as Keeps the close scrutiny grounded in the boarding encounter.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established dim evening ambience with controlled facial contrast and no change of lighting at the closer distance.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old boat remains moored for maintenance. Charlie's large blanket conceals his damaged body, and his travel backpack has not been discarded. 크리스: He is beside the old boat, where he has been carrying out maintenance.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 크리스 right now, so 크리스's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 크리스: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 크리스 (한국인 남성, 40대 초반, 중년 초입의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dusk.\n\nSHOT TEXT (authoritative, Korean): 용접기를 내린 채 차가운 눈빛으로 현우와 모포 쓴 찰리를 빤히 주시하는 크리스의 얼굴 클로즈업.\n\nLOCATION (lock): At the outdoor repair spot beside an old boat moored in the harbor, where its captain pauses work at dusk. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old boat (Undergoing maintenance) — Only an oblique fragment of its dock-facing side is visible behind 크리스; used as Keeps the close scrutiny grounded in the boarding encounter.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established dim evening ambience with controlled facial contrast and no change of lighting at the closer distance.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old boat remains moored for maintenance. Charlie's large blanket conceals his damaged body, and his travel backpack has not been discarded. 크리스: He is beside the old boat, where he has been carrying out maintenance.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 크리스 right now, so 크리스's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 크리스: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 크리스 (한국인 남성, 40대 초반, 중년 초입의 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S78sh13__bgfirst_bg.png",
     "asset_id": "640ad084-58c6-44a6-b7fc-1706ec522e95",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S78sh13.png",
     "asset_id": "4effe083-7d46-4dd7-a28a-33b44efc4e7b",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 크리스: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1199325>",
     "asset_id": "303f2ae8-9f48-4964-8a8b-c24edec958ba",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_harbor_repair_berth_41df90.png",
     "asset_id": "7aa8b28b-26f7-41d5-b9ce-185137890bc3",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 크리스: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1199325>",
     "asset_id": "303f2ae8-9f48-4964-8a8b-c24edec958ba",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "크리스의 시선은 화면 좌측 전경에 서 있는 두 인물(현우와 찰리)을 정확히 향하고 있음.",
    "built_space": "항구 야외 작업장이며, 우측 뒤편으로 낡은 배의 선체가 비스듬하게 올바른 비율로 배치되어 있음.",
    "entities": "크리스(레퍼런스 외모 일치), 전경의 뒷모습 두 인물(한 명은 모포를 두름), 손에 든 용접기 모두 확인됨.",
    "hard_violations": [
     "[gpt-high] 크리스만을 가시 인물로 허용한 인물 제한을 어기고, 왼쪽 전경에 남성 한 명과 모포를 쓴 인물 한 명을 추가했다."
    ],
    "physics": "크리스가 한 손에 용접기를 들고 서 있는 자세 및 전경 인물들의 서 있는 모습에서 물리적 지지 및 무게 중심이 자연스러움."
   },
   {
    "label": "B",
    "direction": "크리스의 시선이 지시된 대상이 아닌 카메라 렌즈를 직접 향하고 있음.",
    "built_space": "낡은 배가 배경 전체를 차지하고 있으며, 인물 앞쪽에 정체불명의 금속판이 배치되어 있음.",
    "entities": "크리스(레퍼런스 일치)와 장갑 낀 손에 든 용접기는 보이나, 현우와 찰리는 프레임 내에 없음.",
    "hard_violations": [],
    "physics": "크리스가 금속판 위에 턱을 괴고 있으며, 지시와 달리 용접기를 밑으로 내리지 않고 얼굴 앞까지 끌어올려 판 위에 얹어둔 상태임."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "요구된 얼굴 클로즈업보다 프레임이 다소 넓어졌으나, 내린 용접기와 전경의 인물들을 향한 시선 등 프롬프트의 상황을 매우 충실하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "클로즈업 프레이밍에 가깝지만, 용접기를 얼굴 높이로 들어 올리고 렌즈를 정면 응시하여 '내린 용접기'와 주시 대상에 대한 지시를 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "크리스의 시선은 화면 좌측 전경에 서 있는 두 인물(현우와 찰리)을 정확히 향하고 있음.",
        "built_space": "항구 야외 작업장이며, 우측 뒤편으로 낡은 배의 선체가 비스듬하게 올바른 비율로 배치되어 있음.",
        "entities": "크리스(레퍼런스 외모 일치), 전경의 뒷모습 두 인물(한 명은 모포를 두름), 손에 든 용접기 모두 확인됨.",
        "hard_violations": [],
        "physics": "크리스가 한 손에 용접기를 들고 서 있는 자세 및 전경 인물들의 서 있는 모습에서 물리적 지지 및 무게 중심이 자연스러움."
       },
       {
        "label": "B",
        "direction": "크리스의 시선이 지시된 대상이 아닌 카메라 렌즈를 직접 향하고 있음.",
        "built_space": "낡은 배가 배경 전체를 차지하고 있으며, 인물 앞쪽에 정체불명의 금속판이 배치되어 있음.",
        "entities": "크리스(레퍼런스 일치)와 장갑 낀 손에 든 용접기는 보이나, 현우와 찰리는 프레임 내에 없음.",
        "hard_violations": [],
        "physics": "크리스가 금속판 위에 턱을 괴고 있으며, 지시와 달리 용접기를 밑으로 내리지 않고 얼굴 앞까지 끌어올려 판 위에 얹어둔 상태임."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "요구된 얼굴 클로즈업보다 프레임이 다소 넓어졌으나, 내린 용접기와 전경의 인물들을 향한 시선 등 프롬프트의 상황을 매우 충실하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "클로즈업 프레이밍에 가깝지만, 용접기를 얼굴 높이로 들어 올리고 렌즈를 정면 응시하여 '내린 용접기'와 주시 대상에 대한 지시를 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "크리스의 시선은 화면 좌측 전경에 서 있는 두 인물(현우와 찰리)을 정확히 향하고 있음.",
        "built_space": "항구 야외 작업장이며, 우측 뒤편으로 낡은 배의 선체가 비스듬하게 올바른 비율로 배치되어 있음.",
        "entities": "크리스(레퍼런스 외모 일치), 전경의 뒷모습 두 인물(한 명은 모포를 두름), 손에 든 용접기 모두 확인됨.",
        "hard_violations": [],
        "physics": "크리스가 한 손에 용접기를 들고 서 있는 자세 및 전경 인물들의 서 있는 모습에서 물리적 지지 및 무게 중심이 자연스러움."
       },
       {
        "label": "B",
        "direction": "크리스의 시선이 지시된 대상이 아닌 카메라 렌즈를 직접 향하고 있음.",
        "built_space": "낡은 배가 배경 전체를 차지하고 있으며, 인물 앞쪽에 정체불명의 금속판이 배치되어 있음.",
        "entities": "크리스(레퍼런스 일치)와 장갑 낀 손에 든 용접기는 보이나, 현우와 찰리는 프레임 내에 없음.",
        "hard_violations": [],
        "physics": "크리스가 금속판 위에 턱을 괴고 있으며, 지시와 달리 용접기를 밑으로 내리지 않고 얼굴 앞까지 끌어올려 판 위에 얹어둔 상태임."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "얼굴 클로즈업, 어두운 저녁빛, 손으로 잡아 내려놓은 용접기는 요구에 가깝지만, 렌즈를 향한 시선과 작업면에 지나치게 가까운 얼굴 때문에 두 방문자를 주시하는 순간은 덜 명확하다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "방문자를 향한 차가운 시선과 항구 장소는 분명하지만, 두 인물을 추가하고 상반신·항구 전경으로 넓혀 크리스만의 얼굴 클로즈업이라는 핵심 구도를 어겼다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "크리스의 얼굴과 두 눈은 거의 렌즈 정면을 향한다. 현우와 찰리는 화면 밖이므로 실제 시선이 두 사람에게 닿는지는 확인되지 않는다. 용접 토치 끝은 화면 오른쪽 아래의 철판을 향하며 사람을 겨누지 않는다.",
        "built_space": "뒤에는 녹슨 흰색 선체와 파란 띠, 왼쪽의 원형 창 하나, 상단 난간 일부가 비스듬하게 보인다. 장소 사진의 배 재질과 색은 이어지며 배 전체나 항구 전경으로 확대하지 않았다. 전경에는 금속 작업면 하나와 그 위의 작은 철판 하나가 있고, 크리스는 얼굴을 그 작업면 높이 가까이 낮추고 있다. 계류 상태와 부두의 나머지 설비는 이 크롭에서 확인되지 않는다.",
        "entities": "보이는 사람은 크리스 한 명이다. 검은 머리의 중년 초입 동아시아계 남성으로, 참고 인물의 얼굴 윤곽과 눈·코 형태에 대체로 가깝다. 눈은 정상적인 사람의 눈이며 표정은 굳어 있다. 남색 계열 옷이지만 참고의 흰 셔츠와 니트 조합보다는 작업복처럼 보인다. 장갑 낀 손에 용접 토치가 있고, 현우·찰리·모포·배낭은 얼굴 중심 구도 밖이다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "용접 토치 손잡이는 크리스의 장갑 낀 손이 잡고 있고, 노즐과 철판은 금속 작업면에 놓여 있다. 손과 팔은 화면 왼쪽의 소매로 연결되어 소유 관계가 자연스럽다. 얼굴을 작업면에 가깝게 낮춘 자세는 가능하지만 방문자를 살피는 동작으로는 다소 어색하다. 하체가 잘렸다는 이유로 부유한다고 볼 근거는 없다."
       },
       {
        "label": "B",
        "direction": "크리스는 화면 왼쪽 전경의 검은 머리 남성을 향해 눈을 돌리고 있다. 그 옆에는 모포를 뒤집어쓴 인물이 있어 방문자들을 살피는 관계가 읽힌다. 다만 모포 속 인물을 직접 응시하는 순간은 아니다. 허리 가까이 잡은 토치의 노즐은 왼쪽 위를 향하며 방문자들의 얼굴을 직접 겨누지는 않는다.",
        "built_space": "오른쪽에 녹슨 흰색 배, 파란 띠, 원형 창 두 개, 난간, 기둥과 켜진 등 하나가 보인다. 중앙 부두에는 계선주 하나와 계류 밧줄이 있고, 뒤로 바다·방파제·등대 하나가 보인다. 장소 사진의 구조는 잘 이어지지만, 선체의 작은 사선 조각만 남기는 대신 항구와 부두를 넓게 보여준다. 크리스는 오른쪽 배 옆에 서 있고 두 방문자는 왼쪽 전경을 차지한다.",
        "entities": "크리스는 참고와 닮은 검은 머리의 중년 초입 동아시아계 남성이며, 회녹색 작업 재킷과 밝은 속옷을 입어 참고 의상과 다르다. 맨손에 용접 토치를 들고 있다. 추가로 뒷머리와 어깨가 보이는 남성 한 명, 회색 모포로 머리와 몸을 덮은 인물 한 명이 있다. 전경 남성에게 배낭 끈처럼 보이는 부분은 있으나 찰리의 배낭인지는 확인되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "크리스만을 가시 인물로 허용한 인물 제한을 어기고, 왼쪽 전경에 남성 한 명과 모포를 쓴 인물 한 명을 추가했다."
        ],
        "physics": "크리스의 손가락이 토치 손잡이를 감싸고 손목과 소매가 자연스럽게 이어져 도구의 지지가 분명하다. 토치는 작동하지 않고 낮게 들려 있다. 인물들의 하체는 프레임 밖이지만 부두에 서 있는 상체 배치로 읽히며 부유나 불가능한 관절은 보이지 않는다. 모포는 안쪽 신체를 따라 어깨에서 아래로 드리워진다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "얼굴 클로즈업, 어두운 저녁빛, 손으로 잡아 내려놓은 용접기는 요구에 가깝지만, 렌즈를 향한 시선과 작업면에 지나치게 가까운 얼굴 때문에 두 방문자를 주시하는 순간은 덜 명확하다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "방문자를 향한 차가운 시선과 항구 장소는 분명하지만, 두 인물을 추가하고 상반신·항구 전경으로 넓혀 크리스만의 얼굴 클로즈업이라는 핵심 구도를 어겼다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "크리스의 얼굴과 두 눈은 거의 렌즈 정면을 향한다. 현우와 찰리는 화면 밖이므로 실제 시선이 두 사람에게 닿는지는 확인되지 않는다. 용접 토치 끝은 화면 오른쪽 아래의 철판을 향하며 사람을 겨누지 않는다.",
        "built_space": "뒤에는 녹슨 흰색 선체와 파란 띠, 왼쪽의 원형 창 하나, 상단 난간 일부가 비스듬하게 보인다. 장소 사진의 배 재질과 색은 이어지며 배 전체나 항구 전경으로 확대하지 않았다. 전경에는 금속 작업면 하나와 그 위의 작은 철판 하나가 있고, 크리스는 얼굴을 그 작업면 높이 가까이 낮추고 있다. 계류 상태와 부두의 나머지 설비는 이 크롭에서 확인되지 않는다.",
        "entities": "보이는 사람은 크리스 한 명이다. 검은 머리의 중년 초입 동아시아계 남성으로, 참고 인물의 얼굴 윤곽과 눈·코 형태에 대체로 가깝다. 눈은 정상적인 사람의 눈이며 표정은 굳어 있다. 남색 계열 옷이지만 참고의 흰 셔츠와 니트 조합보다는 작업복처럼 보인다. 장갑 낀 손에 용접 토치가 있고, 현우·찰리·모포·배낭은 얼굴 중심 구도 밖이다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "용접 토치 손잡이는 크리스의 장갑 낀 손이 잡고 있고, 노즐과 철판은 금속 작업면에 놓여 있다. 손과 팔은 화면 왼쪽의 소매로 연결되어 소유 관계가 자연스럽다. 얼굴을 작업면에 가깝게 낮춘 자세는 가능하지만 방문자를 살피는 동작으로는 다소 어색하다. 하체가 잘렸다는 이유로 부유한다고 볼 근거는 없다."
       },
       {
        "label": "A",
        "direction": "크리스는 화면 왼쪽 전경의 검은 머리 남성을 향해 눈을 돌리고 있다. 그 옆에는 모포를 뒤집어쓴 인물이 있어 방문자들을 살피는 관계가 읽힌다. 다만 모포 속 인물을 직접 응시하는 순간은 아니다. 허리 가까이 잡은 토치의 노즐은 왼쪽 위를 향하며 방문자들의 얼굴을 직접 겨누지는 않는다.",
        "built_space": "오른쪽에 녹슨 흰색 배, 파란 띠, 원형 창 두 개, 난간, 기둥과 켜진 등 하나가 보인다. 중앙 부두에는 계선주 하나와 계류 밧줄이 있고, 뒤로 바다·방파제·등대 하나가 보인다. 장소 사진의 구조는 잘 이어지지만, 선체의 작은 사선 조각만 남기는 대신 항구와 부두를 넓게 보여준다. 크리스는 오른쪽 배 옆에 서 있고 두 방문자는 왼쪽 전경을 차지한다.",
        "entities": "크리스는 참고와 닮은 검은 머리의 중년 초입 동아시아계 남성이며, 회녹색 작업 재킷과 밝은 속옷을 입어 참고 의상과 다르다. 맨손에 용접 토치를 들고 있다. 추가로 뒷머리와 어깨가 보이는 남성 한 명, 회색 모포로 머리와 몸을 덮은 인물 한 명이 있다. 전경 남성에게 배낭 끈처럼 보이는 부분은 있으나 찰리의 배낭인지는 확인되지 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "크리스만을 가시 인물로 허용한 인물 제한을 어기고, 왼쪽 전경에 남성 한 명과 모포를 쓴 인물 한 명을 추가했다."
        ],
        "physics": "크리스의 손가락이 토치 손잡이를 감싸고 손목과 소매가 자연스럽게 이어져 도구의 지지가 분명하다. 토치는 작동하지 않고 낮게 들려 있다. 인물들의 하체는 프레임 밖이지만 부두에 서 있는 상체 배치로 읽히며 부유나 불가능한 관절은 보이지 않는다. 모포는 안쪽 신체를 따라 어깨에서 아래로 드리워진다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.429,
    "B": 1.571
   },
   "adjusted": {
    "A": 1.179,
    "B": 1.571
   },
   "violations": {
    "A": [
     "[gpt-high] 크리스만을 가시 인물로 허용한 인물 제한을 어기고, 왼쪽 전경에 남성 한 명과 모포를 쓴 인물 한 명을 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1179,
   "B": 1571
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1179,
    "verdict_ko": "요구된 얼굴 클로즈업보다 프레임이 다소 넓어졌으나, 내린 용접기와 전경의 인물들을 향한 시선 등 프롬프트의 상황을 매우 충실하게 구현했습니다.  ★위반: [gpt-high] 크리스만을 가시 인물로 허용한 인물 제한을 어기고, 왼쪽 전경에 남성 한 명과 모포를 쓴 인물 한 명을 추가했다."
   },
   {
    "label": "B",
    "score": 1571,
    "verdict_ko": "클로즈업 프레이밍에 가깝지만, 용접기를 얼굴 높이로 들어 올리고 렌즈를 정면 응시하여 '내린 용접기'와 주시 대상에 대한 지시를 위반했습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_harbor_repair_berth_41df90.png",
    "asset_id": "7aa8b28b-26f7-41d5-b9ce-185137890bc3",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 크리스: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1199325>",
    "asset_id": "303f2ae8-9f48-4964-8a8b-c24edec958ba",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-cb9e-712a-912e-7736e1a45a04",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S78sh13__bgfirst_bg.png",
   "bg_asset_id": "640ad084-58c6-44a6-b7fc-1706ec522e95",
   "bg_record_key": "S78sh13::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "harbor_repair_berth",
   "groupbg_asset_id": "7aa8b28b-26f7-41d5-b9ce-185137890bc3"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S78sh13::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T13:32:35.525077+00:00",
  "fingerprint": "d434835cb30aa30bafd6297afc894fe5b598318a8ddcebbcce54675a8a2d2275",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S78sh13_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S78sh13_sel.png",
  "source_sha256": "a1a8ea4b4190de3815c6de2df799147cc4f74674e528a5867b4554e502b4df6b",
  "file": "S78sh13_cine.png",
  "staged_sha256": "48e3a6b7546bb37ee141a797d692a974028fa050556930e37ebb4832efe0af5d",
  "latency_ms": 12043
 },
 "S78sh17::signage": {
  "fp": "8bc706f713ce35a3",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S78sh17": {
  "input_fingerprint": "d5e4a5bba528fc14",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dusk.\n\nSHOT TEXT (authoritative, Korean): 어두운 바다를 가르는 하얀 물살을 남긴 채, 부둣가에서 멀어진 위치에 떠 있는 낡은 배의 뒷모습 풀샷.\n\nLOCATION (lock): On the dark harbor water just beyond the quay, where the small departing boat leaves a white wake. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Departing boat in the upper-center of the frame, background, moves toward open sea beyond the upper frame; Trailing wake in the lower-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Departing old boat (Moving away from the harbor) — Stern and one side remain visible in rear three-quarter view; used as Small, complete subject against the surrounding sea; Wake (White disturbed water trailing behind the boat); used as Connects the distant stern to the lower foreground; Sea (Dark, with water disturbed along the departure path); used as Provides open space around the diminishing vessel.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The dark evening sea and white wake provide restrained tonal separation around the receding boat.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The small motorboat is leaving the harbor under power, churning a wake. Charlie is aboard, still covered by the blanket, with his damaged body and travel backpack unchanged.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dusk.\n\nSHOT TEXT (authoritative, Korean): 어두운 바다를 가르는 하얀 물살을 남긴 채, 부둣가에서 멀어진 위치에 떠 있는 낡은 배의 뒷모습 풀샷.\n\nLOCATION (lock): On the dark harbor water just beyond the quay, where the small departing boat leaves a white wake. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Departing boat in the upper-center of the frame, background, moves toward open sea beyond the upper frame; Trailing wake in the lower-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Departing old boat (Moving away from the harbor) — Stern and one side remain visible in rear three-quarter view; used as Small, complete subject against the surrounding sea; Wake (White disturbed water trailing behind the boat); used as Connects the distant stern to the lower foreground; Sea (Dark, with water disturbed along the departure path); used as Provides open space around the diminishing vessel.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The dark evening sea and white wake provide restrained tonal separation around the receding boat.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The small motorboat is leaving the harbor under power, churning a wake. Charlie is aboard, still covered by the blanket, with his damaged body and travel backpack unchanged.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): dusk.\n\nSHOT TEXT (authoritative, Korean): 어두운 바다를 가르는 하얀 물살을 남긴 채, 부둣가에서 멀어진 위치에 떠 있는 낡은 배의 뒷모습 풀샷.\n\nLOCATION (lock): On the dark harbor water just beyond the quay, where the small departing boat leaves a white wake. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Departing boat in the upper-center of the frame, background, moves toward open sea beyond the upper frame; Trailing wake in the lower-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Departing old boat (Moving away from the harbor) — Stern and one side remain visible in rear three-quarter view; used as Small, complete subject against the surrounding sea; Wake (White disturbed water trailing behind the boat); used as Connects the distant stern to the lower foreground; Sea (Dark, with water disturbed along the departure path); used as Provides open space around the diminishing vessel.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The dark evening sea and white wake provide restrained tonal separation around the receding boat.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The small motorboat is leaving the harbor under power, churning a wake. Charlie is aboard, still covered by the blanket, with his damaged body and travel backpack unchanged.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "배의 선수가 화면 상단의 탁 트인 먼 바다를 향해 올바르게 나아감.",
    "built_space": "인공 구조물 없이 어두운 바다와 배만 단독으로 배치됨.",
    "entities": "낡은 배, 하얀 물살, 어두운 바다가 프롬프트대로 묘사되었고 인물은 없음.",
    "hard_violations": [],
    "physics": "배가 수면 위에 안정적으로 떠서 항해하며 자연스러운 항적을 만듦."
   },
   {
    "label": "B",
    "direction": "배가 화면 상단으로 향하지만 뱃머리 방향에 육지와 항구가 가로막고 있어 먼 바다 방향과 모순됨.",
    "built_space": "화면 상단 배경에 방파제, 건물, 항구 조명 시설이 위치함.",
    "entities": "낡은 배와 하얀 항적이 있으나, 선실 부근에 사람의 실루엣이 나타남.",
    "hard_violations": [
     "[gemini-pro] 명시적인 인물 등장 금지(NO PEOPLE IN THIS SHOT) 규칙을 위반하고 배 안에 인물 생성",
     "[gemini-pro] 배가 먼 바다가 아닌 항구를 향하는 구도적 모순",
     "[gpt-high] 선내에 앉아 있는 인물과 조타실의 인물 실루엣이 노출되어, 어떤 사람 형상도 등장시키지 말라는 명시적 조건을 위반한다.",
     "[gpt-high] 장소 설명에 없는 다수의 항만 건물·정박선·탑형 시설을 배경에 추가하여 장소 요소를 임의로 발명하지 말라는 제한을 위반한다."
    ],
    "physics": "배가 물 위에 떠서 이동하며 물살을 일으킴."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "인물 배제 규칙을 정확히 따랐으며, 명시된 대로 화면 상단의 먼 바다를 향해 나아가는 낡은 배의 구도를 성공적으로 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "명시된 인물 배제 규칙을 어기고 선상에 사람을 묘사했으며, 배가 먼 바다가 아닌 항구 육지 쪽을 향하고 있어 치명적인 오류를 범함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "배의 선수가 화면 상단의 탁 트인 먼 바다를 향해 올바르게 나아감.",
        "built_space": "인공 구조물 없이 어두운 바다와 배만 단독으로 배치됨.",
        "entities": "낡은 배, 하얀 물살, 어두운 바다가 프롬프트대로 묘사되었고 인물은 없음.",
        "hard_violations": [],
        "physics": "배가 수면 위에 안정적으로 떠서 항해하며 자연스러운 항적을 만듦."
       },
       {
        "label": "B",
        "direction": "배가 화면 상단으로 향하지만 뱃머리 방향에 육지와 항구가 가로막고 있어 먼 바다 방향과 모순됨.",
        "built_space": "화면 상단 배경에 방파제, 건물, 항구 조명 시설이 위치함.",
        "entities": "낡은 배와 하얀 항적이 있으나, 선실 부근에 사람의 실루엣이 나타남.",
        "hard_violations": [
         "명시적인 인물 등장 금지(NO PEOPLE IN THIS SHOT) 규칙을 위반하고 배 안에 인물 생성",
         "배가 먼 바다가 아닌 항구를 향하는 구도적 모순"
        ],
        "physics": "배가 물 위에 떠서 이동하며 물살을 일으킴."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "인물 배제 규칙을 정확히 따랐으며, 명시된 대로 화면 상단의 먼 바다를 향해 나아가는 낡은 배의 구도를 성공적으로 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "명시된 인물 배제 규칙을 어기고 선상에 사람을 묘사했으며, 배가 먼 바다가 아닌 항구 육지 쪽을 향하고 있어 치명적인 오류를 범함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "배의 선수가 화면 상단의 탁 트인 먼 바다를 향해 올바르게 나아감.",
        "built_space": "인공 구조물 없이 어두운 바다와 배만 단독으로 배치됨.",
        "entities": "낡은 배, 하얀 물살, 어두운 바다가 프롬프트대로 묘사되었고 인물은 없음.",
        "hard_violations": [],
        "physics": "배가 수면 위에 안정적으로 떠서 항해하며 자연스러운 항적을 만듦."
       },
       {
        "label": "B",
        "direction": "배가 화면 상단으로 향하지만 뱃머리 방향에 육지와 항구가 가로막고 있어 먼 바다 방향과 모순됨.",
        "built_space": "화면 상단 배경에 방파제, 건물, 항구 조명 시설이 위치함.",
        "entities": "낡은 배와 하얀 항적이 있으나, 선실 부근에 사람의 실루엣이 나타남.",
        "hard_violations": [
         "명시적인 인물 등장 금지(NO PEOPLE IN THIS SHOT) 규칙을 위반하고 배 안에 인물 생성",
         "배가 먼 바다가 아닌 항구를 향하는 구도적 모순"
        ],
        "physics": "배가 물 위에 떠서 이동하며 물살을 일으킴."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "선미와 흰 항적의 방향은 맞지만, 선내 인물이 노출되어 인물 금지 조건을 어기고 배도 요구된 원경의 작은 피사체보다 크게 보입니다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "인물 없이 상단 중앙의 낡은 배를 후방 사선으로 보여주며, 어두운 바다와 하단 전경으로 이어지는 흰 항적이 지정된 구도를 충실히 구현합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "선수는 화면 위쪽에서 약간 왼쪽으로 향하고, 선미와 한쪽 측면이 보인다. 흰 항적은 선미에서 하단 중앙 전경으로 이어져 카메라에서 멀어지는 운항 방향과 맞는다. 선내 인물의 시선은 어두워 판별하기 어렵다.",
        "built_space": "배에는 조타실 하나와 뒤쪽 갑판을 덮는 차양 하나가 보인다. 차양 아래 뒤쪽에 담요를 두른 인물 형상이 있고, 그 앞 조타실에도 별도의 어두운 인물 실루엣이 보인다. 상단 배경에는 부두 건물 여러 동, 정박선 여러 척, 방파제와 탑형 시설이 추가되어 있으며, 배와 바다 중심의 제한된 장소 묘사보다 많은 구조물을 도입했다.",
        "entities": "낡은 소형 동력선 한 척, 어두운 바다, 흰 항적과 해질녘의 낮은 조도는 일치한다. 그러나 담요를 두른 사람과 조타실의 사람 형상이 보여 인물 금지 조건에 맞지 않는다. 얼굴·연령·민족성·신체 손상과 배낭 상태는 확인할 수 없다. 선미의 표식은 선명하게 읽히지 않는다.",
        "hard_violations": [
         "선내에 앉아 있는 인물과 조타실의 인물 실루엣이 노출되어, 어떤 사람 형상도 등장시키지 말라는 명시적 조건을 위반한다.",
         "장소 설명에 없는 다수의 항만 건물·정박선·탑형 시설을 배경에 추가하여 장소 요소를 임의로 발명하지 말라는 제한을 위반한다."
        ],
        "physics": "선체는 수면에 잠겨 부력으로 지지되고, 선미 바로 뒤의 거품과 양옆으로 퍼지는 물결은 동력 운항으로 설명된다. 뒤쪽 인물은 갑판의 좌석 부근에 앉아 있고 담요는 몸 위에 걸쳐져 있다. 조타실 인물은 하체가 가려져 있으나 공중에 떠 있는 정황은 없다."
       },
       {
        "label": "B",
        "direction": "선수는 화면 위쪽에서 약간 왼쪽의 열린 바다로 향한다. 카메라에는 선미와 한쪽 측면이 함께 보이며, 항적은 선미에서 하단 중앙과 오른쪽 전경으로 퍼져 멀어지는 배의 진행 방향과 일치한다.",
        "built_space": "배에는 조타실 하나, 선미의 열린 갑판 하나와 측면 난간이 보인다. 배 전체가 상단 중앙에 작게 들어오고 주위는 바다로 비어 있다. 별도의 부두 건물이나 정박선을 추가하지 않았으며, 사람이 서거나 앉은 모습과 부자연스러운 반사는 보이지 않는다.",
        "entities": "마모된 소형 동력선 한 척, 어두운 바다와 선미부터 이어지는 흰 항적이 확인된다. 살아 있는 사람이나 얼굴은 보이지 않는다. 찰리의 담요·손상된 몸·배낭은 화면에서 식별되지 않으며, 인물을 노출하지 않는 이 구도에서는 이를 불일치로 볼 근거가 없다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "배는 수면에 정상적으로 잠겨 부력으로 지지된다. 선미에 붙어 시작하는 거품과 뒤로 넓어지는 항적은 추진하며 떠나는 동력선의 움직임에 부합한다. 조타실과 난간은 선체에 고정되어 있고, 지지 없이 떠 있는 물체나 인체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "선미와 흰 항적의 방향은 맞지만, 선내 인물이 노출되어 인물 금지 조건을 어기고 배도 요구된 원경의 작은 피사체보다 크게 보입니다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "인물 없이 상단 중앙의 낡은 배를 후방 사선으로 보여주며, 어두운 바다와 하단 전경으로 이어지는 흰 항적이 지정된 구도를 충실히 구현합니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "선수는 화면 위쪽에서 약간 왼쪽으로 향하고, 선미와 한쪽 측면이 보인다. 흰 항적은 선미에서 하단 중앙 전경으로 이어져 카메라에서 멀어지는 운항 방향과 맞는다. 선내 인물의 시선은 어두워 판별하기 어렵다.",
        "built_space": "배에는 조타실 하나와 뒤쪽 갑판을 덮는 차양 하나가 보인다. 차양 아래 뒤쪽에 담요를 두른 인물 형상이 있고, 그 앞 조타실에도 별도의 어두운 인물 실루엣이 보인다. 상단 배경에는 부두 건물 여러 동, 정박선 여러 척, 방파제와 탑형 시설이 추가되어 있으며, 배와 바다 중심의 제한된 장소 묘사보다 많은 구조물을 도입했다.",
        "entities": "낡은 소형 동력선 한 척, 어두운 바다, 흰 항적과 해질녘의 낮은 조도는 일치한다. 그러나 담요를 두른 사람과 조타실의 사람 형상이 보여 인물 금지 조건에 맞지 않는다. 얼굴·연령·민족성·신체 손상과 배낭 상태는 확인할 수 없다. 선미의 표식은 선명하게 읽히지 않는다.",
        "hard_violations": [
         "선내에 앉아 있는 인물과 조타실의 인물 실루엣이 노출되어, 어떤 사람 형상도 등장시키지 말라는 명시적 조건을 위반한다.",
         "장소 설명에 없는 다수의 항만 건물·정박선·탑형 시설을 배경에 추가하여 장소 요소를 임의로 발명하지 말라는 제한을 위반한다."
        ],
        "physics": "선체는 수면에 잠겨 부력으로 지지되고, 선미 바로 뒤의 거품과 양옆으로 퍼지는 물결은 동력 운항으로 설명된다. 뒤쪽 인물은 갑판의 좌석 부근에 앉아 있고 담요는 몸 위에 걸쳐져 있다. 조타실 인물은 하체가 가려져 있으나 공중에 떠 있는 정황은 없다."
       },
       {
        "label": "A",
        "direction": "선수는 화면 위쪽에서 약간 왼쪽의 열린 바다로 향한다. 카메라에는 선미와 한쪽 측면이 함께 보이며, 항적은 선미에서 하단 중앙과 오른쪽 전경으로 퍼져 멀어지는 배의 진행 방향과 일치한다.",
        "built_space": "배에는 조타실 하나, 선미의 열린 갑판 하나와 측면 난간이 보인다. 배 전체가 상단 중앙에 작게 들어오고 주위는 바다로 비어 있다. 별도의 부두 건물이나 정박선을 추가하지 않았으며, 사람이 서거나 앉은 모습과 부자연스러운 반사는 보이지 않는다.",
        "entities": "마모된 소형 동력선 한 척, 어두운 바다와 선미부터 이어지는 흰 항적이 확인된다. 살아 있는 사람이나 얼굴은 보이지 않는다. 찰리의 담요·손상된 몸·배낭은 화면에서 식별되지 않으며, 인물을 노출하지 않는 이 구도에서는 이를 불일치로 볼 근거가 없다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "배는 수면에 정상적으로 잠겨 부력으로 지지된다. 선미에 붙어 시작하는 거품과 뒤로 넓어지는 항적은 추진하며 떠나는 동력선의 움직임에 부합한다. 조타실과 난간은 선체에 고정되어 있고, 지지 없이 떠 있는 물체나 인체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.651
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.401
   },
   "violations": {
    "B": [
     "[gemini-pro] 명시적인 인물 등장 금지(NO PEOPLE IN THIS SHOT) 규칙을 위반하고 배 안에 인물 생성",
     "[gemini-pro] 배가 먼 바다가 아닌 항구를 향하는 구도적 모순",
     "[gpt-high] 선내에 앉아 있는 인물과 조타실의 인물 실루엣이 노출되어, 어떤 사람 형상도 등장시키지 말라는 명시적 조건을 위반한다.",
     "[gpt-high] 장소 설명에 없는 다수의 항만 건물·정박선·탑형 시설을 배경에 추가하여 장소 요소를 임의로 발명하지 말라는 제한을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 401
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "인물 배제 규칙을 정확히 따랐으며, 명시된 대로 화면 상단의 먼 바다를 향해 나아가는 낡은 배의 구도를 성공적으로 구현함."
   },
   {
    "label": "B",
    "score": 401,
    "verdict_ko": "명시된 인물 배제 규칙을 어기고 선상에 사람을 묘사했으며, 배가 먼 바다가 아닌 항구 육지 쪽을 향하고 있어 치명적인 오류를 범함.  ★위반: [gemini-pro] 명시적인 인물 등장 금지(NO PEOPLE IN THIS SHOT) 규칙을 위반하고 배 안에 인물 생성 / [gemini-pro] 배가 먼 바다가 아닌 항구를 향하는 구도적 모순 / [gpt-high] 선내에 앉아 있는 인물과 조타실의 인물 실루엣이 노출되어, 어떤 사람 형상도 등장시키지 말라는 명시적 조건을 위반한다. / [gpt-high] 장소 설명에 없는 다수의 항만 건물·정박선·탑형 시설을 배경에 추가하여 장소 요소를 임의로 발명하지 말라는 제한을 위반한다."
   }
  ],
  "refs": [],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-d07f-73e0-b4ed-d68439b48722",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S78sh17::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:52:44.466196+00:00",
  "fingerprint": "3bfc1aeb855b481a554a465f8647a1253dee84f964eba36c55f956985bbbb523",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S78sh17_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S78sh17_sel.png",
  "source_sha256": "4b146f4966b201e9eb79bda26cdd6b094b7e1c64bf4d9a8e11b55fa7ec889813",
  "file": "S78sh17_cine.png",
  "staged_sha256": "986982dba03a996a80b3be4cc33055bc82e4640ac21f5c1e72d296f29c8ff8b9",
  "latency_ms": 9225
 },
 "S79sh10::signage": {
  "fp": "8ad1272ac5edb949",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S79sh10": {
  "input_fingerprint": "ed64c145871b79bd",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 현우의 물음에 복잡한 표정의 디지털 눈 이모티콘을 띄운 찰리의 낡은 얼굴 클로즈업.\n\nLOCATION (lock): At the sofa inside the boat's small two-person crew cabin, under modest nighttime cabin lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Sofa (Supporting both reclining figures) — A narrow portion behind 찰리 and beneath 현우's shoulder remains visible; used as Maintains the shared resting position without competing with 찰리's face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient illumination and gentle tonal separation keep the digital expression readable without adding an unsupported light source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The small two-person cabin door is closed, and the sofa remains the resting place. Charlie reclines there with his accumulated body damage; the blanket and travel backpack have no stated removal.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 현우의 물음에 복잡한 표정의 디지털 눈 이모티콘을 띄운 찰리의 낡은 얼굴 클로즈업.\n\nLOCATION (lock): At the sofa inside the boat's small two-person crew cabin, under modest nighttime cabin lighting. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Sofa (Supporting both reclining figures) — A narrow portion behind 찰리 and beneath 현우's shoulder remains visible; used as Maintains the shared resting position without competing with 찰리's face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient illumination and gentle tonal separation keep the digital expression readable without adding an unsupported light source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The small two-person cabin door is closed, and the sofa remains the resting place. Charlie reclines there with his accumulated body damage; the blanket and travel backpack have no stated removal.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 현우의 물음에 복잡한 표정의 디지털 눈 이모티콘을 띄운 찰리의 낡은 얼굴 클로즈업.\n\nLOCATION (lock): At the sofa inside the boat's small two-person crew cabin, under modest nighttime cabin lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Sofa (Supporting both reclining figures) — A narrow portion behind 찰리 and beneath 현우's shoulder remains visible; used as Maintains the shared resting position without competing with 찰리's face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient illumination and gentle tonal separation keep the digital expression readable without adding an unsupported light source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The small two-person cabin door is closed, and the sofa remains the resting place. Charlie reclines there with his accumulated body damage; the blanket and travel backpack have no stated removal.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S79sh10__bgfirst_bg.png",
     "asset_id": "9ca43f94-b62c-4185-bbd5-d874b7069c57",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S79sh10.png",
     "asset_id": "471ec738-cab9-44f7-b95c-4082e4ab2709",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L263B02.png",
     "asset_id": "a3a7948a-4ed6-4034-bdbc-66720a85c15b",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "찰리의 시선이 왼쪽 전경에 위치한 현우를 향함.",
    "built_space": "좁은 선실 내부로 벽면에 지도와 뒷편 창문이 보이며 찰리 뒤로 소파가 위치함.",
    "entities": "낡은 질감의 찰리 외형이 일치하나 눈은 단순 점선 눈동자 형태임. 현우의 뒷모습 일부가 보임.",
    "hard_violations": [],
    "physics": "인물이 소파에 앉아 안정적으로 지지받고 있음."
   },
   {
    "label": "B",
    "direction": "찰리의 시선이 오른쪽 전경의 현우를 향함.",
    "built_space": "선실 내 소파에 위치하며 뒤로 창문과 선반 등 구조물이 확인됨.",
    "entities": "찰리의 외형이 레퍼런스와 일치하며 두 눈에 명확한 텍스트 형태의 이모티콘이 표시됨. 현우의 어깨가 전경에 존재함.",
    "hard_violations": [],
    "physics": "소파에 비스듬히 기댄 자세가 자연스럽게 지지됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "눈에 표시된 텍스트 기호 형태가 '복잡한 표정의 디지털 눈 이모티콘'이라는 지시사항을 가장 직관적이고 정확하게 구현함."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "전반적인 구도와 배경은 준수하나, 눈의 표현이 이모티콘보다는 단순한 형태의 LED 안구에 가까워 묘사가 다소 부족함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 시선이 왼쪽 전경에 위치한 현우를 향함.",
        "built_space": "좁은 선실 내부로 벽면에 지도와 뒷편 창문이 보이며 찰리 뒤로 소파가 위치함.",
        "entities": "낡은 질감의 찰리 외형이 일치하나 눈은 단순 점선 눈동자 형태임. 현우의 뒷모습 일부가 보임.",
        "hard_violations": [],
        "physics": "인물이 소파에 앉아 안정적으로 지지받고 있음."
       },
       {
        "label": "B",
        "direction": "찰리의 시선이 오른쪽 전경의 현우를 향함.",
        "built_space": "선실 내 소파에 위치하며 뒤로 창문과 선반 등 구조물이 확인됨.",
        "entities": "찰리의 외형이 레퍼런스와 일치하며 두 눈에 명확한 텍스트 형태의 이모티콘이 표시됨. 현우의 어깨가 전경에 존재함.",
        "hard_violations": [],
        "physics": "소파에 비스듬히 기댄 자세가 자연스럽게 지지됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "눈에 표시된 텍스트 기호 형태가 '복잡한 표정의 디지털 눈 이모티콘'이라는 지시사항을 가장 직관적이고 정확하게 구현함."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "전반적인 구도와 배경은 준수하나, 눈의 표현이 이모티콘보다는 단순한 형태의 LED 안구에 가까워 묘사가 다소 부족함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 시선이 왼쪽 전경에 위치한 현우를 향함.",
        "built_space": "좁은 선실 내부로 벽면에 지도와 뒷편 창문이 보이며 찰리 뒤로 소파가 위치함.",
        "entities": "낡은 질감의 찰리 외형이 일치하나 눈은 단순 점선 눈동자 형태임. 현우의 뒷모습 일부가 보임.",
        "hard_violations": [],
        "physics": "인물이 소파에 앉아 안정적으로 지지받고 있음."
       },
       {
        "label": "B",
        "direction": "찰리의 시선이 오른쪽 전경의 현우를 향함.",
        "built_space": "선실 내 소파에 위치하며 뒤로 창문과 선반 등 구조물이 확인됨.",
        "entities": "찰리의 외형이 레퍼런스와 일치하며 두 눈에 명확한 텍스트 형태의 이모티콘이 표시됨. 현우의 어깨가 전경에 존재함.",
        "hard_violations": [],
        "physics": "소파에 비스듬히 기댄 자세가 자연스럽게 지지됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "얼굴 비중이 더 크고 비대칭 디지털 표정과 소파에 기댄 자세가 명확하지만, 가슴과 팔까지 넓게 담아 요구된 얼굴 클로즈업에는 못 미친다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "찰리의 마모된 외형과 야간 선실은 부합하지만, 현우의 큰 전경과 찰리의 상체가 얼굴 클로즈업을 밀어내고 디지털 눈의 복잡한 감정도 덜 분명하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 머리를 화면 오른쪽의 현우 쪽으로 기울이고 얼굴을 그쪽으로 돌리고 있다. 디지털 눈은 서로 다른 점·선 표정이어서 정확한 동공 시선은 특정하기 어렵지만, 대화 상대를 향한 반응으로 읽힌다. 오른쪽 전경의 현우도 찰리를 향한다. 조준하거나 이동하는 물체는 없다.",
        "built_space": "회색 소파 한 개의 등받이와 좌면 일부가 찰리 뒤와 아래에 보인다. 배경에는 지도 한 장, 책이 꽂힌 선반 구역 하나, 작은 조명 하나, 커튼이 있는 선창 하나, 오른쪽의 닫힌 문 하나가 보이며 참고 선실의 주요 배치와 대체로 맞는다. 다만 소파와 선실 배경이 요구된 좁은 배경 조각보다 넓게 드러난다. 현우의 어깨 아래 지지면은 전경에 가려져 확인되지 않는다. 반사는 없다.",
        "entities": "찰리 한 명과 현우로 읽히는 부분 인물 한 명이 보인다. 찰리의 샌드 베이지 장갑판, 흰 각진 마스크, 안테나, 노출된 목 기계부와 긁힌 표면은 참고 정체성에 부합한다. 눈 내부의 흰 점·선 표시는 물리적 디스플레이에 들어가 있으며 서로 다른 난처한 표정을 나타낸다. 현우는 짧은 검은 머리와 어두운 옷을 입은 성인 남성으로 보이나, 흐린 뒷모습만으로 세부 정체성은 확인할 수 없다. 담요와 여행 배낭은 이 구도에서 확인되지 않으며 제거되었다고 단정할 수 없다. 명확히 읽히는 문구나 화면 위 자막은 없다.",
        "hard_violations": [],
        "physics": "찰리의 기울어진 상체는 소파 등받이에 기대고 있으며 하체 쪽은 좌면으로 이어진다. 머리는 기계식 목에 연결되고 팔은 몸 옆으로 내려와 있어 떠 있는 부품은 없다. 현우의 하체는 프레임 밖이므로 지지 자세와 함께 누운 상태는 확인되지 않지만, 공중에 떠 있다고 볼 근거도 없다."
       },
       {
        "label": "B",
        "direction": "찰리의 얼굴은 화면 왼쪽 전경의 현우를 향하고, 디지털 눈도 왼쪽 상대에게 반응하는 인상을 준다. 현우는 찰리를 바라보는 뒷모습이다. 두 눈의 높이와 모양에 차이는 있지만, 복잡한 감정보다는 처지거나 의문을 품은 표정에 가깝다. 무기나 이동 물체는 없다.",
        "built_space": "찰리 뒤와 현우 어깨 옆에 회색 소파 등받이 한 개가 보인다. 배경에는 지도 한 장, 책장 구역 하나, 벽 조명 하나, 커튼이 달린 선창 하나, 조리대와 하부 수납장 일부가 보인다. 참고 장소의 재료와 시설 종류는 대체로 유지되지만 선창·조리대까지 크게 노출되어 얼굴보다 공간 설명의 비중이 높다. 출입문은 프레임 밖이므로 닫힘 여부를 판단할 수 없다. 불가능한 반사는 보이지 않는다.",
        "entities": "찰리 한 명과 현우로 읽히는 부분 인물 한 명이 있다. 찰리의 베이지색 대형 어깨 장갑, 흰 마스크, 안테나와 기계식 목은 참고와 부합하며 얼굴의 마모도 뚜렷하다. 눈은 점 배열 디스플레이지만 감정의 복합성이 A보다 약하다. 왼쪽 인물은 짧은 검은 머리의 성인 동아시아계 남성으로 보이며 회색 상의를 입었다. 담요와 배낭은 확인되지 않으나 프레임 밖일 수 있다. 배경 책등과 지도에 작은 인쇄 흔적이 있지만 확실히 판독되는 문구는 확인하기 어렵다.",
        "hard_violations": [],
        "physics": "찰리의 등 뒤에 소파 등받이가 있어 상체를 받칠 수 있으며, 머리와 양팔은 목 및 어깨 관절에 정상적으로 연결된다. 다만 자세가 비교적 세워져 있어 편히 누워 쉬는 상태는 A보다 약하게 전달된다. 현우의 어깨와 몸통은 프레임 아래로 이어지고 지지면은 가려져 있다. 지지 없이 떠 있는 인물이나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "얼굴 비중이 더 크고 비대칭 디지털 표정과 소파에 기댄 자세가 명확하지만, 가슴과 팔까지 넓게 담아 요구된 얼굴 클로즈업에는 못 미친다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "찰리의 마모된 외형과 야간 선실은 부합하지만, 현우의 큰 전경과 찰리의 상체가 얼굴 클로즈업을 밀어내고 디지털 눈의 복잡한 감정도 덜 분명하다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "찰리는 머리를 화면 오른쪽의 현우 쪽으로 기울이고 얼굴을 그쪽으로 돌리고 있다. 디지털 눈은 서로 다른 점·선 표정이어서 정확한 동공 시선은 특정하기 어렵지만, 대화 상대를 향한 반응으로 읽힌다. 오른쪽 전경의 현우도 찰리를 향한다. 조준하거나 이동하는 물체는 없다.",
        "built_space": "회색 소파 한 개의 등받이와 좌면 일부가 찰리 뒤와 아래에 보인다. 배경에는 지도 한 장, 책이 꽂힌 선반 구역 하나, 작은 조명 하나, 커튼이 있는 선창 하나, 오른쪽의 닫힌 문 하나가 보이며 참고 선실의 주요 배치와 대체로 맞는다. 다만 소파와 선실 배경이 요구된 좁은 배경 조각보다 넓게 드러난다. 현우의 어깨 아래 지지면은 전경에 가려져 확인되지 않는다. 반사는 없다.",
        "entities": "찰리 한 명과 현우로 읽히는 부분 인물 한 명이 보인다. 찰리의 샌드 베이지 장갑판, 흰 각진 마스크, 안테나, 노출된 목 기계부와 긁힌 표면은 참고 정체성에 부합한다. 눈 내부의 흰 점·선 표시는 물리적 디스플레이에 들어가 있으며 서로 다른 난처한 표정을 나타낸다. 현우는 짧은 검은 머리와 어두운 옷을 입은 성인 남성으로 보이나, 흐린 뒷모습만으로 세부 정체성은 확인할 수 없다. 담요와 여행 배낭은 이 구도에서 확인되지 않으며 제거되었다고 단정할 수 없다. 명확히 읽히는 문구나 화면 위 자막은 없다.",
        "hard_violations": [],
        "physics": "찰리의 기울어진 상체는 소파 등받이에 기대고 있으며 하체 쪽은 좌면으로 이어진다. 머리는 기계식 목에 연결되고 팔은 몸 옆으로 내려와 있어 떠 있는 부품은 없다. 현우의 하체는 프레임 밖이므로 지지 자세와 함께 누운 상태는 확인되지 않지만, 공중에 떠 있다고 볼 근거도 없다."
       },
       {
        "label": "A",
        "direction": "찰리의 얼굴은 화면 왼쪽 전경의 현우를 향하고, 디지털 눈도 왼쪽 상대에게 반응하는 인상을 준다. 현우는 찰리를 바라보는 뒷모습이다. 두 눈의 높이와 모양에 차이는 있지만, 복잡한 감정보다는 처지거나 의문을 품은 표정에 가깝다. 무기나 이동 물체는 없다.",
        "built_space": "찰리 뒤와 현우 어깨 옆에 회색 소파 등받이 한 개가 보인다. 배경에는 지도 한 장, 책장 구역 하나, 벽 조명 하나, 커튼이 달린 선창 하나, 조리대와 하부 수납장 일부가 보인다. 참고 장소의 재료와 시설 종류는 대체로 유지되지만 선창·조리대까지 크게 노출되어 얼굴보다 공간 설명의 비중이 높다. 출입문은 프레임 밖이므로 닫힘 여부를 판단할 수 없다. 불가능한 반사는 보이지 않는다.",
        "entities": "찰리 한 명과 현우로 읽히는 부분 인물 한 명이 있다. 찰리의 베이지색 대형 어깨 장갑, 흰 마스크, 안테나와 기계식 목은 참고와 부합하며 얼굴의 마모도 뚜렷하다. 눈은 점 배열 디스플레이지만 감정의 복합성이 A보다 약하다. 왼쪽 인물은 짧은 검은 머리의 성인 동아시아계 남성으로 보이며 회색 상의를 입었다. 담요와 배낭은 확인되지 않으나 프레임 밖일 수 있다. 배경 책등과 지도에 작은 인쇄 흔적이 있지만 확실히 판독되는 문구는 확인하기 어렵다.",
        "hard_violations": [],
        "physics": "찰리의 등 뒤에 소파 등받이가 있어 상체를 받칠 수 있으며, 머리와 양팔은 목 및 어깨 관절에 정상적으로 연결된다. 다만 자세가 비교적 세워져 있어 편히 누워 쉬는 상태는 A보다 약하게 전달된다. 현우의 어깨와 몸통은 프레임 아래로 이어지고 지지면은 가려져 있다. 지지 없이 떠 있는 인물이나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.571,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.571,
    "B": 2.0
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1571
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "눈에 표시된 텍스트 기호 형태가 '복잡한 표정의 디지털 눈 이모티콘'이라는 지시사항을 가장 직관적이고 정확하게 구현함."
   },
   {
    "label": "A",
    "score": 1571,
    "verdict_ko": "전반적인 구도와 배경은 준수하나, 눈의 표현이 이모티콘보다는 단순한 형태의 LED 안구에 가까워 묘사가 다소 부족함."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L263B02.png",
    "asset_id": "a3a7948a-4ed6-4034-bdbc-66720a85c15b",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-d220-775f-9ed3-8347f9a4897e",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S79sh10__bgfirst_bg.png",
   "bg_asset_id": "9ca43f94-b62c-4185-bbd5-d874b7069c57",
   "bg_record_key": "S79sh10::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S79sh10::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T12:02:32.910900+00:00",
  "fingerprint": "3f11e93aad345a750f8c8c9354680ac76e30f8b3128cae2d701e569751eabb14",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S79sh10_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S79sh10_sel.png",
  "source_sha256": "743797096d641d4f07c284271c3efd4eb11f2b0845e09da248235346a865968e",
  "file": "S79sh10_cine.png",
  "staged_sha256": "f4ba9e4b4b1449b7c6f3d4b3e6ea64d546e8b3c6f93a1b96b186b9d209642c7d",
  "latency_ms": 12807
 },
 "S79sh12::signage": {
  "fp": "80220a2ca3440b3f",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S79sh12": {
  "input_fingerprint": "cff0a43a72b32763",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 낮게 울리는 진동 속, 깜짝 놀란 표정으로 두 눈을 번쩍 뜬 채 굳어 있는 현우의 고요한 얼굴.\n\nLOCATION (lock): On the sofa inside the small crew cabin aboard the boat, in low nighttime cabin light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Sofa (Still supporting 현우 as he wakes) — A cropped supporting section sits beneath and behind his head; used as Anchors the arrested face to his previous sleeping position.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the cabin's ambient illumination unchanged across the temporal cut, with restrained contrast preserving the stillness of 현우's startled face.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The cabin door remains closed and the sofa remains in place. Charlie is still in the cabin with his damaged body, blanket and travel backpack. 현우: He has fallen asleep on the sofa and is now waking; his dirty clothing and accumulated injuries remain unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 낮게 울리는 진동 속, 깜짝 놀란 표정으로 두 눈을 번쩍 뜬 채 굳어 있는 현우의 고요한 얼굴.\n\nLOCATION (lock): On the sofa inside the small crew cabin aboard the boat, in low nighttime cabin light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Sofa (Still supporting 현우 as he wakes) — A cropped supporting section sits beneath and behind his head; used as Anchors the arrested face to his previous sleeping position.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the cabin's ambient illumination unchanged across the temporal cut, with restrained contrast preserving the stillness of 현우's startled face.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The cabin door remains closed and the sofa remains in place. Charlie is still in the cabin with his damaged body, blanket and travel backpack. 현우: He has fallen asleep on the sofa and is now waking; his dirty clothing and accumulated injuries remain unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 낮게 울리는 진동 속, 깜짝 놀란 표정으로 두 눈을 번쩍 뜬 채 굳어 있는 현우의 고요한 얼굴.\n\nLOCATION (lock): On the sofa inside the small crew cabin aboard the boat, in low nighttime cabin light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Sofa (Still supporting 현우 as he wakes) — A cropped supporting section sits beneath and behind his head; used as Anchors the arrested face to his previous sleeping position.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Keep the cabin's ambient illumination unchanged across the temporal cut, with restrained contrast preserving the stillness of 현우's startled face.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The cabin door remains closed and the sofa remains in place. Charlie is still in the cabin with his damaged body, blanket and travel backpack. 현우: He has fallen asleep on the sofa and is now waking; his dirty clothing and accumulated injuries remain unchanged.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "시선은 정면 위쪽을 향함.",
    "built_space": "선실 내부. 전경의 소파 외에 배경 왼쪽에도 동일한 회색 소파가 추가로 존재함.",
    "entities": "현우(18세 아시아계 남성)의 외모가 일치함. 지시된 오염된 의상과 놀란 표정을 정확히 구현함.",
    "hard_violations": [
     "[gemini-pro] 배경 왼쪽에 소파 중복 생성 (duplicated fitting)"
    ],
    "physics": "등과 머리가 소파 등받이에 기대어 지탱됨."
   },
   {
    "label": "B",
    "direction": "시선은 카메라 정면을 향함.",
    "built_space": "선실 내부. 전경 소파, 배경 왼쪽 나무 책상, 오른쪽 하얀색 캐비닛 구조가 레퍼런스와 일치함.",
    "entities": "현우(18세 아시아계 남성)의 외모가 일치함. 놀란 표정은 잘 연출되었으나, 지시와 달리 깨끗한 남색 옷을 입고 있음.",
    "hard_violations": [
     "[gpt-high] 잠금된 선실 배치에 없는 탁자와 물컵을 소파 뒤 중경에 추가했다."
    ],
    "physics": "등과 머리가 소파에 안정적으로 지탱됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "레퍼런스와 일치하는 공간 구조를 유지했으나, 의상의 오염 상태 지시를 누락하여 감점됨."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "의상의 오염 상태는 잘 표현했으나, 배경에 소파가 중복 생성되는 치명적 공간 오류로 실격됨."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 정면 위쪽을 향함.",
        "built_space": "선실 내부. 전경의 소파 외에 배경 왼쪽에도 동일한 회색 소파가 추가로 존재함.",
        "entities": "현우(18세 아시아계 남성)의 외모가 일치함. 지시된 오염된 의상과 놀란 표정을 정확히 구현함.",
        "hard_violations": [
         "배경 왼쪽에 소파 중복 생성 (duplicated fitting)"
        ],
        "physics": "등과 머리가 소파 등받이에 기대어 지탱됨."
       },
       {
        "label": "B",
        "direction": "시선은 카메라 정면을 향함.",
        "built_space": "선실 내부. 전경 소파, 배경 왼쪽 나무 책상, 오른쪽 하얀색 캐비닛 구조가 레퍼런스와 일치함.",
        "entities": "현우(18세 아시아계 남성)의 외모가 일치함. 놀란 표정은 잘 연출되었으나, 지시와 달리 깨끗한 남색 옷을 입고 있음.",
        "hard_violations": [],
        "physics": "등과 머리가 소파에 안정적으로 지탱됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "레퍼런스와 일치하는 공간 구조를 유지했으나, 의상의 오염 상태 지시를 누락하여 감점됨."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "의상의 오염 상태는 잘 표현했으나, 배경에 소파가 중복 생성되는 치명적 공간 오류로 실격됨."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 정면 위쪽을 향함.",
        "built_space": "선실 내부. 전경의 소파 외에 배경 왼쪽에도 동일한 회색 소파가 추가로 존재함.",
        "entities": "현우(18세 아시아계 남성)의 외모가 일치함. 지시된 오염된 의상과 놀란 표정을 정확히 구현함.",
        "hard_violations": [
         "배경 왼쪽에 소파 중복 생성 (duplicated fitting)"
        ],
        "physics": "등과 머리가 소파 등받이에 기대어 지탱됨."
       },
       {
        "label": "B",
        "direction": "시선은 카메라 정면을 향함.",
        "built_space": "선실 내부. 전경 소파, 배경 왼쪽 나무 책상, 오른쪽 하얀색 캐비닛 구조가 레퍼런스와 일치함.",
        "entities": "현우(18세 아시아계 남성)의 외모가 일치함. 놀란 표정은 잘 연출되었으나, 지시와 달리 깨끗한 남색 옷을 입고 있음.",
        "hard_violations": [],
        "physics": "등과 머리가 소파에 안정적으로 지탱됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "얼굴 중심의 밀착 구도와 남색 의상은 잘 맞지만, 선실에 탁자와 물컵을 추가했고 잠들었던 자세보다는 상체를 세운 정면 반응으로 보인다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "소파에 머리를 기댄 채 눈을 번쩍 뜬 정지 순간과 선실의 공간 연속성이 더 충실하지만, 얼굴 비중이 다소 작고 티셔츠 색은 인물 참조와 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "두 눈을 크게 뜨고 카메라 쪽 정면을 응시한다. 화면 안에 시선의 대상은 없으며, 원문도 특정 대상을 지정하지 않는다. 입을 조금 벌린 놀람은 분명하지만 시선이 렌즈에 직접 고정된 인상이 강하다. 무기나 방향을 확인할 휴대 소품은 없다.",
        "built_space": "앞쪽 회색 소파 등받이가 어깨와 머리 뒤를 차지하고, 왼쪽 뒤에도 회색 좌석 일부가 보인다. 왼쪽 벽 선반, 따뜻한 조명 한 개, 오른쪽 닫힌 문 한 개가 보인다. 참조의 좌석 주변에는 없던 탁자 한 개와 투명 물컵 한 개가 왼쪽 중경에 추가되어 공간 구성이 바뀌었다. 반사상은 없다.",
        "entities": "인물은 한 명이며 앳된 동아시아계 남성, 헝클어진 검은 머리, 남색 둥근목 티셔츠가 현우 참조와 대체로 맞는다. 국적은 외관만으로 확인할 수 없다. 얼굴에 땀과 피부 자국은 있지만 누적 부상과 옷의 오염은 뚜렷하지 않다. 회색 직물 소파는 참조와 유사하다. 찰리, 담요, 배낭은 이 얼굴 구도 밖이므로 부재로 감점할 사항이 아니다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "잠금된 선실 배치에 없는 탁자와 물컵을 소파 뒤 중경에 추가했다."
        ],
        "physics": "등과 어깨는 소파 등받이에 기대고 있으며 목이 머리를 지탱한다. 머리 뒤에도 소파가 있어 공중에 떠 있는 신체는 아니다. 다만 얼굴과 목이 거의 수직이어서 누워 자다가 그대로 깨어난 자세보다는 일으켜 앉은 자세에 가깝다. 물컵은 탁자 위에 놓여 있어 지지는 정상이다."
       },
       {
        "label": "B",
        "direction": "두 눈을 크게 뜨고 카메라보다 위쪽의 화면 밖을 바라본다. 특정 대상은 보이지 않지만 원문에 지정된 시선 대상도 없다. 뒤로 기댄 상태에서 갑자기 깨어 위를 보는 방향이 자연스럽다. 무기나 지향성 소품은 없다.",
        "built_space": "머리와 어깨 바로 뒤에 회색 소파의 낮은 받침 부분이 있고, 오른쪽에는 높은 등받이 부분, 왼쪽 뒤에는 이어지는 좌석과 쿠션 한 개가 보인다. 왼쪽 벽 선반, 따뜻한 조명 한 개, 뒤쪽 작은 창 한 개, 오른쪽 닫힌 문 한 개가 참조의 선실 구성과 대체로 연결된다. 얼굴 주변보다 선실과 소파가 차지하는 면적이 다소 크다. 불가능한 반사나 명백한 설비 중복은 보이지 않는다.",
        "entities": "한 명의 앳된 동아시아계 남성이 등장하며 검은 헝클어진 머리와 얼굴 특징은 현우 참조에 대체로 부합한다. 눈은 정상적인 홍채와 동공을 유지하면서 크게 떠 있다. 티셔츠에 오염은 분명하지만 회갈색으로, 참조의 남색 의상과 다르다. 뚜렷한 부상은 확인하기 어렵다. 회색 직물 소파와 따뜻하고 어두운 선실 조명은 참조에 가깝다. 찰리와 그의 소지품은 프레임 밖이며, 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "뒤통수와 목 뒤가 소파 받침에 닿고 어깨와 상체도 소파에 실려 있어 체중 지지가 명확하다. 누운 자세를 유지한 채 눈만 크게 뜨고 입을 살짝 벌린 상태로, 잠에서 놀라 깨어 굳는 행동으로 가능한 자세다. 지지 없이 떠 있는 신체나 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "얼굴 중심의 밀착 구도와 남색 의상은 잘 맞지만, 선실에 탁자와 물컵을 추가했고 잠들었던 자세보다는 상체를 세운 정면 반응으로 보인다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "소파에 머리를 기댄 채 눈을 번쩍 뜬 정지 순간과 선실의 공간 연속성이 더 충실하지만, 얼굴 비중이 다소 작고 티셔츠 색은 인물 참조와 다르다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "두 눈을 크게 뜨고 카메라 쪽 정면을 응시한다. 화면 안에 시선의 대상은 없으며, 원문도 특정 대상을 지정하지 않는다. 입을 조금 벌린 놀람은 분명하지만 시선이 렌즈에 직접 고정된 인상이 강하다. 무기나 방향을 확인할 휴대 소품은 없다.",
        "built_space": "앞쪽 회색 소파 등받이가 어깨와 머리 뒤를 차지하고, 왼쪽 뒤에도 회색 좌석 일부가 보인다. 왼쪽 벽 선반, 따뜻한 조명 한 개, 오른쪽 닫힌 문 한 개가 보인다. 참조의 좌석 주변에는 없던 탁자 한 개와 투명 물컵 한 개가 왼쪽 중경에 추가되어 공간 구성이 바뀌었다. 반사상은 없다.",
        "entities": "인물은 한 명이며 앳된 동아시아계 남성, 헝클어진 검은 머리, 남색 둥근목 티셔츠가 현우 참조와 대체로 맞는다. 국적은 외관만으로 확인할 수 없다. 얼굴에 땀과 피부 자국은 있지만 누적 부상과 옷의 오염은 뚜렷하지 않다. 회색 직물 소파는 참조와 유사하다. 찰리, 담요, 배낭은 이 얼굴 구도 밖이므로 부재로 감점할 사항이 아니다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "잠금된 선실 배치에 없는 탁자와 물컵을 소파 뒤 중경에 추가했다."
        ],
        "physics": "등과 어깨는 소파 등받이에 기대고 있으며 목이 머리를 지탱한다. 머리 뒤에도 소파가 있어 공중에 떠 있는 신체는 아니다. 다만 얼굴과 목이 거의 수직이어서 누워 자다가 그대로 깨어난 자세보다는 일으켜 앉은 자세에 가깝다. 물컵은 탁자 위에 놓여 있어 지지는 정상이다."
       },
       {
        "label": "A",
        "direction": "두 눈을 크게 뜨고 카메라보다 위쪽의 화면 밖을 바라본다. 특정 대상은 보이지 않지만 원문에 지정된 시선 대상도 없다. 뒤로 기댄 상태에서 갑자기 깨어 위를 보는 방향이 자연스럽다. 무기나 지향성 소품은 없다.",
        "built_space": "머리와 어깨 바로 뒤에 회색 소파의 낮은 받침 부분이 있고, 오른쪽에는 높은 등받이 부분, 왼쪽 뒤에는 이어지는 좌석과 쿠션 한 개가 보인다. 왼쪽 벽 선반, 따뜻한 조명 한 개, 뒤쪽 작은 창 한 개, 오른쪽 닫힌 문 한 개가 참조의 선실 구성과 대체로 연결된다. 얼굴 주변보다 선실과 소파가 차지하는 면적이 다소 크다. 불가능한 반사나 명백한 설비 중복은 보이지 않는다.",
        "entities": "한 명의 앳된 동아시아계 남성이 등장하며 검은 헝클어진 머리와 얼굴 특징은 현우 참조에 대체로 부합한다. 눈은 정상적인 홍채와 동공을 유지하면서 크게 떠 있다. 티셔츠에 오염은 분명하지만 회갈색으로, 참조의 남색 의상과 다르다. 뚜렷한 부상은 확인하기 어렵다. 회색 직물 소파와 따뜻하고 어두운 선실 조명은 참조에 가깝다. 찰리와 그의 소지품은 프레임 밖이며, 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "뒤통수와 목 뒤가 소파 받침에 닿고 어깨와 상체도 소파에 실려 있어 체중 지지가 명확하다. 누운 자세를 유지한 채 눈만 크게 뜨고 입을 살짝 벌린 상태로, 잠에서 놀라 깨어 굳는 행동으로 가능한 자세다. 지지 없이 떠 있는 신체나 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.5,
    "B": 1.5
   },
   "adjusted": {
    "A": 1.25,
    "B": 1.25
   },
   "violations": {
    "A": [
     "[gemini-pro] 배경 왼쪽에 소파 중복 생성 (duplicated fitting)"
    ],
    "B": [
     "[gpt-high] 잠금된 선실 배치에 없는 탁자와 물컵을 소파 뒤 중경에 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1250,
   "A": 1250
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1250,
    "verdict_ko": "레퍼런스와 일치하는 공간 구조를 유지했으나, 의상의 오염 상태 지시를 누락하여 감점됨.  ★위반: [gpt-high] 잠금된 선실 배치에 없는 탁자와 물컵을 소파 뒤 중경에 추가했다."
   },
   {
    "label": "A",
    "score": 1250,
    "verdict_ko": "의상의 오염 상태는 잘 표현했으나, 배경에 소파가 중복 생성되는 치명적 공간 오류로 실격됨.  ★위반: [gemini-pro] 배경 왼쪽에 소파 중복 생성 (duplicated fitting)"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S79sh10_sel.png",
    "asset_id": "1488d608-8806-42bd-b372-f0a8c3977d3c",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-d572-7170-8f84-d865aaf46f82",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S79sh10"
  }
 },
 "S79sh12::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T12:03:54.521042+00:00",
  "fingerprint": "6046501320878409c784f0f8d9293035e3369e4404abe89ce88d368f075e37f1",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S79sh12_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S79sh12_sel.png",
  "source_sha256": "6e2b5215a70908d3c49033bfb4b066b317e6cf2e16af63d561d948acbce0df71",
  "file": "S79sh12_cine.png",
  "staged_sha256": "334edd6c630d998a5e636e333ade801e1d1322e147b3bd93859e11297cca22d9",
  "latency_ms": 10781
 },
 "S79sh14::signage": {
  "fp": "795724a02daeef62",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S79sh14": {
  "input_fingerprint": "13a6a8be17e1a898",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 기울어진 바닥을 타고 선원실 벽 쪽으로 거칠게 미끄러지는 도중, 몸이 아래 방향으로 강하게 쏠린 mid-action 자세의 현우와 찰리 역동적인 전신.\n\nLOCATION (lock): Inside the boat's small crew cabin, along the floor sloping toward the wall as the vessel lists in nighttime cabin light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Sofa at the start of the slide in the upper-left of the frame, background; Destination cabin wall in the lower-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Sofa (Behind the pair as they slide away) — Viewed obliquely from above at the uphill side of the composition; used as Marks the origin of their displacement; Cabin floor (Sharply tilted with the ship) — Its visible plane slopes through the composition toward lower right; used as Makes the direction of involuntary movement legible; Cabin wall (Ahead of the sliding pair, before contact) — The interior face is visible at the downhill end of their trajectory; used as Defines the imminent collision and remaining clearance.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the cabin's existing ambient illumination and controlled contrast as the physical tilt disrupts the previously quiet composition.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The vessel and cabin are now sharply heeled to one side, with the cabin door still closed unless subsequently opened. Charlie's damaged body is thrown toward the wall; no removal of his blanket or travel backpack is established. 현우: He is thrown toward the cabin wall as the floor tilts, still bearing his earlier injuries and dirty clothing. The cabin is visibly tilted by the vessel's violent roll.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 기울어진 바닥을 타고 선원실 벽 쪽으로 거칠게 미끄러지는 도중, 몸이 아래 방향으로 강하게 쏠린 mid-action 자세의 현우와 찰리 역동적인 전신.\n\nLOCATION (lock): Inside the boat's small crew cabin, along the floor sloping toward the wall as the vessel lists in nighttime cabin light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Sofa at the start of the slide in the upper-left of the frame, background; Destination cabin wall in the lower-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Sofa (Behind the pair as they slide away) — Viewed obliquely from above at the uphill side of the composition; used as Marks the origin of their displacement; Cabin floor (Sharply tilted with the ship) — Its visible plane slopes through the composition toward lower right; used as Makes the direction of involuntary movement legible; Cabin wall (Ahead of the sliding pair, before contact) — The interior face is visible at the downhill end of their trajectory; used as Defines the imminent collision and remaining clearance.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the cabin's existing ambient illumination and controlled contrast as the physical tilt disrupts the previously quiet composition.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The vessel and cabin are now sharply heeled to one side, with the cabin door still closed unless subsequently opened. Charlie's damaged body is thrown toward the wall; no removal of his blanket or travel backpack is established. 현우: He is thrown toward the cabin wall as the floor tilts, still bearing his earlier injuries and dirty clothing. The cabin is visibly tilted by the vessel's violent roll.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 기울어진 바닥을 타고 선원실 벽 쪽으로 거칠게 미끄러지는 도중, 몸이 아래 방향으로 강하게 쏠린 mid-action 자세의 현우와 찰리 역동적인 전신.\n\nLOCATION (lock): Inside the boat's small crew cabin, along the floor sloping toward the wall as the vessel lists in nighttime cabin light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Sofa at the start of the slide in the upper-left of the frame, background; Destination cabin wall in the lower-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Sofa (Behind the pair as they slide away) — Viewed obliquely from above at the uphill side of the composition; used as Marks the origin of their displacement; Cabin floor (Sharply tilted with the ship) — Its visible plane slopes through the composition toward lower right; used as Makes the direction of involuntary movement legible; Cabin wall (Ahead of the sliding pair, before contact) — The interior face is visible at the downhill end of their trajectory; used as Defines the imminent collision and remaining clearance.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the cabin's existing ambient illumination and controlled contrast as the physical tilt disrupts the previously quiet composition.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The vessel and cabin are now sharply heeled to one side, with the cabin door still closed unless subsequently opened. Charlie's damaged body is thrown toward the wall; no removal of his blanket or travel backpack is established. 현우: He is thrown toward the cabin wall as the floor tilts, still bearing his earlier injuries and dirty clothing. The cabin is visibly tilted by the vessel's violent roll.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "두 인물 모두 기울어진 우측 방향을 향해 쏠려 있음.",
    "built_space": "좌측 상단에 소파, 우측에 벽이 위치하며 바닥이 우측으로 기울어짐.",
    "entities": "현우의 얼굴과 옷차림, 찰리의 로봇 형태 및 배낭/담요는 레퍼런스를 반영함.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 해부학적 구조: 현우의 왼쪽 다리와 흰 운동화가 골반/엉덩이에서 비정상적인 각도로 튀어나와 있음."
    ],
    "physics": "찰리는 두 발로 바닥을 딛고 서서 버티는 자세이며, 현우는 손과 무릎을 바닥과 벽에 대고 있으나 하체 구조가 무너져 지지 상태가 성립하지 않음."
   },
   {
    "label": "B",
    "direction": "두 인물의 몸과 움직임이 우측 하단의 벽을 향해 강하게 쏠려 있으며, 현우는 다가오는 벽을 응시함.",
    "built_space": "좌측 상단에 소파, 우측 하단에 벽이 배치되어 있으며 바닥이 심하게 우측 하단으로 기울어진 선원실 내부.",
    "entities": "현우는 오염된 옷을 입고 레퍼런스와 일치하는 외형임. 찰리는 고릴라형 로봇 형태를 띠며 등에는 배낭과 담요를 메고 있음.",
    "hard_violations": [],
    "physics": "현우는 한 손은 바닥, 한 손은 벽을 짚고 미끄러짐을 지탱 중. 찰리는 낮은 자세로 바닥에 손발을 딛고 미끄러지는 관성을 현실감 있게 버티고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "프롬프트가 요구한 구도와 역동적인 미끄러짐 액션을 정확하게 구현했으며, 두 캐릭터의 외형 및 소품도 레퍼런스와 잘 일치합니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "현우의 왼쪽 다리가 신체 구조상 불가능한 위치에 달려 있는 심각한 해부학적 오류가 있어 감점되었습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "두 인물의 몸과 움직임이 우측 하단의 벽을 향해 강하게 쏠려 있으며, 현우는 다가오는 벽을 응시함.",
        "built_space": "좌측 상단에 소파, 우측 하단에 벽이 배치되어 있으며 바닥이 심하게 우측 하단으로 기울어진 선원실 내부.",
        "entities": "현우는 오염된 옷을 입고 레퍼런스와 일치하는 외형임. 찰리는 고릴라형 로봇 형태를 띠며 등에는 배낭과 담요를 메고 있음.",
        "hard_violations": [],
        "physics": "현우는 한 손은 바닥, 한 손은 벽을 짚고 미끄러짐을 지탱 중. 찰리는 낮은 자세로 바닥에 손발을 딛고 미끄러지는 관성을 현실감 있게 버티고 있음."
       },
       {
        "label": "A",
        "direction": "두 인물 모두 기울어진 우측 방향을 향해 쏠려 있음.",
        "built_space": "좌측 상단에 소파, 우측에 벽이 위치하며 바닥이 우측으로 기울어짐.",
        "entities": "현우의 얼굴과 옷차림, 찰리의 로봇 형태 및 배낭/담요는 레퍼런스를 반영함.",
        "hard_violations": [
         "물리적으로 불가능한 해부학적 구조: 현우의 왼쪽 다리와 흰 운동화가 골반/엉덩이에서 비정상적인 각도로 튀어나와 있음."
        ],
        "physics": "찰리는 두 발로 바닥을 딛고 서서 버티는 자세이며, 현우는 손과 무릎을 바닥과 벽에 대고 있으나 하체 구조가 무너져 지지 상태가 성립하지 않음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "프롬프트가 요구한 구도와 역동적인 미끄러짐 액션을 정확하게 구현했으며, 두 캐릭터의 외형 및 소품도 레퍼런스와 잘 일치합니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "현우의 왼쪽 다리가 신체 구조상 불가능한 위치에 달려 있는 심각한 해부학적 오류가 있어 감점되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "두 인물의 몸과 움직임이 우측 하단의 벽을 향해 강하게 쏠려 있으며, 현우는 다가오는 벽을 응시함.",
        "built_space": "좌측 상단에 소파, 우측 하단에 벽이 배치되어 있으며 바닥이 심하게 우측 하단으로 기울어진 선원실 내부.",
        "entities": "현우는 오염된 옷을 입고 레퍼런스와 일치하는 외형임. 찰리는 고릴라형 로봇 형태를 띠며 등에는 배낭과 담요를 메고 있음.",
        "hard_violations": [],
        "physics": "현우는 한 손은 바닥, 한 손은 벽을 짚고 미끄러짐을 지탱 중. 찰리는 낮은 자세로 바닥에 손발을 딛고 미끄러지는 관성을 현실감 있게 버티고 있음."
       },
       {
        "label": "A",
        "direction": "두 인물 모두 기울어진 우측 방향을 향해 쏠려 있음.",
        "built_space": "좌측 상단에 소파, 우측에 벽이 위치하며 바닥이 우측으로 기울어짐.",
        "entities": "현우의 얼굴과 옷차림, 찰리의 로봇 형태 및 배낭/담요는 레퍼런스를 반영함.",
        "hard_violations": [
         "물리적으로 불가능한 해부학적 구조: 현우의 왼쪽 다리와 흰 운동화가 골반/엉덩이에서 비정상적인 각도로 튀어나와 있음."
        ],
        "physics": "찰리는 두 발로 바닥을 딛고 서서 버티는 자세이며, 현우는 손과 무릎을 바닥과 벽에 대고 있으나 하체 구조가 무너져 지지 상태가 성립하지 않음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "두 인물의 전신과 오른쪽 아래로 무너지는 찰리의 자세가 미끄러지는 순간에 더 충실하며, 현우의 시선과 벽까지의 짧은 여유는 아쉽다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "공간의 하강 방향은 맞지만 현우가 이미 벽을 양손으로 짚었고 찰리는 버티며 걷는 모습에 가까워, 충돌 전 두 몸이 거칠게 미끄러지는 순간과 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "소파는 왼쪽 위에 있고 바닥의 선들은 오른쪽 아래 벽 쪽으로 내려간다. 찰리는 상체와 머리가 오른쪽 아래로 기울고 다리는 소파 쪽에 뒤처져, 벽으로 쏠리는 방향이 읽힌다. 현우도 같은 벽 쪽으로 내려와 오른손을 내밀지만, 얼굴과 시선은 목적지 벽보다 왼쪽의 찰리 쪽을 향한다. 무기나 조준하는 물체는 없다.",
        "built_space": "회색 소파 한 조가 왼쪽, 여러 칸의 목재 선반이 그 뒤, 커튼 달린 창 하나가 후면에 보인다. 오른쪽에는 닫힌 문과 밝은 색 벽면이 있으며, 켜진 벽등과 천장등이 보인다. 회색 직물, 낡은 바닥, 선반과 실내 조명은 이전 장면의 공간과 대체로 이어진다. 두 인물은 소파를 벗어나 바닥에 있고 목적지 벽은 오른쪽에 노출되어 있다. 다만 소파의 가까운 팔걸이가 프레임 왼쪽을 크게 차지하고, 현우와 벽 사이의 남은 공간은 매우 좁다.",
        "entities": "현우는 앳된 동아시아계 남성으로, 헝클어진 검은 머리와 더러운 회갈색 반팔이 이전 장면과 부합한다. 얼굴과 팔의 상처·오염도 보인다. 한국계 미국인이라는 국적은 외관만으로 확인할 수 없다. 찰리는 긴 육중한 팔, 상대적으로 짧은 다리, 샌드 베이지 장갑판, 흰 기계식 마스크 얼굴을 갖춰 캐릭터 참조와 부합한다. 등에는 여행 배낭이 있고 회색 담요도 남아 있다. 등장 인물은 둘뿐이다. 바닥에는 물병과 종이류가 추가로 보이지만 읽을 수 있는 문구는 없다.",
        "hard_violations": [],
        "physics": "현우는 접힌 다리와 바닥에 댄 왼손으로 몸을 지탱하면서 오른손을 벽 쪽으로 뻗는다. 찰리는 낮아진 발끝과 발 가장자리가 바닥에 닿아 있고, 상체가 그 지지점보다 오른쪽으로 크게 넘어가 균형을 잃는 순간으로 읽힌다. 두 몸이 근거 없이 공중에 떠 있는 모습은 아니다. 배낭은 찰리의 등에 붙어 있고 담요는 어깨와 배낭에 걸려 지지된다. 정지한 자세보다는 기울어진 바닥에서 미끄러지며 버티려는 동작에 가깝다."
       },
       {
        "label": "B",
        "direction": "소파에서 오른쪽 아래 벽으로 이어지는 하강 방향은 분명하다. 현우의 얼굴과 양팔은 오른쪽 벽을 향하고 찰리의 얼굴도 현우와 그 앞쪽 바닥을 향한다. 다만 현우의 두 손이 이미 벽면에 도달해, 벽을 향해 이동하는 충돌 직전보다는 도착 후 제동하는 순간으로 보인다. 찰리는 상체를 비교적 세운 채 그 뒤를 따라오는 방향이다.",
        "built_space": "왼쪽 위에 회색 소파 한 조와 뒤쪽 목재 선반, 후면에 커튼 달린 창 하나와 벽등이 보인다. 오른쪽에는 닫힌 밝은 색 문 및 가까운 금속 문면이 있고 긴 손잡이 두 개가 보인다. 기존 선원실의 재료와 조명은 대체로 유지되지만 가까운 금속 문과 손잡이가 이전 장면보다 강하게 부각된다. 두 인물의 전신은 포함되며 소파도 뒤에 있으나, 현우가 목적지 벽을 이미 짚어 요구된 충돌 여유가 사라졌다.",
        "entities": "현우의 젊은 동아시아계 남성 외모, 검은 머리, 오염된 회갈색 반팔과 어두운 바지는 참조와 대체로 일치한다. 얼굴은 숙여져 있어 세부 동일성 판단이 A보다 어렵다. 찰리는 베이지색 장갑, 흰 마스크 얼굴, 긴 팔과 짧은 다리를 유지하며 회색 담요와 등에 멘 배낭도 보인다. 둘 외의 인물은 없다. 왼쪽 앞 바닥에는 별도의 갈색 가방이 보이며, 이는 지정된 찰리의 배낭과는 다른 부가 소품이다. 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "현우의 접힌 다리와 무릎이 바닥에 닿고 양손은 벽을 짚어 몸을 지지한다. 따라서 자세 자체는 가능하지만 이미 벽에서 충격을 받아 버티는 동작이다. 찰리는 벌린 발로 바닥을 지탱하고 무릎을 굽힌 채 상체를 상당히 세우고 있어, 손상된 몸이 강하게 쓸려 내려가는 것보다는 보행하거나 균형을 잡는 모습이다. 배낭과 담요는 등과 어깨에 지지되며, 바닥의 별도 가방도 바닥에 놓여 있다. 지지 없는 부유는 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "두 인물의 전신과 오른쪽 아래로 무너지는 찰리의 자세가 미끄러지는 순간에 더 충실하며, 현우의 시선과 벽까지의 짧은 여유는 아쉽다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "공간의 하강 방향은 맞지만 현우가 이미 벽을 양손으로 짚었고 찰리는 버티며 걷는 모습에 가까워, 충돌 전 두 몸이 거칠게 미끄러지는 순간과 다르다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "소파는 왼쪽 위에 있고 바닥의 선들은 오른쪽 아래 벽 쪽으로 내려간다. 찰리는 상체와 머리가 오른쪽 아래로 기울고 다리는 소파 쪽에 뒤처져, 벽으로 쏠리는 방향이 읽힌다. 현우도 같은 벽 쪽으로 내려와 오른손을 내밀지만, 얼굴과 시선은 목적지 벽보다 왼쪽의 찰리 쪽을 향한다. 무기나 조준하는 물체는 없다.",
        "built_space": "회색 소파 한 조가 왼쪽, 여러 칸의 목재 선반이 그 뒤, 커튼 달린 창 하나가 후면에 보인다. 오른쪽에는 닫힌 문과 밝은 색 벽면이 있으며, 켜진 벽등과 천장등이 보인다. 회색 직물, 낡은 바닥, 선반과 실내 조명은 이전 장면의 공간과 대체로 이어진다. 두 인물은 소파를 벗어나 바닥에 있고 목적지 벽은 오른쪽에 노출되어 있다. 다만 소파의 가까운 팔걸이가 프레임 왼쪽을 크게 차지하고, 현우와 벽 사이의 남은 공간은 매우 좁다.",
        "entities": "현우는 앳된 동아시아계 남성으로, 헝클어진 검은 머리와 더러운 회갈색 반팔이 이전 장면과 부합한다. 얼굴과 팔의 상처·오염도 보인다. 한국계 미국인이라는 국적은 외관만으로 확인할 수 없다. 찰리는 긴 육중한 팔, 상대적으로 짧은 다리, 샌드 베이지 장갑판, 흰 기계식 마스크 얼굴을 갖춰 캐릭터 참조와 부합한다. 등에는 여행 배낭이 있고 회색 담요도 남아 있다. 등장 인물은 둘뿐이다. 바닥에는 물병과 종이류가 추가로 보이지만 읽을 수 있는 문구는 없다.",
        "hard_violations": [],
        "physics": "현우는 접힌 다리와 바닥에 댄 왼손으로 몸을 지탱하면서 오른손을 벽 쪽으로 뻗는다. 찰리는 낮아진 발끝과 발 가장자리가 바닥에 닿아 있고, 상체가 그 지지점보다 오른쪽으로 크게 넘어가 균형을 잃는 순간으로 읽힌다. 두 몸이 근거 없이 공중에 떠 있는 모습은 아니다. 배낭은 찰리의 등에 붙어 있고 담요는 어깨와 배낭에 걸려 지지된다. 정지한 자세보다는 기울어진 바닥에서 미끄러지며 버티려는 동작에 가깝다."
       },
       {
        "label": "A",
        "direction": "소파에서 오른쪽 아래 벽으로 이어지는 하강 방향은 분명하다. 현우의 얼굴과 양팔은 오른쪽 벽을 향하고 찰리의 얼굴도 현우와 그 앞쪽 바닥을 향한다. 다만 현우의 두 손이 이미 벽면에 도달해, 벽을 향해 이동하는 충돌 직전보다는 도착 후 제동하는 순간으로 보인다. 찰리는 상체를 비교적 세운 채 그 뒤를 따라오는 방향이다.",
        "built_space": "왼쪽 위에 회색 소파 한 조와 뒤쪽 목재 선반, 후면에 커튼 달린 창 하나와 벽등이 보인다. 오른쪽에는 닫힌 밝은 색 문 및 가까운 금속 문면이 있고 긴 손잡이 두 개가 보인다. 기존 선원실의 재료와 조명은 대체로 유지되지만 가까운 금속 문과 손잡이가 이전 장면보다 강하게 부각된다. 두 인물의 전신은 포함되며 소파도 뒤에 있으나, 현우가 목적지 벽을 이미 짚어 요구된 충돌 여유가 사라졌다.",
        "entities": "현우의 젊은 동아시아계 남성 외모, 검은 머리, 오염된 회갈색 반팔과 어두운 바지는 참조와 대체로 일치한다. 얼굴은 숙여져 있어 세부 동일성 판단이 A보다 어렵다. 찰리는 베이지색 장갑, 흰 마스크 얼굴, 긴 팔과 짧은 다리를 유지하며 회색 담요와 등에 멘 배낭도 보인다. 둘 외의 인물은 없다. 왼쪽 앞 바닥에는 별도의 갈색 가방이 보이며, 이는 지정된 찰리의 배낭과는 다른 부가 소품이다. 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "현우의 접힌 다리와 무릎이 바닥에 닿고 양손은 벽을 짚어 몸을 지지한다. 따라서 자세 자체는 가능하지만 이미 벽에서 충격을 받아 버티는 동작이다. 찰리는 벌린 발로 바닥을 지탱하고 무릎을 굽힌 채 상체를 상당히 세우고 있어, 손상된 몸이 강하게 쓸려 내려가는 것보다는 보행하거나 균형을 잡는 모습이다. 배낭과 담요는 등과 어깨에 지지되며, 바닥의 별도 가방도 바닥에 놓여 있다. 지지 없는 부유는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.0,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.75,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 물리적으로 불가능한 해부학적 구조: 현우의 왼쪽 다리와 흰 운동화가 골반/엉덩이에서 비정상적인 각도로 튀어나와 있음."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 750
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "프롬프트가 요구한 구도와 역동적인 미끄러짐 액션을 정확하게 구현했으며, 두 캐릭터의 외형 및 소품도 레퍼런스와 잘 일치합니다."
   },
   {
    "label": "A",
    "score": 750,
    "verdict_ko": "현우의 왼쪽 다리가 신체 구조상 불가능한 위치에 달려 있는 심각한 해부학적 오류가 있어 감점되었습니다.  ★위반: [gemini-pro] 물리적으로 불가능한 해부학적 구조: 현우의 왼쪽 다리와 흰 운동화가 골반/엉덩이에서 비정상적인 각도로 튀어나와 있음."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S79sh12_sel.png",
    "asset_id": "57ee8966-e86a-4ae6-af1e-bf79a9f36c10",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-d722-7984-826d-d193cad7ed59",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S79sh12"
  }
 },
 "S79sh14::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T12:05:12.012243+00:00",
  "fingerprint": "1734c6457628c83b06d9f1047710830d47a21ceee8516f956ddd3219803f2e97",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S79sh14_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S79sh14_sel.png",
  "source_sha256": "805020cec4b7b51e12fef54092327fe5aa6d06a54146986c3b23eedca9d53bd0",
  "file": "S79sh14_cine.png",
  "staged_sha256": "150ea145f36705601a98a61754d10809eaa8e5a0d4d88580223d5e5d448ea047",
  "latency_ms": 10498
 },
 "S80sh9::signage": {
  "fp": "6fb1eaa78cf58749",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::a955c71a887aa880": {
  "subjects": [],
  "subject_text": "크리스의 선박 갑판\n바다에 노출된 작은 선박의 낡은 갑판 공간. 바닥 위로 지지 기둥과 짐이 배치되고 실내로 연결되는 출입구가 있다.",
  "identity": "canonical",
  "scope_id": "L261",
  "scope_role": "location_exterior",
  "scope_sha": "2edff8e6cc4e90f6"
 },
 "S80sh9::bgfirst_bg": {
  "input_fingerprint": "42e829411c948878",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 허공으로 떨어지는 현우의 손목을 낚아채듯 꽉 움켜쥔 찰리의 다른 금속 손 클로즈업.\n\nLOCATION (lock): At the exposed edge of the boat's storm-lashed deck, beside the upright support gripped by the robot.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Ship's edge (Beside 현우's interrupted fall during the storm) — Seen diagonally from the deck side beneath the extending arms; used as Separates the supported deck side from the space of the fall; Supporting post (Serving as 찰리's anchor) — A small oblique portion is visible toward upper left behind his forearm; used as Explains the direction of resistance without competing with the grip.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Nighttime storm ambience and torrential rain reduce visibility while restrained local contrast keeps the human wrist and metal grip distinguishable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 허공으로 떨어지는 현우의 손목을 낚아채듯 꽉 움켜쥔 찰리의 다른 금속 손 클로즈업.\n\nLOCATION (lock): At the exposed edge of the boat's storm-lashed deck, beside the upright support gripped by the robot.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Ship's edge (Beside 현우's interrupted fall during the storm) — Seen diagonally from the deck side beneath the extending arms; used as Separates the supported deck side from the space of the fall; Supporting post (Serving as 찰리's anchor) — A small oblique portion is visible toward upper left behind his forearm; used as Explains the direction of resistance without competing with the grip.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Nighttime storm ambience and torrential rain reduce visibility while restrained local contrast keeps the human wrist and metal grip distinguishable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S80sh9__bgfirst_bg.png",
  "asset_id": "308f16b6-d6a7-43a8-94fa-ab3e5aa531be",
  "input_asset_ids": [
   "7e2e9a21-fd2a-469c-b057-4ecfe185964a",
   "a7ba2d1e-a8a6-411b-a6de-e6dbb6a4b132"
  ]
 },
 "S80sh9": {
  "input_fingerprint": "e5f740b5704328e9",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 허공으로 떨어지는 현우의 손목을 낚아채듯 꽉 움켜쥔 찰리의 다른 금속 손 클로즈업.\n\nLOCATION (lock): At the exposed edge of the boat's storm-lashed deck, beside the upright support gripped by the robot. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Ship's edge (Beside 현우's interrupted fall during the storm) — Seen diagonally from the deck side beneath the extending arms; used as Separates the supported deck side from the space of the fall; Supporting post (Serving as 찰리's anchor) — A small oblique portion is visible toward upper left behind his forearm; used as Explains the direction of resistance without competing with the grip.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Nighttime storm ambience and torrential rain reduce visibility while restrained local contrast keeps the human wrist and metal grip distinguishable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Torrential rain, typhoon winds and huge waves batter the ship at night. Charlie's already damaged body is rain-soaked, with one hand braced around a shipboard post and the other extended over the edge. 현우: He is soaked and suspended at the ship's edge with one arm stretched upward, his earlier injuries still present.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 허공으로 떨어지는 현우의 손목을 낚아채듯 꽉 움켜쥔 찰리의 다른 금속 손 클로즈업.\n\nLOCATION (lock): At the exposed edge of the boat's storm-lashed deck, beside the upright support gripped by the robot. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Ship's edge (Beside 현우's interrupted fall during the storm) — Seen diagonally from the deck side beneath the extending arms; used as Separates the supported deck side from the space of the fall; Supporting post (Serving as 찰리's anchor) — A small oblique portion is visible toward upper left behind his forearm; used as Explains the direction of resistance without competing with the grip.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Nighttime storm ambience and torrential rain reduce visibility while restrained local contrast keeps the human wrist and metal grip distinguishable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Torrential rain, typhoon winds and huge waves batter the ship at night. Charlie's already damaged body is rain-soaked, with one hand braced around a shipboard post and the other extended over the edge. 현우: He is soaked and suspended at the ship's edge with one arm stretched upward, his earlier injuries still present.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 허공으로 떨어지는 현우의 손목을 낚아채듯 꽉 움켜쥔 찰리의 다른 금속 손 클로즈업.\n\nLOCATION (lock): At the exposed edge of the boat's storm-lashed deck, beside the upright support gripped by the robot. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Ship's edge (Beside 현우's interrupted fall during the storm) — Seen diagonally from the deck side beneath the extending arms; used as Separates the supported deck side from the space of the fall; Supporting post (Serving as 찰리's anchor) — A small oblique portion is visible toward upper left behind his forearm; used as Explains the direction of resistance without competing with the grip.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Nighttime storm ambience and torrential rain reduce visibility while restrained local contrast keeps the human wrist and metal grip distinguishable.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Torrential rain, typhoon winds and huge waves batter the ship at night. Charlie's already damaged body is rain-soaked, with one hand braced around a shipboard post and the other extended over the edge. 현우: He is soaked and suspended at the ship's edge with one arm stretched upward, his earlier injuries still present.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S80sh9__bgfirst_bg.png",
     "asset_id": "308f16b6-d6a7-43a8-94fa-ab3e5aa531be",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S80sh9.png",
     "asset_id": "7e2e9a21-fd2a-469c-b057-4ecfe185964a",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L261B01.png",
     "asset_id": "a7ba2d1e-a8a6-411b-a6de-e6dbb6a4b132",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "로봇 팔은 우측 하단으로, 사람 팔은 좌측 상단으로 향해 서로 교차함.",
    "built_space": "갑판 위의 난간과 거친 바다, 그리고 좌측에 전체적인 형태가 드러난 하얀색 지지 기둥 1개가 있음.",
    "entities": "샌드 베이지색의 로봇 팔과 비에 젖은 사람 팔이 존재함.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 인체 해부학적 구조 및 맞물림 (로봇 손가락이 손목을 감싸는 고리 형태로 변형되고 사람의 손과 융합됨)"
    ],
    "physics": "두 손이 서로 엉켜 있으나 파지법이 물리적으로 성립하지 않는 기괴한 구조로 인해 지지 상태를 명확히 설명할 수 없음."
   },
   {
    "label": "B",
    "direction": "로봇의 팔이 우측으로 뻗어 사람의 손목을 향하고 있으며, 사람의 손은 허공에서 위를 향해 뻗어 있음.",
    "built_space": "배의 가장자리 난간과 거친 바다가 보이며, 좌측 상단 로봇 팔 뒤편으로 하얀색 지지 기둥의 일부가 정확한 위치에 묘사됨.",
    "entities": "찰리의 샌드 베이지색 금속 팔과 현우의 비에 젖은 맨팔이 정확히 묘사됨.",
    "hard_violations": [],
    "physics": "로봇의 금속 손이 사람의 손목을 단단히 쥐어 허공으로 떨어지려는 몸을 지탱하고 있으며, 빗방울과 그립의 상호작용이 자연스러움."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "요구된 클로즈업 앵글과 손목을 움켜쥔 동작을 사실적으로 구현했으며, 지정된 배경 요소의 배치도 지시문과 정확히 일치합니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "금속 손과 사람의 손이 기괴하게 융합되어 물리적으로 불가능한 구조를 이루는 치명적인 형태 오류가 있습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "로봇의 팔이 우측으로 뻗어 사람의 손목을 향하고 있으며, 사람의 손은 허공에서 위를 향해 뻗어 있음.",
        "built_space": "배의 가장자리 난간과 거친 바다가 보이며, 좌측 상단 로봇 팔 뒤편으로 하얀색 지지 기둥의 일부가 정확한 위치에 묘사됨.",
        "entities": "찰리의 샌드 베이지색 금속 팔과 현우의 비에 젖은 맨팔이 정확히 묘사됨.",
        "hard_violations": [],
        "physics": "로봇의 금속 손이 사람의 손목을 단단히 쥐어 허공으로 떨어지려는 몸을 지탱하고 있으며, 빗방울과 그립의 상호작용이 자연스러움."
       },
       {
        "label": "A",
        "direction": "로봇 팔은 우측 하단으로, 사람 팔은 좌측 상단으로 향해 서로 교차함.",
        "built_space": "갑판 위의 난간과 거친 바다, 그리고 좌측에 전체적인 형태가 드러난 하얀색 지지 기둥 1개가 있음.",
        "entities": "샌드 베이지색의 로봇 팔과 비에 젖은 사람 팔이 존재함.",
        "hard_violations": [
         "물리적으로 불가능한 인체 해부학적 구조 및 맞물림 (로봇 손가락이 손목을 감싸는 고리 형태로 변형되고 사람의 손과 융합됨)"
        ],
        "physics": "두 손이 서로 엉켜 있으나 파지법이 물리적으로 성립하지 않는 기괴한 구조로 인해 지지 상태를 명확히 설명할 수 없음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "요구된 클로즈업 앵글과 손목을 움켜쥔 동작을 사실적으로 구현했으며, 지정된 배경 요소의 배치도 지시문과 정확히 일치합니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "금속 손과 사람의 손이 기괴하게 융합되어 물리적으로 불가능한 구조를 이루는 치명적인 형태 오류가 있습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "로봇의 팔이 우측으로 뻗어 사람의 손목을 향하고 있으며, 사람의 손은 허공에서 위를 향해 뻗어 있음.",
        "built_space": "배의 가장자리 난간과 거친 바다가 보이며, 좌측 상단 로봇 팔 뒤편으로 하얀색 지지 기둥의 일부가 정확한 위치에 묘사됨.",
        "entities": "찰리의 샌드 베이지색 금속 팔과 현우의 비에 젖은 맨팔이 정확히 묘사됨.",
        "hard_violations": [],
        "physics": "로봇의 금속 손이 사람의 손목을 단단히 쥐어 허공으로 떨어지려는 몸을 지탱하고 있으며, 빗방울과 그립의 상호작용이 자연스러움."
       },
       {
        "label": "A",
        "direction": "로봇 팔은 우측 하단으로, 사람 팔은 좌측 상단으로 향해 서로 교차함.",
        "built_space": "갑판 위의 난간과 거친 바다, 그리고 좌측에 전체적인 형태가 드러난 하얀색 지지 기둥 1개가 있음.",
        "entities": "샌드 베이지색의 로봇 팔과 비에 젖은 사람 팔이 존재함.",
        "hard_violations": [
         "물리적으로 불가능한 인체 해부학적 구조 및 맞물림 (로봇 손가락이 손목을 감싸는 고리 형태로 변형되고 사람의 손과 융합됨)"
        ],
        "physics": "두 손이 서로 엉켜 있으나 파지법이 물리적으로 성립하지 않는 기괴한 구조로 인해 지지 상태를 명확히 설명할 수 없음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "금속 손이 현우의 손목을 꽉 붙든 순간을 크게 잡아 핵심 클로즈업에 더 충실하지만, 배경 지지대는 요구한 작은 일부보다 많이 드러난다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "손목 지지와 선박 가장자리의 관계는 명확하지만, 갑판과 파도까지 넓게 보여 손의 클로즈업이라는 최우선 구도 지시에서 멀어진다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 팔은 화면 왼쪽에서 오른쪽으로 뻗어 현우의 손목을 붙잡는다. 현우의 팔은 아래에서 위로 올라오며, 펼쳐진 손가락은 오른쪽 위를 향한다. 붙잡는 대상은 손바닥이 아니라 손목이다. 얼굴과 시선은 프레임 밖이다.",
        "built_space": "녹슨 흰색 지지대 하나와 그에 매달린 줄이 팔 뒤에 보인다. 선측 벽과 상단 난간은 팔 아래로 대각선을 이루며 왼쪽 갑판과 오른쪽 바다를 구분한다. 장소의 재질과 배치는 참조에 부합하지만, 지지대는 왼쪽 위의 작은 조각이 아니라 화면 높이 대부분에 걸쳐 보인다.",
        "entities": "보이는 인물 부분은 찰리의 베이지색 금속 팔과 손, 현우의 맨팔과 손이다. 찰리의 각진 장갑판과 검은 관절, 마모 흔적은 참조와 잘 맞는다. 현우의 손은 젊은 사람의 손으로 읽히지만 손만으로 정확한 나이·성별·민족적 정체성을 확인할 수는 없다. 얼굴과 의상은 잘려 있어 평가 대상이 아니다. 팔에 뚜렷한 기존 부상은 식별되지 않는다. 비와 어두운 바다가 보이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "금속 손가락과 엄지가 현우의 손목 양쪽을 감싸 직접 지지한다. 아래로 이어지는 현우의 팔과 왼쪽으로 이어지는 찰리의 팔이 서로 반대 방향의 하중을 전달하는 형태라 추락을 멈추는 동작으로 성립한다. 찰리의 다른 손과 지지대의 접촉은 프레임 밖이므로 확인할 수 없지만, 보이는 손목이 아무 지지 없이 떠 있지는 않다."
       },
       {
        "label": "B",
        "direction": "찰리의 팔은 왼쪽 위에서 오른쪽 아래로 뻗어 현우의 손목을 감싼다. 현우의 팔은 오른쪽 아래에서 위로 뻗고 손은 찰리 쪽으로 굽혀져 있어 서로 매달려 붙드는 동작으로 읽힌다. 얼굴과 시선은 보이지 않는다.",
        "built_space": "왼쪽에 굵은 흰색 지지대 하나와 감긴 줄이 있고, 선측 벽과 상단 난간이 뒤쪽 중앙에서 오른쪽 아래로 이어진다. 왼쪽 갑판에는 그물 덮인 적재물들이, 먼 배경에는 문 하나와 사다리 하나 및 점등된 조명이 보인다. 참조 장소의 요소들은 잘 맞지만, 지지대의 밑동과 넓은 갑판까지 드러나 요구한 작은 배경 조각 및 손 중심 클로즈업보다 훨씬 넓다.",
        "entities": "찰리의 베이지색 장갑 팔과 관절식 손, 현우의 젖은 맨팔과 손 및 짙은 소매 끝이 보인다. 금속 재질과 색은 찰리 참조에 부합하나 장갑판 형태는 다소 다르다. 현우의 팔에는 긁힌 듯한 붉은 자국이 있다. 얼굴이 없으므로 정확한 나이와 정체성은 확인할 수 없다. 폭우와 큰 파도, 야간 조명은 요구와 맞으며 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "찰리의 굽힌 금속 손가락이 현우의 손목을 둘러싸고, 현우의 손도 금속 손 쪽에 접촉한다. 현우의 팔은 선측 바깥 아래쪽으로 이어져 손목을 통해 체중을 지지받는 관계가 성립한다. 찰리의 팔은 왼쪽 위 프레임 밖으로 이어지며, 지지대를 잡은 다른 손은 보이지 않는다. 보이는 범위에 지지 없는 부유나 불가능한 관절 꺾임은 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "금속 손이 현우의 손목을 꽉 붙든 순간을 크게 잡아 핵심 클로즈업에 더 충실하지만, 배경 지지대는 요구한 작은 일부보다 많이 드러난다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "손목 지지와 선박 가장자리의 관계는 명확하지만, 갑판과 파도까지 넓게 보여 손의 클로즈업이라는 최우선 구도 지시에서 멀어진다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 팔은 화면 왼쪽에서 오른쪽으로 뻗어 현우의 손목을 붙잡는다. 현우의 팔은 아래에서 위로 올라오며, 펼쳐진 손가락은 오른쪽 위를 향한다. 붙잡는 대상은 손바닥이 아니라 손목이다. 얼굴과 시선은 프레임 밖이다.",
        "built_space": "녹슨 흰색 지지대 하나와 그에 매달린 줄이 팔 뒤에 보인다. 선측 벽과 상단 난간은 팔 아래로 대각선을 이루며 왼쪽 갑판과 오른쪽 바다를 구분한다. 장소의 재질과 배치는 참조에 부합하지만, 지지대는 왼쪽 위의 작은 조각이 아니라 화면 높이 대부분에 걸쳐 보인다.",
        "entities": "보이는 인물 부분은 찰리의 베이지색 금속 팔과 손, 현우의 맨팔과 손이다. 찰리의 각진 장갑판과 검은 관절, 마모 흔적은 참조와 잘 맞는다. 현우의 손은 젊은 사람의 손으로 읽히지만 손만으로 정확한 나이·성별·민족적 정체성을 확인할 수는 없다. 얼굴과 의상은 잘려 있어 평가 대상이 아니다. 팔에 뚜렷한 기존 부상은 식별되지 않는다. 비와 어두운 바다가 보이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "금속 손가락과 엄지가 현우의 손목 양쪽을 감싸 직접 지지한다. 아래로 이어지는 현우의 팔과 왼쪽으로 이어지는 찰리의 팔이 서로 반대 방향의 하중을 전달하는 형태라 추락을 멈추는 동작으로 성립한다. 찰리의 다른 손과 지지대의 접촉은 프레임 밖이므로 확인할 수 없지만, 보이는 손목이 아무 지지 없이 떠 있지는 않다."
       },
       {
        "label": "A",
        "direction": "찰리의 팔은 왼쪽 위에서 오른쪽 아래로 뻗어 현우의 손목을 감싼다. 현우의 팔은 오른쪽 아래에서 위로 뻗고 손은 찰리 쪽으로 굽혀져 있어 서로 매달려 붙드는 동작으로 읽힌다. 얼굴과 시선은 보이지 않는다.",
        "built_space": "왼쪽에 굵은 흰색 지지대 하나와 감긴 줄이 있고, 선측 벽과 상단 난간이 뒤쪽 중앙에서 오른쪽 아래로 이어진다. 왼쪽 갑판에는 그물 덮인 적재물들이, 먼 배경에는 문 하나와 사다리 하나 및 점등된 조명이 보인다. 참조 장소의 요소들은 잘 맞지만, 지지대의 밑동과 넓은 갑판까지 드러나 요구한 작은 배경 조각 및 손 중심 클로즈업보다 훨씬 넓다.",
        "entities": "찰리의 베이지색 장갑 팔과 관절식 손, 현우의 젖은 맨팔과 손 및 짙은 소매 끝이 보인다. 금속 재질과 색은 찰리 참조에 부합하나 장갑판 형태는 다소 다르다. 현우의 팔에는 긁힌 듯한 붉은 자국이 있다. 얼굴이 없으므로 정확한 나이와 정체성은 확인할 수 없다. 폭우와 큰 파도, 야간 조명은 요구와 맞으며 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "찰리의 굽힌 금속 손가락이 현우의 손목을 둘러싸고, 현우의 손도 금속 손 쪽에 접촉한다. 현우의 팔은 선측 바깥 아래쪽으로 이어져 손목을 통해 체중을 지지받는 관계가 성립한다. 찰리의 팔은 왼쪽 위 프레임 밖으로 이어지며, 지지대를 잡은 다른 손은 보이지 않는다. 보이는 범위에 지지 없는 부유나 불가능한 관절 꺾임은 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.179,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.929,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 물리적으로 불가능한 인체 해부학적 구조 및 맞물림 (로봇 손가락이 손목을 감싸는 고리 형태로 변형되고 사람의 손과 융합됨)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 929
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "요구된 클로즈업 앵글과 손목을 움켜쥔 동작을 사실적으로 구현했으며, 지정된 배경 요소의 배치도 지시문과 정확히 일치합니다."
   },
   {
    "label": "A",
    "score": 929,
    "verdict_ko": "금속 손과 사람의 손이 기괴하게 융합되어 물리적으로 불가능한 구조를 이루는 치명적인 형태 오류가 있습니다.  ★위반: [gemini-pro] 물리적으로 불가능한 인체 해부학적 구조 및 맞물림 (로봇 손가락이 손목을 감싸는 고리 형태로 변형되고 사람의 손과 융합됨)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L261B01.png",
    "asset_id": "a7ba2d1e-a8a6-411b-a6de-e6dbb6a4b132",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-d8df-704d-9f83-4c10cf784a3c",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S80sh9__bgfirst_bg.png",
   "bg_asset_id": "308f16b6-d6a7-43a8-94fa-ab3e5aa531be",
   "bg_record_key": "S80sh9::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S80sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:57:32.030105+00:00",
  "fingerprint": "1def057aadcacecd9c4332d62d2a276e4f79486d963e2d23babc4519590ca0b2",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S80sh9_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S80sh9_sel.png",
  "source_sha256": "5336bc40f5cf7b166abfeb700ba041fae06af379b11d30792379f07a18119b59",
  "file": "S80sh9_cine.png",
  "staged_sha256": "e693e80fe61178d17dcb9bb560e4fde4e30fb50fd290899085d9404eb68e2d5c",
  "latency_ms": 11638
 },
 "S80sh15::signage": {
  "fp": "4420b6e6498fab5b",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S80sh15": {
  "input_fingerprint": "a2dfa970ba8d101a",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 파도의 엄청난 힘에 의해 찰리의 금속 손아귀에서 거칠게 뜯겨져 나간 현우의 손 클로즈업.\n\nLOCATION (lock): At the open deck edge of the boat during the nighttime storm, where a breaking wave tears the two apart. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Ship's edge (Beside the fall as the wave breaks the grip) — Retains the same diagonal deck-side view as the earlier hand close-up; used as Provides an unchanged reference for the separation and impending downward tilt; Wave water (Striking the ship and pulling the pair apart); used as Partially interrupts the surrounding space without concealing the gap between the hands.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established nighttime storm illumination and rain-obscured contrast through the release, without adding a flash or new source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The ship is being struck by another immense wave amid the ongoing typhoon and torrential rain. Charlie's wet, damaged body remains at the ship's edge as the outstretched grip breaks. 현우: He is fully soaked and being swept away from the ship, no longer held by his extended hand.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 파도의 엄청난 힘에 의해 찰리의 금속 손아귀에서 거칠게 뜯겨져 나간 현우의 손 클로즈업.\n\nLOCATION (lock): At the open deck edge of the boat during the nighttime storm, where a breaking wave tears the two apart. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Ship's edge (Beside the fall as the wave breaks the grip) — Retains the same diagonal deck-side view as the earlier hand close-up; used as Provides an unchanged reference for the separation and impending downward tilt; Wave water (Striking the ship and pulling the pair apart); used as Partially interrupts the surrounding space without concealing the gap between the hands.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established nighttime storm illumination and rain-obscured contrast through the release, without adding a flash or new source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The ship is being struck by another immense wave amid the ongoing typhoon and torrential rain. Charlie's wet, damaged body remains at the ship's edge as the outstretched grip breaks. 현우: He is fully soaked and being swept away from the ship, no longer held by his extended hand.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 파도의 엄청난 힘에 의해 찰리의 금속 손아귀에서 거칠게 뜯겨져 나간 현우의 손 클로즈업.\n\nLOCATION (lock): At the open deck edge of the boat during the nighttime storm, where a breaking wave tears the two apart. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Ship's edge (Beside the fall as the wave breaks the grip) — Retains the same diagonal deck-side view as the earlier hand close-up; used as Provides an unchanged reference for the separation and impending downward tilt; Wave water (Striking the ship and pulling the pair apart); used as Partially interrupts the surrounding space without concealing the gap between the hands.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established nighttime storm illumination and rain-obscured contrast through the release, without adding a flash or new source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The ship is being struck by another immense wave amid the ongoing typhoon and torrential rain. Charlie's wet, damaged body remains at the ship's edge as the outstretched grip breaks. 현우: He is fully soaked and being swept away from the ship, no longer held by his extended hand.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "로봇 손이 현우의 손목을 강하게 쥐고 있음.",
    "built_space": "갑판 난간과 기둥이 참조 이미지와 동일하게 배치됨.",
    "entities": "찰리의 금속 팔과 현우의 팔이 지시대로 뚜렷하게 나타남.",
    "hard_violations": [],
    "physics": "현우의 팔은 로봇 손에 의해 물리적으로 지탱되며 화면 밖의 몸으로 이어짐."
   },
   {
    "label": "B",
    "direction": "로봇 손이 펴져 있고 현우의 손은 떨어져 파도 쪽을 향함.",
    "built_space": "갑판 난간과 기둥이 참조에 맞게 배치됨.",
    "entities": "찰리의 금속 팔은 명확하나, 현우의 손은 손목 아래 팔 부분이 소실됨.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 해부학 (신체와 절단된 손)",
     "[gemini-pro] 아무것도 지탱하지 않는 허공에 뜬 신체 부위"
    ],
    "physics": "현우의 손이 신체와의 연결이나 다른 지지대 없이 허공의 물보라 위에 둥둥 떠 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "두 손이 분리되는 순간을 요구한 지시와 달리 여전히 손을 꽉 잡고 있어 핵심 액션 구현에 실패함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "손이 떨어지는 액션은 반영했으나, 현우의 손이 팔과 단절되어 허공에 떠 있는 치명적 해부학 오류가 있음."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "로봇 손이 현우의 손목을 강하게 쥐고 있음.",
        "built_space": "갑판 난간과 기둥이 참조 이미지와 동일하게 배치됨.",
        "entities": "찰리의 금속 팔과 현우의 팔이 지시대로 뚜렷하게 나타남.",
        "hard_violations": [],
        "physics": "현우의 팔은 로봇 손에 의해 물리적으로 지탱되며 화면 밖의 몸으로 이어짐."
       },
       {
        "label": "B",
        "direction": "로봇 손이 펴져 있고 현우의 손은 떨어져 파도 쪽을 향함.",
        "built_space": "갑판 난간과 기둥이 참조에 맞게 배치됨.",
        "entities": "찰리의 금속 팔은 명확하나, 현우의 손은 손목 아래 팔 부분이 소실됨.",
        "hard_violations": [
         "물리적으로 불가능한 해부학 (신체와 절단된 손)",
         "아무것도 지탱하지 않는 허공에 뜬 신체 부위"
        ],
        "physics": "현우의 손이 신체와의 연결이나 다른 지지대 없이 허공의 물보라 위에 둥둥 떠 있음."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "두 손이 분리되는 순간을 요구한 지시와 달리 여전히 손을 꽉 잡고 있어 핵심 액션 구현에 실패함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "손이 떨어지는 액션은 반영했으나, 현우의 손이 팔과 단절되어 허공에 떠 있는 치명적 해부학 오류가 있음."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "로봇 손이 현우의 손목을 강하게 쥐고 있음.",
        "built_space": "갑판 난간과 기둥이 참조 이미지와 동일하게 배치됨.",
        "entities": "찰리의 금속 팔과 현우의 팔이 지시대로 뚜렷하게 나타남.",
        "hard_violations": [],
        "physics": "현우의 팔은 로봇 손에 의해 물리적으로 지탱되며 화면 밖의 몸으로 이어짐."
       },
       {
        "label": "B",
        "direction": "로봇 손이 펴져 있고 현우의 손은 떨어져 파도 쪽을 향함.",
        "built_space": "갑판 난간과 기둥이 참조에 맞게 배치됨.",
        "entities": "찰리의 금속 팔은 명확하나, 현우의 손은 손목 아래 팔 부분이 소실됨.",
        "hard_violations": [
         "물리적으로 불가능한 해부학 (신체와 절단된 손)",
         "아무것도 지탱하지 않는 허공에 뜬 신체 부위"
        ],
        "physics": "현우의 손이 신체와의 연결이나 다른 지지대 없이 허공의 물보라 위에 둥둥 떠 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "두 손 사이의 분명한 간격과 선체 바깥 파도로 휩쓸리는 현우의 손을 클로즈업하여, 손아귀가 풀린 정확한 순간을 구현했다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "손의 크기와 금속 질감은 충실하지만 찰리가 여전히 현우의 손목을 움켜쥐고 있어, 뜯겨 나간 손과 두 손 사이 간격이라는 핵심 지시를 충족하지 못한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 팔은 화면 왼쪽에서 오른쪽의 현우 손을 향하고, 벌어진 금속 손가락 앞에 명확한 빈 간격이 있다. 현우의 손가락은 오른쪽 위로 뻗고 팔은 난간 바깥 아래쪽 물속으로 이어져, 선체에서 멀어지는 상태로 읽힌다. 얼굴과 시선은 보이지 않는다.",
        "built_space": "녹슨 굵은 수직 기둥 하나, 그 옆에 매달린 줄, 대각선으로 이어지는 선측 벽과 상부 난간 한 줄 및 여러 짧은 지지대가 보인다. 왼쪽은 젖은 갑판, 오른쪽은 파도이며 이전 장면의 재질과 공간 관계가 유지된다. 찰리는 갑판 쪽, 현우의 손과 팔은 난간 바깥에 있다. 반사상이나 중복된 주요 설비는 없다.",
        "entities": "현우의 젖은 맨손 하나와 전완 일부, 찰리의 샌드 베이지 장갑 팔과 금속 손 하나가 보인다. 현우의 손은 젊은 인물의 손으로 자연스럽지만 얼굴·머리·민족적 정체성은 이 크롭에서 확인할 수 없다. 찰리의 마모된 각진 장갑과 검은 관절은 참조에 부합한다. 폭우, 부서지는 파도, 야간의 차가운 조명이 있으며 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "찰리의 손은 손목 관절과 왼쪽 프레임 밖으로 이어지는 팔에 연결되어 있다. 현우의 손도 전완에 연결되며 전완 아래쪽은 파도에 가려진다. 몸 전체가 보이지 않는 손 클로즈업이므로 이를 공중에 뜬 무지지 신체로 볼 근거는 없다. 선체 바깥을 덮치는 파도가 휩쓸림의 원인으로 보이고, 놓친 금속 손과 멀어진 맨손의 관계도 물리적으로 가능하다."
       },
       {
        "label": "B",
        "direction": "찰리의 팔은 왼쪽에서 오른쪽으로 뻗어 현우의 손목을 감싸고 있다. 현우의 손가락은 오른쪽 위를 향하지만 두 손 사이에 분리 간격이 없으며, 파도가 덮쳐도 아직 붙잡힌 순간으로 읽힌다. 얼굴이나 시선은 보이지 않는다.",
        "built_space": "녹슨 굵은 수직 기둥 하나와 매달린 줄, 대각선 선측 벽 및 상부 난간이 보이며 갑판과 바다의 좌우 관계는 참조와 같다. 왼쪽 먼 통로에는 밝은 조명 세 점과 통로 구조가 드러나 이전 장면보다 배경 광원이 두드러진다. 찰리는 갑판 쪽에서 선측 바깥의 현우 팔을 붙잡고 있다. 불가능한 반사나 주요 설비의 명백한 중복은 없다.",
        "entities": "현우의 젖은 맨손과 아래 프레임으로 이어지는 전완, 찰리의 베이지색 장갑 팔과 금속 손이 보인다. 피부와 기계 장갑의 외관은 이전 장면과 가깝다. 얼굴·머리·옷은 크롭 밖이라 신원 세부를 확인할 수 없다. 야간 폭우와 큰 파도는 구현되어 있고 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우의 팔은 아래 프레임 밖으로 연결되고 손목은 찰리의 금속 손가락에 실제로 붙잡혀 있다. 찰리의 손과 팔 역시 관절로 연결되어 있어 떠 있는 물체는 없다. 파도가 선측에 부딪혀 솟는 모습은 가능하지만, 손목을 감싼 접촉이 유지되어 파도의 힘으로 손이 뜯겨 나간 결과는 나타나지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "두 손 사이의 분명한 간격과 선체 바깥 파도로 휩쓸리는 현우의 손을 클로즈업하여, 손아귀가 풀린 정확한 순간을 구현했다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "손의 크기와 금속 질감은 충실하지만 찰리가 여전히 현우의 손목을 움켜쥐고 있어, 뜯겨 나간 손과 두 손 사이 간격이라는 핵심 지시를 충족하지 못한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 팔은 화면 왼쪽에서 오른쪽의 현우 손을 향하고, 벌어진 금속 손가락 앞에 명확한 빈 간격이 있다. 현우의 손가락은 오른쪽 위로 뻗고 팔은 난간 바깥 아래쪽 물속으로 이어져, 선체에서 멀어지는 상태로 읽힌다. 얼굴과 시선은 보이지 않는다.",
        "built_space": "녹슨 굵은 수직 기둥 하나, 그 옆에 매달린 줄, 대각선으로 이어지는 선측 벽과 상부 난간 한 줄 및 여러 짧은 지지대가 보인다. 왼쪽은 젖은 갑판, 오른쪽은 파도이며 이전 장면의 재질과 공간 관계가 유지된다. 찰리는 갑판 쪽, 현우의 손과 팔은 난간 바깥에 있다. 반사상이나 중복된 주요 설비는 없다.",
        "entities": "현우의 젖은 맨손 하나와 전완 일부, 찰리의 샌드 베이지 장갑 팔과 금속 손 하나가 보인다. 현우의 손은 젊은 인물의 손으로 자연스럽지만 얼굴·머리·민족적 정체성은 이 크롭에서 확인할 수 없다. 찰리의 마모된 각진 장갑과 검은 관절은 참조에 부합한다. 폭우, 부서지는 파도, 야간의 차가운 조명이 있으며 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "찰리의 손은 손목 관절과 왼쪽 프레임 밖으로 이어지는 팔에 연결되어 있다. 현우의 손도 전완에 연결되며 전완 아래쪽은 파도에 가려진다. 몸 전체가 보이지 않는 손 클로즈업이므로 이를 공중에 뜬 무지지 신체로 볼 근거는 없다. 선체 바깥을 덮치는 파도가 휩쓸림의 원인으로 보이고, 놓친 금속 손과 멀어진 맨손의 관계도 물리적으로 가능하다."
       },
       {
        "label": "A",
        "direction": "찰리의 팔은 왼쪽에서 오른쪽으로 뻗어 현우의 손목을 감싸고 있다. 현우의 손가락은 오른쪽 위를 향하지만 두 손 사이에 분리 간격이 없으며, 파도가 덮쳐도 아직 붙잡힌 순간으로 읽힌다. 얼굴이나 시선은 보이지 않는다.",
        "built_space": "녹슨 굵은 수직 기둥 하나와 매달린 줄, 대각선 선측 벽 및 상부 난간이 보이며 갑판과 바다의 좌우 관계는 참조와 같다. 왼쪽 먼 통로에는 밝은 조명 세 점과 통로 구조가 드러나 이전 장면보다 배경 광원이 두드러진다. 찰리는 갑판 쪽에서 선측 바깥의 현우 팔을 붙잡고 있다. 불가능한 반사나 주요 설비의 명백한 중복은 없다.",
        "entities": "현우의 젖은 맨손과 아래 프레임으로 이어지는 전완, 찰리의 베이지색 장갑 팔과 금속 손이 보인다. 피부와 기계 장갑의 외관은 이전 장면과 가깝다. 얼굴·머리·옷은 크롭 밖이라 신원 세부를 확인할 수 없다. 야간 폭우와 큰 파도는 구현되어 있고 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우의 팔은 아래 프레임 밖으로 연결되고 손목은 찰리의 금속 손가락에 실제로 붙잡혀 있다. 찰리의 손과 팔 역시 관절로 연결되어 있어 떠 있는 물체는 없다. 파도가 선측에 부딪혀 솟는 모습은 가능하지만, 손목을 감싼 접촉이 유지되어 파도의 힘으로 손이 뜯겨 나간 결과는 나타나지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.444,
    "B": 1.75
   },
   "adjusted": {
    "A": 1.444,
    "B": 1.5
   },
   "violations": {
    "B": [
     "[gemini-pro] 물리적으로 불가능한 해부학 (신체와 절단된 손)",
     "[gemini-pro] 아무것도 지탱하지 않는 허공에 뜬 신체 부위"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1444,
   "B": 1500
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1444,
    "verdict_ko": "두 손이 분리되는 순간을 요구한 지시와 달리 여전히 손을 꽉 잡고 있어 핵심 액션 구현에 실패함."
   },
   {
    "label": "B",
    "score": 1500,
    "verdict_ko": "손이 떨어지는 액션은 반영했으나, 현우의 손이 팔과 단절되어 허공에 떠 있는 치명적 해부학 오류가 있음.  ★위반: [gemini-pro] 물리적으로 불가능한 해부학 (신체와 절단된 손) / [gemini-pro] 아무것도 지탱하지 않는 허공에 뜬 신체 부위"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리, 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S80sh9_sel.png",
    "asset_id": "0a522b96-308d-4e05-b6cf-88eef5d9e825",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-dc42-7bd7-9e09-51c107488e4c",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S80sh9"
  }
 },
 "S80sh15::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T08:58:50.990883+00:00",
  "fingerprint": "d8a22938f0f2227ce92ad0a16f02a067e9a60d4591d45757750eb2c3920e4464",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S80sh15_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S80sh15_sel.png",
  "source_sha256": "e0ba8f3c25540809b9ed4034af550bbfede9caee9188f77aa60dee7029c2aab3",
  "file": "S80sh15_cine.png",
  "staged_sha256": "06d5daebffa1e9ae77529f52195723a9e39f5fb4b639dc1949ef389b856a17ef",
  "latency_ms": 9622
 },
 "S80sh20::signage": {
  "fp": "4905fe22411e7df2",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S80sh20::bgfirst_bg": {
  "input_fingerprint": "50ca4ed39ec1fd85",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 칠흑 같은 심해 속으로 축 늘어진 채 나란히 떠 있는 현우와 찰리의 거대한 실루엣.\n\nLOCATION (lock): Deep underwater beneath the storm-struck boat, in near-total darkness among sinking cargo.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Deep surrounding water (Increasingly dark as the pair descend); used as Provides unbroken negative space around the paired silhouettes without introducing a seabed or visible surface.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Progressively diminishing underwater visibility leaves only restrained silhouette separation before the image fades into darkness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 칠흑 같은 심해 속으로 축 늘어진 채 나란히 떠 있는 현우와 찰리의 거대한 실루엣.\n\nLOCATION (lock): Deep underwater beneath the storm-struck boat, in near-total darkness among sinking cargo.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Deep surrounding water (Increasingly dark as the pair descend); used as Provides unbroken negative space around the paired silhouettes without introducing a seabed or visible surface.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Progressively diminishing underwater visibility leaves only restrained silhouette separation before the image fades into darkness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S80sh20__bgfirst_bg.png",
  "asset_id": "e9739767-c664-4b57-931d-69841d4b4382",
  "input_asset_ids": [
   "fee007fc-1173-489b-af5a-181096534991",
   "dac9a97a-00da-4766-aa32-41c0c8abe71c"
  ]
 },
 "S80sh20": {
  "input_fingerprint": "4425354470b5e679",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 칠흑 같은 심해 속으로 축 늘어진 채 나란히 떠 있는 현우와 찰리의 거대한 실루엣.\n\nLOCATION (lock): Deep underwater beneath the storm-struck boat, in near-total darkness among sinking cargo. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Deep surrounding water (Increasingly dark as the pair descend); used as Provides unbroken negative space around the paired silhouettes without introducing a seabed or visible surface.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Progressively diminishing underwater visibility leaves only restrained silhouette separation before the image fades into darkness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Hyunwoo is unconscious and sinking beneath the sea beside Charlie, his body limp and unsupported in the water. The source does not establish the direction of his head, the orientation of his torso, or the arrangement of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Heavy shipboard cargo sinks through increasingly dark water. Charlie is submerged and descending with his preexisting body damage still present. 현우: He is unconscious and sinking deeper underwater with his body limp.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 칠흑 같은 심해 속으로 축 늘어진 채 나란히 떠 있는 현우와 찰리의 거대한 실루엣.\n\nLOCATION (lock): Deep underwater beneath the storm-struck boat, in near-total darkness among sinking cargo. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Deep surrounding water (Increasingly dark as the pair descend); used as Provides unbroken negative space around the paired silhouettes without introducing a seabed or visible surface.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Progressively diminishing underwater visibility leaves only restrained silhouette separation before the image fades into darkness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Hyunwoo is unconscious and sinking beneath the sea beside Charlie, his body limp and unsupported in the water. The source does not establish the direction of his head, the orientation of his torso, or the arrangement of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Heavy shipboard cargo sinks through increasingly dark water. Charlie is submerged and descending with his preexisting body damage still present. 현우: He is unconscious and sinking deeper underwater with his body limp.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 칠흑 같은 심해 속으로 축 늘어진 채 나란히 떠 있는 현우와 찰리의 거대한 실루엣.\n\nLOCATION (lock): Deep underwater beneath the storm-struck boat, in near-total darkness among sinking cargo. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Deep surrounding water (Increasingly dark as the pair descend); used as Provides unbroken negative space around the paired silhouettes without introducing a seabed or visible surface.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Progressively diminishing underwater visibility leaves only restrained silhouette separation before the image fades into darkness.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Hyunwoo is unconscious and sinking beneath the sea beside Charlie, his body limp and unsupported in the water. The source does not establish the direction of his head, the orientation of his torso, or the arrangement of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Heavy shipboard cargo sinks through increasingly dark water. Charlie is submerged and descending with his preexisting body damage still present. 현우: He is unconscious and sinking deeper underwater with his body limp.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S80sh20__bgfirst_bg.png",
     "asset_id": "e9739767-c664-4b57-931d-69841d4b4382",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S80sh20.png",
     "asset_id": "fee007fc-1173-489b-af5a-181096534991",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L261B02.png",
     "asset_id": "dac9a97a-00da-4766-aa32-41c0c8abe71c",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "카메라는 아래에서 위를 향해(로우 앵글) 물속으로 가라앉는 두 캐릭터를 올려다보고 있습니다.",
    "built_space": "수면의 빛이 닿는 깊은 물속 배경이지만, 지시문에 요구된 가라앉는 화물이나 배의 모습은 보이지 않고 물의 질감과 기포만 존재합니다.",
    "entities": "현우(흰 반팔 티셔츠와 바지)와 찰리(샌드 베이지색 기계)가 묘사되었으나, 완전한 실루엣이 아닌 질감과 색상이 꽤 명확하게 드러납니다.",
    "hard_violations": [],
    "physics": "두 캐릭터는 지지대 없이 물속에서 하강하는 모습이며, 현우는 팔다리가 축 늘어져 기절한 상태의 중력을 올바르게 반영하고 있습니다."
   },
   {
    "label": "B",
    "direction": "카메라는 정면에서 약간 아래를 향해 심해로 가라앉는 두 캐릭터와 주변 사물들을 바라보고 있습니다.",
    "built_space": "어두운 심해 배경으로, 상단 수면에 배의 실루엣이 보이며, 두 캐릭터 주변으로 나무 상자 등 가라앉는 화물들이 흩어져 있습니다.",
    "entities": "검은 실루엣으로 처리된 현우와 찰리(고릴라형 기계)가 나란히 떠 있으며, 나무 상자와 금속 상자 형태의 화물, 그리고 상단의 배가 확인됩니다.",
    "hard_violations": [],
    "physics": "현우와 찰리, 그리고 주변의 화물들은 모두 물속에 떠서 중력에 의해 자연스럽게 가라앉는 상태로, 수중 부력과 중력 외에는 아무런 지지대 없이 물리적으로 올바르게 묘사되었습니다. 현우의 몸은 축 늘어진 무의식 상태를 잘 보여줍니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "지시문에 명시된 가라앉는 화물과 수면 위의 배 윤곽을 모두 포함하여, 칠흑 같은 심해 속 두 캐릭터의 거대한 실루엣과 축 늘어진 자세를 완벽하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "두 인물이 나란히 가라앉는 구도는 좋으나, 필수 요소인 가라앉는 화물과 배가 생략되었고, 실루엣으로 보이기에는 디테일이 너무 밝게 표현되었습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "카메라는 정면에서 약간 아래를 향해 심해로 가라앉는 두 캐릭터와 주변 사물들을 바라보고 있습니다.",
        "built_space": "어두운 심해 배경으로, 상단 수면에 배의 실루엣이 보이며, 두 캐릭터 주변으로 나무 상자 등 가라앉는 화물들이 흩어져 있습니다.",
        "entities": "검은 실루엣으로 처리된 현우와 찰리(고릴라형 기계)가 나란히 떠 있으며, 나무 상자와 금속 상자 형태의 화물, 그리고 상단의 배가 확인됩니다.",
        "hard_violations": [],
        "physics": "현우와 찰리, 그리고 주변의 화물들은 모두 물속에 떠서 중력에 의해 자연스럽게 가라앉는 상태로, 수중 부력과 중력 외에는 아무런 지지대 없이 물리적으로 올바르게 묘사되었습니다. 현우의 몸은 축 늘어진 무의식 상태를 잘 보여줍니다."
       },
       {
        "label": "A",
        "direction": "카메라는 아래에서 위를 향해(로우 앵글) 물속으로 가라앉는 두 캐릭터를 올려다보고 있습니다.",
        "built_space": "수면의 빛이 닿는 깊은 물속 배경이지만, 지시문에 요구된 가라앉는 화물이나 배의 모습은 보이지 않고 물의 질감과 기포만 존재합니다.",
        "entities": "현우(흰 반팔 티셔츠와 바지)와 찰리(샌드 베이지색 기계)가 묘사되었으나, 완전한 실루엣이 아닌 질감과 색상이 꽤 명확하게 드러납니다.",
        "hard_violations": [],
        "physics": "두 캐릭터는 지지대 없이 물속에서 하강하는 모습이며, 현우는 팔다리가 축 늘어져 기절한 상태의 중력을 올바르게 반영하고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "지시문에 명시된 가라앉는 화물과 수면 위의 배 윤곽을 모두 포함하여, 칠흑 같은 심해 속 두 캐릭터의 거대한 실루엣과 축 늘어진 자세를 완벽하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "두 인물이 나란히 가라앉는 구도는 좋으나, 필수 요소인 가라앉는 화물과 배가 생략되었고, 실루엣으로 보이기에는 디테일이 너무 밝게 표현되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "카메라는 정면에서 약간 아래를 향해 심해로 가라앉는 두 캐릭터와 주변 사물들을 바라보고 있습니다.",
        "built_space": "어두운 심해 배경으로, 상단 수면에 배의 실루엣이 보이며, 두 캐릭터 주변으로 나무 상자 등 가라앉는 화물들이 흩어져 있습니다.",
        "entities": "검은 실루엣으로 처리된 현우와 찰리(고릴라형 기계)가 나란히 떠 있으며, 나무 상자와 금속 상자 형태의 화물, 그리고 상단의 배가 확인됩니다.",
        "hard_violations": [],
        "physics": "현우와 찰리, 그리고 주변의 화물들은 모두 물속에 떠서 중력에 의해 자연스럽게 가라앉는 상태로, 수중 부력과 중력 외에는 아무런 지지대 없이 물리적으로 올바르게 묘사되었습니다. 현우의 몸은 축 늘어진 무의식 상태를 잘 보여줍니다."
       },
       {
        "label": "A",
        "direction": "카메라는 아래에서 위를 향해(로우 앵글) 물속으로 가라앉는 두 캐릭터를 올려다보고 있습니다.",
        "built_space": "수면의 빛이 닿는 깊은 물속 배경이지만, 지시문에 요구된 가라앉는 화물이나 배의 모습은 보이지 않고 물의 질감과 기포만 존재합니다.",
        "entities": "현우(흰 반팔 티셔츠와 바지)와 찰리(샌드 베이지색 기계)가 묘사되었으나, 완전한 실루엣이 아닌 질감과 색상이 꽤 명확하게 드러납니다.",
        "hard_violations": [],
        "physics": "두 캐릭터는 지지대 없이 물속에서 하강하는 모습이며, 현우는 팔다리가 축 늘어져 기절한 상태의 중력을 올바르게 반영하고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "넓은 암흑의 여백 안에 나란히 잠긴 두 전신과 침강 화물을 배치해 더 충실하지만, 위쪽 선체와 찰리의 다소 정돈된 자세는 아쉽다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "축 처진 자세는 잘 보이나, 밝고 선명한 수면과 크게 잡힌 인물들이 심해의 어두운 와이드 구도를 벗어나며 현우의 의상도 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 현우는 머리를 화면 오른쪽 위로 젖히고 발을 왼쪽 아래로 늘어뜨렸다. 얼굴은 위쪽을 향하며 특정 대상을 응시하지 않는다. 오른쪽 찰리는 얼굴을 약간 오른쪽으로 돌리고 두 팔과 발을 아래로 내린다. 둘 다 수영하거나 목표물을 향해 이동하는 동작 없이 나란히 침강하는 모습이다.",
        "built_space": "고정 구조물이나 해저는 없고, 중앙의 두 전신 주위로 어두운 물이 넓게 이어진다. 화물은 왼쪽 위와 아래, 오른쪽 위와 아래에 각각 하나씩 총 네 개가 보인다. 상단 중앙에는 멀리 선체로 읽히는 윤곽이 있어 완전히 끊김 없는 물속 여백은 아니다. 뚜렷한 수면이나 반사는 없다.",
        "entities": "현우는 헝클어진 검은 머리와 앳된 동아시아계 남성의 외형, 어두운 남색 반팔 상의로 참조와 대체로 일치한다. 어두운 얼굴 때문에 정확한 동일인 여부와 나이는 판별하기 어렵다. 찰리는 흰 마스크형 얼굴, 각진 베이지 계열 장갑, 육중한 긴 팔과 짧은 다리로 참조의 특징을 유지한다. 장갑의 구체적인 기존 손상은 어둠 때문에 확인하기 어렵다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 몸과 화물은 공중이 아니라 물속에 있으며, 중력에 따른 침강과 물의 부력·저항으로 설명되는 배치다. 현우의 손목과 팔, 다리는 느슨하고 머리카락은 물속에서 퍼져 있어 의식 없는 몸으로 읽힌다. 찰리의 팔은 아래로 늘어지지만 몸통과 굽힌 다리는 다소 정돈된 자세다. 화물에도 바닥이나 줄은 없으나 수중 침강이라는 물리적 상황이 성립한다."
       },
       {
        "label": "B",
        "direction": "현우는 고개를 아래로 떨구고 얼굴을 화면 오른쪽 아래로 향한다. 찰리도 머리를 오른쪽 아래로 숙이며 특정 대상을 응시하지 않는다. 두 몸은 발이 카메라에 더 가까운 저각도에서 보이고, 팔과 다리는 아래로 늘어져 있다. 추진 동작 없이 나란히 가라앉는 방향으로 읽힌다.",
        "built_space": "화면 상단에 물결치는 밝은 수면이 명확하게 보이며, 해저나 고정 시설은 없다. 두 전신이 화면 높이 대부분을 차지해 주변의 연속된 암흑 여백이 적다. 침강하는 화물은 확인되지 않는다. 수면이 보이는 얕은 공간감은 수면을 노출하지 않는 심해 구도와 맞지 않는다.",
        "entities": "현우는 검은 헝클어진 머리의 젊은 동아시아계 남성으로 보이지만, 밝은 회색 반팔 상의는 참조의 남색 상의와 다르다. 숙인 얼굴은 어두워 정확한 얼굴 일치를 확인하기 어렵다. 찰리는 베이지 장갑과 흰 기계식 마스크, 안테나 및 장갑의 긁힘을 갖췄으나 참조보다 다리가 길고 팔의 고릴라형 비례가 약하다. 두 대상 외의 인물과 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우의 목과 손목이 떨어지고 무릎도 느슨하게 굽어 있어 의식 없는 수중 자세로 성립한다. 찰리 역시 머리와 팔을 늘어뜨린다. 몸을 받치는 고체 표면은 없지만 물의 부력과 저항을 받으며 침강하는 상황이므로 불가능한 공중 부양은 아니다. 카메라 쪽으로 돌출된 발은 저각도 원근으로 설명된다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "넓은 암흑의 여백 안에 나란히 잠긴 두 전신과 침강 화물을 배치해 더 충실하지만, 위쪽 선체와 찰리의 다소 정돈된 자세는 아쉽다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "축 처진 자세는 잘 보이나, 밝고 선명한 수면과 크게 잡힌 인물들이 심해의 어두운 와이드 구도를 벗어나며 현우의 의상도 다르다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽 현우는 머리를 화면 오른쪽 위로 젖히고 발을 왼쪽 아래로 늘어뜨렸다. 얼굴은 위쪽을 향하며 특정 대상을 응시하지 않는다. 오른쪽 찰리는 얼굴을 약간 오른쪽으로 돌리고 두 팔과 발을 아래로 내린다. 둘 다 수영하거나 목표물을 향해 이동하는 동작 없이 나란히 침강하는 모습이다.",
        "built_space": "고정 구조물이나 해저는 없고, 중앙의 두 전신 주위로 어두운 물이 넓게 이어진다. 화물은 왼쪽 위와 아래, 오른쪽 위와 아래에 각각 하나씩 총 네 개가 보인다. 상단 중앙에는 멀리 선체로 읽히는 윤곽이 있어 완전히 끊김 없는 물속 여백은 아니다. 뚜렷한 수면이나 반사는 없다.",
        "entities": "현우는 헝클어진 검은 머리와 앳된 동아시아계 남성의 외형, 어두운 남색 반팔 상의로 참조와 대체로 일치한다. 어두운 얼굴 때문에 정확한 동일인 여부와 나이는 판별하기 어렵다. 찰리는 흰 마스크형 얼굴, 각진 베이지 계열 장갑, 육중한 긴 팔과 짧은 다리로 참조의 특징을 유지한다. 장갑의 구체적인 기존 손상은 어둠 때문에 확인하기 어렵다. 추가 인물이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "두 몸과 화물은 공중이 아니라 물속에 있으며, 중력에 따른 침강과 물의 부력·저항으로 설명되는 배치다. 현우의 손목과 팔, 다리는 느슨하고 머리카락은 물속에서 퍼져 있어 의식 없는 몸으로 읽힌다. 찰리의 팔은 아래로 늘어지지만 몸통과 굽힌 다리는 다소 정돈된 자세다. 화물에도 바닥이나 줄은 없으나 수중 침강이라는 물리적 상황이 성립한다."
       },
       {
        "label": "A",
        "direction": "현우는 고개를 아래로 떨구고 얼굴을 화면 오른쪽 아래로 향한다. 찰리도 머리를 오른쪽 아래로 숙이며 특정 대상을 응시하지 않는다. 두 몸은 발이 카메라에 더 가까운 저각도에서 보이고, 팔과 다리는 아래로 늘어져 있다. 추진 동작 없이 나란히 가라앉는 방향으로 읽힌다.",
        "built_space": "화면 상단에 물결치는 밝은 수면이 명확하게 보이며, 해저나 고정 시설은 없다. 두 전신이 화면 높이 대부분을 차지해 주변의 연속된 암흑 여백이 적다. 침강하는 화물은 확인되지 않는다. 수면이 보이는 얕은 공간감은 수면을 노출하지 않는 심해 구도와 맞지 않는다.",
        "entities": "현우는 검은 헝클어진 머리의 젊은 동아시아계 남성으로 보이지만, 밝은 회색 반팔 상의는 참조의 남색 상의와 다르다. 숙인 얼굴은 어두워 정확한 얼굴 일치를 확인하기 어렵다. 찰리는 베이지 장갑과 흰 기계식 마스크, 안테나 및 장갑의 긁힘을 갖췄으나 참조보다 다리가 길고 팔의 고릴라형 비례가 약하다. 두 대상 외의 인물과 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우의 목과 손목이 떨어지고 무릎도 느슨하게 굽어 있어 의식 없는 수중 자세로 성립한다. 찰리 역시 머리와 팔을 늘어뜨린다. 몸을 받치는 고체 표면은 없지만 물의 부력과 저항을 받으며 침강하는 상황이므로 불가능한 공중 부양은 아니다. 카메라 쪽으로 돌출된 발은 저각도 원근으로 설명된다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.292,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.292,
    "B": 2.0
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1292
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "지시문에 명시된 가라앉는 화물과 수면 위의 배 윤곽을 모두 포함하여, 칠흑 같은 심해 속 두 캐릭터의 거대한 실루엣과 축 늘어진 자세를 완벽하게 구현했습니다."
   },
   {
    "label": "A",
    "score": 1292,
    "verdict_ko": "두 인물이 나란히 가라앉는 구도는 좋으나, 필수 요소인 가라앉는 화물과 배가 생략되었고, 실루엣으로 보이기에는 디테일이 너무 밝게 표현되었습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L261B02.png",
    "asset_id": "dac9a97a-00da-4766-aa32-41c0c8abe71c",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-de00-77ee-98a6-98105e3bf606",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S80sh20__bgfirst_bg.png",
   "bg_asset_id": "e9739767-c664-4b57-931d-69841d4b4382",
   "bg_record_key": "S80sh20::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S80sh20::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T09:00:23.432614+00:00",
  "fingerprint": "91b2d5fd5f15da74da98a7530eb398b49f1061f4c4f9c81d62898bc7160ce2e0",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S80sh20_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S80sh20_sel.png",
  "source_sha256": "6c269515bb232bf1fe253c255af5144b86ba25fa061b047f59a0e440b908c5ec",
  "file": "S80sh20_cine.png",
  "staged_sha256": "87c59b408f7c1f1a65a6fd21e7e9f291eb32c3d225a590ea8dcb2f6d44c4a406",
  "latency_ms": 9174
 },
 "S81sh5::signage": {
  "fp": "b9ac00a4e5c29977",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::8a4521a0f910a4bf": {
  "subjects": [],
  "subject_text": "제주도 연구소 임시보호소 병실\n흰색 침대와 밝은 형광등이 있는 병실. 창밖으로 에메랄드빛 바다가 넓게 펼쳐지고 실내에는 통화 장치가 마련되어 있다.",
  "identity": "canonical",
  "scope_id": "L266",
  "scope_role": "location_interior",
  "scope_sha": "5e6010fbf1088b41"
 },
 "S81sh5::bgfirst_bg": {
  "input_fingerprint": "3905670d577532fa",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 병실 문가에 서서 다정하게 미소 짓는 서지민의 전신.\n\nLOCATION (lock): At the doorway inside an island research facility's temporary-care room, lit by strong fluorescent lights and daylight from a sea-facing window.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Room doorway (Open, with 서지민 at the threshold) — Seen obliquely from the room's interior; used as Frames her full body while leaving space toward 현우; White bed (A small corner remains visible beside the camera position) — Only a cropped near corner appears at lower left; used as Anchors the view to 현우's bedside without showing him.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Strong fluorescent illumination established inside the room is rendered with controlled highlights and readable facial detail.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 병실 문가에 서서 다정하게 미소 짓는 서지민의 전신.\n\nLOCATION (lock): At the doorway inside an island research facility's temporary-care room, lit by strong fluorescent lights and daylight from a sea-facing window.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Room doorway (Open, with 서지민 at the threshold) — Seen obliquely from the room's interior; used as Frames her full body while leaving space toward 현우; White bed (A small corner remains visible beside the camera position) — Only a cropped near corner appears at lower left; used as Anchors the view to 현우's bedside without showing him.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Strong fluorescent illumination established inside the room is rendered with controlled highlights and readable facial detail.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S81sh5__bgfirst_bg.png",
  "asset_id": "0ebb2820-cefc-4db9-9c50-9e74728ec4d6",
  "input_asset_ids": [
   "87280629-3bc9-4dc2-9c89-4967995f760f",
   "78924c8c-f3a0-4a9f-b31d-34a3fce75e4a"
  ]
 },
 "S81sh5": {
  "input_fingerprint": "1bd21d480a7d0580",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 병실 문가에 서서 다정하게 미소 짓는 서지민의 전신.\n\nLOCATION (lock): At the doorway inside an island research facility's temporary-care room, lit by strong fluorescent lights and daylight from a sea-facing window. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Room doorway (Open, with 서지민 at the threshold) — Seen obliquely from the room's interior; used as Frames her full body while leaving space toward 현우; White bed (A small corner remains visible beside the camera position) — Only a cropped near corner appears at lower left; used as Anchors the view to 현우's bedside without showing him.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Strong fluorescent illumination established inside the room is rendered with controlled highlights and readable facial detail.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A white bed is lit by strong fluorescent lighting in the temporary shelter, with an emerald-blue sea visible through the window. The storm-dark shipboard setting has ended. 서지민: She wears research clothing and has her researcher access card available.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 서지민 (한국인 여성, 20세, 앳된 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 병실 문가에 서서 다정하게 미소 짓는 서지민의 전신.\n\nLOCATION (lock): At the doorway inside an island research facility's temporary-care room, lit by strong fluorescent lights and daylight from a sea-facing window. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Room doorway (Open, with 서지민 at the threshold) — Seen obliquely from the room's interior; used as Frames her full body while leaving space toward 현우; White bed (A small corner remains visible beside the camera position) — Only a cropped near corner appears at lower left; used as Anchors the view to 현우's bedside without showing him.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Strong fluorescent illumination established inside the room is rendered with controlled highlights and readable facial detail.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A white bed is lit by strong fluorescent lighting in the temporary shelter, with an emerald-blue sea visible through the window. The storm-dark shipboard setting has ended. 서지민: She wears research clothing and has her researcher access card available.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 서지민 (한국인 여성, 20세, 앳된 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 병실 문가에 서서 다정하게 미소 짓는 서지민의 전신.\n\nLOCATION (lock): At the doorway inside an island research facility's temporary-care room, lit by strong fluorescent lights and daylight from a sea-facing window. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Room doorway (Open, with 서지민 at the threshold) — Seen obliquely from the room's interior; used as Frames her full body while leaving space toward 현우; White bed (A small corner remains visible beside the camera position) — Only a cropped near corner appears at lower left; used as Anchors the view to 현우's bedside without showing him.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Strong fluorescent illumination established inside the room is rendered with controlled highlights and readable facial detail.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): A white bed is lit by strong fluorescent lighting in the temporary shelter, with an emerald-blue sea visible through the window. The storm-dark shipboard setting has ended. 서지민: She wears research clothing and has her researcher access card available.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 서지민 (한국인 여성, 20세, 앳된 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S81sh5__bgfirst_bg.png",
     "asset_id": "0ebb2820-cefc-4db9-9c50-9e74728ec4d6",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S81sh5.png",
     "asset_id": "87280629-3bc9-4dc2-9c89-4967995f760f",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 서지민: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1147284>",
     "asset_id": "a529b479-7055-42c6-be7d-19d14b5e000f",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L266B02.png",
     "asset_id": "78924c8c-f3a0-4a9f-b31d-34a3fce75e4a",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 서지민: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1147284>",
     "asset_id": "a529b479-7055-42c6-be7d-19d14b5e000f",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "지민은 방 안쪽을 향해 미소 지으며 시선을 둠.",
    "built_space": "카메라가 방 내부에 위치하고 좌측 하단에 침대 모서리가 있는 점은 지시와 일치하나, 마주봐야 할 창문(좌측)과 출입문(우측)이 한 화면에 인접해 보이도록 방의 구조가 물리적으로 불가능하게 왜곡됨.",
    "entities": "서지민의 인상착의와 연구원 복장, 사원증을 갖추었으나, 실사 인물이 아닌 2D 일러스트레이션으로 렌더링됨.",
    "hard_violations": [
     "[gemini-pro] invented people/objects (실사 영화 스틸컷이 아닌 2D 애니메이션 스타일의 일러스트 생성)",
     "[gemini-pro] physically impossible staging (레퍼런스의 방 구조를 무시하고 창문과 문의 위치를 임의로 왜곡)",
     "[gpt-high] 인물의 의복과 신체가 뚜렷한 만화식 윤곽선과 평면 음영으로 표현되어, 실사 공간에 삽화를 붙인 듯한 그래픽 인물이 되었습니다.",
     "[gpt-high] 지시되지 않은 회색 파일을 인물의 소지품으로 추가했습니다."
    ],
    "physics": "두 발로 바닥을 딛고 걸어 들어오는 자세."
   },
   {
    "label": "B",
    "direction": "지민은 카메라(프레임 정면)를 향해 미소 지으며 바라봄.",
    "built_space": "카메라는 방 바깥에서 안쪽을 들여다보는 구도로, 방 내부에서 촬영하라는 지시를 위반함. 전경 좌측에 침대의 일부가 보이고, 방 안쪽에 또 다른 침대가 전체 크기로 존재하여 침대가 2개로 중복됨.",
    "entities": "서지민의 외모(한국인 여성, 검은 머리)와 일치하는 실사 인물. 연구원 복장과 사원증을 올바르게 착용함.",
    "hard_violations": [
     "[gemini-pro] duplicated fitting (레퍼런스에 하나인 침대가 방 안팎으로 2개 생성됨)",
     "[gemini-pro] camera position violation (방 내부에서 밖을 향하는 앵글이 아닌, 방 바깥에서 안을 들여다보는 구도)",
     "[gpt-high] 기준 장소의 침대 한 개를 전경과 문 너머의 두 개로 중복시켰습니다.",
     "[gpt-high] 지정된 병실 내부에서 출입문을 보는 배치가 아니라, 별도 침상 공간에서 바다 창이 있는 병실을 문 너머로 바라보도록 공간을 재구성했습니다."
    ],
    "physics": "두 발로 바닥을 안정적으로 딛고 서 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "실사 렌더링에는 성공했으나, 지정된 카메라 위치(방 내부)를 위반하고 공간에 침대를 중복 생성하는 치명적인 오류가 있습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "실사 영화 스틸컷 지시를 완전히 무시하고 2D 일러스트로 생성되었으며, 원본 공간의 구조를 임의로 왜곡했습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "지민은 카메라(프레임 정면)를 향해 미소 지으며 바라봄.",
        "built_space": "카메라는 방 바깥에서 안쪽을 들여다보는 구도로, 방 내부에서 촬영하라는 지시를 위반함. 전경 좌측에 침대의 일부가 보이고, 방 안쪽에 또 다른 침대가 전체 크기로 존재하여 침대가 2개로 중복됨.",
        "entities": "서지민의 외모(한국인 여성, 검은 머리)와 일치하는 실사 인물. 연구원 복장과 사원증을 올바르게 착용함.",
        "hard_violations": [
         "duplicated fitting (레퍼런스에 하나인 침대가 방 안팎으로 2개 생성됨)",
         "camera position violation (방 내부에서 밖을 향하는 앵글이 아닌, 방 바깥에서 안을 들여다보는 구도)"
        ],
        "physics": "두 발로 바닥을 안정적으로 딛고 서 있음."
       },
       {
        "label": "A",
        "direction": "지민은 방 안쪽을 향해 미소 지으며 시선을 둠.",
        "built_space": "카메라가 방 내부에 위치하고 좌측 하단에 침대 모서리가 있는 점은 지시와 일치하나, 마주봐야 할 창문(좌측)과 출입문(우측)이 한 화면에 인접해 보이도록 방의 구조가 물리적으로 불가능하게 왜곡됨.",
        "entities": "서지민의 인상착의와 연구원 복장, 사원증을 갖추었으나, 실사 인물이 아닌 2D 일러스트레이션으로 렌더링됨.",
        "hard_violations": [
         "invented people/objects (실사 영화 스틸컷이 아닌 2D 애니메이션 스타일의 일러스트 생성)",
         "physically impossible staging (레퍼런스의 방 구조를 무시하고 창문과 문의 위치를 임의로 왜곡)"
        ],
        "physics": "두 발로 바닥을 딛고 걸어 들어오는 자세."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "실사 렌더링에는 성공했으나, 지정된 카메라 위치(방 내부)를 위반하고 공간에 침대를 중복 생성하는 치명적인 오류가 있습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "실사 영화 스틸컷 지시를 완전히 무시하고 2D 일러스트로 생성되었으며, 원본 공간의 구조를 임의로 왜곡했습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "지민은 카메라(프레임 정면)를 향해 미소 지으며 바라봄.",
        "built_space": "카메라는 방 바깥에서 안쪽을 들여다보는 구도로, 방 내부에서 촬영하라는 지시를 위반함. 전경 좌측에 침대의 일부가 보이고, 방 안쪽에 또 다른 침대가 전체 크기로 존재하여 침대가 2개로 중복됨.",
        "entities": "서지민의 외모(한국인 여성, 검은 머리)와 일치하는 실사 인물. 연구원 복장과 사원증을 올바르게 착용함.",
        "hard_violations": [
         "duplicated fitting (레퍼런스에 하나인 침대가 방 안팎으로 2개 생성됨)",
         "camera position violation (방 내부에서 밖을 향하는 앵글이 아닌, 방 바깥에서 안을 들여다보는 구도)"
        ],
        "physics": "두 발로 바닥을 안정적으로 딛고 서 있음."
       },
       {
        "label": "A",
        "direction": "지민은 방 안쪽을 향해 미소 지으며 시선을 둠.",
        "built_space": "카메라가 방 내부에 위치하고 좌측 하단에 침대 모서리가 있는 점은 지시와 일치하나, 마주봐야 할 창문(좌측)과 출입문(우측)이 한 화면에 인접해 보이도록 방의 구조가 물리적으로 불가능하게 왜곡됨.",
        "entities": "서지민의 인상착의와 연구원 복장, 사원증을 갖추었으나, 실사 인물이 아닌 2D 일러스트레이션으로 렌더링됨.",
        "hard_violations": [
         "invented people/objects (실사 영화 스틸컷이 아닌 2D 애니메이션 스타일의 일러스트 생성)",
         "physically impossible staging (레퍼런스의 방 구조를 무시하고 창문과 문의 위치를 임의로 왜곡)"
        ],
        "physics": "두 발로 바닥을 딛고 걸어 들어오는 자세."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "실사 인물과 미소는 부합하지만, 침대가 두 개로 늘어나고 문 너머에 바다 창이 있는 병실을 배치해 지정된 병실 내부 시점과 공간 구조를 위반합니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "병상 쪽에서 문가의 전신을 보는 구도는 더 정확하지만, 인물이 윤곽선 있는 삽화처럼 표현되고 지시되지 않은 파일을 들고 있어 최종 실사 컷으로는 부적합합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "서지민은 거의 카메라 정면을 보며 웃습니다. 카메라 옆 병상의 현우를 향한 시선으로 해석할 여지는 있지만, 왼쪽 병상 쪽으로 시선이 명확하게 향하지는 않습니다. 한쪽 팔은 오른쪽 문틀 방향으로 뻗어 있습니다.",
        "built_space": "왼쪽 전경에 침대 하나, 문 너머 왼쪽에 별도의 침대 하나가 보여 총 두 개입니다. 중앙의 열린 출입구 너머에 바다 창 하나, 오른쪽 벽에 모니터 하나와 제어 장치가 있고 천장에 긴 조명과 환기구가 보입니다. 서지민은 출입구에 서 있지만, 카메라는 기준 사진의 병실 안에서 출입문을 돌아보는 대신 별도의 침상 공간에서 그 병실 안을 바라보는 배치입니다. 전경 침대도 작은 모서리보다 훨씬 넓게 노출됩니다.",
        "entities": "보이는 사람은 검은 어깨 길이 머리의 젊은 동아시아계 여성 한 명으로, 기준 인물의 얼굴과 체형에 대체로 가깝습니다. 짙은 남색 연구복 형태의 옷과 목걸이형 출입증이 있으며 다정한 미소와 전신이 보입니다. 흰 침대, 낮의 바다, 창과 의료시설 장치가 있지만 침대 수는 맞지 않습니다. 현우나 다른 사람은 보이지 않으며 확실히 판독되는 문구는 확인되지 않습니다.",
        "hard_violations": [
         "기준 장소의 침대 한 개를 전경과 문 너머의 두 개로 중복시켰습니다.",
         "지정된 병실 내부에서 출입문을 보는 배치가 아니라, 별도 침상 공간에서 바다 창이 있는 병실을 문 너머로 바라보도록 공간을 재구성했습니다."
        ],
        "physics": "양쪽 신발이 바닥에 닿아 체중을 지지합니다. 뻗은 손은 문틀에 가려 접촉이 명확하지 않지만 서 있는 데 추가 지지는 필요하지 않습니다. 출입증은 목줄에 매달리고 침대는 다리와 바퀴로 지지됩니다. 떠 있는 신체나 물체는 없습니다."
       },
       {
        "label": "B",
        "direction": "서지민의 얼굴과 시선은 화면 왼쪽의 병상 및 카메라 옆 공간을 향하며 웃고 있어, 화면 밖 현우를 대하는 관계가 더 명확합니다. 한 손은 오른쪽 문틀 쪽으로 뻗고 다른 팔은 회색 판형 파일을 몸에 끌어안습니다.",
        "built_space": "왼쪽 전경에 흰 침대 하나의 일부, 왼쪽 벽에 바다 창 하나, 중앙 벽에 모니터 하나, 오른쪽에 열린 출입구 하나가 보입니다. 출입구 뒤에는 복도와 다른 문틀이 보이고, 실내와 복도 천장에 조명이 있습니다. 서지민은 문턱에 서 있으며 병실 안에서 출입문을 비스듬히 보는 관계는 맞습니다. 다만 전경 침대가 작은 모서리에 그치지 않고 화면 하단을 크게 차지하며, 창과 벽면 장치의 세부 형태는 기준 사진과 다릅니다.",
        "entities": "사람은 젊은 동아시아계 여성 한 명이며 검은 단발성 머리와 미소가 보입니다. 그러나 얼굴과 체형은 기준 인물보다 양식화되어 있고, 특히 옷과 팔다리에 검은 윤곽선과 평면적인 음영이 있어 실사 배우로 읽히지 않습니다. 밝은 재킷과 바지, 출입증을 착용하며 지시되지 않은 회색 파일도 들고 있습니다. 흰 침대와 낮의 바다는 보이고 다른 사람이나 명확히 읽히는 글자는 없습니다.",
        "hard_violations": [
         "인물의 의복과 신체가 뚜렷한 만화식 윤곽선과 평면 음영으로 표현되어, 실사 공간에 삽화를 붙인 듯한 그래픽 인물이 되었습니다.",
         "지시되지 않은 회색 파일을 인물의 소지품으로 추가했습니다."
        ],
        "physics": "앞쪽 신발과 뒤쪽 신발이 문턱 주변 바닥에 닿아 몸을 지지합니다. 다리를 엇갈린 자세는 가능한 체중 이동이며 공중에 떠 있지는 않습니다. 회색 파일은 팔과 손으로 몸에 붙여 지지하고, 출입증은 목줄에 매달려 있습니다. 뻗은 손은 문틀 뒤로 일부 가려지지만 불가능한 지지 관계는 보이지 않습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "실사 인물과 미소는 부합하지만, 침대가 두 개로 늘어나고 문 너머에 바다 창이 있는 병실을 배치해 지정된 병실 내부 시점과 공간 구조를 위반합니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "병상 쪽에서 문가의 전신을 보는 구도는 더 정확하지만, 인물이 윤곽선 있는 삽화처럼 표현되고 지시되지 않은 파일을 들고 있어 최종 실사 컷으로는 부적합합니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "서지민은 거의 카메라 정면을 보며 웃습니다. 카메라 옆 병상의 현우를 향한 시선으로 해석할 여지는 있지만, 왼쪽 병상 쪽으로 시선이 명확하게 향하지는 않습니다. 한쪽 팔은 오른쪽 문틀 방향으로 뻗어 있습니다.",
        "built_space": "왼쪽 전경에 침대 하나, 문 너머 왼쪽에 별도의 침대 하나가 보여 총 두 개입니다. 중앙의 열린 출입구 너머에 바다 창 하나, 오른쪽 벽에 모니터 하나와 제어 장치가 있고 천장에 긴 조명과 환기구가 보입니다. 서지민은 출입구에 서 있지만, 카메라는 기준 사진의 병실 안에서 출입문을 돌아보는 대신 별도의 침상 공간에서 그 병실 안을 바라보는 배치입니다. 전경 침대도 작은 모서리보다 훨씬 넓게 노출됩니다.",
        "entities": "보이는 사람은 검은 어깨 길이 머리의 젊은 동아시아계 여성 한 명으로, 기준 인물의 얼굴과 체형에 대체로 가깝습니다. 짙은 남색 연구복 형태의 옷과 목걸이형 출입증이 있으며 다정한 미소와 전신이 보입니다. 흰 침대, 낮의 바다, 창과 의료시설 장치가 있지만 침대 수는 맞지 않습니다. 현우나 다른 사람은 보이지 않으며 확실히 판독되는 문구는 확인되지 않습니다.",
        "hard_violations": [
         "기준 장소의 침대 한 개를 전경과 문 너머의 두 개로 중복시켰습니다.",
         "지정된 병실 내부에서 출입문을 보는 배치가 아니라, 별도 침상 공간에서 바다 창이 있는 병실을 문 너머로 바라보도록 공간을 재구성했습니다."
        ],
        "physics": "양쪽 신발이 바닥에 닿아 체중을 지지합니다. 뻗은 손은 문틀에 가려 접촉이 명확하지 않지만 서 있는 데 추가 지지는 필요하지 않습니다. 출입증은 목줄에 매달리고 침대는 다리와 바퀴로 지지됩니다. 떠 있는 신체나 물체는 없습니다."
       },
       {
        "label": "A",
        "direction": "서지민의 얼굴과 시선은 화면 왼쪽의 병상 및 카메라 옆 공간을 향하며 웃고 있어, 화면 밖 현우를 대하는 관계가 더 명확합니다. 한 손은 오른쪽 문틀 쪽으로 뻗고 다른 팔은 회색 판형 파일을 몸에 끌어안습니다.",
        "built_space": "왼쪽 전경에 흰 침대 하나의 일부, 왼쪽 벽에 바다 창 하나, 중앙 벽에 모니터 하나, 오른쪽에 열린 출입구 하나가 보입니다. 출입구 뒤에는 복도와 다른 문틀이 보이고, 실내와 복도 천장에 조명이 있습니다. 서지민은 문턱에 서 있으며 병실 안에서 출입문을 비스듬히 보는 관계는 맞습니다. 다만 전경 침대가 작은 모서리에 그치지 않고 화면 하단을 크게 차지하며, 창과 벽면 장치의 세부 형태는 기준 사진과 다릅니다.",
        "entities": "사람은 젊은 동아시아계 여성 한 명이며 검은 단발성 머리와 미소가 보입니다. 그러나 얼굴과 체형은 기준 인물보다 양식화되어 있고, 특히 옷과 팔다리에 검은 윤곽선과 평면적인 음영이 있어 실사 배우로 읽히지 않습니다. 밝은 재킷과 바지, 출입증을 착용하며 지시되지 않은 회색 파일도 들고 있습니다. 흰 침대와 낮의 바다는 보이고 다른 사람이나 명확히 읽히는 글자는 없습니다.",
        "hard_violations": [
         "인물의 의복과 신체가 뚜렷한 만화식 윤곽선과 평면 음영으로 표현되어, 실사 공간에 삽화를 붙인 듯한 그래픽 인물이 되었습니다.",
         "지시되지 않은 회색 파일을 인물의 소지품으로 추가했습니다."
        ],
        "physics": "앞쪽 신발과 뒤쪽 신발이 문턱 주변 바닥에 닿아 몸을 지지합니다. 다리를 엇갈린 자세는 가능한 체중 이동이며 공중에 떠 있지는 않습니다. 회색 파일은 팔과 손으로 몸에 붙여 지지하고, 출입증은 목줄에 매달려 있습니다. 뻗은 손은 문틀 뒤로 일부 가려지지만 불가능한 지지 관계는 보이지 않습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.75,
    "B": 1.667
   },
   "adjusted": {
    "A": 1.5,
    "B": 1.417
   },
   "violations": {
    "B": [
     "[gemini-pro] duplicated fitting (레퍼런스에 하나인 침대가 방 안팎으로 2개 생성됨)",
     "[gemini-pro] camera position violation (방 내부에서 밖을 향하는 앵글이 아닌, 방 바깥에서 안을 들여다보는 구도)",
     "[gpt-high] 기준 장소의 침대 한 개를 전경과 문 너머의 두 개로 중복시켰습니다.",
     "[gpt-high] 지정된 병실 내부에서 출입문을 보는 배치가 아니라, 별도 침상 공간에서 바다 창이 있는 병실을 문 너머로 바라보도록 공간을 재구성했습니다."
    ],
    "A": [
     "[gemini-pro] invented people/objects (실사 영화 스틸컷이 아닌 2D 애니메이션 스타일의 일러스트 생성)",
     "[gemini-pro] physically impossible staging (레퍼런스의 방 구조를 무시하고 창문과 문의 위치를 임의로 왜곡)",
     "[gpt-high] 인물의 의복과 신체가 뚜렷한 만화식 윤곽선과 평면 음영으로 표현되어, 실사 공간에 삽화를 붙인 듯한 그래픽 인물이 되었습니다.",
     "[gpt-high] 지시되지 않은 회색 파일을 인물의 소지품으로 추가했습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1417,
   "A": 1500
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1417,
    "verdict_ko": "실사 렌더링에는 성공했으나, 지정된 카메라 위치(방 내부)를 위반하고 공간에 침대를 중복 생성하는 치명적인 오류가 있습니다.  ★위반: [gemini-pro] duplicated fitting (레퍼런스에 하나인 침대가 방 안팎으로 2개 생성됨) / [gemini-pro] camera position violation (방 내부에서 밖을 향하는 앵글이 아닌, 방 바깥에서 안을 들여다보는 구도) / [gpt-high] 기준 장소의 침대 한 개를 전경과 문 너머의 두 개로 중복시켰습니다. / [gpt-high] 지정된 병실 내부에서 출입문을 보는 배치가 아니라, 별도 침상 공간에서 바다 창이 있는 병실을 문 너머로 바라보도록 공간을 재구성했습니다."
   },
   {
    "label": "A",
    "score": 1500,
    "verdict_ko": "실사 영화 스틸컷 지시를 완전히 무시하고 2D 일러스트로 생성되었으며, 원본 공간의 구조를 임의로 왜곡했습니다.  ★위반: [gemini-pro] invented people/objects (실사 영화 스틸컷이 아닌 2D 애니메이션 스타일의 일러스트 생성) / [gemini-pro] physically impossible staging (레퍼런스의 방 구조를 무시하고 창문과 문의 위치를 임의로 왜곡) / [gpt-high] 인물의 의복과 신체가 뚜렷한 만화식 윤곽선과 평면 음영으로 표현되어, 실사 공간에 삽화를 붙인 듯한 그래픽 인물이 되었습니다. / [gpt-high] 지시되지 않은 회색 파일을 인물의 소지품으로 추가했습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L266B02.png",
    "asset_id": "78924c8c-f3a0-4a9f-b31d-34a3fce75e4a",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 서지민: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1147284>",
    "asset_id": "a529b479-7055-42c6-be7d-19d14b5e000f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-e165-7c1c-b30d-54ec6c810f8e",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S81sh5__bgfirst_bg.png",
   "bg_asset_id": "0ebb2820-cefc-4db9-9c50-9e74728ec4d6",
   "bg_record_key": "S81sh5::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S81sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T09:02:19.711795+00:00",
  "fingerprint": "36cbf52e4d44c2b8edc24ffd40a7a3d0eb53ab1c68ed25a0e1d416a005d26135",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S81sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S81sh5_sel.png",
  "source_sha256": "d4c79127a37704c478ae84c0d19bd40ec22dbfb422e29dff4367396d059270bd",
  "file": "S81sh5_cine.png",
  "staged_sha256": "7ab6bb0b1076a8cb93b4d0da0855c4fe6e76c1fad75b3311fce5beadd143ce03",
  "latency_ms": 9768
 },
 "S81sh13::signage": {
  "fp": "8d90e05dd031158a",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S81sh13::bgfirst_bg": {
  "input_fingerprint": "947b3aa49300d01c",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 멈춰 선 채 두 눈이 동그라져 앞을 빤히 주시하는 현우의 놀란 얼굴 클로즈업.\n\nLOCATION (lock): In the corridor connecting the temporary-care rooms to the island research facility, under ordinary corridor lighting.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Research laboratory corridor (Visible behind 현우 as he stops following 서지민) — The corridor recedes along the established walking line; used as Soft background depth and open space along his eyeline.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient light and restrained contrast preserve the immediacy of his startled expression without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 멈춰 선 채 두 눈이 동그라져 앞을 빤히 주시하는 현우의 놀란 얼굴 클로즈업.\n\nLOCATION (lock): In the corridor connecting the temporary-care rooms to the island research facility, under ordinary corridor lighting.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Research laboratory corridor (Visible behind 현우 as he stops following 서지민) — The corridor recedes along the established walking line; used as Soft background depth and open space along his eyeline.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient light and restrained contrast preserve the immediacy of his startled expression without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S81sh13__bgfirst_bg.png",
  "asset_id": "2596dc8d-81c7-4a2d-aa17-c63234623f09",
  "input_asset_ids": [
   "2d7ff4e0-834c-440a-92c5-40e9da371422",
   "4bcfae29-00ac-4c7b-a17f-ba83dae38ae3"
  ]
 },
 "S81sh13": {
  "input_fingerprint": "cb2a21f49e18cf66",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 멈춰 선 채 두 눈이 동그라져 앞을 빤히 주시하는 현우의 놀란 얼굴 클로즈업.\n\nLOCATION (lock): In the corridor connecting the temporary-care rooms to the island research facility, under ordinary corridor lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Research laboratory corridor (Visible behind 현우 as he stops following 서지민) — The corridor recedes along the established walking line; used as Soft background depth and open space along his eyeline.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient light and restrained contrast preserve the immediacy of his startled expression without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): 현우: He is wearing the fresh white clothes provided during his recovery and has stopped in the research-facility corridor.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 멈춰 선 채 두 눈이 동그라져 앞을 빤히 주시하는 현우의 놀란 얼굴 클로즈업.\n\nLOCATION (lock): In the corridor connecting the temporary-care rooms to the island research facility, under ordinary corridor lighting. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Research laboratory corridor (Visible behind 현우 as he stops following 서지민) — The corridor recedes along the established walking line; used as Soft background depth and open space along his eyeline.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient light and restrained contrast preserve the immediacy of his startled expression without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): 현우: He is wearing the fresh white clothes provided during his recovery and has stopped in the research-facility corridor.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 멈춰 선 채 두 눈이 동그라져 앞을 빤히 주시하는 현우의 놀란 얼굴 클로즈업.\n\nLOCATION (lock): In the corridor connecting the temporary-care rooms to the island research facility, under ordinary corridor lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Research laboratory corridor (Visible behind 현우 as he stops following 서지민) — The corridor recedes along the established walking line; used as Soft background depth and open space along his eyeline.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient light and restrained contrast preserve the immediacy of his startled expression without introducing a new light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): 현우: He is wearing the fresh white clothes provided during his recovery and has stopped in the research-facility corridor.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S81sh13__bgfirst_bg.png",
     "asset_id": "2596dc8d-81c7-4a2d-aa17-c63234623f09",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S81sh13.png",
     "asset_id": "2d7ff4e0-834c-440a-92c5-40e9da371422",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L266B01.png",
     "asset_id": "4bcfae29-00ac-4c7b-a17f-ba83dae38ae3",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "시선은 프레임 우측 안쪽으로 뻗어 있는 복도를 향하고 있음.",
    "built_space": "금속 프레임과 유리문이 있는 연구소 복도가 우측 뒤로 깊이감 있게 뻗어 있으며, 식별 가능한 글자는 보이지 않음.",
    "entities": "현우(10대 후반, 동양인 남성, 흰색 티셔츠, 헝클어진 머리)가 놀란 표정으로 클로즈업됨.",
    "hard_violations": [],
    "physics": "바닥에 안정적으로 서서 이동을 멈춘 상태를 자연스럽게 유지함."
   },
   {
    "label": "B",
    "direction": "시선이 카메라 렌즈(정면)를 똑바로 응시하고 있음.",
    "built_space": "연구소 복도가 인물 뒤로 대칭을 이루며 뻗어 있고, 우측 벽면에 '연구센터'라는 글자가 선명하게 배치됨.",
    "entities": "현우(10대 후반, 동양인 남성, 흰색 티셔츠)가 클로즈업되어 정면을 바라봄.",
    "hard_violations": [
     "[gpt-high] 오른쪽 벽 표지에서 ‘연구센터’라는 문자를 읽을 수 있어 읽을 수 있는 글자 금지 조건을 위반한다."
    ],
    "physics": "바닥에 대칭적으로 올곧게 서 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "시선 방향을 따라 복도의 깊이감을 배치하라는 구도 지시를 훌륭하게 구현했으며, 읽을 수 있는 문자가 없어야 한다는 부정 프롬프트도 완벽히 충족했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "복도가 시선 방향이 아닌 인물 뒤로 배치되어 구도 지시를 어겼으며, 우측 벽면에 '연구센터' 문자가 선명하게 노출되어 문자 배제 조건을 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 프레임 우측 안쪽으로 뻗어 있는 복도를 향하고 있음.",
        "built_space": "금속 프레임과 유리문이 있는 연구소 복도가 우측 뒤로 깊이감 있게 뻗어 있으며, 식별 가능한 글자는 보이지 않음.",
        "entities": "현우(10대 후반, 동양인 남성, 흰색 티셔츠, 헝클어진 머리)가 놀란 표정으로 클로즈업됨.",
        "hard_violations": [],
        "physics": "바닥에 안정적으로 서서 이동을 멈춘 상태를 자연스럽게 유지함."
       },
       {
        "label": "B",
        "direction": "시선이 카메라 렌즈(정면)를 똑바로 응시하고 있음.",
        "built_space": "연구소 복도가 인물 뒤로 대칭을 이루며 뻗어 있고, 우측 벽면에 '연구센터'라는 글자가 선명하게 배치됨.",
        "entities": "현우(10대 후반, 동양인 남성, 흰색 티셔츠)가 클로즈업되어 정면을 바라봄.",
        "hard_violations": [],
        "physics": "바닥에 대칭적으로 올곧게 서 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "시선 방향을 따라 복도의 깊이감을 배치하라는 구도 지시를 훌륭하게 구현했으며, 읽을 수 있는 문자가 없어야 한다는 부정 프롬프트도 완벽히 충족했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "복도가 시선 방향이 아닌 인물 뒤로 배치되어 구도 지시를 어겼으며, 우측 벽면에 '연구센터' 문자가 선명하게 노출되어 문자 배제 조건을 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 프레임 우측 안쪽으로 뻗어 있는 복도를 향하고 있음.",
        "built_space": "금속 프레임과 유리문이 있는 연구소 복도가 우측 뒤로 깊이감 있게 뻗어 있으며, 식별 가능한 글자는 보이지 않음.",
        "entities": "현우(10대 후반, 동양인 남성, 흰색 티셔츠, 헝클어진 머리)가 놀란 표정으로 클로즈업됨.",
        "hard_violations": [],
        "physics": "바닥에 안정적으로 서서 이동을 멈춘 상태를 자연스럽게 유지함."
       },
       {
        "label": "B",
        "direction": "시선이 카메라 렌즈(정면)를 똑바로 응시하고 있음.",
        "built_space": "연구소 복도가 인물 뒤로 대칭을 이루며 뻗어 있고, 우측 벽면에 '연구센터'라는 글자가 선명하게 배치됨.",
        "entities": "현우(10대 후반, 동양인 남성, 흰색 티셔츠)가 클로즈업되어 정면을 바라봄.",
        "hard_violations": [],
        "physics": "바닥에 대칭적으로 올곧게 서 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "놀란 표정과 장소는 맞지만 얼굴 클로즈업보다 넓은 가슴 위 구도이며, 오른쪽의 ‘연구센터’ 표지가 읽혀 문자 금지 조건을 위반한다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "현우의 얼굴을 크게 잡고 정상적인 눈을 동그랗게 뜬 놀람과 전방 응시를 구현했으며, 시선 쪽 여백과 연구시설 복도의 깊이를 함께 살렸다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 얼굴과 두 눈을 거의 카메라 정면으로 향하고 렌즈 부근의 화면 밖 전방을 응시한다. 구체적인 응시 대상은 보이지 않으며, 눈을 크게 뜨고 입을 조금 벌린 모습은 앞을 빤히 보는 놀람에 부합한다. 무기나 방향을 확인할 휴대 물체는 없다.",
        "built_space": "현우가 중앙의 큰 금속 출입구 앞에 있고 그 뒤로 복도가 이어진다. 전경 출입구 한 벌, 뒤쪽 출입구 한 벌, 왼쪽 운반 카트 한 대, 오른쪽 실내 창 한 개와 그 아래 적재 상자들이 보인다. 회색 바닥과 금속 문틀, 벽 하부 보호판은 장소 참조와 잘 맞는다. 오른쪽 벽 표지 한 개에는 글자가 남아 있다. 얼굴보다 주변 공간과 상체가 많이 보여 요구한 클로즈업보다 넓다.",
        "entities": "보이는 사람은 현우 한 명뿐이다. 앳된 동아시아계 남성의 외형, 헝클어진 검은 머리, 얼굴 윤곽과 피부 질감은 인물 참조에 가깝고 흰색 티셔츠는 회복 중 지급받은 흰옷 조건에 맞는다. 국적은 외형만으로 확인할 수 없다. 눈은 정상적인 홍채와 동공을 유지한다. 오른쪽 표지의 ‘연구센터’는 흐리지만 읽을 수 있다.",
        "hard_violations": [
         "오른쪽 벽 표지에서 ‘연구센터’라는 문자를 읽을 수 있어 읽을 수 있는 글자 금지 조건을 위반한다."
        ],
        "physics": "머리는 목과 상체에 자연스럽게 연결되고 흰 티셔츠는 어깨에 걸쳐 아래로 드리워진다. 발은 구도 밖이므로 지면 접촉이나 체중 이동은 확인할 수 없지만, 상체에 부유나 불가능한 자세는 없다. 카트와 상자들은 바닥 위에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "현우의 얼굴과 두 눈은 화면 오른쪽의 화면 밖 전방을 향한다. 특정 대상은 보이지 않지만 시선이 한 지점에 고정되어 있고, 그 방향으로 화면 여백이 확보되어 있다. 두 눈을 크게 뜨고 입술을 벌린 표정은 멈춰 서서 앞을 주시하는 놀람으로 읽힌다.",
        "built_space": "현우의 얼굴이 왼쪽 전경을 크게 차지하고 오른쪽 뒤로 복도가 깊게 이어진다. 왼쪽 가장자리의 가까운 금속 문틀 일부, 중경의 열린 양문 출입구 한 벌, 왼쪽 운반 카트 한 대, 오른쪽 실내 창 한 개, 끝의 외부 창 한 면이 보인다. 천장에는 사각 조명들이 연속 배치되어 있다. 금속 프레임, 회색 반사 바닥, 밝은 벽과 하부 보호판이 참조 장소와 일치하는 성격이며, 창과 바닥의 빛도 공간상 자연스럽다.",
        "entities": "인물은 현우 한 명이며 다른 사람이나 신체 일부는 없다. 앳된 동아시아계 남성의 얼굴, 헝클어진 검은 머리, 코와 입술 및 피부 질감이 참조 인물에 가깝다. 한국계 미국인이라는 국적은 시각적으로 확인할 수 없다. 보이는 옷은 흰색 티셔츠로 지정된 의상과 맞는다. 눈의 놀람은 눈꺼풀과 표정으로 표현되며 안구의 비정상적 변형은 없다. 배경 표지에서는 읽을 수 있는 문자가 보이지 않는다.",
        "hard_violations": [],
        "physics": "머리와 목, 어깨의 연결이 자연스럽고 정지한 상체 자세가 가능하다. 발과 손은 클로즈업 밖이므로 지지 발이나 중단된 보행 동작까지 확인할 수는 없다. 몸이 공중에 떠 있다는 단서는 없고 옷은 어깨와 몸통에 의해 지지된다. 뒤쪽 카트는 바닥에 놓여 있으며 다른 부유 물체도 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "놀란 표정과 장소는 맞지만 얼굴 클로즈업보다 넓은 가슴 위 구도이며, 오른쪽의 ‘연구센터’ 표지가 읽혀 문자 금지 조건을 위반한다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "현우의 얼굴을 크게 잡고 정상적인 눈을 동그랗게 뜬 놀람과 전방 응시를 구현했으며, 시선 쪽 여백과 연구시설 복도의 깊이를 함께 살렸다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 얼굴과 두 눈을 거의 카메라 정면으로 향하고 렌즈 부근의 화면 밖 전방을 응시한다. 구체적인 응시 대상은 보이지 않으며, 눈을 크게 뜨고 입을 조금 벌린 모습은 앞을 빤히 보는 놀람에 부합한다. 무기나 방향을 확인할 휴대 물체는 없다.",
        "built_space": "현우가 중앙의 큰 금속 출입구 앞에 있고 그 뒤로 복도가 이어진다. 전경 출입구 한 벌, 뒤쪽 출입구 한 벌, 왼쪽 운반 카트 한 대, 오른쪽 실내 창 한 개와 그 아래 적재 상자들이 보인다. 회색 바닥과 금속 문틀, 벽 하부 보호판은 장소 참조와 잘 맞는다. 오른쪽 벽 표지 한 개에는 글자가 남아 있다. 얼굴보다 주변 공간과 상체가 많이 보여 요구한 클로즈업보다 넓다.",
        "entities": "보이는 사람은 현우 한 명뿐이다. 앳된 동아시아계 남성의 외형, 헝클어진 검은 머리, 얼굴 윤곽과 피부 질감은 인물 참조에 가깝고 흰색 티셔츠는 회복 중 지급받은 흰옷 조건에 맞는다. 국적은 외형만으로 확인할 수 없다. 눈은 정상적인 홍채와 동공을 유지한다. 오른쪽 표지의 ‘연구센터’는 흐리지만 읽을 수 있다.",
        "hard_violations": [
         "오른쪽 벽 표지에서 ‘연구센터’라는 문자를 읽을 수 있어 읽을 수 있는 글자 금지 조건을 위반한다."
        ],
        "physics": "머리는 목과 상체에 자연스럽게 연결되고 흰 티셔츠는 어깨에 걸쳐 아래로 드리워진다. 발은 구도 밖이므로 지면 접촉이나 체중 이동은 확인할 수 없지만, 상체에 부유나 불가능한 자세는 없다. 카트와 상자들은 바닥 위에 놓여 있다."
       },
       {
        "label": "A",
        "direction": "현우의 얼굴과 두 눈은 화면 오른쪽의 화면 밖 전방을 향한다. 특정 대상은 보이지 않지만 시선이 한 지점에 고정되어 있고, 그 방향으로 화면 여백이 확보되어 있다. 두 눈을 크게 뜨고 입술을 벌린 표정은 멈춰 서서 앞을 주시하는 놀람으로 읽힌다.",
        "built_space": "현우의 얼굴이 왼쪽 전경을 크게 차지하고 오른쪽 뒤로 복도가 깊게 이어진다. 왼쪽 가장자리의 가까운 금속 문틀 일부, 중경의 열린 양문 출입구 한 벌, 왼쪽 운반 카트 한 대, 오른쪽 실내 창 한 개, 끝의 외부 창 한 면이 보인다. 천장에는 사각 조명들이 연속 배치되어 있다. 금속 프레임, 회색 반사 바닥, 밝은 벽과 하부 보호판이 참조 장소와 일치하는 성격이며, 창과 바닥의 빛도 공간상 자연스럽다.",
        "entities": "인물은 현우 한 명이며 다른 사람이나 신체 일부는 없다. 앳된 동아시아계 남성의 얼굴, 헝클어진 검은 머리, 코와 입술 및 피부 질감이 참조 인물에 가깝다. 한국계 미국인이라는 국적은 시각적으로 확인할 수 없다. 보이는 옷은 흰색 티셔츠로 지정된 의상과 맞는다. 눈의 놀람은 눈꺼풀과 표정으로 표현되며 안구의 비정상적 변형은 없다. 배경 표지에서는 읽을 수 있는 문자가 보이지 않는다.",
        "hard_violations": [],
        "physics": "머리와 목, 어깨의 연결이 자연스럽고 정지한 상체 자세가 가능하다. 발과 손은 클로즈업 밖이므로 지지 발이나 중단된 보행 동작까지 확인할 수는 없다. 몸이 공중에 떠 있다는 단서는 없고 옷은 어깨와 몸통에 의해 지지된다. 뒤쪽 카트는 바닥에 놓여 있으며 다른 부유 물체도 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.762
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.512
   },
   "violations": {
    "B": [
     "[gpt-high] 오른쪽 벽 표지에서 ‘연구센터’라는 문자를 읽을 수 있어 읽을 수 있는 글자 금지 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 512
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "시선 방향을 따라 복도의 깊이감을 배치하라는 구도 지시를 훌륭하게 구현했으며, 읽을 수 있는 문자가 없어야 한다는 부정 프롬프트도 완벽히 충족했습니다."
   },
   {
    "label": "B",
    "score": 512,
    "verdict_ko": "복도가 시선 방향이 아닌 인물 뒤로 배치되어 구도 지시를 어겼으며, 우측 벽면에 '연구센터' 문자가 선명하게 노출되어 문자 배제 조건을 위반했습니다.  ★위반: [gpt-high] 오른쪽 벽 표지에서 ‘연구센터’라는 문자를 읽을 수 있어 읽을 수 있는 글자 금지 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L266B01.png",
    "asset_id": "4bcfae29-00ac-4c7b-a17f-ba83dae38ae3",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-e4bc-784b-9f81-6b105393452c",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S81sh13__bgfirst_bg.png",
   "bg_asset_id": "2596dc8d-81c7-4a2d-aa17-c63234623f09",
   "bg_record_key": "S81sh13::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S81sh13::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T09:04:01.533219+00:00",
  "fingerprint": "0fe34fe8fa55ef59c131e2460a85e9cab9db629f3a48a9d34bba0d0f96d9f3a2",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S81sh13_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S81sh13_sel.png",
  "source_sha256": "2057287911b3f504c86fed86f715e0a90b1d24b4583cfa28ad453315da45f691",
  "file": "S81sh13_cine.png",
  "staged_sha256": "9649d3ff2f89b2f5047ae6cb45b07f18543b33565ea579a0bf534ed72f758a82",
  "latency_ms": 9791
 },
 "S81sh16::signage": {
  "fp": "27e27e766769817c",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S81sh16": {
  "input_fingerprint": "0f267a6a4eaec1f5",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 두꺼운 자동문이 좌우로 스르륵 열리고 있는 틈새 너머로 두 눈을 뜬 현우의 상체.\n\nLOCATION (lock): At the access-controlled doorway between the research corridor and main center, under the facility's interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Left separating door panel in the middle-left of the frame, foreground; Right separating door panel in the middle-right of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Automatic door panels (Thick panels sliding apart to either side) — Their room-facing surfaces and inner edges flank the view of 현우 across the threshold; used as Moving lateral frame that progressively reveals his upper body; Corridor beyond the threshold (Visible behind 현우 through the opening) — Seen from inside the room looking back toward the corridor; used as Maintains spatial continuity across the reverse-position cut.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained, setting-appropriate ambient illumination across the threshold without inventing a contrast between the two spaces.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The access-controlled door is opening after a researcher card has been presented to the wall reader. 현우: He remains dressed in white recovery clothes at the entrance to the research center.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 두꺼운 자동문이 좌우로 스르륵 열리고 있는 틈새 너머로 두 눈을 뜬 현우의 상체.\n\nLOCATION (lock): At the access-controlled doorway between the research corridor and main center, under the facility's interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Left separating door panel in the middle-left of the frame, foreground; Right separating door panel in the middle-right of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Automatic door panels (Thick panels sliding apart to either side) — Their room-facing surfaces and inner edges flank the view of 현우 across the threshold; used as Moving lateral frame that progressively reveals his upper body; Corridor beyond the threshold (Visible behind 현우 through the opening) — Seen from inside the room looking back toward the corridor; used as Maintains spatial continuity across the reverse-position cut.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained, setting-appropriate ambient illumination across the threshold without inventing a contrast between the two spaces.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The access-controlled door is opening after a researcher card has been presented to the wall reader. 현우: He remains dressed in white recovery clothes at the entrance to the research center.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 두꺼운 자동문이 좌우로 스르륵 열리고 있는 틈새 너머로 두 눈을 뜬 현우의 상체.\n\nLOCATION (lock): At the access-controlled doorway between the research corridor and main center, under the facility's interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: Left separating door panel in the middle-left of the frame, foreground; Right separating door panel in the middle-right of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Automatic door panels (Thick panels sliding apart to either side) — Their room-facing surfaces and inner edges flank the view of 현우 across the threshold; used as Moving lateral frame that progressively reveals his upper body; Corridor beyond the threshold (Visible behind 현우 through the opening) — Seen from inside the room looking back toward the corridor; used as Maintains spatial continuity across the reverse-position cut.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained, setting-appropriate ambient illumination across the threshold without inventing a contrast between the two spaces.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The access-controlled door is opening after a researcher card has been presented to the wall reader. 현우: He remains dressed in white recovery clothes at the entrance to the research center.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선은 정면을 향해 있습니다.",
    "built_space": "배경은 이전 샷의 복도를 반영하고 있으나, 전경에 위치한 자동문에 거대한 가로 금속 막대가 가로지르고 있어 좌우로 열리는 문 패널의 구조를 알 수 없게 만듭니다.",
    "entities": "현우의 얼굴, 헝클어진 검은 머리, 흰색 환자복은 레퍼런스와 일치합니다.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 구조 (좌우로 열리는 자동문 앞을 가로막고 있는 정체불명의 거대한 가로 금속 패널)"
    ],
    "physics": "현우는 바닥에 안정적으로 서 있으나, 전경의 문 구조가 슬라이딩 도어로서 물리적으로 성립하지 않습니다."
   },
   {
    "label": "B",
    "direction": "현우의 시선은 정면(카메라 렌즈)을 뚜렷하게 응시하고 있습니다.",
    "built_space": "프롬프트의 지시대로 방 안에서 복도를 바라보는 카메라 위치이며, 두꺼운 자동문 패널이 좌우로 열리며 중앙에 틈새를 만들고 있습니다. 배경의 복도는 레퍼런스의 엘리베이터와 유리벽 위치 등을 정확히 유지하고 있습니다.",
    "entities": "현우의 얼굴 특징, 나이대, 머리스타일 및 흰색 옷차림이 캐릭터 레퍼런스와 훌륭하게 일치합니다.",
    "hard_violations": [],
    "physics": "좌우로 스르륵 열리는 문의 형태와 손잡이가 물리적으로 자연스러우며, 현우 역시 틈새 너머 바닥에 체중을 싣고 온전하게 서 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "프롬프트가 요구한 두꺼운 자동문이 좌우로 열리는 구도와 배경 복도의 연속성을 완벽하게 구현했으며 현우의 외모도 일치합니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "전경의 문 구조에 가로로 뻗은 거대한 금속 막대가 있어 좌우로 열리는 슬라이딩 도어의 물리적 형태를 심각하게 훼손했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선은 정면을 향해 있습니다.",
        "built_space": "배경은 이전 샷의 복도를 반영하고 있으나, 전경에 위치한 자동문에 거대한 가로 금속 막대가 가로지르고 있어 좌우로 열리는 문 패널의 구조를 알 수 없게 만듭니다.",
        "entities": "현우의 얼굴, 헝클어진 검은 머리, 흰색 환자복은 레퍼런스와 일치합니다.",
        "hard_violations": [
         "물리적으로 불가능한 구조 (좌우로 열리는 자동문 앞을 가로막고 있는 정체불명의 거대한 가로 금속 패널)"
        ],
        "physics": "현우는 바닥에 안정적으로 서 있으나, 전경의 문 구조가 슬라이딩 도어로서 물리적으로 성립하지 않습니다."
       },
       {
        "label": "B",
        "direction": "현우의 시선은 정면(카메라 렌즈)을 뚜렷하게 응시하고 있습니다.",
        "built_space": "프롬프트의 지시대로 방 안에서 복도를 바라보는 카메라 위치이며, 두꺼운 자동문 패널이 좌우로 열리며 중앙에 틈새를 만들고 있습니다. 배경의 복도는 레퍼런스의 엘리베이터와 유리벽 위치 등을 정확히 유지하고 있습니다.",
        "entities": "현우의 얼굴 특징, 나이대, 머리스타일 및 흰색 옷차림이 캐릭터 레퍼런스와 훌륭하게 일치합니다.",
        "hard_violations": [],
        "physics": "좌우로 스르륵 열리는 문의 형태와 손잡이가 물리적으로 자연스러우며, 현우 역시 틈새 너머 바닥에 체중을 싣고 온전하게 서 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "프롬프트가 요구한 두꺼운 자동문이 좌우로 열리는 구도와 배경 복도의 연속성을 완벽하게 구현했으며 현우의 외모도 일치합니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "전경의 문 구조에 가로로 뻗은 거대한 금속 막대가 있어 좌우로 열리는 슬라이딩 도어의 물리적 형태를 심각하게 훼손했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선은 정면을 향해 있습니다.",
        "built_space": "배경은 이전 샷의 복도를 반영하고 있으나, 전경에 위치한 자동문에 거대한 가로 금속 막대가 가로지르고 있어 좌우로 열리는 문 패널의 구조를 알 수 없게 만듭니다.",
        "entities": "현우의 얼굴, 헝클어진 검은 머리, 흰색 환자복은 레퍼런스와 일치합니다.",
        "hard_violations": [
         "물리적으로 불가능한 구조 (좌우로 열리는 자동문 앞을 가로막고 있는 정체불명의 거대한 가로 금속 패널)"
        ],
        "physics": "현우는 바닥에 안정적으로 서 있으나, 전경의 문 구조가 슬라이딩 도어로서 물리적으로 성립하지 않습니다."
       },
       {
        "label": "B",
        "direction": "현우의 시선은 정면(카메라 렌즈)을 뚜렷하게 응시하고 있습니다.",
        "built_space": "프롬프트의 지시대로 방 안에서 복도를 바라보는 카메라 위치이며, 두꺼운 자동문 패널이 좌우로 열리며 중앙에 틈새를 만들고 있습니다. 배경의 복도는 레퍼런스의 엘리베이터와 유리벽 위치 등을 정확히 유지하고 있습니다.",
        "entities": "현우의 얼굴 특징, 나이대, 머리스타일 및 흰색 옷차림이 캐릭터 레퍼런스와 훌륭하게 일치합니다.",
        "hard_violations": [],
        "physics": "좌우로 스르륵 열리는 문의 형태와 손잡이가 물리적으로 자연스러우며, 현우 역시 틈새 너머 바닥에 체중을 싣고 온전하게 서 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "좌우 전경의 두꺼운 문 사이로 현우의 상체와 뒤쪽 복도가 드러나는 미디엄 숏을 더 충실히 구현하지만, 문 손잡이와 불투명 유리의 연속성은 불확실하다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "인물과 시설 분위기는 유지했지만, 문틈이 지나치게 좁아 어깨와 상체 대부분을 가리고 두 문 가장자리도 중앙에 몰려 지정된 구도가 약해진다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 두 눈을 뜨고 문 너머 실내의 카메라 쪽을 정면으로 바라본다. 특정 상대는 보이지 않는다. 두 문짝 사이에는 좌우로 벌어진 틈이 있으나 정지 화면만으로 이동 방향 자체를 확인할 수는 없다. 무기나 방향성 있는 휴대 물체는 없다.",
        "built_space": "전경에 문짝 두 개가 있고 각각의 두꺼운 금속 안쪽 테두리가 화면 중간 왼쪽과 오른쪽을 차지한다. 각 문짝에 세로 손잡이 하나씩, 총 두 개가 보이며 바깥쪽 유리 하단은 불투명하다. 현우는 문 뒤 문턱 부근에 서 있고 그 뒤로 천장등이 이어지는 복도와 먼 창이 보인다. 금속 문틀, 밝은 벽, 오른쪽 실내 유리창은 참고 장소와 유사하지만 전경 문의 손잡이와 불투명 처리까지 같은 설비인지는 확인되지 않는다. 불가능한 거울 반사는 보이지 않는다.",
        "entities": "사람은 현우 한 명뿐이다. 앳된 동아시아계 남성의 외모, 헝클어진 검은 머리, 얼굴 윤곽과 피부 질감이 참고 인물에 가깝다. 정확한 나이와 국적은 외모만으로 확인할 수 없다. 이전 숏과 같은 흰색 둥근 목 상의를 입었고 두 눈의 홍채와 동공은 정상적으로 보인다. 연구원 카드와 판독기는 뚜렷이 보이지 않지만 카드 제시 이후의 순간과 모순되지는 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우의 하체와 발은 프레임 밖이지만 몸통은 문 뒤에서 자연스럽게 직립하며 공중에 떠 있다는 징후는 없다. 문짝은 건축 문틀 안에 설치되어 지지되는 것으로 보이고 손잡이도 문에 고정되어 있다. 옷은 어깨와 몸통에 걸려 자연스럽게 주름진다. 다만 자세는 다소 정면으로 고정된 느낌이다."
       },
       {
        "label": "B",
        "direction": "현우는 정상적인 두 눈을 뜨고 문 너머 카메라 쪽을 바라본다. 문짝 두 개는 중앙에서 조금 떨어져 있으며 좌우로 열리는 초기 순간으로 해석할 수 있다. 별도의 시선 대상이나 손에 든 방향성 물체는 없다.",
        "built_space": "전경에는 금속 테두리의 유리 문짝 두 개가 있고 각 문짝에 넓은 가로 금속 띠가 하나씩 있다. 손잡이는 보이지 않는다. 두 안쪽 테두리가 중앙에 가까이 몰려 매우 좁은 틈을 만들며, 그 뒤 현우의 양어깨와 몸통 상당 부분을 가린다. 뒤쪽에는 왼쪽 금속 출입문, 복도의 유리 문틀, 오른쪽 실내창, 반복되는 천장등과 먼 창이 보인다. 참고 장소의 재료와 복도 분위기는 가깝지만 전경의 넓은 금속 띠는 참고에서 확인되지 않는 차이다. 유리 너머 공간과 반사에 명백한 광학적 모순은 없다.",
        "entities": "보이는 사람은 현우 한 명이다. 젊은 동아시아계 남성의 얼굴, 검은 헝클어진 머리와 흰색 둥근 목 상의가 참고와 대체로 일치한다. 정확한 나이와 국적은 시각적으로 확정할 수 없다. 양쪽 눈은 열려 있고 비정상적인 변형은 없다. 카드와 판독기는 명확히 식별되지 않으며 추가 인물이나 읽을 수 있는 문구는 없다.",
        "hard_violations": [],
        "physics": "현우는 문 뒤에서 수직으로 서 있으며 발은 화면 밖이다. 오른쪽 유리 아래로 보이는 팔 일부도 같은 인물의 위치와 양립한다. 부유하거나 신체가 문을 관통하는 명백한 징후는 없다. 유리와 금속 띠는 문틀에 결합되어 지지되고, 문짝도 출입구 구조 안에 놓여 있다. 자세는 다소 경직되어 있지만 물리적으로 불가능하지는 않다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "좌우 전경의 두꺼운 문 사이로 현우의 상체와 뒤쪽 복도가 드러나는 미디엄 숏을 더 충실히 구현하지만, 문 손잡이와 불투명 유리의 연속성은 불확실하다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "인물과 시설 분위기는 유지했지만, 문틈이 지나치게 좁아 어깨와 상체 대부분을 가리고 두 문 가장자리도 중앙에 몰려 지정된 구도가 약해진다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 두 눈을 뜨고 문 너머 실내의 카메라 쪽을 정면으로 바라본다. 특정 상대는 보이지 않는다. 두 문짝 사이에는 좌우로 벌어진 틈이 있으나 정지 화면만으로 이동 방향 자체를 확인할 수는 없다. 무기나 방향성 있는 휴대 물체는 없다.",
        "built_space": "전경에 문짝 두 개가 있고 각각의 두꺼운 금속 안쪽 테두리가 화면 중간 왼쪽과 오른쪽을 차지한다. 각 문짝에 세로 손잡이 하나씩, 총 두 개가 보이며 바깥쪽 유리 하단은 불투명하다. 현우는 문 뒤 문턱 부근에 서 있고 그 뒤로 천장등이 이어지는 복도와 먼 창이 보인다. 금속 문틀, 밝은 벽, 오른쪽 실내 유리창은 참고 장소와 유사하지만 전경 문의 손잡이와 불투명 처리까지 같은 설비인지는 확인되지 않는다. 불가능한 거울 반사는 보이지 않는다.",
        "entities": "사람은 현우 한 명뿐이다. 앳된 동아시아계 남성의 외모, 헝클어진 검은 머리, 얼굴 윤곽과 피부 질감이 참고 인물에 가깝다. 정확한 나이와 국적은 외모만으로 확인할 수 없다. 이전 숏과 같은 흰색 둥근 목 상의를 입었고 두 눈의 홍채와 동공은 정상적으로 보인다. 연구원 카드와 판독기는 뚜렷이 보이지 않지만 카드 제시 이후의 순간과 모순되지는 않는다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우의 하체와 발은 프레임 밖이지만 몸통은 문 뒤에서 자연스럽게 직립하며 공중에 떠 있다는 징후는 없다. 문짝은 건축 문틀 안에 설치되어 지지되는 것으로 보이고 손잡이도 문에 고정되어 있다. 옷은 어깨와 몸통에 걸려 자연스럽게 주름진다. 다만 자세는 다소 정면으로 고정된 느낌이다."
       },
       {
        "label": "A",
        "direction": "현우는 정상적인 두 눈을 뜨고 문 너머 카메라 쪽을 바라본다. 문짝 두 개는 중앙에서 조금 떨어져 있으며 좌우로 열리는 초기 순간으로 해석할 수 있다. 별도의 시선 대상이나 손에 든 방향성 물체는 없다.",
        "built_space": "전경에는 금속 테두리의 유리 문짝 두 개가 있고 각 문짝에 넓은 가로 금속 띠가 하나씩 있다. 손잡이는 보이지 않는다. 두 안쪽 테두리가 중앙에 가까이 몰려 매우 좁은 틈을 만들며, 그 뒤 현우의 양어깨와 몸통 상당 부분을 가린다. 뒤쪽에는 왼쪽 금속 출입문, 복도의 유리 문틀, 오른쪽 실내창, 반복되는 천장등과 먼 창이 보인다. 참고 장소의 재료와 복도 분위기는 가깝지만 전경의 넓은 금속 띠는 참고에서 확인되지 않는 차이다. 유리 너머 공간과 반사에 명백한 광학적 모순은 없다.",
        "entities": "보이는 사람은 현우 한 명이다. 젊은 동아시아계 남성의 얼굴, 검은 헝클어진 머리와 흰색 둥근 목 상의가 참고와 대체로 일치한다. 정확한 나이와 국적은 시각적으로 확정할 수 없다. 양쪽 눈은 열려 있고 비정상적인 변형은 없다. 카드와 판독기는 명확히 식별되지 않으며 추가 인물이나 읽을 수 있는 문구는 없다.",
        "hard_violations": [],
        "physics": "현우는 문 뒤에서 수직으로 서 있으며 발은 화면 밖이다. 오른쪽 유리 아래로 보이는 팔 일부도 같은 인물의 위치와 양립한다. 부유하거나 신체가 문을 관통하는 명백한 징후는 없다. 유리와 금속 띠는 문틀에 결합되어 지지되고, 문짝도 출입구 구조 안에 놓여 있다. 자세는 다소 경직되어 있지만 물리적으로 불가능하지는 않다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.208,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.958,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 물리적으로 불가능한 구조 (좌우로 열리는 자동문 앞을 가로막고 있는 정체불명의 거대한 가로 금속 패널)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 958
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "프롬프트가 요구한 두꺼운 자동문이 좌우로 열리는 구도와 배경 복도의 연속성을 완벽하게 구현했으며 현우의 외모도 일치합니다."
   },
   {
    "label": "A",
    "score": 958,
    "verdict_ko": "전경의 문 구조에 가로로 뻗은 거대한 금속 막대가 있어 좌우로 열리는 슬라이딩 도어의 물리적 형태를 심각하게 훼손했습니다.  ★위반: [gemini-pro] 물리적으로 불가능한 구조 (좌우로 열리는 자동문 앞을 가로막고 있는 정체불명의 거대한 가로 금속 패널)"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S81sh13_sel.png",
    "asset_id": "071967d2-d700-420c-8a22-a7667a02ec5b",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-e80a-79ef-9558-26df6d6c5f35",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S81sh13"
  }
 },
 "S81sh16::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T09:04:56.254259+00:00",
  "fingerprint": "e15e50d5d3adeac804c7172409815a7f8800149f04dfa3babbf3f9e6d6ca67fa",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S81sh16_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S81sh16_sel.png",
  "source_sha256": "45cee1cb8d11a41d287542d40237ffc81a078d00203456f571f8516351bb55b9",
  "file": "S81sh16_cine.png",
  "staged_sha256": "3f30f7fe84cb1b10cd89451ae2deeb2036b3e65b683b5ed548e49f94c3071705",
  "latency_ms": 9099
 },
 "S82sh14::signage": {
  "fp": "1b297fe9484dc4df",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S82sh14": {
  "input_fingerprint": "519903354cb5438a",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 실험실 중앙의 차가운 스테인리스 침대 위로 미동 없이 축 늘어져 누운 찰리의 거대한 전신.\n\nLOCATION (lock): Inside the circular glass enclosure of the main center's adjoining laboratory, on a stainless-steel examination bed under cool laboratory lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Circular glass enclosure (Separates the central bed from the surrounding researchers) — 찰리, the bed, and the scanning arms are visible through the near section of glass; used as Establishes the physical barrier before the camera approaches 현우; Stainless-steel bed (Supports 찰리's motionless body) — Its long side runs diagonally across the elevated view; used as Central support and scale reference; Artificial-intelligence scanning arms (Scanning different areas of 찰리's body) — Articulated sections approach the body from different positions around the bed; used as Surround the still figure with purposeful mechanical activity; Peripheral computer workstations (In use by roughly ten researchers, each engaged at a separate computer) — Seen obliquely around the outer laboratory, without emphasis on screen contents; used as Peripheral scale and asynchronous background activity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral laboratory ambient illumination and controlled tonal contrast reveal the stainless-steel bed and scanning equipment without theatrical highlights.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent and motionless on the stainless-steel examination bed inside the circular glass enclosure, his body supported by the bed while robotic arms scan him. The source does not specify his head's direction, his torso's upward-facing surface, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie lies inactive on a stainless-steel bed inside a circular glass enclosure, with hardware damaged by prolonged immersion and multiple wires already attached to the body. Robotic arms scan the body amid computers and research equipment.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 실험실 중앙의 차가운 스테인리스 침대 위로 미동 없이 축 늘어져 누운 찰리의 거대한 전신.\n\nLOCATION (lock): Inside the circular glass enclosure of the main center's adjoining laboratory, on a stainless-steel examination bed under cool laboratory lighting. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Circular glass enclosure (Separates the central bed from the surrounding researchers) — 찰리, the bed, and the scanning arms are visible through the near section of glass; used as Establishes the physical barrier before the camera approaches 현우; Stainless-steel bed (Supports 찰리's motionless body) — Its long side runs diagonally across the elevated view; used as Central support and scale reference; Artificial-intelligence scanning arms (Scanning different areas of 찰리's body) — Articulated sections approach the body from different positions around the bed; used as Surround the still figure with purposeful mechanical activity; Peripheral computer workstations (In use by roughly ten researchers, each engaged at a separate computer) — Seen obliquely around the outer laboratory, without emphasis on screen contents; used as Peripheral scale and asynchronous background activity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral laboratory ambient illumination and controlled tonal contrast reveal the stainless-steel bed and scanning equipment without theatrical highlights.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent and motionless on the stainless-steel examination bed inside the circular glass enclosure, his body supported by the bed while robotic arms scan him. The source does not specify his head's direction, his torso's upward-facing surface, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie lies inactive on a stainless-steel bed inside a circular glass enclosure, with hardware damaged by prolonged immersion and multiple wires already attached to the body. Robotic arms scan the body amid computers and research equipment.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 실험실 중앙의 차가운 스테인리스 침대 위로 미동 없이 축 늘어져 누운 찰리의 거대한 전신.\n\nLOCATION (lock): Inside the circular glass enclosure of the main center's adjoining laboratory, on a stainless-steel examination bed under cool laboratory lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Circular glass enclosure (Separates the central bed from the surrounding researchers) — 찰리, the bed, and the scanning arms are visible through the near section of glass; used as Establishes the physical barrier before the camera approaches 현우; Stainless-steel bed (Supports 찰리's motionless body) — Its long side runs diagonally across the elevated view; used as Central support and scale reference; Artificial-intelligence scanning arms (Scanning different areas of 찰리's body) — Articulated sections approach the body from different positions around the bed; used as Surround the still figure with purposeful mechanical activity; Peripheral computer workstations (In use by roughly ten researchers, each engaged at a separate computer) — Seen obliquely around the outer laboratory, without emphasis on screen contents; used as Peripheral scale and asynchronous background activity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral laboratory ambient illumination and controlled tonal contrast reveal the stainless-steel bed and scanning equipment without theatrical highlights.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent and motionless on the stainless-steel examination bed inside the circular glass enclosure, his body supported by the bed while robotic arms scan him. The source does not specify his head's direction, his torso's upward-facing surface, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie lies inactive on a stainless-steel bed inside a circular glass enclosure, with hardware damaged by prolonged immersion and multiple wires already attached to the body. Robotic arms scan the body amid computers and research equipment.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S82sh14__bgfirst_bg.png",
     "asset_id": "09c1e32d-ad58-4bd4-a51a-d8b7432ecb22",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S82sh14.png",
     "asset_id": "90735080-f327-4c19-9101-6d98f10c3031",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L268B01.png",
     "asset_id": "a491e625-4c37-4571-b3ba-85a3daf996f9",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "로봇 팔들이 침대 위 찰리를 향해 있고, 주변 연구원들은 각자의 모니터를 바라봄.",
    "built_space": "원형 유리 구조물을 위에서 내려다보는 부감(elevated view). 스테인리스 침대가 대각선으로 놓였고 가장자리에 약 10명의 연구원이 배치됨.",
    "entities": "샌드 베이지색 장갑과 마스크를 쓴 찰리, 로봇 팔, 연구원과 컴퓨터 등 프롬프트의 모든 요소가 일치함.",
    "hard_violations": [
     "[gemini-pro] 화면 곳곳(모니터와 인물 머리 근처)에 정체불명의 흰색 화살표 마커(데이터셋 잔재)가 허공에 떠 있음.",
     "[gpt-high] 연구원들의 손과 작업대 부근 여러 곳에 물리적 물체에 속하지 않는 흰 화살표·포인터 모양 표식이 남아 있어 도식 표식 및 오버레이 금지 조건을 위반한다."
    ],
    "physics": "찰리의 거대한 몸체는 침대에 완전히 지탱되어 누워 있으며, 연구원들과 로봇 팔도 바닥과 의자에 정상적으로 지지를 받음."
   },
   {
    "label": "B",
    "direction": "로봇 팔들이 찰리를 향하고 배경의 연구원들은 작업을 진행 중임.",
    "built_space": "명시된 부감이 아닌 지면 높이의 샷이며, 침대 방향이 가로이고 원형 유리 구조가 뚜렷하지 않음.",
    "entities": "찰리의 외형은 기준과 일치하며, 로봇 팔과 일부 연구원이 확인됨.",
    "hard_violations": [
     "[gemini-pro] 프롬프트가 요구한 부감(elevated view)을 무시하고 평면적인 지면 시점으로 렌더링됨.",
     "[gemini-pro] 우측 로봇 팔이 유리벽을 물리적으로 불가능하게 통과하여 뻗어 있음.",
     "[gpt-high] 찰리의 상체와 머리가 평평한 침대에서 상당히 일어나 있지만, 이를 받치는 등받이·받침대·장비가 보이지 않는다. 무의식 상태로 완전히 지지되어 누워 있어야 하는 몸이 스스로 자세를 유지하는 형태다.",
     "[gpt-high] 오른쪽 끝 배경 모니터에 읽을 수 있는 영문 경고 글자가 남아 있어 글자 금지 조건을 위반한다."
    ],
    "physics": "찰리의 몸은 침대가 지탱하며 우측 팔은 중력에 의해 가장자리 아래로 자연스럽게 늘어져 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "지정된 부감과 대각선 구도를 완벽히 구현했으나, 화면 곳곳에 허공을 가리키는 흰색 화살표(마커)가 노출되어 치명적인 오류를 범했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "요구된 부감을 무시하고 지면 시점으로 연출했으며, 우측 로봇 팔이 유리를 통과하는 물리적 구조 오류가 있습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "로봇 팔들이 침대 위 찰리를 향해 있고, 주변 연구원들은 각자의 모니터를 바라봄.",
        "built_space": "원형 유리 구조물을 위에서 내려다보는 부감(elevated view). 스테인리스 침대가 대각선으로 놓였고 가장자리에 약 10명의 연구원이 배치됨.",
        "entities": "샌드 베이지색 장갑과 마스크를 쓴 찰리, 로봇 팔, 연구원과 컴퓨터 등 프롬프트의 모든 요소가 일치함.",
        "hard_violations": [
         "화면 곳곳(모니터와 인물 머리 근처)에 정체불명의 흰색 화살표 마커(데이터셋 잔재)가 허공에 떠 있음."
        ],
        "physics": "찰리의 거대한 몸체는 침대에 완전히 지탱되어 누워 있으며, 연구원들과 로봇 팔도 바닥과 의자에 정상적으로 지지를 받음."
       },
       {
        "label": "B",
        "direction": "로봇 팔들이 찰리를 향하고 배경의 연구원들은 작업을 진행 중임.",
        "built_space": "명시된 부감이 아닌 지면 높이의 샷이며, 침대 방향이 가로이고 원형 유리 구조가 뚜렷하지 않음.",
        "entities": "찰리의 외형은 기준과 일치하며, 로봇 팔과 일부 연구원이 확인됨.",
        "hard_violations": [
         "프롬프트가 요구한 부감(elevated view)을 무시하고 평면적인 지면 시점으로 렌더링됨.",
         "우측 로봇 팔이 유리벽을 물리적으로 불가능하게 통과하여 뻗어 있음."
        ],
        "physics": "찰리의 몸은 침대가 지탱하며 우측 팔은 중력에 의해 가장자리 아래로 자연스럽게 늘어져 있음."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "지정된 부감과 대각선 구도를 완벽히 구현했으나, 화면 곳곳에 허공을 가리키는 흰색 화살표(마커)가 노출되어 치명적인 오류를 범했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "요구된 부감을 무시하고 지면 시점으로 연출했으며, 우측 로봇 팔이 유리를 통과하는 물리적 구조 오류가 있습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "로봇 팔들이 침대 위 찰리를 향해 있고, 주변 연구원들은 각자의 모니터를 바라봄.",
        "built_space": "원형 유리 구조물을 위에서 내려다보는 부감(elevated view). 스테인리스 침대가 대각선으로 놓였고 가장자리에 약 10명의 연구원이 배치됨.",
        "entities": "샌드 베이지색 장갑과 마스크를 쓴 찰리, 로봇 팔, 연구원과 컴퓨터 등 프롬프트의 모든 요소가 일치함.",
        "hard_violations": [
         "화면 곳곳(모니터와 인물 머리 근처)에 정체불명의 흰색 화살표 마커(데이터셋 잔재)가 허공에 떠 있음."
        ],
        "physics": "찰리의 거대한 몸체는 침대에 완전히 지탱되어 누워 있으며, 연구원들과 로봇 팔도 바닥과 의자에 정상적으로 지지를 받음."
       },
       {
        "label": "B",
        "direction": "로봇 팔들이 찰리를 향하고 배경의 연구원들은 작업을 진행 중임.",
        "built_space": "명시된 부감이 아닌 지면 높이의 샷이며, 침대 방향이 가로이고 원형 유리 구조가 뚜렷하지 않음.",
        "entities": "찰리의 외형은 기준과 일치하며, 로봇 팔과 일부 연구원이 확인됨.",
        "hard_violations": [
         "프롬프트가 요구한 부감(elevated view)을 무시하고 평면적인 지면 시점으로 렌더링됨.",
         "우측 로봇 팔이 유리벽을 물리적으로 불가능하게 통과하여 뻗어 있음."
        ],
        "physics": "찰리의 몸은 침대가 지탱하며 우측 팔은 중력에 의해 가장자리 아래로 자연스럽게 늘어져 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "찰리의 육중한 체형은 가깝지만, 받침 없이 상체를 일으킨 자세가 축 늘어진 부동 상태를 위반하며 주변부의 읽히는 경고 글자도 부적합하다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "유리 너머의 높은 와이드 시점과 침대에 전신을 맡긴 자세는 더 충실하지만, 흰 화살표 표식이 남아 있고 찰리의 체형과 장비 배치는 참조와 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 얼굴은 위쪽을 향하며 특정 대상을 응시하지 않는다. 뒤쪽 스캐닝 팔은 머리와 상부 몸통 쪽으로, 오른쪽 장비 끝은 어깨 쪽으로 향한다. 배경 연구원들은 대체로 각자 앞의 모니터를 보고 있으며 화면도 작업자 쪽을 향한다.",
        "built_space": "원형 유리벽 안에 금속 침대 한 대와 양쪽의 스캐닝 장비 두 대가 보인다. 침대의 긴 축은 왼쪽 아래에서 오른쪽 위로 놓인다. 유리 밖에는 약 일곱 명의 연구원과 여러 작업대가 보인다. 참조의 창문, 벽면 모니터, 밝은 실험실 재료는 비교적 잘 유지되지만, 침대와 몸이 화면을 크게 채워 원형 구획 전체와 주변 작업대의 관계는 덜 드러난다. 명백히 불가능한 반사는 보이지 않는다.",
        "entities": "찰리는 흰 마스크형 얼굴, 각진 샌드 베이지 장갑, 넓은 몸통, 육중한 긴 팔과 짧은 다리를 갖추어 참조의 정체성에 가깝다. 장갑의 오염과 마모, 몸에 연결된 여러 케이블도 보인다. 금속 침대와 관절식 스캐너는 실물 장비처럼 표현된다. 배경에는 성인 남녀 연구원들이 보이며 일부는 동아시아계 외모로 보인다. 오른쪽 끝 배경 화면에는 영문 경고 글자의 일부가 읽힌다.",
        "hard_violations": [
         "찰리의 상체와 머리가 평평한 침대에서 상당히 일어나 있지만, 이를 받치는 등받이·받침대·장비가 보이지 않는다. 무의식 상태로 완전히 지지되어 누워 있어야 하는 몸이 스스로 자세를 유지하는 형태다.",
         "오른쪽 끝 배경 모니터에 읽을 수 있는 영문 경고 글자가 남아 있어 글자 금지 조건을 위반한다."
        ],
        "physics": "침대는 바닥의 금속 지지부에 놓이고 골반과 다리를 받친다. 화면 오른쪽 팔과 손은 어깨에서 아래로 늘어져 있어 그 자체는 중력에 맞는다. 그러나 들어 올려진 등·어깨·머리를 지지하는 것은 보이지 않는다. 스캐닝 팔은 기계 관절과 받침에 연결되어 있으며 케이블은 몸과 장비 사이에서 처지거나 바닥에 놓인다."
       },
       {
        "label": "B",
        "direction": "찰리의 얼굴은 천장을 향하고 몸은 움직임 없이 누워 있다. 머리 뒤쪽 스캐너 끝은 머리와 상부 몸통 쪽으로 내려오고, 오른쪽 스캐너는 몸통 측면으로 향한다. 왼쪽 앞 스캐너 끝은 가까운 팔 옆의 침대 가장자리 쪽을 향해 실제 신체 부위에 대한 조준은 덜 명확하다. 연구원들은 모니터나 손에 든 판형 기기를 보고 있으며 기능 면도 사용자 쪽을 향한다.",
        "built_space": "전경의 곡면 유리를 통해 원형 구획, 중앙 침대 한 대, 독립 받침대에 설치된 스캐닝 팔 세 대가 보인다. 높은 사선 시점에서 침대의 긴 변이 화면을 대각선으로 가로지르며 전신이 들어온다. 유리 밖에는 전경 인물을 포함해 약 열 명이 보이고, 대부분 개별 컴퓨터 앞에 앉아 있으며 한 명은 서서 기기를 본다. 참조의 원형 유리벽과 바닥 환형 홈은 유지하지만, 참조에서 보이는 스캐너 두 대가 세 대로 바뀌고 벽면의 대형 화면 배열도 생략됐다. 명백히 불가능한 반사는 보이지 않는다.",
        "entities": "찰리는 베이지 장갑과 흰 기계식 얼굴을 갖췄지만, 참조보다 어깨와 팔이 가늘고 다리가 길어 고릴라형의 거대한 체격이 약하다. 장갑의 손상과 오염, 몸 주변에 연결된 여러 선은 보인다. 스테인리스 침대, 관절식 스캐너, 컴퓨터 작업대가 모두 식별된다. 연구원들은 주로 동아시아계 외모의 성인 남녀로 보인다. 여러 연구원과 작업대 주변에 작은 흰 화살표 모양 표식이 떠 있다.",
        "hard_violations": [
         "연구원들의 손과 작업대 부근 여러 곳에 물리적 물체에 속하지 않는 흰 화살표·포인터 모양 표식이 남아 있어 도식 표식 및 오버레이 금지 조건을 위반한다."
        ],
        "physics": "찰리의 머리, 등, 골반, 팔과 다리가 침대 표면에 놓이고 손도 몸 옆의 침대에 내려져 있다. 발은 발뒤꿈치 쪽으로 지지되며, 상체를 스스로 일으키거나 지지 없이 떠 있는 부분은 보이지 않는다. 침대와 세 스캐너는 각각 바닥에 닿는 받침을 갖고, 케이블은 장비에 연결되어 자연스럽게 처진다. 연구원들은 의자에 앉거나 바닥에 서 있고, 서 있는 인물의 판형 기기는 손으로 지지된다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "찰리의 육중한 체형은 가깝지만, 받침 없이 상체를 일으킨 자세가 축 늘어진 부동 상태를 위반하며 주변부의 읽히는 경고 글자도 부적합하다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "유리 너머의 높은 와이드 시점과 침대에 전신을 맡긴 자세는 더 충실하지만, 흰 화살표 표식이 남아 있고 찰리의 체형과 장비 배치는 참조와 다르다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 얼굴은 위쪽을 향하며 특정 대상을 응시하지 않는다. 뒤쪽 스캐닝 팔은 머리와 상부 몸통 쪽으로, 오른쪽 장비 끝은 어깨 쪽으로 향한다. 배경 연구원들은 대체로 각자 앞의 모니터를 보고 있으며 화면도 작업자 쪽을 향한다.",
        "built_space": "원형 유리벽 안에 금속 침대 한 대와 양쪽의 스캐닝 장비 두 대가 보인다. 침대의 긴 축은 왼쪽 아래에서 오른쪽 위로 놓인다. 유리 밖에는 약 일곱 명의 연구원과 여러 작업대가 보인다. 참조의 창문, 벽면 모니터, 밝은 실험실 재료는 비교적 잘 유지되지만, 침대와 몸이 화면을 크게 채워 원형 구획 전체와 주변 작업대의 관계는 덜 드러난다. 명백히 불가능한 반사는 보이지 않는다.",
        "entities": "찰리는 흰 마스크형 얼굴, 각진 샌드 베이지 장갑, 넓은 몸통, 육중한 긴 팔과 짧은 다리를 갖추어 참조의 정체성에 가깝다. 장갑의 오염과 마모, 몸에 연결된 여러 케이블도 보인다. 금속 침대와 관절식 스캐너는 실물 장비처럼 표현된다. 배경에는 성인 남녀 연구원들이 보이며 일부는 동아시아계 외모로 보인다. 오른쪽 끝 배경 화면에는 영문 경고 글자의 일부가 읽힌다.",
        "hard_violations": [
         "찰리의 상체와 머리가 평평한 침대에서 상당히 일어나 있지만, 이를 받치는 등받이·받침대·장비가 보이지 않는다. 무의식 상태로 완전히 지지되어 누워 있어야 하는 몸이 스스로 자세를 유지하는 형태다.",
         "오른쪽 끝 배경 모니터에 읽을 수 있는 영문 경고 글자가 남아 있어 글자 금지 조건을 위반한다."
        ],
        "physics": "침대는 바닥의 금속 지지부에 놓이고 골반과 다리를 받친다. 화면 오른쪽 팔과 손은 어깨에서 아래로 늘어져 있어 그 자체는 중력에 맞는다. 그러나 들어 올려진 등·어깨·머리를 지지하는 것은 보이지 않는다. 스캐닝 팔은 기계 관절과 받침에 연결되어 있으며 케이블은 몸과 장비 사이에서 처지거나 바닥에 놓인다."
       },
       {
        "label": "A",
        "direction": "찰리의 얼굴은 천장을 향하고 몸은 움직임 없이 누워 있다. 머리 뒤쪽 스캐너 끝은 머리와 상부 몸통 쪽으로 내려오고, 오른쪽 스캐너는 몸통 측면으로 향한다. 왼쪽 앞 스캐너 끝은 가까운 팔 옆의 침대 가장자리 쪽을 향해 실제 신체 부위에 대한 조준은 덜 명확하다. 연구원들은 모니터나 손에 든 판형 기기를 보고 있으며 기능 면도 사용자 쪽을 향한다.",
        "built_space": "전경의 곡면 유리를 통해 원형 구획, 중앙 침대 한 대, 독립 받침대에 설치된 스캐닝 팔 세 대가 보인다. 높은 사선 시점에서 침대의 긴 변이 화면을 대각선으로 가로지르며 전신이 들어온다. 유리 밖에는 전경 인물을 포함해 약 열 명이 보이고, 대부분 개별 컴퓨터 앞에 앉아 있으며 한 명은 서서 기기를 본다. 참조의 원형 유리벽과 바닥 환형 홈은 유지하지만, 참조에서 보이는 스캐너 두 대가 세 대로 바뀌고 벽면의 대형 화면 배열도 생략됐다. 명백히 불가능한 반사는 보이지 않는다.",
        "entities": "찰리는 베이지 장갑과 흰 기계식 얼굴을 갖췄지만, 참조보다 어깨와 팔이 가늘고 다리가 길어 고릴라형의 거대한 체격이 약하다. 장갑의 손상과 오염, 몸 주변에 연결된 여러 선은 보인다. 스테인리스 침대, 관절식 스캐너, 컴퓨터 작업대가 모두 식별된다. 연구원들은 주로 동아시아계 외모의 성인 남녀로 보인다. 여러 연구원과 작업대 주변에 작은 흰 화살표 모양 표식이 떠 있다.",
        "hard_violations": [
         "연구원들의 손과 작업대 부근 여러 곳에 물리적 물체에 속하지 않는 흰 화살표·포인터 모양 표식이 남아 있어 도식 표식 및 오버레이 금지 조건을 위반한다."
        ],
        "physics": "찰리의 머리, 등, 골반, 팔과 다리가 침대 표면에 놓이고 손도 몸 옆의 침대에 내려져 있다. 발은 발뒤꿈치 쪽으로 지지되며, 상체를 스스로 일으키거나 지지 없이 떠 있는 부분은 보이지 않는다. 침대와 세 스캐너는 각각 바닥에 닿는 받침을 갖고, 케이블은 장비에 연결되어 자연스럽게 처진다. 연구원들은 의자에 앉거나 바닥에 서 있고, 서 있는 인물의 판형 기기는 손으로 지지된다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.5
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.25
   },
   "violations": {
    "A": [
     "[gemini-pro] 화면 곳곳(모니터와 인물 머리 근처)에 정체불명의 흰색 화살표 마커(데이터셋 잔재)가 허공에 떠 있음.",
     "[gpt-high] 연구원들의 손과 작업대 부근 여러 곳에 물리적 물체에 속하지 않는 흰 화살표·포인터 모양 표식이 남아 있어 도식 표식 및 오버레이 금지 조건을 위반한다."
    ],
    "B": [
     "[gemini-pro] 프롬프트가 요구한 부감(elevated view)을 무시하고 평면적인 지면 시점으로 렌더링됨.",
     "[gemini-pro] 우측 로봇 팔이 유리벽을 물리적으로 불가능하게 통과하여 뻗어 있음.",
     "[gpt-high] 찰리의 상체와 머리가 평평한 침대에서 상당히 일어나 있지만, 이를 받치는 등받이·받침대·장비가 보이지 않는다. 무의식 상태로 완전히 지지되어 누워 있어야 하는 몸이 스스로 자세를 유지하는 형태다.",
     "[gpt-high] 오른쪽 끝 배경 모니터에 읽을 수 있는 영문 경고 글자가 남아 있어 글자 금지 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 1250
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "지정된 부감과 대각선 구도를 완벽히 구현했으나, 화면 곳곳에 허공을 가리키는 흰색 화살표(마커)가 노출되어 치명적인 오류를 범했습니다.  ★위반: [gemini-pro] 화면 곳곳(모니터와 인물 머리 근처)에 정체불명의 흰색 화살표 마커(데이터셋 잔재)가 허공에 떠 있음. / [gpt-high] 연구원들의 손과 작업대 부근 여러 곳에 물리적 물체에 속하지 않는 흰 화살표·포인터 모양 표식이 남아 있어 도식 표식 및 오버레이 금지 조건을 위반한다."
   },
   {
    "label": "B",
    "score": 1250,
    "verdict_ko": "요구된 부감을 무시하고 지면 시점으로 연출했으며, 우측 로봇 팔이 유리를 통과하는 물리적 구조 오류가 있습니다.  ★위반: [gemini-pro] 프롬프트가 요구한 부감(elevated view)을 무시하고 평면적인 지면 시점으로 렌더링됨. / [gemini-pro] 우측 로봇 팔이 유리벽을 물리적으로 불가능하게 통과하여 뻗어 있음. / [gpt-high] 찰리의 상체와 머리가 평평한 침대에서 상당히 일어나 있지만, 이를 받치는 등받이·받침대·장비가 보이지 않는다. 무의식 상태로 완전히 지지되어 누워 있어야 하는 몸이 스스로 자세를 유지하는 형태다. / [gpt-high] 오른쪽 끝 배경 모니터에 읽을 수 있는 영문 경고 글자가 남아 있어 글자 금지 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L268B01.png",
    "asset_id": "a491e625-4c37-4571-b3ba-85a3daf996f9",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-e9b7-7b28-b174-2e4dcef1e696",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S82sh14__bgfirst_bg.png",
   "bg_asset_id": "09c1e32d-ad58-4bd4-a51a-d8b7432ecb22",
   "bg_record_key": "S82sh14::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S82sh14::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T12:07:05.972362+00:00",
  "fingerprint": "36ab4f80f49b433316470c730b84a7a583ab715cc78a87df68573645936b5f1c",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S82sh14_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S82sh14_sel.png",
  "source_sha256": "8cc39508d3d64184ccd5a2544c7e7f4948c9e9524fa458c69d4f47a8031b5540",
  "file": "S82sh14_cine.png",
  "staged_sha256": "dfcbe4d46936280b39f26eea2099bee00f1ac3cbc4365c5aea5263cdb4209400",
  "latency_ms": 10330
 },
 "S82sh17::signage": {
  "fp": "6c736eeca5a08bea",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S82sh17": {
  "input_fingerprint": "56ce7abf43c98429",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 차가운 유리 벽에 양손을 바짝 대고 찰리를 뚫어져라 내려다보는 현우의 절박한 얼굴.\n\nLOCATION (lock): At the observation side of the circular glass laboratory enclosure, with cool light over the examination bed beyond. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Glass barrier between 현우 and 찰리 in the middle-center of the frame, midground; Bed supporting 찰리 beyond the glass in the lower-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Circular glass wall (Separates 현우's pressed hands from 찰리) — Seen obliquely along its curvature, with 찰리 and part of the bed visible through it; used as Physical separation within the intimate composition; Stainless-steel bed (Partially visible beneath 찰리) — A small portion of the near side appears beyond the glass at lower right; used as Anchors the downward eyeline and the body's depth.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the laboratory's ambient tonal balance, allowing clear transmission through the glass rather than emphasizing reflected light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the stainless-steel bed, circular glass enclosure, scanning equipment, and cool laboratory lighting. Exclude furnishings from the separate shelter bedroom.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent and motionless on the stainless-steel examination bed inside the circular glass enclosure, his body supported by the bed while robotic arms scan him. The source does not specify his head's direction, his torso's upward-facing surface, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie remains inactive on the stainless-steel bed within the circular glass enclosure, with immersion-damaged hardware and multiple attached wires. The surrounding robotic arms continue the scanning procedure. 현우: He remains in white recovery clothes, close to the laboratory's glass enclosure and visibly worried.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 차가운 유리 벽에 양손을 바짝 대고 찰리를 뚫어져라 내려다보는 현우의 절박한 얼굴.\n\nLOCATION (lock): At the observation side of the circular glass laboratory enclosure, with cool light over the examination bed beyond. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Glass barrier between 현우 and 찰리 in the middle-center of the frame, midground; Bed supporting 찰리 beyond the glass in the lower-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Circular glass wall (Separates 현우's pressed hands from 찰리) — Seen obliquely along its curvature, with 찰리 and part of the bed visible through it; used as Physical separation within the intimate composition; Stainless-steel bed (Partially visible beneath 찰리) — A small portion of the near side appears beyond the glass at lower right; used as Anchors the downward eyeline and the body's depth.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the laboratory's ambient tonal balance, allowing clear transmission through the glass rather than emphasizing reflected light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the stainless-steel bed, circular glass enclosure, scanning equipment, and cool laboratory lighting. Exclude furnishings from the separate shelter bedroom.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent and motionless on the stainless-steel examination bed inside the circular glass enclosure, his body supported by the bed while robotic arms scan him. The source does not specify his head's direction, his torso's upward-facing surface, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie remains inactive on the stainless-steel bed within the circular glass enclosure, with immersion-damaged hardware and multiple attached wires. The surrounding robotic arms continue the scanning procedure. 현우: He remains in white recovery clothes, close to the laboratory's glass enclosure and visibly worried.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 차가운 유리 벽에 양손을 바짝 대고 찰리를 뚫어져라 내려다보는 현우의 절박한 얼굴.\n\nLOCATION (lock): At the observation side of the circular glass laboratory enclosure, with cool light over the examination bed beyond. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: Glass barrier between 현우 and 찰리 in the middle-center of the frame, midground; Bed supporting 찰리 beyond the glass in the lower-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Circular glass wall (Separates 현우's pressed hands from 찰리) — Seen obliquely along its curvature, with 찰리 and part of the bed visible through it; used as Physical separation within the intimate composition; Stainless-steel bed (Partially visible beneath 찰리) — A small portion of the near side appears beyond the glass at lower right; used as Anchors the downward eyeline and the body's depth.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the laboratory's ambient tonal balance, allowing clear transmission through the glass rather than emphasizing reflected light.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the stainless-steel bed, circular glass enclosure, scanning equipment, and cool laboratory lighting. Exclude furnishings from the separate shelter bedroom.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent and motionless on the stainless-steel examination bed inside the circular glass enclosure, his body supported by the bed while robotic arms scan him. The source does not specify his head's direction, his torso's upward-facing surface, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie remains inactive on the stainless-steel bed within the circular glass enclosure, with immersion-damaged hardware and multiple attached wires. The surrounding robotic arms continue the scanning procedure. 현우: He remains in white recovery clothes, close to the laboratory's glass enclosure and visibly worried.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "현우는 유리벽 너머 하단에 누워 있는 찰리를 향해 시선을 정확히 내리깔고 있음.",
    "built_space": "둥근 유리벽이 현우와 찰리를 분리하고 있으며 스틸 침대는 우측 원경에 위치함. 하지만 지시되지 않은 인물 2명이 배경의 컴퓨터 앞에 앉아 있음.",
    "entities": "현우는 레퍼런스의 외모와 일치하나 지정된 흰색 회복복 대신 남색 티셔츠를 입고 있음. 찰리와 침대는 올바르게 등장함. 텍스트에 없는 엑스트라 2명이 배경에 존재함.",
    "hard_violations": [
     "[gemini-pro] invented people (프롬프트에서 금지한 배경 인물 2명이 포함됨)",
     "[gpt-high] 이 샷에서 제외해야 하는 작업자들을 왼쪽 가장자리, 뒤쪽 중앙, 오른쪽 배경에 추가했다."
    ],
    "physics": "현우의 양손은 유리에 자연스럽게 밀착되어 지탱되고 있으며, 찰리 역시 침대 위에 중력에 맞게 누워 있음."
   },
   {
    "label": "B",
    "direction": "현우는 유리벽 너머의 찰리를 향해 시선을 향하고 있음.",
    "built_space": "유리벽이 수직의 둥근 벽이 아닌 구형의 거품이나 둥근 창틀처럼 비정상적인 형태로 왜곡되어 있음. 배경의 컴퓨터 앞에 지시되지 않은 인물 1명이 앉아 있음.",
    "entities": "현우는 남색 티셔츠를 입고 있음. 찰리가 현우의 얼굴 크기보다도 작은 미니어처 장난감처럼 극단적으로 작게 축소되어 등장함. 텍스트에 없는 배경 인물 1명이 존재함.",
    "hard_violations": [
     "[gemini-pro] invented people (프롬프트에서 금지한 배경 인물 1명이 포함됨)",
     "[gemini-pro] physically impossible anatomy (왼쪽 손의 손가락이 6개임)",
     "[gemini-pro] physically impossible staging (찰리와 침대가 미니어처처럼 물리적으로 불가능한 비율로 왜곡 및 축소됨)",
     "[gpt-high] 이전 샷의 사람들을 제외하라는 지시와 달리, 뒤쪽 작업대에 앉은 작업자 한 명을 추가했다."
    ],
    "physics": "왼쪽 손(화면 우측)에 6개의 손가락이 묘사되는 해부학적 오류가 발생함. 찰리와 침대의 크기 비율이 물리적인 원근법과 전혀 맞지 않음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "프롬프트에서 엄격히 금지한 배경 인물을 추가하는 치명적 규칙 위반을 범했고 지정된 흰색 의상도 무시했으나, 공간과 피사체의 스케일은 B보다 자연스럽습니다."
       },
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "금지된 배경 인물 포함뿐만 아니라, 찰리가 미니어처처럼 축소된 스케일 오류와 손가락 개수 오류 등 치명적인 결함이 다수 존재하여 사용할 수 없습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 유리벽 너머 하단에 누워 있는 찰리를 향해 시선을 정확히 내리깔고 있음.",
        "built_space": "둥근 유리벽이 현우와 찰리를 분리하고 있으며 스틸 침대는 우측 원경에 위치함. 하지만 지시되지 않은 인물 2명이 배경의 컴퓨터 앞에 앉아 있음.",
        "entities": "현우는 레퍼런스의 외모와 일치하나 지정된 흰색 회복복 대신 남색 티셔츠를 입고 있음. 찰리와 침대는 올바르게 등장함. 텍스트에 없는 엑스트라 2명이 배경에 존재함.",
        "hard_violations": [
         "invented people (프롬프트에서 금지한 배경 인물 2명이 포함됨)"
        ],
        "physics": "현우의 양손은 유리에 자연스럽게 밀착되어 지탱되고 있으며, 찰리 역시 침대 위에 중력에 맞게 누워 있음."
       },
       {
        "label": "B",
        "direction": "현우는 유리벽 너머의 찰리를 향해 시선을 향하고 있음.",
        "built_space": "유리벽이 수직의 둥근 벽이 아닌 구형의 거품이나 둥근 창틀처럼 비정상적인 형태로 왜곡되어 있음. 배경의 컴퓨터 앞에 지시되지 않은 인물 1명이 앉아 있음.",
        "entities": "현우는 남색 티셔츠를 입고 있음. 찰리가 현우의 얼굴 크기보다도 작은 미니어처 장난감처럼 극단적으로 작게 축소되어 등장함. 텍스트에 없는 배경 인물 1명이 존재함.",
        "hard_violations": [
         "invented people (프롬프트에서 금지한 배경 인물 1명이 포함됨)",
         "physically impossible anatomy (왼쪽 손의 손가락이 6개임)",
         "physically impossible staging (찰리와 침대가 미니어처처럼 물리적으로 불가능한 비율로 왜곡 및 축소됨)"
        ],
        "physics": "왼쪽 손(화면 우측)에 6개의 손가락이 묘사되는 해부학적 오류가 발생함. 찰리와 침대의 크기 비율이 물리적인 원근법과 전혀 맞지 않음."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "프롬프트에서 엄격히 금지한 배경 인물을 추가하는 치명적 규칙 위반을 범했고 지정된 흰색 의상도 무시했으나, 공간과 피사체의 스케일은 B보다 자연스럽습니다."
       },
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "금지된 배경 인물 포함뿐만 아니라, 찰리가 미니어처처럼 축소된 스케일 오류와 손가락 개수 오류 등 치명적인 결함이 다수 존재하여 사용할 수 없습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "현우는 유리벽 너머 하단에 누워 있는 찰리를 향해 시선을 정확히 내리깔고 있음.",
        "built_space": "둥근 유리벽이 현우와 찰리를 분리하고 있으며 스틸 침대는 우측 원경에 위치함. 하지만 지시되지 않은 인물 2명이 배경의 컴퓨터 앞에 앉아 있음.",
        "entities": "현우는 레퍼런스의 외모와 일치하나 지정된 흰색 회복복 대신 남색 티셔츠를 입고 있음. 찰리와 침대는 올바르게 등장함. 텍스트에 없는 엑스트라 2명이 배경에 존재함.",
        "hard_violations": [
         "invented people (프롬프트에서 금지한 배경 인물 2명이 포함됨)"
        ],
        "physics": "현우의 양손은 유리에 자연스럽게 밀착되어 지탱되고 있으며, 찰리 역시 침대 위에 중력에 맞게 누워 있음."
       },
       {
        "label": "B",
        "direction": "현우는 유리벽 너머의 찰리를 향해 시선을 향하고 있음.",
        "built_space": "유리벽이 수직의 둥근 벽이 아닌 구형의 거품이나 둥근 창틀처럼 비정상적인 형태로 왜곡되어 있음. 배경의 컴퓨터 앞에 지시되지 않은 인물 1명이 앉아 있음.",
        "entities": "현우는 남색 티셔츠를 입고 있음. 찰리가 현우의 얼굴 크기보다도 작은 미니어처 장난감처럼 극단적으로 작게 축소되어 등장함. 텍스트에 없는 배경 인물 1명이 존재함.",
        "hard_violations": [
         "invented people (프롬프트에서 금지한 배경 인물 1명이 포함됨)",
         "physically impossible anatomy (왼쪽 손의 손가락이 6개임)",
         "physically impossible staging (찰리와 침대가 미니어처처럼 물리적으로 불가능한 비율로 왜곡 및 축소됨)"
        ],
        "physics": "왼쪽 손(화면 우측)에 6개의 손가락이 묘사되는 해부학적 오류가 발생함. 찰리와 침대의 크기 비율이 물리적인 원근법과 전혀 맞지 않음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "금지된 배경 인물 때문에 실격이지만, 현우의 얼굴 클로즈업과 양손 밀착, 오른쪽 아래 찰리를 향한 시선은 B보다 충실하다. 흰 회복복 대신 남색 티셔츠를 입었다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "금지된 배경 인물들이 등장하며, 상반신과 침대를 크게 보여주는 넓은 구도와 찰리보다 아래로 떨어지는 시선이 절박한 얼굴 클로즈업 지시에서 벗어난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 두 눈은 오른쪽 아래 침대 위 찰리 쪽을 향한다. 양손은 손바닥을 유리에 대고 펼쳐져 있다. 오른쪽 위 스캐너 끝은 찰리의 머리를 향하고, 나머지 장치는 침대 양옆을 향한다.",
        "built_space": "중앙의 곡면 유리 너머 오른쪽 아래에 검사 침대 한 개와 찰리가 있고, 로봇 팔은 부분 노출을 포함해 세 개가 보인다. 뒤쪽에는 원형 작업대와 모니터, 앉은 작업자 한 명이 보인다. 차가운 금속 재질과 조명은 장소 참조에 가깝지만, 유리 테두리의 세로 방향 굴곡이 과장되어 있으며 침대도 지시한 작은 일부보다 훨씬 많이 노출된다. 손 주변의 약한 반사는 유리 접촉면에서 가능한 범위다.",
        "entities": "현우는 참조와 유사한 앳된 동아시아계 남성으로, 헝클어진 검은 머리와 얼굴 특징이 대체로 맞는다. 국적은 외관만으로 확인할 수 없다. 옷은 지정된 흰 회복복이 아니라 남색 티셔츠다. 찰리는 참조처럼 손상되고 때 묻은 밝은색 기계 몸체이며 금속 침대, 연결선, 스캔 장치가 보인다. 배경의 흰옷 작업자 한 명은 이 샷에 허용되지 않은 인물이다. 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [
         "이전 샷의 사람들을 제외하라는 지시와 달리, 뒤쪽 작업대에 앉은 작업자 한 명을 추가했다."
        ],
        "physics": "현우의 양손은 유리에 닿아 있고 팔은 어깨에서 자연스럽게 이어진다. 하체는 프레임 밖이므로 발의 지지는 확인되지 않지만 몸이 공중에 떠 있다는 증거는 없다. 찰리의 머리, 몸통, 팔과 다리는 침대에 놓여 있으며 발은 뒤꿈치 쪽으로 지지된다. 스캐너는 관절 팔과 받침대에 연결되고 전선은 침대 가장자리에서 아래로 처져 있다."
       },
       {
        "label": "B",
        "direction": "현우의 얼굴과 눈은 오른쪽 아래를 향하지만, 찰리는 그 시선보다 높은 오른쪽 배경에 놓여 있어 찰리를 뚫어져라 내려다보는 관계가 약하다. 양손은 유리를 향해 펼쳐져 있으나 손바닥 전체가 바짝 눌린 접촉은 A보다 불명확하다. 위쪽 스캐너 끝은 찰리의 머리 부근을 향한다.",
        "built_space": "중앙에 세로 이음부가 있는 유리벽, 오른쪽에 검사 침대 한 개, 중앙과 상단에 로봇 팔 두 개가 보인다. 원형 바닥 경계와 뒤쪽 곡선 작업대는 참조 장소와 연결된다. 현우의 허리 부근까지, 침대의 넓은 상판과 하부까지 보여 얼굴 클로즈업 및 침대 일부만 노출하라는 구도보다 넓다. 왼쪽 가장자리와 뒤쪽 중앙, 오른쪽에 작업자들이 보인다. 명백히 불가능한 반사는 보이지 않는다.",
        "entities": "현우의 앳된 동아시아계 남성 외형과 검은 머리는 참조와 대체로 맞지만, 지정된 흰 회복복 대신 남색 반소매 티셔츠를 입었다. 찰리는 참조와 같은 밝은색 기계 몸체로 침대에 누워 있고, 연결선과 스캔 장치도 존재한다. 배경에는 부분적으로 보이는 인물을 포함해 작업자 세 명이 있어 허용된 등장인물 구성을 어긴다. 읽을 수 있는 글자는 확인되지 않는다.",
        "hard_violations": [
         "이 샷에서 제외해야 하는 작업자들을 왼쪽 가장자리, 뒤쪽 중앙, 오른쪽 배경에 추가했다."
        ],
        "physics": "현우는 상체를 앞으로 기울이고 팔을 들어 유리에 손을 대려는 물리적으로 가능한 자세다. 손바닥 밀착 정도는 불분명하지만 팔이나 손이 독립적으로 떠 있지는 않다. 찰리의 몸통과 팔다리는 침대에 지지되고, 머리도 침대 쪽에 놓여 있다. 로봇 팔은 받침대에 연결되며 전선은 중력 방향으로 늘어진다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "금지된 배경 인물 때문에 실격이지만, 현우의 얼굴 클로즈업과 양손 밀착, 오른쪽 아래 찰리를 향한 시선은 B보다 충실하다. 흰 회복복 대신 남색 티셔츠를 입었다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "금지된 배경 인물들이 등장하며, 상반신과 침대를 크게 보여주는 넓은 구도와 찰리보다 아래로 떨어지는 시선이 절박한 얼굴 클로즈업 지시에서 벗어난다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 두 눈은 오른쪽 아래 침대 위 찰리 쪽을 향한다. 양손은 손바닥을 유리에 대고 펼쳐져 있다. 오른쪽 위 스캐너 끝은 찰리의 머리를 향하고, 나머지 장치는 침대 양옆을 향한다.",
        "built_space": "중앙의 곡면 유리 너머 오른쪽 아래에 검사 침대 한 개와 찰리가 있고, 로봇 팔은 부분 노출을 포함해 세 개가 보인다. 뒤쪽에는 원형 작업대와 모니터, 앉은 작업자 한 명이 보인다. 차가운 금속 재질과 조명은 장소 참조에 가깝지만, 유리 테두리의 세로 방향 굴곡이 과장되어 있으며 침대도 지시한 작은 일부보다 훨씬 많이 노출된다. 손 주변의 약한 반사는 유리 접촉면에서 가능한 범위다.",
        "entities": "현우는 참조와 유사한 앳된 동아시아계 남성으로, 헝클어진 검은 머리와 얼굴 특징이 대체로 맞는다. 국적은 외관만으로 확인할 수 없다. 옷은 지정된 흰 회복복이 아니라 남색 티셔츠다. 찰리는 참조처럼 손상되고 때 묻은 밝은색 기계 몸체이며 금속 침대, 연결선, 스캔 장치가 보인다. 배경의 흰옷 작업자 한 명은 이 샷에 허용되지 않은 인물이다. 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [
         "이전 샷의 사람들을 제외하라는 지시와 달리, 뒤쪽 작업대에 앉은 작업자 한 명을 추가했다."
        ],
        "physics": "현우의 양손은 유리에 닿아 있고 팔은 어깨에서 자연스럽게 이어진다. 하체는 프레임 밖이므로 발의 지지는 확인되지 않지만 몸이 공중에 떠 있다는 증거는 없다. 찰리의 머리, 몸통, 팔과 다리는 침대에 놓여 있으며 발은 뒤꿈치 쪽으로 지지된다. 스캐너는 관절 팔과 받침대에 연결되고 전선은 침대 가장자리에서 아래로 처져 있다."
       },
       {
        "label": "A",
        "direction": "현우의 얼굴과 눈은 오른쪽 아래를 향하지만, 찰리는 그 시선보다 높은 오른쪽 배경에 놓여 있어 찰리를 뚫어져라 내려다보는 관계가 약하다. 양손은 유리를 향해 펼쳐져 있으나 손바닥 전체가 바짝 눌린 접촉은 A보다 불명확하다. 위쪽 스캐너 끝은 찰리의 머리 부근을 향한다.",
        "built_space": "중앙에 세로 이음부가 있는 유리벽, 오른쪽에 검사 침대 한 개, 중앙과 상단에 로봇 팔 두 개가 보인다. 원형 바닥 경계와 뒤쪽 곡선 작업대는 참조 장소와 연결된다. 현우의 허리 부근까지, 침대의 넓은 상판과 하부까지 보여 얼굴 클로즈업 및 침대 일부만 노출하라는 구도보다 넓다. 왼쪽 가장자리와 뒤쪽 중앙, 오른쪽에 작업자들이 보인다. 명백히 불가능한 반사는 보이지 않는다.",
        "entities": "현우의 앳된 동아시아계 남성 외형과 검은 머리는 참조와 대체로 맞지만, 지정된 흰 회복복 대신 남색 반소매 티셔츠를 입었다. 찰리는 참조와 같은 밝은색 기계 몸체로 침대에 누워 있고, 연결선과 스캔 장치도 존재한다. 배경에는 부분적으로 보이는 인물을 포함해 작업자 세 명이 있어 허용된 등장인물 구성을 어긴다. 읽을 수 있는 글자는 확인되지 않는다.",
        "hard_violations": [
         "이 샷에서 제외해야 하는 작업자들을 왼쪽 가장자리, 뒤쪽 중앙, 오른쪽 배경에 추가했다."
        ],
        "physics": "현우는 상체를 앞으로 기울이고 팔을 들어 유리에 손을 대려는 물리적으로 가능한 자세다. 손바닥 밀착 정도는 불분명하지만 팔이나 손이 독립적으로 떠 있지는 않다. 찰리의 몸통과 팔다리는 침대에 지지되고, 머리도 침대 쪽에 놓여 있다. 로봇 팔은 받침대에 연결되며 전선은 중력 방향으로 늘어진다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.667,
    "B": 1.333
   },
   "adjusted": {
    "A": 1.417,
    "B": 1.083
   },
   "violations": {
    "A": [
     "[gemini-pro] invented people (프롬프트에서 금지한 배경 인물 2명이 포함됨)",
     "[gpt-high] 이 샷에서 제외해야 하는 작업자들을 왼쪽 가장자리, 뒤쪽 중앙, 오른쪽 배경에 추가했다."
    ],
    "B": [
     "[gemini-pro] invented people (프롬프트에서 금지한 배경 인물 1명이 포함됨)",
     "[gemini-pro] physically impossible anatomy (왼쪽 손의 손가락이 6개임)",
     "[gemini-pro] physically impossible staging (찰리와 침대가 미니어처처럼 물리적으로 불가능한 비율로 왜곡 및 축소됨)",
     "[gpt-high] 이전 샷의 사람들을 제외하라는 지시와 달리, 뒤쪽 작업대에 앉은 작업자 한 명을 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1417,
   "B": 1083
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1417,
    "verdict_ko": "프롬프트에서 엄격히 금지한 배경 인물을 추가하는 치명적 규칙 위반을 범했고 지정된 흰색 의상도 무시했으나, 공간과 피사체의 스케일은 B보다 자연스럽습니다.  ★위반: [gemini-pro] invented people (프롬프트에서 금지한 배경 인물 2명이 포함됨) / [gpt-high] 이 샷에서 제외해야 하는 작업자들을 왼쪽 가장자리, 뒤쪽 중앙, 오른쪽 배경에 추가했다."
   },
   {
    "label": "B",
    "score": 1083,
    "verdict_ko": "금지된 배경 인물 포함뿐만 아니라, 찰리가 미니어처처럼 축소된 스케일 오류와 손가락 개수 오류 등 치명적인 결함이 다수 존재하여 사용할 수 없습니다.  ★위반: [gemini-pro] invented people (프롬프트에서 금지한 배경 인물 1명이 포함됨) / [gemini-pro] physically impossible anatomy (왼쪽 손의 손가락이 6개임) / [gemini-pro] physically impossible staging (찰리와 침대가 미니어처처럼 물리적으로 불가능한 비율로 왜곡 및 축소됨) / [gpt-high] 이전 샷의 사람들을 제외하라는 지시와 달리, 뒤쪽 작업대에 앉은 작업자 한 명을 추가했다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S82sh14_sel.png",
    "asset_id": "73257111-705a-49ad-9ee9-a8e062fa4e86",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-ed0e-7054-82cf-42469fe88ed6",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S82sh14"
  }
 },
 "S82sh17::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T12:08:19.569770+00:00",
  "fingerprint": "da4cf35c81a8f3135eacd71b454015105189e33668512a6be94c61ed56a83923",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S82sh17_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S82sh17_sel.png",
  "source_sha256": "09029a5d1216584bae788f1eb41af3841b537feab82de892624ab5f9d1d168f7",
  "file": "S82sh17_cine.png",
  "staged_sha256": "85a49bb8525b222632d251c0d3923f626218640aca8db820ba3db8471cff8343",
  "latency_ms": 9524
 },
 "S82sh26::signage": {
  "fp": "a4c961976af9b64e",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S82sh26": {
  "input_fingerprint": "f064e08b3d40153d",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 푸른 조명이 감도는 유리 벽 앞에 나란히 서서 찰리를 내려다보는 현우와 지소영의 굳은 뒷모습 풀샷.\n\nLOCATION (lock): In the observation area outside the laboratory's circular glass wall, bathed in the blue light specified in the shot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Bed visible between the observers beyond the glass in the middle-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Circular glass wall (Remains between the observers and 찰리) — The bed and 찰리 are seen through the section in front of the two observers; used as Sustains the separation across the widened composition; Stainless-steel bed (Occupied by the motionless 찰리) — Seen beyond the interval between 현우 and 지소영; used as Shared eyeline destination, kept small enough to preserve laboratory space; Scanning arms (Continue scanning 찰리) — Visible around the bed beyond the observers; used as Ongoing treatment contrasts with the observers' stillness.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The described blue laboratory illumination remains subdued, preserving separation between the two backs, the glass, and 찰리 beyond.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same glass enclosure, stainless-steel bed, surrounding research equipment, and cool lighting. Exclude shelter beds and domestic furnishings from the temporary accommodation.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent and motionless on the stainless-steel examination bed inside the circular glass enclosure, his body supported by the bed while robotic arms scan him. The source does not specify his head's direction, his torso's upward-facing surface, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie remains on the stainless-steel bed behind the circular glass wall, inactive and connected to multiple wires during treatment. Robotic scanning arms and the surrounding computer stations remain in place. 현우: He is wearing white recovery clothes and stands at the glass wall with a worried expression. 지소영: She stands at the glass wall wearing her neat research coat.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 푸른 조명이 감도는 유리 벽 앞에 나란히 서서 찰리를 내려다보는 현우와 지소영의 굳은 뒷모습 풀샷.\n\nLOCATION (lock): In the observation area outside the laboratory's circular glass wall, bathed in the blue light specified in the shot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Bed visible between the observers beyond the glass in the middle-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Circular glass wall (Remains between the observers and 찰리) — The bed and 찰리 are seen through the section in front of the two observers; used as Sustains the separation across the widened composition; Stainless-steel bed (Occupied by the motionless 찰리) — Seen beyond the interval between 현우 and 지소영; used as Shared eyeline destination, kept small enough to preserve laboratory space; Scanning arms (Continue scanning 찰리) — Visible around the bed beyond the observers; used as Ongoing treatment contrasts with the observers' stillness.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The described blue laboratory illumination remains subdued, preserving separation between the two backs, the glass, and 찰리 beyond.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same glass enclosure, stainless-steel bed, surrounding research equipment, and cool lighting. Exclude shelter beds and domestic furnishings from the temporary accommodation.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent and motionless on the stainless-steel examination bed inside the circular glass enclosure, his body supported by the bed while robotic arms scan him. The source does not specify his head's direction, his torso's upward-facing surface, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie remains on the stainless-steel bed behind the circular glass wall, inactive and connected to multiple wires during treatment. Robotic scanning arms and the surrounding computer stations remain in place. 현우: He is wearing white recovery clothes and stands at the glass wall with a worried expression. 지소영: She stands at the glass wall wearing her neat research coat.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 푸른 조명이 감도는 유리 벽 앞에 나란히 서서 찰리를 내려다보는 현우와 지소영의 굳은 뒷모습 풀샷.\n\nLOCATION (lock): In the observation area outside the laboratory's circular glass wall, bathed in the blue light specified in the shot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Bed visible between the observers beyond the glass in the middle-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Circular glass wall (Remains between the observers and 찰리) — The bed and 찰리 are seen through the section in front of the two observers; used as Sustains the separation across the widened composition; Stainless-steel bed (Occupied by the motionless 찰리) — Seen beyond the interval between 현우 and 지소영; used as Shared eyeline destination, kept small enough to preserve laboratory space; Scanning arms (Continue scanning 찰리) — Visible around the bed beyond the observers; used as Ongoing treatment contrasts with the observers' stillness.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The described blue laboratory illumination remains subdued, preserving separation between the two backs, the glass, and 찰리 beyond.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same glass enclosure, stainless-steel bed, surrounding research equipment, and cool lighting. Exclude shelter beds and domestic furnishings from the temporary accommodation.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent and motionless on the stainless-steel examination bed inside the circular glass enclosure, his body supported by the bed while robotic arms scan him. The source does not specify his head's direction, his torso's upward-facing surface, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie remains on the stainless-steel bed behind the circular glass wall, inactive and connected to multiple wires during treatment. Robotic scanning arms and the surrounding computer stations remain in place. 현우: He is wearing white recovery clothes and stands at the glass wall with a worried expression. 지소영: She stands at the glass wall wearing her neat research coat.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "현우와 지소영은 유리 너머 침대 위 찰리를 향하고 있으며, 로봇 팔들도 찰리를 조준함.",
    "built_space": "둥근 유리 벽 뒤로 침대가 중앙에 있고 배경에 모니터와 장비들이 배치됨.",
    "entities": "현우(흰색 긴팔), 지소영(연구 가운), 찰리(로봇), 명시되지 않은 배경 연구원 1명.",
    "hard_violations": [
     "[gemini-pro] 발명된 인물 (배경 연구원)",
     "[gemini-pro] 물리적으로 불가능한 구조 (지지대 없이 허공에 떠 있는 로봇 팔)",
     "[gpt-high] 침대 뒤 작업대에 출연이 허용되지 않은 연구원이 남아 있다."
    ],
    "physics": "인물들은 바닥에 잘 서 있으나, 왼쪽 로봇 팔의 기저부가 허공에 끊겨 있어 아무것도 이를 지지하지 않음."
   },
   {
    "label": "B",
    "direction": "현우와 지소영은 유리 벽 너머 찰리를 응시하고, 로봇 팔들이 찰리를 향해 있음.",
    "built_space": "둥근 유리 벽 너머 중앙에 침대가 있고, 주변에 컴퓨터 스테이션이 정렬됨.",
    "entities": "현우(회색 반팔), 지소영(연구 가운), 찰리(로봇), 명시되지 않은 배경 연구원 1명.",
    "hard_violations": [
     "[gemini-pro] 발명된 인물 (배경 연구원)",
     "[gpt-high] 출연이 허용되지 않은 배경 연구원 두 명이 유리 안쪽 작업대에 남아 있다."
    ],
    "physics": "인물들은 바닥에 서 있고 찰리는 침대에 지탱되며, 로봇 팔도 장비 본체에 안정적으로 연결되어 지지됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "프롬프트에 없는 배경 인물이 포함되어 치명적인 위반이나, 전체적인 공간 구조와 물리적 안정성은 A보다 낫습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "프롬프트에 없는 인물이 추가되었을 뿐만 아니라, 왼쪽 로봇 팔이 허공에 떠 있는 물리적 오류가 있어 심각한 위반입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우와 지소영은 유리 너머 침대 위 찰리를 향하고 있으며, 로봇 팔들도 찰리를 조준함.",
        "built_space": "둥근 유리 벽 뒤로 침대가 중앙에 있고 배경에 모니터와 장비들이 배치됨.",
        "entities": "현우(흰색 긴팔), 지소영(연구 가운), 찰리(로봇), 명시되지 않은 배경 연구원 1명.",
        "hard_violations": [
         "발명된 인물 (배경 연구원)",
         "물리적으로 불가능한 구조 (지지대 없이 허공에 떠 있는 로봇 팔)"
        ],
        "physics": "인물들은 바닥에 잘 서 있으나, 왼쪽 로봇 팔의 기저부가 허공에 끊겨 있어 아무것도 이를 지지하지 않음."
       },
       {
        "label": "B",
        "direction": "현우와 지소영은 유리 벽 너머 찰리를 응시하고, 로봇 팔들이 찰리를 향해 있음.",
        "built_space": "둥근 유리 벽 너머 중앙에 침대가 있고, 주변에 컴퓨터 스테이션이 정렬됨.",
        "entities": "현우(회색 반팔), 지소영(연구 가운), 찰리(로봇), 명시되지 않은 배경 연구원 1명.",
        "hard_violations": [
         "발명된 인물 (배경 연구원)"
        ],
        "physics": "인물들은 바닥에 서 있고 찰리는 침대에 지탱되며, 로봇 팔도 장비 본체에 안정적으로 연결되어 지지됨."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "프롬프트에 없는 배경 인물이 포함되어 치명적인 위반이나, 전체적인 공간 구조와 물리적 안정성은 A보다 낫습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "프롬프트에 없는 인물이 추가되었을 뿐만 아니라, 왼쪽 로봇 팔이 허공에 떠 있는 물리적 오류가 있어 심각한 위반입니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "현우와 지소영은 유리 너머 침대 위 찰리를 향하고 있으며, 로봇 팔들도 찰리를 조준함.",
        "built_space": "둥근 유리 벽 뒤로 침대가 중앙에 있고 배경에 모니터와 장비들이 배치됨.",
        "entities": "현우(흰색 긴팔), 지소영(연구 가운), 찰리(로봇), 명시되지 않은 배경 연구원 1명.",
        "hard_violations": [
         "발명된 인물 (배경 연구원)",
         "물리적으로 불가능한 구조 (지지대 없이 허공에 떠 있는 로봇 팔)"
        ],
        "physics": "인물들은 바닥에 잘 서 있으나, 왼쪽 로봇 팔의 기저부가 허공에 끊겨 있어 아무것도 이를 지지하지 않음."
       },
       {
        "label": "B",
        "direction": "현우와 지소영은 유리 벽 너머 찰리를 응시하고, 로봇 팔들이 찰리를 향해 있음.",
        "built_space": "둥근 유리 벽 너머 중앙에 침대가 있고, 주변에 컴퓨터 스테이션이 정렬됨.",
        "entities": "현우(회색 반팔), 지소영(연구 가운), 찰리(로봇), 명시되지 않은 배경 연구원 1명.",
        "hard_violations": [
         "발명된 인물 (배경 연구원)"
        ],
        "physics": "인물들은 바닥에 서 있고 찰리는 침대에 지탱되며, 로봇 팔도 장비 본체에 안정적으로 연결되어 지지됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "유리 너머 중앙 침대를 바라보는 배치는 맞지만, 제외해야 할 연구원 두 명이 남아 있고 두 관찰자의 발까지 담는 뒷모습 풀샷도 아니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "중앙 침대·전선·스캔 장치와 흰 회복복은 요구에 가깝지만, 배경 연구원이 남아 있어 탈락 사유가 있으며 풀샷과 이전 장면의 의상 연속성도 충족하지 못한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 왼쪽, 지소영은 오른쪽에서 등을 보이며 중앙 침대의 찰리 쪽으로 몸과 머리를 향한다. 눈은 보이지 않아 정확한 시선은 확인할 수 없고, 고개를 아래로 숙인 정도는 약하다. 침대 위 두 스캔 암의 끝은 찰리의 머리와 상체 쪽을 향한다.",
        "built_space": "곡면 유리벽 한 겹이 관찰자들과 침대 사이를 가로막고, 왼쪽에는 금속 문틀이 있다. 중앙에는 금속 검사 침대 한 대, 그 위쪽에는 주요 스캔 암 두 개가 보이며 오른쪽에는 관절식 장치와 이동형 장비가 모여 있다. 벽면 작업대와 여러 모니터가 원형 공간을 둘러싼다. 관찰자들은 유리 바깥에 나란히 서 있어 공간 관계는 맞지만, 하체가 화면 아래에서 잘려 풀샷은 아니다. 왼쪽과 중앙 뒤 작업대에는 각각 연구원 한 명이 보인다.",
        "entities": "전경에는 검은 헝클어진 머리의 젊은 남성과 단정한 검은 단발머리의 여성이 있다. 뒷모습이므로 한국계 정체성이나 정확한 나이·얼굴 일치는 확인할 수 없다. 지소영의 연구복은 맞는다. 현우의 옅은 회색 반팔 상의는 이전 장면의 남색 반팔과 색이 다르고, 명시된 흰 회복복과도 차이가 있다. 찰리는 참조처럼 밝은색 장갑 외형을 지닌 채 침대에 누워 있다. 추가 연구원 두 명은 이 장면에서 제외하라는 지시와 어긋난다. 판독 가능한 문구는 뚜렷하지 않다.",
        "hard_violations": [
         "출연이 허용되지 않은 배경 연구원 두 명이 유리 안쪽 작업대에 남아 있다."
        ],
        "physics": "찰리의 머리·몸통·다리와 양팔은 침대 위에 놓여 있으며, 들린 손이나 지지 없이 떠 있는 신체는 보이지 않는다. 침대는 하부 금속 받침으로 바닥에 지지된다. 스캔 암은 관절과 장비 지지대에 연결되어 있다. 두 관찰자의 발은 프레임 밖이지만 몸이 공중에 떠 있다는 징후는 없으며, 팔은 중력 방향으로 내려가 있다."
       },
       {
        "label": "B",
        "direction": "두 관찰자는 유리 너머 사이 공간에 놓인 찰리를 향해 등을 보이고 서 있다. 현우의 머리는 약간 안쪽으로 돌아가 있고 지소영도 침대 방향을 향하지만, 눈이 가려져 정확한 하향 시선은 확인할 수 없다. 위쪽 스캔 암 두 개는 찰리의 머리와 상체를 향하며 오른쪽의 낮은 관절식 장치도 몸통 쪽으로 접근한다.",
        "built_space": "왼쪽 출입문과 금속 기둥, 원형 유리벽, 바닥의 원형 경계가 보인다. 관찰자들은 유리 밖에 있고 중앙 검사 침대 한 대는 두 사람 사이로 보인다. 침대 주변에는 위쪽 스캔 암 두 개와 오른쪽 낮은 관절식 장치가 있으며, 뒤쪽 벽에는 작업대·모니터·의자가 배치되어 있다. 참조의 원형 실험실 재료와 구조는 대체로 이어진다. 다만 두 관찰자의 다리 아래와 발이 잘려 요청된 풀샷이 아니며, 침대 뒤 작업대에 앉은 연구원이 보인다.",
        "entities": "현우는 검은 헝클어진 머리와 젊은 남성의 체격, 지소영은 단정한 검은 단발과 연구복으로 표현되어 있다. 얼굴이 보이지 않으므로 정확한 나이와 얼굴 정체성은 검증할 수 없다. 현우의 흰 긴소매 회복복은 의상 문구에는 맞지만 이전 장면에서 고정한 남색 반팔과는 다르다. 찰리는 참조의 장갑 외형을 유지하며 금속 침대에 누워 있고 주변 전선도 보인다. 중앙 뒤편의 추가 연구원은 허용된 인물이 아니다. 뚜렷하게 읽히는 글자는 없다.",
        "hard_violations": [
         "침대 뒤 작업대에 출연이 허용되지 않은 연구원이 남아 있다."
        ],
        "physics": "찰리는 머리와 몸통, 팔과 다리를 침대에 맡긴 누운 자세이며, 오른쪽 팔 일부는 장치에 가려져 있다. 보이는 손과 발이 근육 힘으로 허공에 들린 모습은 없다. 침대는 바닥에 닿은 금속 하부 구조로 지지되고, 스캔 암은 기계 관절과 받침에 연결된다. 전선은 장비와 침대 주변으로 늘어진다. 관찰자들의 발은 잘렸지만 서 있는 몸의 무게 관계는 자연스럽고, 지지 없이 떠 있는 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "유리 너머 중앙 침대를 바라보는 배치는 맞지만, 제외해야 할 연구원 두 명이 남아 있고 두 관찰자의 발까지 담는 뒷모습 풀샷도 아니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "중앙 침대·전선·스캔 장치와 흰 회복복은 요구에 가깝지만, 배경 연구원이 남아 있어 탈락 사유가 있으며 풀샷과 이전 장면의 의상 연속성도 충족하지 못한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 왼쪽, 지소영은 오른쪽에서 등을 보이며 중앙 침대의 찰리 쪽으로 몸과 머리를 향한다. 눈은 보이지 않아 정확한 시선은 확인할 수 없고, 고개를 아래로 숙인 정도는 약하다. 침대 위 두 스캔 암의 끝은 찰리의 머리와 상체 쪽을 향한다.",
        "built_space": "곡면 유리벽 한 겹이 관찰자들과 침대 사이를 가로막고, 왼쪽에는 금속 문틀이 있다. 중앙에는 금속 검사 침대 한 대, 그 위쪽에는 주요 스캔 암 두 개가 보이며 오른쪽에는 관절식 장치와 이동형 장비가 모여 있다. 벽면 작업대와 여러 모니터가 원형 공간을 둘러싼다. 관찰자들은 유리 바깥에 나란히 서 있어 공간 관계는 맞지만, 하체가 화면 아래에서 잘려 풀샷은 아니다. 왼쪽과 중앙 뒤 작업대에는 각각 연구원 한 명이 보인다.",
        "entities": "전경에는 검은 헝클어진 머리의 젊은 남성과 단정한 검은 단발머리의 여성이 있다. 뒷모습이므로 한국계 정체성이나 정확한 나이·얼굴 일치는 확인할 수 없다. 지소영의 연구복은 맞는다. 현우의 옅은 회색 반팔 상의는 이전 장면의 남색 반팔과 색이 다르고, 명시된 흰 회복복과도 차이가 있다. 찰리는 참조처럼 밝은색 장갑 외형을 지닌 채 침대에 누워 있다. 추가 연구원 두 명은 이 장면에서 제외하라는 지시와 어긋난다. 판독 가능한 문구는 뚜렷하지 않다.",
        "hard_violations": [
         "출연이 허용되지 않은 배경 연구원 두 명이 유리 안쪽 작업대에 남아 있다."
        ],
        "physics": "찰리의 머리·몸통·다리와 양팔은 침대 위에 놓여 있으며, 들린 손이나 지지 없이 떠 있는 신체는 보이지 않는다. 침대는 하부 금속 받침으로 바닥에 지지된다. 스캔 암은 관절과 장비 지지대에 연결되어 있다. 두 관찰자의 발은 프레임 밖이지만 몸이 공중에 떠 있다는 징후는 없으며, 팔은 중력 방향으로 내려가 있다."
       },
       {
        "label": "A",
        "direction": "두 관찰자는 유리 너머 사이 공간에 놓인 찰리를 향해 등을 보이고 서 있다. 현우의 머리는 약간 안쪽으로 돌아가 있고 지소영도 침대 방향을 향하지만, 눈이 가려져 정확한 하향 시선은 확인할 수 없다. 위쪽 스캔 암 두 개는 찰리의 머리와 상체를 향하며 오른쪽의 낮은 관절식 장치도 몸통 쪽으로 접근한다.",
        "built_space": "왼쪽 출입문과 금속 기둥, 원형 유리벽, 바닥의 원형 경계가 보인다. 관찰자들은 유리 밖에 있고 중앙 검사 침대 한 대는 두 사람 사이로 보인다. 침대 주변에는 위쪽 스캔 암 두 개와 오른쪽 낮은 관절식 장치가 있으며, 뒤쪽 벽에는 작업대·모니터·의자가 배치되어 있다. 참조의 원형 실험실 재료와 구조는 대체로 이어진다. 다만 두 관찰자의 다리 아래와 발이 잘려 요청된 풀샷이 아니며, 침대 뒤 작업대에 앉은 연구원이 보인다.",
        "entities": "현우는 검은 헝클어진 머리와 젊은 남성의 체격, 지소영은 단정한 검은 단발과 연구복으로 표현되어 있다. 얼굴이 보이지 않으므로 정확한 나이와 얼굴 정체성은 검증할 수 없다. 현우의 흰 긴소매 회복복은 의상 문구에는 맞지만 이전 장면에서 고정한 남색 반팔과는 다르다. 찰리는 참조의 장갑 외형을 유지하며 금속 침대에 누워 있고 주변 전선도 보인다. 중앙 뒤편의 추가 연구원은 허용된 인물이 아니다. 뚜렷하게 읽히는 글자는 없다.",
        "hard_violations": [
         "침대 뒤 작업대에 출연이 허용되지 않은 연구원이 남아 있다."
        ],
        "physics": "찰리는 머리와 몸통, 팔과 다리를 침대에 맡긴 누운 자세이며, 오른쪽 팔 일부는 장치에 가려져 있다. 보이는 손과 발이 근육 힘으로 허공에 들린 모습은 없다. 침대는 바닥에 닿은 금속 하부 구조로 지지되고, 스캔 암은 기계 관절과 받침에 연결된다. 전선은 장비와 침대 주변으로 늘어진다. 관찰자들의 발은 잘렸지만 서 있는 몸의 무게 관계는 자연스럽고, 지지 없이 떠 있는 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.75,
    "B": 1.667
   },
   "adjusted": {
    "A": 1.5,
    "B": 1.417
   },
   "violations": {
    "A": [
     "[gemini-pro] 발명된 인물 (배경 연구원)",
     "[gemini-pro] 물리적으로 불가능한 구조 (지지대 없이 허공에 떠 있는 로봇 팔)",
     "[gpt-high] 침대 뒤 작업대에 출연이 허용되지 않은 연구원이 남아 있다."
    ],
    "B": [
     "[gemini-pro] 발명된 인물 (배경 연구원)",
     "[gpt-high] 출연이 허용되지 않은 배경 연구원 두 명이 유리 안쪽 작업대에 남아 있다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1417,
   "A": 1500
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1417,
    "verdict_ko": "프롬프트에 없는 배경 인물이 포함되어 치명적인 위반이나, 전체적인 공간 구조와 물리적 안정성은 A보다 낫습니다.  ★위반: [gemini-pro] 발명된 인물 (배경 연구원) / [gpt-high] 출연이 허용되지 않은 배경 연구원 두 명이 유리 안쪽 작업대에 남아 있다."
   },
   {
    "label": "A",
    "score": 1500,
    "verdict_ko": "프롬프트에 없는 인물이 추가되었을 뿐만 아니라, 왼쪽 로봇 팔이 허공에 떠 있는 물리적 오류가 있어 심각한 위반입니다.  ★위반: [gemini-pro] 발명된 인물 (배경 연구원) / [gemini-pro] 물리적으로 불가능한 구조 (지지대 없이 허공에 떠 있는 로봇 팔) / [gpt-high] 침대 뒤 작업대에 출연이 허용되지 않은 연구원이 남아 있다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S82sh17_sel.png",
    "asset_id": "2c72afdf-43d4-46d6-9ef5-8f1465a4e8dc",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 지소영: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1243508>",
    "asset_id": "c7496f13-cfcf-44a5-976d-96c783d20580",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-eebe-7fd0-af93-b421c838b31a",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S82sh17"
  }
 },
 "S82sh26::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T12:09:51.953139+00:00",
  "fingerprint": "9d4283945d98164712aa4b36bc337a32e96e6ff66e2c56011ff48cde6d18fb80",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S82sh26_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S82sh26_sel.png",
  "source_sha256": "f80f0c77a0d1a84722ca6eae210908e1e03bf8af9f7961164eddab6fb3002373",
  "file": "S82sh26_cine.png",
  "staged_sha256": "313ba8dc54f0bf9b86f3627380431f5ec8d2b1b9ab0a782a29ff10e1cb837214",
  "latency_ms": 9557
 },
 "S83sh5::signage": {
  "fp": "a888e4ab561146ab",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S83sh5": {
  "input_fingerprint": "ca0e75bb15e37fe2",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 모니터 화면 속, 활짝 웃고 있는 앰버와 라울의 얼굴 클로즈업.\n\nLOCATION (lock): On a video-call monitor inside the research facility's temporary-care room, showing two remote callers without an identifiable background location. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Video-call monitor (Displaying 앰버 and 라울 smiling during the live call) — The image-bearing front is seen obliquely, with its boundary retained around the displayed faces; used as Mediates the close view and distinguishes remote people from local space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the tonal separation between the electronic call image and the local ambient surroundings without adding an unsupported colored glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The call is now in video mode, showing Amber after her completed operation and Raul appearing in the live image.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 모니터 화면 속, 활짝 웃고 있는 앰버와 라울의 얼굴 클로즈업.\n\nLOCATION (lock): On a video-call monitor inside the research facility's temporary-care room, showing two remote callers without an identifiable background location. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Video-call monitor (Displaying 앰버 and 라울 smiling during the live call) — The image-bearing front is seen obliquely, with its boundary retained around the displayed faces; used as Mediates the close view and distinguishes remote people from local space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the tonal separation between the electronic call image and the local ambient surroundings without adding an unsupported colored glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The call is now in video mode, showing Amber after her completed operation and Raul appearing in the live image.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 모니터 화면 속, 활짝 웃고 있는 앰버와 라울의 얼굴 클로즈업.\n\nLOCATION (lock): On a video-call monitor inside the research facility's temporary-care room, showing two remote callers without an identifiable background location. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Video-call monitor (Displaying 앰버 and 라울 smiling during the live call) — The image-bearing front is seen obliquely, with its boundary retained around the displayed faces; used as Mediates the close view and distinguishes remote people from local space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the tonal separation between the electronic call image and the local ambient surroundings without adding an unsupported colored glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The call is now in video mode, showing Amber after her completed operation and Raul appearing in the live image.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽); 라울 (10세 남자아이, 라틴계 흑인 혼혈, 어린 얼굴, 뒤로 묶은 꽁지머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S83sh5__bgfirst_bg.png",
     "asset_id": "3bc12182-b8a5-4b6b-bbfb-1fc47da3fb74",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S83sh5.png",
     "asset_id": "481275cf-ea10-437e-8452-befc99e03a7f",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163202>",
     "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L266B02.png",
     "asset_id": "78924c8c-f3a0-4a9f-b31d-34a3fce75e4a",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1163202>",
     "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "화면 속 두 인물이 정면을 바라보며 미소 짓고 있음.",
    "built_space": "클로즈업 지시를 무시하고 병실 전체(침대, 창문 포함)를 보여주며, 벽에 걸린 모니터는 공간의 원근법을 무시하고 가로로 기형적으로 길게 늘어나 있음.",
    "entities": "화면 속에 앰버와 라울이 지정된 인상착의로 등장함.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 무대 연출 (원근과 비율이 심각하게 왜곡된 모니터 형태)"
    ],
    "physics": "모니터가 벽에 걸려 있으나 형태 왜곡으로 인해 지지 구조가 비현실적으로 보임."
   },
   {
    "label": "B",
    "direction": "화면 속 두 인물이 정면을 바라보며 미소 짓고 있음.",
    "built_space": "벽걸이 모니터와 장착 패널을 가깝게 잡은 클로즈업 구도이며, 카메라와 모니터의 배치가 자연스러움.",
    "entities": "화면 속에 앰버와 라울이 지정된 인상착의(금발 머리, 꽁지머리 등)로 정확히 등장함.",
    "hard_violations": [
     "[gemini-pro] 읽을 수 있는 텍스트 생성 (우측 벽면 안내문에 글자가 노출됨)",
     "[gpt-high] 오른쪽 벽 안내문의 한글 및 단자의 읽을 수 있는 영문 표기가 남아 있어, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
    ],
    "physics": "모니터가 벽면 패널에 안정적으로 부착되어 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "지시된 클로즈업 프레이밍과 화면 속 인물들의 표정을 정확히 구현했으나, 우측 벽면에 금지된 텍스트가 생성된 점이 감점 요인입니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "지시된 클로즈업을 무시하고 방 전체를 보여주는 와이드 샷으로 렌더링했으며, 모니터의 가로 비율과 원근이 비현실적으로 왜곡되었습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "화면 속 두 인물이 정면을 바라보며 미소 짓고 있음.",
        "built_space": "벽걸이 모니터와 장착 패널을 가깝게 잡은 클로즈업 구도이며, 카메라와 모니터의 배치가 자연스러움.",
        "entities": "화면 속에 앰버와 라울이 지정된 인상착의(금발 머리, 꽁지머리 등)로 정확히 등장함.",
        "hard_violations": [
         "읽을 수 있는 텍스트 생성 (우측 벽면 안내문에 글자가 노출됨)"
        ],
        "physics": "모니터가 벽면 패널에 안정적으로 부착되어 있음."
       },
       {
        "label": "A",
        "direction": "화면 속 두 인물이 정면을 바라보며 미소 짓고 있음.",
        "built_space": "클로즈업 지시를 무시하고 병실 전체(침대, 창문 포함)를 보여주며, 벽에 걸린 모니터는 공간의 원근법을 무시하고 가로로 기형적으로 길게 늘어나 있음.",
        "entities": "화면 속에 앰버와 라울이 지정된 인상착의로 등장함.",
        "hard_violations": [
         "물리적으로 불가능한 무대 연출 (원근과 비율이 심각하게 왜곡된 모니터 형태)"
        ],
        "physics": "모니터가 벽에 걸려 있으나 형태 왜곡으로 인해 지지 구조가 비현실적으로 보임."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "지시된 클로즈업 프레이밍과 화면 속 인물들의 표정을 정확히 구현했으나, 우측 벽면에 금지된 텍스트가 생성된 점이 감점 요인입니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "지시된 클로즈업을 무시하고 방 전체를 보여주는 와이드 샷으로 렌더링했으며, 모니터의 가로 비율과 원근이 비현실적으로 왜곡되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "화면 속 두 인물이 정면을 바라보며 미소 짓고 있음.",
        "built_space": "벽걸이 모니터와 장착 패널을 가깝게 잡은 클로즈업 구도이며, 카메라와 모니터의 배치가 자연스러움.",
        "entities": "화면 속에 앰버와 라울이 지정된 인상착의(금발 머리, 꽁지머리 등)로 정확히 등장함.",
        "hard_violations": [
         "읽을 수 있는 텍스트 생성 (우측 벽면 안내문에 글자가 노출됨)"
        ],
        "physics": "모니터가 벽면 패널에 안정적으로 부착되어 있음."
       },
       {
        "label": "A",
        "direction": "화면 속 두 인물이 정면을 바라보며 미소 짓고 있음.",
        "built_space": "클로즈업 지시를 무시하고 병실 전체(침대, 창문 포함)를 보여주며, 벽에 걸린 모니터는 공간의 원근법을 무시하고 가로로 기형적으로 길게 늘어나 있음.",
        "entities": "화면 속에 앰버와 라울이 지정된 인상착의로 등장함.",
        "hard_violations": [
         "물리적으로 불가능한 무대 연출 (원근과 비율이 심각하게 왜곡된 모니터 형태)"
        ],
        "physics": "모니터가 벽에 걸려 있으나 형태 왜곡으로 인해 지지 구조가 비현실적으로 보임."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "비스듬한 모니터의 전체 테두리와 웃는 두 얼굴은 잘 담았지만, 오른쪽 안내문과 단자의 읽을 수 있는 글자가 문자 금지 조건을 위반한다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "주간 임시보호실의 모니터 안에서 두 아이가 활짝 웃는 순간을 충실히 구현했으나, 화면 상단 테두리가 일부 잘리고 방의 비중이 다소 크다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버와 라울 모두 원격 통화 카메라 쪽을 바라보며 이를 드러내고 웃는다. 모니터의 영상 표시 면은 촬영 카메라에 비스듬히 보이며, 위쪽 웹캠은 방 안쪽을 향한다.",
        "built_space": "흰 벽에 모니터 한 대, 상단 웹캠 한 개, 뒤쪽 금속 설치판 한 개가 보인다. 설치판 오른쪽에는 연결 단자부 두 구획이 있고, 더 오른쪽에는 안내문과 벽 제어판 일부가 보인다. 참고 장소의 벽걸이 장비 구성을 대체로 유지하며 모니터 테두리 전체가 프레임 안에 있다. 화면의 옅은 창빛 반사는 맞은편 쪽 창에서 생길 수 있는 배치다.",
        "entities": "화면 속 인물은 두 명뿐이다. 앰버는 약 10세의 금발 여자아이로 큰 눈, 둥근 얼굴과 남색 상의가 참고와 대체로 맞는다. 라울은 약 10세의 갈색 피부 남자아이로 뒤로 묶은 검은 머리와 남색 상의가 참고와 맞는다. 혼혈 배경 자체는 외모만으로 확정할 수 없지만 두 인물 모두 제시된 외형에 부합한다. 원격 배경은 장소를 특정하기 어려운 밝은 면이다. 오른쪽 안내문의 한글과 단자의 영문 표기가 노출되어 있다.",
        "hard_violations": [
         "오른쪽 벽 안내문의 한글 및 단자의 읽을 수 있는 영문 표기가 남아 있어, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
        ],
        "physics": "모니터는 벽의 설치판에 부착되어 있고 웹캠은 상단 받침에 놓여 있다. 두 아이의 머리는 목과 어깨에 자연스럽게 연결되어 있으며, 화면 밖 하체의 지지는 이 얼굴 클로즈업에서 판단할 필요가 없다. 공중에 떠 있는 신체나 물체는 없다."
       },
       {
        "label": "B",
        "direction": "두 아이는 원격 통화 카메라를 향해 시선을 두고 입을 벌려 활짝 웃는다. 모니터 앞면은 촬영 카메라 쪽에 비스듬히 노출되고 상단 웹캠은 임시보호실 안쪽을 향해 있어 통화 장비의 사용 방향이 자연스럽다.",
        "built_space": "오른쪽 벽의 모니터 한 대와 금속 설치판 한 개, 웹캠 한 개, 오른쪽 세로 제어부 한 개가 보인다. 왼쪽에는 바다가 보이는 창 한 구획과 병상 한 대의 발치 일부가 있으며, 천장 조명과 밝은 벽·회색 바닥도 참고 장소와 부합한다. 얼굴을 둘러싼 모니터의 좌우·아래 테두리는 보이지만 상단 테두리 일부는 프레임 밖으로 잘린다. 방이 차지하는 면적은 얼굴 클로즈업 요구에 비해 다소 크다.",
        "entities": "등장인물은 모니터 속 앰버와 라울 두 명뿐이다. 앰버의 어린 여자아이 외형, 금발, 큰 눈, 둥근 볼과 남색 상의가 참고에 부합한다. 라울의 어린 남자아이 외형, 갈색 피부, 뒤로 묶은 검은 머리와 남색 상의도 부합한다. 두 사람의 혼혈 배경은 영상만으로 확정할 수 없다. 통화 화면 안의 밝은 배경은 특정 장소를 드러내지 않으며, 읽을 수 있는 글자나 자막은 보이지 않는다.",
        "hard_violations": [],
        "physics": "모니터는 벽 부착 설치판에 지지되고 웹캠은 모니터 위 받침에 고정되어 있다. 병상은 보이는 바퀴로 바닥에 지지된다. 두 아이의 얼굴과 목·어깨 연결은 자연스럽고, 웃는 표정에도 해부학적 이상이 없다. 지지 없이 떠 있는 물체나 인물은 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "비스듬한 모니터의 전체 테두리와 웃는 두 얼굴은 잘 담았지만, 오른쪽 안내문과 단자의 읽을 수 있는 글자가 문자 금지 조건을 위반한다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "주간 임시보호실의 모니터 안에서 두 아이가 활짝 웃는 순간을 충실히 구현했으나, 화면 상단 테두리가 일부 잘리고 방의 비중이 다소 크다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "앰버와 라울 모두 원격 통화 카메라 쪽을 바라보며 이를 드러내고 웃는다. 모니터의 영상 표시 면은 촬영 카메라에 비스듬히 보이며, 위쪽 웹캠은 방 안쪽을 향한다.",
        "built_space": "흰 벽에 모니터 한 대, 상단 웹캠 한 개, 뒤쪽 금속 설치판 한 개가 보인다. 설치판 오른쪽에는 연결 단자부 두 구획이 있고, 더 오른쪽에는 안내문과 벽 제어판 일부가 보인다. 참고 장소의 벽걸이 장비 구성을 대체로 유지하며 모니터 테두리 전체가 프레임 안에 있다. 화면의 옅은 창빛 반사는 맞은편 쪽 창에서 생길 수 있는 배치다.",
        "entities": "화면 속 인물은 두 명뿐이다. 앰버는 약 10세의 금발 여자아이로 큰 눈, 둥근 얼굴과 남색 상의가 참고와 대체로 맞는다. 라울은 약 10세의 갈색 피부 남자아이로 뒤로 묶은 검은 머리와 남색 상의가 참고와 맞는다. 혼혈 배경 자체는 외모만으로 확정할 수 없지만 두 인물 모두 제시된 외형에 부합한다. 원격 배경은 장소를 특정하기 어려운 밝은 면이다. 오른쪽 안내문의 한글과 단자의 영문 표기가 노출되어 있다.",
        "hard_violations": [
         "오른쪽 벽 안내문의 한글 및 단자의 읽을 수 있는 영문 표기가 남아 있어, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
        ],
        "physics": "모니터는 벽의 설치판에 부착되어 있고 웹캠은 상단 받침에 놓여 있다. 두 아이의 머리는 목과 어깨에 자연스럽게 연결되어 있으며, 화면 밖 하체의 지지는 이 얼굴 클로즈업에서 판단할 필요가 없다. 공중에 떠 있는 신체나 물체는 없다."
       },
       {
        "label": "A",
        "direction": "두 아이는 원격 통화 카메라를 향해 시선을 두고 입을 벌려 활짝 웃는다. 모니터 앞면은 촬영 카메라 쪽에 비스듬히 노출되고 상단 웹캠은 임시보호실 안쪽을 향해 있어 통화 장비의 사용 방향이 자연스럽다.",
        "built_space": "오른쪽 벽의 모니터 한 대와 금속 설치판 한 개, 웹캠 한 개, 오른쪽 세로 제어부 한 개가 보인다. 왼쪽에는 바다가 보이는 창 한 구획과 병상 한 대의 발치 일부가 있으며, 천장 조명과 밝은 벽·회색 바닥도 참고 장소와 부합한다. 얼굴을 둘러싼 모니터의 좌우·아래 테두리는 보이지만 상단 테두리 일부는 프레임 밖으로 잘린다. 방이 차지하는 면적은 얼굴 클로즈업 요구에 비해 다소 크다.",
        "entities": "등장인물은 모니터 속 앰버와 라울 두 명뿐이다. 앰버의 어린 여자아이 외형, 금발, 큰 눈, 둥근 볼과 남색 상의가 참고에 부합한다. 라울의 어린 남자아이 외형, 갈색 피부, 뒤로 묶은 검은 머리와 남색 상의도 부합한다. 두 사람의 혼혈 배경은 영상만으로 확정할 수 없다. 통화 화면 안의 밝은 배경은 특정 장소를 드러내지 않으며, 읽을 수 있는 글자나 자막은 보이지 않는다.",
        "hard_violations": [],
        "physics": "모니터는 벽 부착 설치판에 지지되고 웹캠은 모니터 위 받침에 고정되어 있다. 병상은 보이는 바퀴로 바닥에 지지된다. 두 아이의 얼굴과 목·어깨 연결은 자연스럽고, 웃는 표정에도 해부학적 이상이 없다. 지지 없이 떠 있는 물체나 인물은 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.667,
    "B": 1.375
   },
   "adjusted": {
    "A": 1.417,
    "B": 1.125
   },
   "violations": {
    "B": [
     "[gemini-pro] 읽을 수 있는 텍스트 생성 (우측 벽면 안내문에 글자가 노출됨)",
     "[gpt-high] 오른쪽 벽 안내문의 한글 및 단자의 읽을 수 있는 영문 표기가 남아 있어, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
    ],
    "A": [
     "[gemini-pro] 물리적으로 불가능한 무대 연출 (원근과 비율이 심각하게 왜곡된 모니터 형태)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1125,
   "A": 1417
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1125,
    "verdict_ko": "지시된 클로즈업 프레이밍과 화면 속 인물들의 표정을 정확히 구현했으나, 우측 벽면에 금지된 텍스트가 생성된 점이 감점 요인입니다.  ★위반: [gemini-pro] 읽을 수 있는 텍스트 생성 (우측 벽면 안내문에 글자가 노출됨) / [gpt-high] 오른쪽 벽 안내문의 한글 및 단자의 읽을 수 있는 영문 표기가 남아 있어, 읽을 수 있는 글자를 전면 금지한 조건을 위반한다."
   },
   {
    "label": "A",
    "score": 1417,
    "verdict_ko": "지시된 클로즈업을 무시하고 방 전체를 보여주는 와이드 샷으로 렌더링했으며, 모니터의 가로 비율과 원근이 비현실적으로 왜곡되었습니다.  ★위반: [gemini-pro] 물리적으로 불가능한 무대 연출 (원근과 비율이 심각하게 왜곡된 모니터 형태)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L266B02.png",
    "asset_id": "78924c8c-f3a0-4a9f-b31d-34a3fce75e4a",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 앰버: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 라울: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1163202>",
    "asset_id": "9bbe1cd0-b296-45a8-882b-1256373d7701",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-f07a-7509-9b14-ea5fe0f17f52",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S83sh5__bgfirst_bg.png",
   "bg_asset_id": "3bc12182-b8a5-4b6b-bbfb-1fc47da3fb74",
   "bg_record_key": "S83sh5::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S83sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T12:11:24.432660+00:00",
  "fingerprint": "7f9b2e861309273c6311f848628f7db5bf0d6d3c5ae866e4844d216df6d0b3ab",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S83sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S83sh5_sel.png",
  "source_sha256": "0c2f8d4f9b0c979c52f20a162c8b24bfc4ac5243fba9ef328df643010c38d6cd",
  "file": "S83sh5_cine.png",
  "staged_sha256": "dbd165e65b1bdeba27be181923f93a2cc382dea81a367a8d288cf38c08093b90",
  "latency_ms": 10968
 },
 "S83sh6::signage": {
  "fp": "b21dae096eda3b23",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S83sh6": {
  "input_fingerprint": "77fb4ff6d11cd93b",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 화면을 뚫어져라 응시한 채 눈시울이 붉어진 현우의 얼굴.\n\nLOCATION (lock): At the video-call station inside the temporary-care bedroom, lit by the monitor and the room's established fluorescent lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Quiet ambient illumination and gentle facial contrast keep the reddened eyes visible without imposing an unmotivated monitor glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same shelter-room surfaces, bed, and daytime lighting. Exclude the doorway figure from the reference and the research laboratory's glass enclosure.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The video call remains connected; Amber is back in control of the phone after Raul's appearance. 현우: He remains at the telephone in his shelter room, watching the video display and becoming choked with emotion.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 화면을 뚫어져라 응시한 채 눈시울이 붉어진 현우의 얼굴.\n\nLOCATION (lock): At the video-call station inside the temporary-care bedroom, lit by the monitor and the room's established fluorescent lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Quiet ambient illumination and gentle facial contrast keep the reddened eyes visible without imposing an unmotivated monitor glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same shelter-room surfaces, bed, and daytime lighting. Exclude the doorway figure from the reference and the research laboratory's glass enclosure.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The video call remains connected; Amber is back in control of the phone after Raul's appearance. 현우: He remains at the telephone in his shelter room, watching the video display and becoming choked with emotion.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 화면을 뚫어져라 응시한 채 눈시울이 붉어진 현우의 얼굴.\n\nLOCATION (lock): At the video-call station inside the temporary-care bedroom, lit by the monitor and the room's established fluorescent lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Quiet ambient illumination and gentle facial contrast keep the reddened eyes visible without imposing an unmotivated monitor glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Use the same shelter-room surfaces, bed, and daytime lighting. Exclude the doorway figure from the reference and the research laboratory's glass enclosure.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The video call remains connected; Amber is back in control of the phone after Raul's appearance. 현우: He remains at the telephone in his shelter room, watching the video display and becoming choked with emotion.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선이 손에 들고 있는 스마트폰을 향하고 있음.",
    "built_space": "창문, 침대 프레임, 벽면 TV가 위치하며 참조된 병실 환경과 일치함.",
    "entities": "현우의 얼굴과 머리 모양이 참조와 일치하며, 눈시울이 붉어진 감정 표현이 잘 묘사됨.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 소품: 스마트폰의 후면(카메라 렌즈 옆)에 화상 통화 화면이 표시됨",
     "[gpt-high] 참고 공간의 침대 외에 별도의 침대가 하나 더 보이는 배치로, 고정 시설의 수가 달라졌다.",
     "[gpt-high] 휴대전화 뒷면에 현우 외 여성의 얼굴을 추가하여 등장인물 제한을 위반한다."
    ],
    "physics": "오른손이 스마트폰의 테두리를 안정적으로 쥐고 있음."
   },
   {
    "label": "B",
    "direction": "현우의 시선이 테이블 위 거치대에 놓인 스마트폰을 향함.",
    "built_space": "침대, 테이블, 벽걸이 TV 및 스위치가 참조 이미지의 병실과 일치하게 배치됨.",
    "entities": "현우의 인상착의가 참조와 일치하며 붉어진 눈과 눈물이 묘사됨.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 소품: 스마트폰의 후면(카메라 렌즈 반대편)에 화면 UI 이미지가 나타남",
     "[gpt-high] 휴대전화 뒷면에 현우 외 인물의 작은 사진이 추가되어, 이 쇼트에 지정되지 않은 얼굴이 나타난다."
    ],
    "physics": "스마트폰이 테이블 위 거치대에 의해 지지되고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "클로즈업 프레이밍과 감정 표현은 우수하나, 스마트폰 뒷면에 화면이 나타나는 치명적인 프롭 오류로 인해 실격됨."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "스마트폰 후면에 UI가 표시되는 물리적 오류가 발생했으며, 프레이밍이 지시된 클로즈업보다 다소 넓음."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선이 손에 들고 있는 스마트폰을 향하고 있음.",
        "built_space": "창문, 침대 프레임, 벽면 TV가 위치하며 참조된 병실 환경과 일치함.",
        "entities": "현우의 얼굴과 머리 모양이 참조와 일치하며, 눈시울이 붉어진 감정 표현이 잘 묘사됨.",
        "hard_violations": [
         "물리적으로 불가능한 소품: 스마트폰의 후면(카메라 렌즈 옆)에 화상 통화 화면이 표시됨"
        ],
        "physics": "오른손이 스마트폰의 테두리를 안정적으로 쥐고 있음."
       },
       {
        "label": "B",
        "direction": "현우의 시선이 테이블 위 거치대에 놓인 스마트폰을 향함.",
        "built_space": "침대, 테이블, 벽걸이 TV 및 스위치가 참조 이미지의 병실과 일치하게 배치됨.",
        "entities": "현우의 인상착의가 참조와 일치하며 붉어진 눈과 눈물이 묘사됨.",
        "hard_violations": [
         "물리적으로 불가능한 소품: 스마트폰의 후면(카메라 렌즈 반대편)에 화면 UI 이미지가 나타남"
        ],
        "physics": "스마트폰이 테이블 위 거치대에 의해 지지되고 있음."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "클로즈업 프레이밍과 감정 표현은 우수하나, 스마트폰 뒷면에 화면이 나타나는 치명적인 프롭 오류로 인해 실격됨."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "스마트폰 후면에 UI가 표시되는 물리적 오류가 발생했으며, 프레이밍이 지시된 클로즈업보다 다소 넓음."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선이 손에 들고 있는 스마트폰을 향하고 있음.",
        "built_space": "창문, 침대 프레임, 벽면 TV가 위치하며 참조된 병실 환경과 일치함.",
        "entities": "현우의 얼굴과 머리 모양이 참조와 일치하며, 눈시울이 붉어진 감정 표현이 잘 묘사됨.",
        "hard_violations": [
         "물리적으로 불가능한 소품: 스마트폰의 후면(카메라 렌즈 옆)에 화상 통화 화면이 표시됨"
        ],
        "physics": "오른손이 스마트폰의 테두리를 안정적으로 쥐고 있음."
       },
       {
        "label": "B",
        "direction": "현우의 시선이 테이블 위 거치대에 놓인 스마트폰을 향함.",
        "built_space": "침대, 테이블, 벽걸이 TV 및 스위치가 참조 이미지의 병실과 일치하게 배치됨.",
        "entities": "현우의 인상착의가 참조와 일치하며 붉어진 눈과 눈물이 묘사됨.",
        "hard_violations": [
         "물리적으로 불가능한 소품: 스마트폰의 후면(카메라 렌즈 반대편)에 화면 UI 이미지가 나타남"
        ],
        "physics": "스마트폰이 테이블 위 거치대에 의해 지지되고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "붉어진 눈시울과 침실의 연속성은 비교적 잘 살렸지만, 얼굴 클로즈업보다 넓고 휴대전화 뒷면에 불필요한 인물 사진이 추가됐다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "화면을 응시하는 시선은 더 명확하지만, 추가 침대와 휴대전화 뒷면의 여성 얼굴이 장소·등장인물 제한을 위반하며 손과 기기가 얼굴 클로즈업을 분산시킨다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 얼굴과 눈은 앞쪽 왼편의 휴대전화 방향을 향한다. 다만 눈동자가 향하는 높이는 낮게 놓인 휴대전화 화면보다 약간 높아 보여, 화면을 뚫어져라 응시한다는 관계가 선명하지 않다. 카메라에는 휴대전화의 후면 카메라가 보이고 표시 면은 현우 쪽이므로 기기의 앞뒤 방향은 적절하다.",
        "built_space": "왼쪽 위에 벽걸이 모니터 한 대와 옆 제어 패널 하나, 오른쪽 벽에 스위치 하나가 보인다. 전경에는 침대 한 개의 침구와 금속 난간, 그 위를 가로지르는 탁자가 있으며 현우 뒤에는 의자 등받이 일부가 보인다. 밝은 벽과 금속 침대는 장소 참고와 대체로 이어지지만, 얼굴뿐 아니라 탁자와 침대까지 크게 포함해 요구된 클로즈업보다 넓다.",
        "entities": "헝클어진 검은 머리, 앳된 동아시아계 남성의 얼굴, 남색 티셔츠는 현우 참고와 대체로 맞는다. 한국계 미국인이라는 국적 배경 자체는 외형으로 확인할 수 없다. 붉은 눈시울과 눈물 자국이 분명하고 안구는 정상적인 사람의 눈이다. 휴대전화 뒷면에는 작은 별도 인물 사진이 보인다. 읽을 수 있는 문자는 없다.",
        "hard_violations": [
         "휴대전화 뒷면에 현우 외 인물의 작은 사진이 추가되어, 이 쇼트에 지정되지 않은 얼굴이 나타난다."
        ],
        "physics": "휴대전화는 탁자 위 거치대가 받치고 있으며 떠 있지 않다. 현우는 뒤에 보이는 의자에 앉은 자세로 읽히고 목과 상체의 연결도 자연스럽다. 하체와 좌면 접점은 프레임 밖이므로 확인할 수 없다. 눈물은 볼 표면을 따라 내려와 물리적으로 자연스럽다."
       },
       {
        "label": "B",
        "direction": "현우의 눈은 오른쪽 앞에 든 휴대전화의 표시 면을 향하며 응시 대상이 명확하다. 후면 렌즈가 카메라 쪽에 있고 주 화면은 현우 쪽에 있어 사용 방향은 맞는다. 카메라에 보이는 여성 얼굴은 주 화면이 아니라 휴대전화 뒷면의 사진처럼 나타난다.",
        "built_space": "뒤쪽에 창 하나, 오른쪽 벽에 모니터 한 대와 제어 패널 하나, 위쪽에 형광등 한 개와 환기구가 보인다. 왼쪽 뒤 침대의 머리·발 쪽 난간 외에 오른쪽 아래에도 별도 침대의 난간과 목재 판이 들어와 침대가 두 개로 읽힌다. 참고의 단일 침대 공간과 달라진다. 얼굴은 A보다 크게 보이지만 상체와 큰 손, 휴대전화까지 포함하는 구도다.",
        "entities": "현우의 검은 머리와 남색 티셔츠, 피부 질감과 얼굴 특징은 인물 참고에 가깝다. 눈시울은 붉고 촉촉하지만 감정 표현은 A보다 절제되어 있다. 휴대전화 뒷면에는 긴 검은 머리의 여성 얼굴이 크게 보이며, 제외하도록 지정한 이전 쇼트의 여성과도 유사하다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "참고 공간의 침대 외에 별도의 침대가 하나 더 보이는 배치로, 고정 시설의 수가 달라졌다.",
         "휴대전화 뒷면에 현우 외 여성의 얼굴을 추가하여 등장인물 제한을 위반한다."
        ],
        "physics": "휴대전화는 손가락과 손바닥이 감싸 지지하고 있으며 손목과 팔의 연결도 가능한 자세다. 떠 있는 물체나 지지 없는 공중 자세는 없다. 몸의 좌면 접점과 발은 프레임 밖이어서 확인할 수 없지만 보이는 상체에는 물리적 모순이 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "붉어진 눈시울과 침실의 연속성은 비교적 잘 살렸지만, 얼굴 클로즈업보다 넓고 휴대전화 뒷면에 불필요한 인물 사진이 추가됐다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "화면을 응시하는 시선은 더 명확하지만, 추가 침대와 휴대전화 뒷면의 여성 얼굴이 장소·등장인물 제한을 위반하며 손과 기기가 얼굴 클로즈업을 분산시킨다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 얼굴과 눈은 앞쪽 왼편의 휴대전화 방향을 향한다. 다만 눈동자가 향하는 높이는 낮게 놓인 휴대전화 화면보다 약간 높아 보여, 화면을 뚫어져라 응시한다는 관계가 선명하지 않다. 카메라에는 휴대전화의 후면 카메라가 보이고 표시 면은 현우 쪽이므로 기기의 앞뒤 방향은 적절하다.",
        "built_space": "왼쪽 위에 벽걸이 모니터 한 대와 옆 제어 패널 하나, 오른쪽 벽에 스위치 하나가 보인다. 전경에는 침대 한 개의 침구와 금속 난간, 그 위를 가로지르는 탁자가 있으며 현우 뒤에는 의자 등받이 일부가 보인다. 밝은 벽과 금속 침대는 장소 참고와 대체로 이어지지만, 얼굴뿐 아니라 탁자와 침대까지 크게 포함해 요구된 클로즈업보다 넓다.",
        "entities": "헝클어진 검은 머리, 앳된 동아시아계 남성의 얼굴, 남색 티셔츠는 현우 참고와 대체로 맞는다. 한국계 미국인이라는 국적 배경 자체는 외형으로 확인할 수 없다. 붉은 눈시울과 눈물 자국이 분명하고 안구는 정상적인 사람의 눈이다. 휴대전화 뒷면에는 작은 별도 인물 사진이 보인다. 읽을 수 있는 문자는 없다.",
        "hard_violations": [
         "휴대전화 뒷면에 현우 외 인물의 작은 사진이 추가되어, 이 쇼트에 지정되지 않은 얼굴이 나타난다."
        ],
        "physics": "휴대전화는 탁자 위 거치대가 받치고 있으며 떠 있지 않다. 현우는 뒤에 보이는 의자에 앉은 자세로 읽히고 목과 상체의 연결도 자연스럽다. 하체와 좌면 접점은 프레임 밖이므로 확인할 수 없다. 눈물은 볼 표면을 따라 내려와 물리적으로 자연스럽다."
       },
       {
        "label": "A",
        "direction": "현우의 눈은 오른쪽 앞에 든 휴대전화의 표시 면을 향하며 응시 대상이 명확하다. 후면 렌즈가 카메라 쪽에 있고 주 화면은 현우 쪽에 있어 사용 방향은 맞는다. 카메라에 보이는 여성 얼굴은 주 화면이 아니라 휴대전화 뒷면의 사진처럼 나타난다.",
        "built_space": "뒤쪽에 창 하나, 오른쪽 벽에 모니터 한 대와 제어 패널 하나, 위쪽에 형광등 한 개와 환기구가 보인다. 왼쪽 뒤 침대의 머리·발 쪽 난간 외에 오른쪽 아래에도 별도 침대의 난간과 목재 판이 들어와 침대가 두 개로 읽힌다. 참고의 단일 침대 공간과 달라진다. 얼굴은 A보다 크게 보이지만 상체와 큰 손, 휴대전화까지 포함하는 구도다.",
        "entities": "현우의 검은 머리와 남색 티셔츠, 피부 질감과 얼굴 특징은 인물 참고에 가깝다. 눈시울은 붉고 촉촉하지만 감정 표현은 A보다 절제되어 있다. 휴대전화 뒷면에는 긴 검은 머리의 여성 얼굴이 크게 보이며, 제외하도록 지정한 이전 쇼트의 여성과도 유사하다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "참고 공간의 침대 외에 별도의 침대가 하나 더 보이는 배치로, 고정 시설의 수가 달라졌다.",
         "휴대전화 뒷면에 현우 외 여성의 얼굴을 추가하여 등장인물 제한을 위반한다."
        ],
        "physics": "휴대전화는 손가락과 손바닥이 감싸 지지하고 있으며 손목과 팔의 연결도 가능한 자세다. 떠 있는 물체나 지지 없는 공중 자세는 없다. 몸의 좌면 접점과 발은 프레임 밖이어서 확인할 수 없지만 보이는 상체에는 물리적 모순이 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.75,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.5,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 물리적으로 불가능한 소품: 스마트폰의 후면(카메라 렌즈 옆)에 화상 통화 화면이 표시됨",
     "[gpt-high] 참고 공간의 침대 외에 별도의 침대가 하나 더 보이는 배치로, 고정 시설의 수가 달라졌다.",
     "[gpt-high] 휴대전화 뒷면에 현우 외 여성의 얼굴을 추가하여 등장인물 제한을 위반한다."
    ],
    "B": [
     "[gemini-pro] 물리적으로 불가능한 소품: 스마트폰의 후면(카메라 렌즈 반대편)에 화면 UI 이미지가 나타남",
     "[gpt-high] 휴대전화 뒷면에 현우 외 인물의 작은 사진이 추가되어, 이 쇼트에 지정되지 않은 얼굴이 나타난다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1500,
   "B": 1750
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1500,
    "verdict_ko": "클로즈업 프레이밍과 감정 표현은 우수하나, 스마트폰 뒷면에 화면이 나타나는 치명적인 프롭 오류로 인해 실격됨.  ★위반: [gemini-pro] 물리적으로 불가능한 소품: 스마트폰의 후면(카메라 렌즈 옆)에 화상 통화 화면이 표시됨 / [gpt-high] 참고 공간의 침대 외에 별도의 침대가 하나 더 보이는 배치로, 고정 시설의 수가 달라졌다. / [gpt-high] 휴대전화 뒷면에 현우 외 여성의 얼굴을 추가하여 등장인물 제한을 위반한다."
   },
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "스마트폰 후면에 UI가 표시되는 물리적 오류가 발생했으며, 프레이밍이 지시된 클로즈업보다 다소 넓음.  ★위반: [gemini-pro] 물리적으로 불가능한 소품: 스마트폰의 후면(카메라 렌즈 반대편)에 화면 UI 이미지가 나타남 / [gpt-high] 휴대전화 뒷면에 현우 외 인물의 작은 사진이 추가되어, 이 쇼트에 지정되지 않은 얼굴이 나타난다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S81sh5_sel.png",
    "asset_id": "74401854-b7f6-443f-85a1-4a195785b1fa",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-f3c0-7368-b874-f75439803e2b",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S81sh5"
  }
 },
 "S83sh6::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T09:10:32.624788+00:00",
  "fingerprint": "af91564d6c12943472f8aa7ead5708f1cdd944e88a5699ec72f7a4a77ee3b28e",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S83sh6_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S83sh6_sel.png",
  "source_sha256": "4cfc1113dab311c09b2aa73c6ed5a56f15aba78a66ebadce20c2734ea6ac5dcc",
  "file": "S83sh6_cine.png",
  "staged_sha256": "2844318c6964bf2cebcca73ba9566987eb3af1379c16671c79b4fa666b574c6a",
  "latency_ms": 9607
 },
 "S83sh7::signage": {
  "fp": "6e2d91105960b25b",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S83sh7": {
  "input_fingerprint": "fbc03df5f7845bda",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 침대 위에 홀로 누운 채 멍하니 천장을 올려다보는 현우의 전신.\n\nLOCATION (lock): On the bed inside the research facility's temporary-care bedroom, beneath fluorescent lighting with daylight outside the sea-facing window. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Bed (Occupied only by 현우 after the call) — The upper surface and footward side are visible from the elevated diagonal position; used as Supports the full-body composition without overwhelming the surrounding space; Room space around the bed (현우 is alone); used as Negative space that increases with the retreat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained room ambience and soft tonal separation, allowing the empty space to feel quiet without inventing a change of time or light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the shelter room, bed, and daytime illumination. Exclude the researcher, who has left, and do not retain an active video-call image after the call ends.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The call has ended and the video image has disappeared. 현우: He is alone in his shelter room, lying on the bed and remaining visibly preoccupied.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 침대 위에 홀로 누운 채 멍하니 천장을 올려다보는 현우의 전신.\n\nLOCATION (lock): On the bed inside the research facility's temporary-care bedroom, beneath fluorescent lighting with daylight outside the sea-facing window. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Bed (Occupied only by 현우 after the call) — The upper surface and footward side are visible from the elevated diagonal position; used as Supports the full-body composition without overwhelming the surrounding space; Room space around the bed (현우 is alone); used as Negative space that increases with the retreat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained room ambience and soft tonal separation, allowing the empty space to feel quiet without inventing a change of time or light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the shelter room, bed, and daytime illumination. Exclude the researcher, who has left, and do not retain an active video-call image after the call ends.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The call has ended and the video image has disappeared. 현우: He is alone in his shelter room, lying on the bed and remaining visibly preoccupied.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 침대 위에 홀로 누운 채 멍하니 천장을 올려다보는 현우의 전신.\n\nLOCATION (lock): On the bed inside the research facility's temporary-care bedroom, beneath fluorescent lighting with daylight outside the sea-facing window. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Bed (Occupied only by 현우 after the call) — The upper surface and footward side are visible from the elevated diagonal position; used as Supports the full-body composition without overwhelming the surrounding space; Room space around the bed (현우 is alone); used as Negative space that increases with the retreat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained room ambience and soft tonal separation, allowing the empty space to feel quiet without inventing a change of time or light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the shelter room, bed, and daytime illumination. Exclude the researcher, who has left, and do not retain an active video-call image after the call ends.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The call has ended and the video image has disappeared. 현우: He is alone in his shelter room, lying on the bed and remaining visibly preoccupied.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "시선은 천장을 향해 위로 고정되어 있음.",
    "built_space": "창문이 왼쪽에 위치한 병실. 카메라는 대각선이 아닌 완전한 측면 프로필 구도를 취함. 침대 뒤쪽 벽에는 TV와 함께 레퍼런스에 1개만 존재하던 패널과 달리 가로형 의료 패널 2개와 세로형 패널이 복제되어 설치됨.",
    "entities": "현우가 짙은 파란색 티셔츠를 입고 홀로 누워 있으며, 얼굴과 착장은 레퍼런스와 일치함.",
    "hard_violations": [
     "[gemini-pro] 레퍼런스에 명시된 공간의 단일 패널 설정을 무시하고 다수의 벽면 의료 패널을 중복/추가 생성함 (Duplicated fittings).",
     "[gemini-pro] 명시된 '높은 대각선 위치(elevated diagonal position)'를 무시하고 평면적인 측면(profile) 뷰로 렌더링함."
    ],
    "physics": "침대 매트리스에 등을 대고 누워 하중이 올바르게 지지되고 있음."
   },
   {
    "label": "B",
    "direction": "시선은 멍하게 천장을 향해 위로 고정되어 있음.",
    "built_space": "침대가 방에 놓여 있고, 카메라는 지시된 높은 대각선 구도(elevated diagonal position)에서 침대 상단면을 내려다봄. 벽면의 구조가 다소 변경되어 TV와 창문이 우측 벽에 배치됨.",
    "entities": "현우가 짙은 파란색 티셔츠를 입고 홀로 누워 있으며, 인물의 인상착의가 레퍼런스와 매우 흡사함.",
    "hard_violations": [],
    "physics": "침대에 누워 몸이 매트리스에 완전히 지지되고 있으며, 왼손이 복부 위에 자연스럽게 놓여 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "명시된 높은 대각선 카메라 구도와 인물의 멍한 표정을 잘 구현했으나, 전신 와이드 샷이 아닌 하체가 잘린 타이트한 프레이밍이 아쉽습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "지시된 대각선 카메라 뷰를 완전히 무시한 측면 샷이며, 레퍼런스에 없는 다수의 벽면 패널을 중복 생성하여 심각한 공간 일관성 오류를 범했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 천장을 향해 위로 고정되어 있음.",
        "built_space": "창문이 왼쪽에 위치한 병실. 카메라는 대각선이 아닌 완전한 측면 프로필 구도를 취함. 침대 뒤쪽 벽에는 TV와 함께 레퍼런스에 1개만 존재하던 패널과 달리 가로형 의료 패널 2개와 세로형 패널이 복제되어 설치됨.",
        "entities": "현우가 짙은 파란색 티셔츠를 입고 홀로 누워 있으며, 얼굴과 착장은 레퍼런스와 일치함.",
        "hard_violations": [
         "레퍼런스에 명시된 공간의 단일 패널 설정을 무시하고 다수의 벽면 의료 패널을 중복/추가 생성함 (Duplicated fittings).",
         "명시된 '높은 대각선 위치(elevated diagonal position)'를 무시하고 평면적인 측면(profile) 뷰로 렌더링함."
        ],
        "physics": "침대 매트리스에 등을 대고 누워 하중이 올바르게 지지되고 있음."
       },
       {
        "label": "B",
        "direction": "시선은 멍하게 천장을 향해 위로 고정되어 있음.",
        "built_space": "침대가 방에 놓여 있고, 카메라는 지시된 높은 대각선 구도(elevated diagonal position)에서 침대 상단면을 내려다봄. 벽면의 구조가 다소 변경되어 TV와 창문이 우측 벽에 배치됨.",
        "entities": "현우가 짙은 파란색 티셔츠를 입고 홀로 누워 있으며, 인물의 인상착의가 레퍼런스와 매우 흡사함.",
        "hard_violations": [],
        "physics": "침대에 누워 몸이 매트리스에 완전히 지지되고 있으며, 왼손이 복부 위에 자연스럽게 놓여 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "명시된 높은 대각선 카메라 구도와 인물의 멍한 표정을 잘 구현했으나, 전신 와이드 샷이 아닌 하체가 잘린 타이트한 프레이밍이 아쉽습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "지시된 대각선 카메라 뷰를 완전히 무시한 측면 샷이며, 레퍼런스에 없는 다수의 벽면 패널을 중복 생성하여 심각한 공간 일관성 오류를 범했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 천장을 향해 위로 고정되어 있음.",
        "built_space": "창문이 왼쪽에 위치한 병실. 카메라는 대각선이 아닌 완전한 측면 프로필 구도를 취함. 침대 뒤쪽 벽에는 TV와 함께 레퍼런스에 1개만 존재하던 패널과 달리 가로형 의료 패널 2개와 세로형 패널이 복제되어 설치됨.",
        "entities": "현우가 짙은 파란색 티셔츠를 입고 홀로 누워 있으며, 얼굴과 착장은 레퍼런스와 일치함.",
        "hard_violations": [
         "레퍼런스에 명시된 공간의 단일 패널 설정을 무시하고 다수의 벽면 의료 패널을 중복/추가 생성함 (Duplicated fittings).",
         "명시된 '높은 대각선 위치(elevated diagonal position)'를 무시하고 평면적인 측면(profile) 뷰로 렌더링함."
        ],
        "physics": "침대 매트리스에 등을 대고 누워 하중이 올바르게 지지되고 있음."
       },
       {
        "label": "B",
        "direction": "시선은 멍하게 천장을 향해 위로 고정되어 있음.",
        "built_space": "침대가 방에 놓여 있고, 카메라는 지시된 높은 대각선 구도(elevated diagonal position)에서 침대 상단면을 내려다봄. 벽면의 구조가 다소 변경되어 TV와 창문이 우측 벽에 배치됨.",
        "entities": "현우가 짙은 파란색 티셔츠를 입고 홀로 누워 있으며, 인물의 인상착의가 레퍼런스와 매우 흡사함.",
        "hard_violations": [],
        "physics": "침대에 누워 몸이 매트리스에 완전히 지지되고 있으며, 왼손이 복부 위에 자연스럽게 놓여 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "높은 대각선 시점과 천장을 보는 현우의 정적인 연기는 맞지만, 발과 침대 발치가 잘려 핵심인 전신 와이드 숏을 충족하지 못한다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "홀로 천장을 보는 모습과 낮의 바다 창은 맞지만, 낮은 측면 시점과 화면을 채우는 신체 구도가 요구된 높은 대각선 전신 와이드 숏에서 더 멀다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 등을 대고 누워 얼굴과 눈을 천장 쪽으로 향한다. 카메라나 꺼진 벽면 화면을 보는 모습은 아니며, 멍하니 위를 보는 행동에 부합한다.",
        "built_space": "침대 하나의 금속 머리판과 가까운 쪽 측면 난간, 흰 침구가 보인다. 왼쪽 벽 상단에는 긴 의료 설비 패널 하나, 맞은편에는 꺼진 화면 하나와 오른쪽 조작부 하나, 그 아래에는 작은 벽면 접속판 세 개가 있다. 오른쪽 창 하나에서 낮빛이 들어온다. 높은 대각선에서 매트리스 윗면을 보지만 침대 발치와 인물의 발은 화면 밖이며, 침대가 화면 대부분을 차지해 주변 여백도 제한적이다. 참고의 밝은 벽과 금속 침대 재질은 이어지지만 바다는 확인되지 않는다.",
        "entities": "현우로 보이는 젊은 동아시아계 남성 한 명만 있다. 앳된 얼굴, 헝클어진 검은 머리, 남색 반팔 상의는 참고와 대체로 맞으며 국적은 외형만으로 확인할 수 없다. 하체는 흰 이불로 덮여 있다. 연구자나 다른 사람, 활성 영상통화 화면은 없다. 이전 숏의 휴대전화와 금속 탁자는 보이지 않는다. 판독 가능한 문구나 합성 표식은 확인되지 않는다.",
        "hard_violations": [],
        "physics": "머리와 등은 베개 및 매트리스에 지지되고, 한 손은 배 위에, 다른 팔과 손은 침구 위에 놓여 있다. 이불은 하체를 따라 부풀고 매트리스 위로 주름져 있어 지지 관계가 자연스럽다. 공중에 떠 있거나 지지 없이 놓인 신체·물체는 없다."
       },
       {
        "label": "B",
        "direction": "현우의 얼굴은 위로 향하고 눈도 천장을 바라본다. 왼쪽 창이나 벽면 화면을 응시하지 않으며, 누운 채 생각에 잠긴 행동은 맞는다.",
        "built_space": "침대 하나에 오른쪽 흰색·목재색 머리판과 왼쪽 금속 난간이 보인다. 왼쪽에는 바다가 보이는 분할 창 하나, 뒤쪽 벽에는 꺼진 화면 하나와 조작부, 가로 의료 설비 패널 두 개가 있으며 오른쪽 가장자리에도 일부 조작 장치가 보인다. 천장에는 직사각형 조명 세 개와 왼쪽 공조 장치 일부가 드러난다. 참고에서 확인되지 않은 설비가 많이 노출되지만 중복 설치라고 단정할 근거는 부족하다. 카메라는 매트리스보다 조금 높은 측면에 가까워 요구된 높은 대각선 시점이 아니며, 발끝도 왼쪽 밖으로 잘린다. 벽과 천장 여백은 있지만 인물과 침대가 전경을 크게 채운다.",
        "entities": "젊은 동아시아계 남성 한 명만 보이며 검은 머리, 남색 반팔 상의와 얼굴 윤곽은 현우 참고에 대체로 부합한다. 하체는 흰 이불 아래에 있다. 연구자나 통화 상대의 영상은 없고 바다와 낮빛은 확인된다. 침대 머리판의 혼합 재질은 참고의 금속 침대 인상과 차이가 있다. 이전 휴대전화는 확인되지 않으며 오른쪽 아래에는 금속 가구 일부가 보인다. 읽을 수 있는 글자는 확인되지 않는다.",
        "hard_violations": [],
        "physics": "머리와 몸통은 베개 및 매트리스에 받쳐져 있고, 보이는 팔은 침대 위에 놓여 손이 이불에 닿는다. 하체를 덮은 이불도 몸과 침대에 지지된다. 누운 자세에서 가능한 접촉 관계이며, 지지 없는 부유나 명백한 해부학적 불가능은 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "높은 대각선 시점과 천장을 보는 현우의 정적인 연기는 맞지만, 발과 침대 발치가 잘려 핵심인 전신 와이드 숏을 충족하지 못한다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "홀로 천장을 보는 모습과 낮의 바다 창은 맞지만, 낮은 측면 시점과 화면을 채우는 신체 구도가 요구된 높은 대각선 전신 와이드 숏에서 더 멀다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 등을 대고 누워 얼굴과 눈을 천장 쪽으로 향한다. 카메라나 꺼진 벽면 화면을 보는 모습은 아니며, 멍하니 위를 보는 행동에 부합한다.",
        "built_space": "침대 하나의 금속 머리판과 가까운 쪽 측면 난간, 흰 침구가 보인다. 왼쪽 벽 상단에는 긴 의료 설비 패널 하나, 맞은편에는 꺼진 화면 하나와 오른쪽 조작부 하나, 그 아래에는 작은 벽면 접속판 세 개가 있다. 오른쪽 창 하나에서 낮빛이 들어온다. 높은 대각선에서 매트리스 윗면을 보지만 침대 발치와 인물의 발은 화면 밖이며, 침대가 화면 대부분을 차지해 주변 여백도 제한적이다. 참고의 밝은 벽과 금속 침대 재질은 이어지지만 바다는 확인되지 않는다.",
        "entities": "현우로 보이는 젊은 동아시아계 남성 한 명만 있다. 앳된 얼굴, 헝클어진 검은 머리, 남색 반팔 상의는 참고와 대체로 맞으며 국적은 외형만으로 확인할 수 없다. 하체는 흰 이불로 덮여 있다. 연구자나 다른 사람, 활성 영상통화 화면은 없다. 이전 숏의 휴대전화와 금속 탁자는 보이지 않는다. 판독 가능한 문구나 합성 표식은 확인되지 않는다.",
        "hard_violations": [],
        "physics": "머리와 등은 베개 및 매트리스에 지지되고, 한 손은 배 위에, 다른 팔과 손은 침구 위에 놓여 있다. 이불은 하체를 따라 부풀고 매트리스 위로 주름져 있어 지지 관계가 자연스럽다. 공중에 떠 있거나 지지 없이 놓인 신체·물체는 없다."
       },
       {
        "label": "A",
        "direction": "현우의 얼굴은 위로 향하고 눈도 천장을 바라본다. 왼쪽 창이나 벽면 화면을 응시하지 않으며, 누운 채 생각에 잠긴 행동은 맞는다.",
        "built_space": "침대 하나에 오른쪽 흰색·목재색 머리판과 왼쪽 금속 난간이 보인다. 왼쪽에는 바다가 보이는 분할 창 하나, 뒤쪽 벽에는 꺼진 화면 하나와 조작부, 가로 의료 설비 패널 두 개가 있으며 오른쪽 가장자리에도 일부 조작 장치가 보인다. 천장에는 직사각형 조명 세 개와 왼쪽 공조 장치 일부가 드러난다. 참고에서 확인되지 않은 설비가 많이 노출되지만 중복 설치라고 단정할 근거는 부족하다. 카메라는 매트리스보다 조금 높은 측면에 가까워 요구된 높은 대각선 시점이 아니며, 발끝도 왼쪽 밖으로 잘린다. 벽과 천장 여백은 있지만 인물과 침대가 전경을 크게 채운다.",
        "entities": "젊은 동아시아계 남성 한 명만 보이며 검은 머리, 남색 반팔 상의와 얼굴 윤곽은 현우 참고에 대체로 부합한다. 하체는 흰 이불 아래에 있다. 연구자나 통화 상대의 영상은 없고 바다와 낮빛은 확인된다. 침대 머리판의 혼합 재질은 참고의 금속 침대 인상과 차이가 있다. 이전 휴대전화는 확인되지 않으며 오른쪽 아래에는 금속 가구 일부가 보인다. 읽을 수 있는 글자는 확인되지 않는다.",
        "hard_violations": [],
        "physics": "머리와 몸통은 베개 및 매트리스에 받쳐져 있고, 보이는 팔은 침대 위에 놓여 손이 이불에 닿는다. 하체를 덮은 이불도 몸과 침대에 지지된다. 누운 자세에서 가능한 접촉 관계이며, 지지 없는 부유나 명백한 해부학적 불가능은 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.333,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.083,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 레퍼런스에 명시된 공간의 단일 패널 설정을 무시하고 다수의 벽면 의료 패널을 중복/추가 생성함 (Duplicated fittings).",
     "[gemini-pro] 명시된 '높은 대각선 위치(elevated diagonal position)'를 무시하고 평면적인 측면(profile) 뷰로 렌더링함."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1083
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "명시된 높은 대각선 카메라 구도와 인물의 멍한 표정을 잘 구현했으나, 전신 와이드 샷이 아닌 하체가 잘린 타이트한 프레이밍이 아쉽습니다."
   },
   {
    "label": "A",
    "score": 1083,
    "verdict_ko": "지시된 대각선 카메라 뷰를 완전히 무시한 측면 샷이며, 레퍼런스에 없는 다수의 벽면 패널을 중복 생성하여 심각한 공간 일관성 오류를 범했습니다.  ★위반: [gemini-pro] 레퍼런스에 명시된 공간의 단일 패널 설정을 무시하고 다수의 벽면 의료 패널을 중복/추가 생성함 (Duplicated fittings). / [gemini-pro] 명시된 '높은 대각선 위치(elevated diagonal position)'를 무시하고 평면적인 측면(profile) 뷰로 렌더링함."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S83sh6_sel.png",
    "asset_id": "d874ef0d-b459-42b2-b02f-3c403ae0f863",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-f566-7afc-996b-8a5f5462c201",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S83sh6"
  }
 },
 "S83sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T09:11:56.861060+00:00",
  "fingerprint": "af40fdf9e4d40b709333a356a70244ef0e65b544d42315731e3122bdbbf640a3",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S83sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S83sh7_sel.png",
  "source_sha256": "b3a3c12f6a0d181c77795b341965473af0bcf80b3f371303eff775f0a5b4e264",
  "file": "S83sh7_cine.png",
  "staged_sha256": "25e1574ae11687e38eb1b51867c49aaa56ce9863efd75b8d228d8734eb48543b",
  "latency_ms": 8220
 },
 "S84sh1::signage": {
  "fp": "f4cb2978eafeee6a",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S84sh1": {
  "input_fingerprint": "5fc2d9ca2eea3397",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 푸른 조명이 감도는 실험실의 차가운 침대 위, 눈을 번쩍 뜬 찰리의 낡은 금속 얼굴 클로즈업.\n\nLOCATION (lock): On the examination bed inside the research laboratory's glass enclosure, under the blue lighting specified in the shot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Laboratory bed (Still supporting 찰리 at the instant of awakening) — Only a narrow portion of its upper surface appears beside his head; used as Grounds the face in its reclining position; Attached wires (Still connected before 찰리 removes them); used as Small lower-edge details that establish his treatment context.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued blue laboratory illumination reveals the worn metal face and newly opened eyes without adding an unsupported eye-light color or flare.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent and motionless on the stainless-steel examination bed inside the circular glass enclosure, his body supported by the bed while robotic arms scan him. The source does not specify his head's direction, his torso's upward-facing surface, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie is still lying on the stainless-steel bed inside the laboratory's glass enclosure, with multiple wires attached across the body. The eyes have just powered on; the wires have not yet been removed.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 푸른 조명이 감도는 실험실의 차가운 침대 위, 눈을 번쩍 뜬 찰리의 낡은 금속 얼굴 클로즈업.\n\nLOCATION (lock): On the examination bed inside the research laboratory's glass enclosure, under the blue lighting specified in the shot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Laboratory bed (Still supporting 찰리 at the instant of awakening) — Only a narrow portion of its upper surface appears beside his head; used as Grounds the face in its reclining position; Attached wires (Still connected before 찰리 removes them); used as Small lower-edge details that establish his treatment context.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued blue laboratory illumination reveals the worn metal face and newly opened eyes without adding an unsupported eye-light color or flare.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent and motionless on the stainless-steel examination bed inside the circular glass enclosure, his body supported by the bed while robotic arms scan him. The source does not specify his head's direction, his torso's upward-facing surface, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie is still lying on the stainless-steel bed inside the laboratory's glass enclosure, with multiple wires attached across the body. The eyes have just powered on; the wires have not yet been removed.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 푸른 조명이 감도는 실험실의 차가운 침대 위, 눈을 번쩍 뜬 찰리의 낡은 금속 얼굴 클로즈업.\n\nLOCATION (lock): On the examination bed inside the research laboratory's glass enclosure, under the blue lighting specified in the shot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Laboratory bed (Still supporting 찰리 at the instant of awakening) — Only a narrow portion of its upper surface appears beside his head; used as Grounds the face in its reclining position; Attached wires (Still connected before 찰리 removes them); used as Small lower-edge details that establish his treatment context.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued blue laboratory illumination reveals the worn metal face and newly opened eyes without adding an unsupported eye-light color or flare.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent and motionless on the stainless-steel examination bed inside the circular glass enclosure, his body supported by the bed while robotic arms scan him. The source does not specify his head's direction, his torso's upward-facing surface, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie is still lying on the stainless-steel bed inside the laboratory's glass enclosure, with multiple wires attached across the body. The eyes have just powered on; the wires have not yet been removed.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "눈동자의 시선이 위쪽 허공을 향하고 있음.",
    "built_space": "배경에 유리 벽면과 모니터 등 푸른빛이 감도는 실험실 환경이 흐릿하게 보임.",
    "entities": "지문대로 샌드 베이지색 금속 헬멧과 점/선이 있는 흰색 마스크를 착용한 찰리가 금속 침대에 누워 있으며, 측면에 여러 가닥의 전선이 연결됨.",
    "hard_violations": [
     "[gpt-high] 찰리만 나와야 하는 장면에 화면 오른쪽의 흰 상의 연구원이 남아 있습니다."
    ],
    "physics": "머리와 전선들이 금속 침대 표면에 물리적으로 안정감 있게 닿아 있음."
   },
   {
    "label": "B",
    "direction": "불빛이 켜진 눈이 위쪽을 향해 있음.",
    "built_space": "배경에 로봇 팔과 실험실의 구조물들이 블러 처리되어 나타남.",
    "entities": "침대 위에 누워 전선이 연결된 로봇 형태이나, 얼굴 전체가 단일한 베이지색으로 처리되어 흰 마스크와 점/선 특징이 부재함.",
    "hard_violations": [
     "[gpt-high] 찰리만 나와야 하는 장면에 흰 상의의 배경 연구원이 남아 있습니다."
    ],
    "physics": "금속 머리가 침대 면에 지지되어 있으며, 연결된 전선들도 중력에 따라 자연스럽게 놓여 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "레퍼런스와 지문에 명시된 '점과 선으로 구성된 흰 마스크' 디테일을 정확히 구현했으며, 요구된 클로즈업 프레이밍과 조명 분위기를 훌륭하게 살렸습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "지문이 요구한 얼굴의 흰색 마스크와 점/선 디테일이 완전히 누락되었고, 금지된 강한 안광 플레어 효과가 추가되어 아쉽습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "눈동자의 시선이 위쪽 허공을 향하고 있음.",
        "built_space": "배경에 유리 벽면과 모니터 등 푸른빛이 감도는 실험실 환경이 흐릿하게 보임.",
        "entities": "지문대로 샌드 베이지색 금속 헬멧과 점/선이 있는 흰색 마스크를 착용한 찰리가 금속 침대에 누워 있으며, 측면에 여러 가닥의 전선이 연결됨.",
        "hard_violations": [],
        "physics": "머리와 전선들이 금속 침대 표면에 물리적으로 안정감 있게 닿아 있음."
       },
       {
        "label": "B",
        "direction": "불빛이 켜진 눈이 위쪽을 향해 있음.",
        "built_space": "배경에 로봇 팔과 실험실의 구조물들이 블러 처리되어 나타남.",
        "entities": "침대 위에 누워 전선이 연결된 로봇 형태이나, 얼굴 전체가 단일한 베이지색으로 처리되어 흰 마스크와 점/선 특징이 부재함.",
        "hard_violations": [],
        "physics": "금속 머리가 침대 면에 지지되어 있으며, 연결된 전선들도 중력에 따라 자연스럽게 놓여 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "레퍼런스와 지문에 명시된 '점과 선으로 구성된 흰 마스크' 디테일을 정확히 구현했으며, 요구된 클로즈업 프레이밍과 조명 분위기를 훌륭하게 살렸습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "지문이 요구한 얼굴의 흰색 마스크와 점/선 디테일이 완전히 누락되었고, 금지된 강한 안광 플레어 효과가 추가되어 아쉽습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "눈동자의 시선이 위쪽 허공을 향하고 있음.",
        "built_space": "배경에 유리 벽면과 모니터 등 푸른빛이 감도는 실험실 환경이 흐릿하게 보임.",
        "entities": "지문대로 샌드 베이지색 금속 헬멧과 점/선이 있는 흰색 마스크를 착용한 찰리가 금속 침대에 누워 있으며, 측면에 여러 가닥의 전선이 연결됨.",
        "hard_violations": [],
        "physics": "머리와 전선들이 금속 침대 표면에 물리적으로 안정감 있게 닿아 있음."
       },
       {
        "label": "B",
        "direction": "불빛이 켜진 눈이 위쪽을 향해 있음.",
        "built_space": "배경에 로봇 팔과 실험실의 구조물들이 블러 처리되어 나타남.",
        "entities": "침대 위에 누워 전선이 연결된 로봇 형태이나, 얼굴 전체가 단일한 베이지색으로 처리되어 흰 마스크와 점/선 특징이 부재함.",
        "hard_violations": [],
        "physics": "금속 머리가 침대 면에 지지되어 있으며, 연결된 전선들도 중력에 따라 자연스럽게 놓여 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "완전한 금속 얼굴과 푸른 조명은 B보다 충실하지만, 배경 연구원이 남아 인물 제한을 위반하고 침대와 배선도 지정된 좁은 범위보다 크게 보입니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "배경 연구원이 남아 인물 제한을 위반하며, 눈 주변의 인간 피부와 발광 눈이 참조의 금속 얼굴 및 절제된 각성 묘사에서 벗어납니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 얼굴과 두 눈은 누운 자세에서 천장 쪽을 향합니다. 특정 인물을 바라보거나 카메라를 응시하는 모습은 아니며, 요구된 각성 순간과 방향상 충돌하지 않습니다. 무기나 겨냥하는 물체는 없습니다.",
        "built_space": "스테인리스 검사대 한 개가 머리 아래에서 화면 양옆으로 넓게 드러납니다. 왼쪽 배경에 로봇 장치 받침 한 개, 뒤쪽에 실험실 작업대와 흐릿한 모니터들이 보입니다. 원형 유리 enclosure의 전체 구조는 이 구도로 확인할 수 없습니다. 상단 중앙 오른쪽에는 흰 상의를 입고 검은 의자에 앉은 연구원의 뒷모습이 흐릿하게 남아 있습니다. 침대 표면은 요구된 머리 옆의 좁은 부분보다 훨씬 넓습니다.",
        "entities": "주인공은 긁히고 도장이 벗겨진 샌드 베이지색 금속 머리와 장갑을 가진 로봇입니다. 인간의 나이·성별·민족성을 판별할 피부는 보이지 않습니다. 참조의 기계적 정체성은 유지하지만, 얼굴은 흰 점·선 마스크보다는 굵직한 각면 장갑 형태입니다. 두 눈은 강한 백색 발광체로 표현되어 절제된 눈 묘사보다 과합니다. 연결된 케이블 여러 가닥과 금속 침대가 있으며, 케이블은 작은 하단 디테일에 그치지 않습니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "찰리만 나와야 하는 장면에 흰 상의의 배경 연구원이 남아 있습니다."
        ],
        "physics": "찰리의 후두부는 침대에 닿아 있고 목은 기계 관절로 몸통에 연결되어 있습니다. 보이는 상체도 침대에 누운 연속적인 자세입니다. 케이블은 머리 옆 단자에 연결되고 나머지는 침대 표면에 놓이거나 가장자리로 처집니다. 지지 없이 떠 있는 신체나 물체는 보이지 않습니다."
       },
       {
        "label": "B",
        "direction": "찰리의 두 눈은 천장과 화면 오른쪽 위를 향합니다. 얼굴 역시 위를 향한 채 누워 있어 시선 방향 자체는 각성 장면에 맞습니다. 특정 목표를 겨냥하는 물체는 없습니다.",
        "built_space": "검사대 한 개의 상판과 테두리가 머리 양옆으로 넓게 보이고, 왼쪽 가장자리에 장치 받침 한 개가 잘려 있습니다. 뒤쪽에는 실험실 작업대와 흐릿한 모니터 두 개가 보입니다. 화면 맨 오른쪽에는 검은 머리와 흰 상의를 가진 연구원이 흐릿하게 남아 있습니다. 유리 enclosure의 원형 구조는 확인하기 어렵고, 침대와 가슴 장갑이 차지하는 면적은 지정된 얼굴 중심 구도보다 큽니다.",
        "entities": "낡은 베이지 장갑, 밝은 마스크형 안면판, 점 모양 구멍과 입선은 보입니다. 그러나 눈 주변에 주름진 인간 피부와 눈꺼풀이 넓게 드러나 참조의 금속 로봇 얼굴을 인간이 쓴 가면처럼 바꿉니다. 이 부분만으로 민족성이나 성별은 확정할 수 없습니다. 두 눈에는 밝은 청백색 발광 무늬가 있습니다. 머리에 연결된 여러 전선과 스테인리스 침대는 존재하며, 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "찰리만 나와야 하는 장면에 화면 오른쪽의 흰 상의 연구원이 남아 있습니다."
        ],
        "physics": "머리 뒤쪽은 침대에 기대고 기계식 목이 머리와 가슴을 연결합니다. 보이는 몸통은 검사대 위에 누워 있으며, 머리가 지지 없이 공중에 떠 있다고 볼 근거는 없습니다. 전선은 측두부 단자에서 내려와 침대 위에 걸쳐 있고 중력에 맞게 휘어 있습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "완전한 금속 얼굴과 푸른 조명은 B보다 충실하지만, 배경 연구원이 남아 인물 제한을 위반하고 침대와 배선도 지정된 좁은 범위보다 크게 보입니다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "배경 연구원이 남아 인물 제한을 위반하며, 눈 주변의 인간 피부와 발광 눈이 참조의 금속 얼굴 및 절제된 각성 묘사에서 벗어납니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 얼굴과 두 눈은 누운 자세에서 천장 쪽을 향합니다. 특정 인물을 바라보거나 카메라를 응시하는 모습은 아니며, 요구된 각성 순간과 방향상 충돌하지 않습니다. 무기나 겨냥하는 물체는 없습니다.",
        "built_space": "스테인리스 검사대 한 개가 머리 아래에서 화면 양옆으로 넓게 드러납니다. 왼쪽 배경에 로봇 장치 받침 한 개, 뒤쪽에 실험실 작업대와 흐릿한 모니터들이 보입니다. 원형 유리 enclosure의 전체 구조는 이 구도로 확인할 수 없습니다. 상단 중앙 오른쪽에는 흰 상의를 입고 검은 의자에 앉은 연구원의 뒷모습이 흐릿하게 남아 있습니다. 침대 표면은 요구된 머리 옆의 좁은 부분보다 훨씬 넓습니다.",
        "entities": "주인공은 긁히고 도장이 벗겨진 샌드 베이지색 금속 머리와 장갑을 가진 로봇입니다. 인간의 나이·성별·민족성을 판별할 피부는 보이지 않습니다. 참조의 기계적 정체성은 유지하지만, 얼굴은 흰 점·선 마스크보다는 굵직한 각면 장갑 형태입니다. 두 눈은 강한 백색 발광체로 표현되어 절제된 눈 묘사보다 과합니다. 연결된 케이블 여러 가닥과 금속 침대가 있으며, 케이블은 작은 하단 디테일에 그치지 않습니다. 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "찰리만 나와야 하는 장면에 흰 상의의 배경 연구원이 남아 있습니다."
        ],
        "physics": "찰리의 후두부는 침대에 닿아 있고 목은 기계 관절로 몸통에 연결되어 있습니다. 보이는 상체도 침대에 누운 연속적인 자세입니다. 케이블은 머리 옆 단자에 연결되고 나머지는 침대 표면에 놓이거나 가장자리로 처집니다. 지지 없이 떠 있는 신체나 물체는 보이지 않습니다."
       },
       {
        "label": "A",
        "direction": "찰리의 두 눈은 천장과 화면 오른쪽 위를 향합니다. 얼굴 역시 위를 향한 채 누워 있어 시선 방향 자체는 각성 장면에 맞습니다. 특정 목표를 겨냥하는 물체는 없습니다.",
        "built_space": "검사대 한 개의 상판과 테두리가 머리 양옆으로 넓게 보이고, 왼쪽 가장자리에 장치 받침 한 개가 잘려 있습니다. 뒤쪽에는 실험실 작업대와 흐릿한 모니터 두 개가 보입니다. 화면 맨 오른쪽에는 검은 머리와 흰 상의를 가진 연구원이 흐릿하게 남아 있습니다. 유리 enclosure의 원형 구조는 확인하기 어렵고, 침대와 가슴 장갑이 차지하는 면적은 지정된 얼굴 중심 구도보다 큽니다.",
        "entities": "낡은 베이지 장갑, 밝은 마스크형 안면판, 점 모양 구멍과 입선은 보입니다. 그러나 눈 주변에 주름진 인간 피부와 눈꺼풀이 넓게 드러나 참조의 금속 로봇 얼굴을 인간이 쓴 가면처럼 바꿉니다. 이 부분만으로 민족성이나 성별은 확정할 수 없습니다. 두 눈에는 밝은 청백색 발광 무늬가 있습니다. 머리에 연결된 여러 전선과 스테인리스 침대는 존재하며, 읽을 수 있는 글자는 없습니다.",
        "hard_violations": [
         "찰리만 나와야 하는 장면에 화면 오른쪽의 흰 상의 연구원이 남아 있습니다."
        ],
        "physics": "머리 뒤쪽은 침대에 기대고 기계식 목이 머리와 가슴을 연결합니다. 보이는 몸통은 검사대 위에 누워 있으며, 머리가 지지 없이 공중에 떠 있다고 볼 근거는 없습니다. 전선은 측두부 단자에서 내려와 침대 위에 걸쳐 있고 중력에 맞게 휘어 있습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.667,
    "B": 1.571
   },
   "adjusted": {
    "A": 1.417,
    "B": 1.321
   },
   "violations": {
    "B": [
     "[gpt-high] 찰리만 나와야 하는 장면에 흰 상의의 배경 연구원이 남아 있습니다."
    ],
    "A": [
     "[gpt-high] 찰리만 나와야 하는 장면에 화면 오른쪽의 흰 상의 연구원이 남아 있습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1417,
   "B": 1321
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1417,
    "verdict_ko": "레퍼런스와 지문에 명시된 '점과 선으로 구성된 흰 마스크' 디테일을 정확히 구현했으며, 요구된 클로즈업 프레이밍과 조명 분위기를 훌륭하게 살렸습니다.  ★위반: [gpt-high] 찰리만 나와야 하는 장면에 화면 오른쪽의 흰 상의 연구원이 남아 있습니다."
   },
   {
    "label": "B",
    "score": 1321,
    "verdict_ko": "지문이 요구한 얼굴의 흰색 마스크와 점/선 디테일이 완전히 누락되었고, 금지된 강한 안광 플레어 효과가 추가되어 아쉽습니다.  ★위반: [gpt-high] 찰리만 나와야 하는 장면에 흰 상의의 배경 연구원이 남아 있습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S82sh14_sel.png",
    "asset_id": "73257111-705a-49ad-9ee9-a8e062fa4e86",
    "role": "prev_still"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-f714-7402-9437-1aa95bffaeb8",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S82sh14"
  },
  "locked_char_refs_excluded": [
   "찰리(C06)"
  ]
 },
 "S84sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T12:12:35.728482+00:00",
  "fingerprint": "8ec59347ba194477177c88a4515103fb6d656fcad7e0f8f2eecad803da6b878e",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S84sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S84sh1_sel.png",
  "source_sha256": "ce0de049e74c4b48e80bcb48c500e4a2a5b95a66e3d83e42b3383666a22c5282",
  "file": "S84sh1_cine.png",
  "staged_sha256": "a0a38c8d9fc983dbaab80d3f3fc24289f53f80d65030855f7bc5cbf0ce091e99",
  "latency_ms": 10426
 },
 "S84sh11::signage": {
  "fp": "928caa1df27ed571",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::e4b209e943d61d96": {
  "subjects": [],
  "subject_text": "제주도 연구소 실험실과 유리벽 관찰 구역\n중앙 원형 유리벽 안에 스테인리스 침대와 기계 팔이 설치된 실험실. 바깥에는 컴퓨터 작업대와 관찰 공간이 둘러져 있다.",
  "identity": "canonical",
  "scope_id": "L269",
  "scope_role": "location_interior",
  "scope_sha": "17b47caf6574618e"
 },
 "S84sh11::bgfirst_bg": {
  "input_fingerprint": "f4e1e745bb3245d1",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 자신 앞까지 다가온 찰리를 올려다보며 온화한 미소를 지은 채 정지해 있는 지소영의 얼굴 클로즈업.\n\nLOCATION (lock): Beside the old piano in the central hall of the research director's suite, under subdued nighttime interior lighting.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old piano (Playing has paused while 지소영 addresses 찰리) — A small oblique section of the keyboard side appears at the lower edge; used as Locates her seated posture and preserves continuity with the interrupted music.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate interior ambience and gentle facial contrast preserve the tenderness of the pause without adding an unsupported warm source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 자신 앞까지 다가온 찰리를 올려다보며 온화한 미소를 지은 채 정지해 있는 지소영의 얼굴 클로즈업.\n\nLOCATION (lock): Beside the old piano in the central hall of the research director's suite, under subdued nighttime interior lighting.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old piano (Playing has paused while 지소영 addresses 찰리) — A small oblique section of the keyboard side appears at the lower edge; used as Locates her seated posture and preserves continuity with the interrupted music.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate interior ambience and gentle facial contrast preserve the tenderness of the pause without adding an unsupported warm source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S84sh11__bgfirst_bg.png",
  "asset_id": "d36617d7-e206-4f23-93f4-d6161a0fbc32",
  "input_asset_ids": [
   "b7d06adc-c73d-44f7-8b3e-12527f45e14c",
   "2d7bcd3a-3137-473f-b939-ac3ee04b6230"
  ]
 },
 "S84sh11": {
  "input_fingerprint": "3916d5e4110ff75b",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 자신 앞까지 다가온 찰리를 올려다보며 온화한 미소를 지은 채 정지해 있는 지소영의 얼굴 클로즈업.\n\nLOCATION (lock): Beside the old piano in the central hall of the research director's suite, under subdued nighttime interior lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old piano (Playing has paused while 지소영 addresses 찰리) — A small oblique section of the keyboard side appears at the lower edge; used as Locates her seated posture and preserves continuity with the interrupted music.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate interior ambience and gentle facial contrast preserve the tenderness of the pause without adding an unsupported warm source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): An old piano stands in the central hall of the modern director's office. Charlie is now mobile beside the piano, with the laboratory wires detached and the glass-enclosure exit left open. 지소영: She remains at the piano in her neat research coat, having paused her playing.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 자신 앞까지 다가온 찰리를 올려다보며 온화한 미소를 지은 채 정지해 있는 지소영의 얼굴 클로즈업.\n\nLOCATION (lock): Beside the old piano in the central hall of the research director's suite, under subdued nighttime interior lighting. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old piano (Playing has paused while 지소영 addresses 찰리) — A small oblique section of the keyboard side appears at the lower edge; used as Locates her seated posture and preserves continuity with the interrupted music.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate interior ambience and gentle facial contrast preserve the tenderness of the pause without adding an unsupported warm source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): An old piano stands in the central hall of the modern director's office. Charlie is now mobile beside the piano, with the laboratory wires detached and the glass-enclosure exit left open. 지소영: She remains at the piano in her neat research coat, having paused her playing.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 자신 앞까지 다가온 찰리를 올려다보며 온화한 미소를 지은 채 정지해 있는 지소영의 얼굴 클로즈업.\n\nLOCATION (lock): Beside the old piano in the central hall of the research director's suite, under subdued nighttime interior lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Old piano (Playing has paused while 지소영 addresses 찰리) — A small oblique section of the keyboard side appears at the lower edge; used as Locates her seated posture and preserves continuity with the interrupted music.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate interior ambience and gentle facial contrast preserve the tenderness of the pause without adding an unsupported warm source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): An old piano stands in the central hall of the modern director's office. Charlie is now mobile beside the piano, with the laboratory wires detached and the glass-enclosure exit left open. 지소영: She remains at the piano in her neat research coat, having paused her playing.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S84sh11__bgfirst_bg.png",
     "asset_id": "d36617d7-e206-4f23-93f4-d6161a0fbc32",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S84sh11.png",
     "asset_id": "b7d06adc-c73d-44f7-8b3e-12527f45e14c",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 지소영: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1243508>",
     "asset_id": "c7496f13-cfcf-44a5-976d-96c783d20580",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L269B03.png",
     "asset_id": "2d7bcd3a-3137-473f-b939-ac3ee04b6230",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 지소영: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1243508>",
     "asset_id": "c7496f13-cfcf-44a5-976d-96c783d20580",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "지소영이 전경에 있는 남성(찰리)을 향해 고개를 들어 시선을 명확히 맞추고 있음.",
    "built_space": "배경에 원형 유리벽과 로봇 팔 등 레퍼런스 사진의 연구실 내부 구조가 정확한 위치에 렌더링됨.",
    "entities": "지소영의 얼굴과 헤어스타일이 레퍼런스와 일치하며, 우측 하단에 피아노 건반 일부가 작게 배치됨.",
    "hard_violations": [
     "[gpt-high] 가시 인물을 지소영으로 한정한 명시적 조건과 달리, 왼쪽 전경에 성인 남성의 상체를 크게 추가했다."
    ],
    "physics": "지소영이 피아노 앞에 앉아 고개를 든 자세와 전경 인물이 서 있는 모습이 물리적으로 자연스러움."
   },
   {
    "label": "B",
    "direction": "지소영이 위를 올려다보고 있으나 시선의 대상(찰리)이 프레임 내에 존재하지 않음.",
    "built_space": "연구실 배경이 흐릿하게 묘사되어 원형 유리 구조 등의 핵심 공간 특징이 잘 드러나지 않음.",
    "entities": "지소영의 인상착의는 일치하나, 거대한 목재 피아노가 화면 우측을 크게 가리고 있어 작게 배치하라는 지시에 어긋남.",
    "hard_violations": [],
    "physics": "자리에 앉아 고개를 든 자세를 지탱하는 신체 균형은 자연스러움."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "찰리의 뒷모습을 전경에 배치해 시선을 맞추는 연출을 훌륭히 살렸으며, 화면 하단에 작게 걸친 피아노와 배경 구조물도 지시에 부합합니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "시선의 대상인 찰리가 등장하지 않아 상호작용이 느껴지지 않으며, 피아노가 지시와 다르게 화면 우측을 과도하게 차지하고 있습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "지소영이 전경에 있는 남성(찰리)을 향해 고개를 들어 시선을 명확히 맞추고 있음.",
        "built_space": "배경에 원형 유리벽과 로봇 팔 등 레퍼런스 사진의 연구실 내부 구조가 정확한 위치에 렌더링됨.",
        "entities": "지소영의 얼굴과 헤어스타일이 레퍼런스와 일치하며, 우측 하단에 피아노 건반 일부가 작게 배치됨.",
        "hard_violations": [],
        "physics": "지소영이 피아노 앞에 앉아 고개를 든 자세와 전경 인물이 서 있는 모습이 물리적으로 자연스러움."
       },
       {
        "label": "B",
        "direction": "지소영이 위를 올려다보고 있으나 시선의 대상(찰리)이 프레임 내에 존재하지 않음.",
        "built_space": "연구실 배경이 흐릿하게 묘사되어 원형 유리 구조 등의 핵심 공간 특징이 잘 드러나지 않음.",
        "entities": "지소영의 인상착의는 일치하나, 거대한 목재 피아노가 화면 우측을 크게 가리고 있어 작게 배치하라는 지시에 어긋남.",
        "hard_violations": [],
        "physics": "자리에 앉아 고개를 든 자세를 지탱하는 신체 균형은 자연스러움."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "찰리의 뒷모습을 전경에 배치해 시선을 맞추는 연출을 훌륭히 살렸으며, 화면 하단에 작게 걸친 피아노와 배경 구조물도 지시에 부합합니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "시선의 대상인 찰리가 등장하지 않아 상호작용이 느껴지지 않으며, 피아노가 지시와 다르게 화면 우측을 과도하게 차지하고 있습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "지소영이 전경에 있는 남성(찰리)을 향해 고개를 들어 시선을 명확히 맞추고 있음.",
        "built_space": "배경에 원형 유리벽과 로봇 팔 등 레퍼런스 사진의 연구실 내부 구조가 정확한 위치에 렌더링됨.",
        "entities": "지소영의 얼굴과 헤어스타일이 레퍼런스와 일치하며, 우측 하단에 피아노 건반 일부가 작게 배치됨.",
        "hard_violations": [],
        "physics": "지소영이 피아노 앞에 앉아 고개를 든 자세와 전경 인물이 서 있는 모습이 물리적으로 자연스러움."
       },
       {
        "label": "B",
        "direction": "지소영이 위를 올려다보고 있으나 시선의 대상(찰리)이 프레임 내에 존재하지 않음.",
        "built_space": "연구실 배경이 흐릿하게 묘사되어 원형 유리 구조 등의 핵심 공간 특징이 잘 드러나지 않음.",
        "entities": "지소영의 인상착의는 일치하나, 거대한 목재 피아노가 화면 우측을 크게 가리고 있어 작게 배치하라는 지시에 어긋남.",
        "hard_violations": [],
        "physics": "자리에 앉아 고개를 든 자세를 지탱하는 신체 균형은 자연스러움."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "지소영의 정체성·연구복·온화한 미소와 올려다보는 동작은 맞지만, 얼굴 클로즈업보다 넓고 피아노가 하단의 작은 일부가 아니라 오른쪽 전경을 크게 차지한다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "찰리를 올려다보는 관계는 명확하지만, 가시 인물을 지소영으로 제한한 조건과 달리 남성을 크게 넣었고 얼굴 클로즈업을 넓은 어깨너머 구도로 바꿨다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "지소영은 턱을 조금 들고 화면 오른쪽 위의 프레임 밖을 바라본다. 다가온 찰리를 올려다보는 방향으로 해석할 수 있으나 상대는 보이지 않아 실제 시선 도착점은 확인할 수 없다. 건반을 내려다보거나 카메라를 응시하지는 않는다.",
        "built_space": "오른쪽에 오래된 목재 업라이트 피아노 한 대와 건반 한 벌이 보인다. 뒤에는 유리 칸막이와 금속 기둥, 실험 장비 및 모니터가 있어 장소 참조의 차가운 연구시설 재질과 이어진다. 지소영은 건반 앞쪽의 연주자 위치에 있으나 좌석은 잘렸다. 피아노 상부까지 오른쪽 화면을 크게 채워, 하단에 건반 측면의 작은 사선 부분만 보여야 한다는 구도와 다르다. 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "보이는 인물은 중년의 한국인 여성으로 제시된 지소영 한 명이며, 얼굴의 연령감과 짧고 단정한 검은 머리가 인물 참조에 대체로 부합한다. 흰 연구복 안에 어두운 상의를 입었고 온화하게 입을 다문 미소를 짓는다. 낡은 목재 피아노는 요청된 소품에 부합한다. 찰리와 분리된 전선, 출입구 상태는 이 구도에서 확인되지 않는다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "목과 어깨, 몸통의 연결과 고개를 들어 올린 자세는 자연스럽고 정지한 순간으로 읽힌다. 손·골반·발과 의자가 프레임 밖이라 착석 접촉은 직접 확인할 수 없지만, 공중에 떠 있다는 증거도 없다. 피아노는 하부가 잘린 정상적인 고정 가구로 보이며 부유하는 물체는 없다."
       },
       {
        "label": "B",
        "direction": "지소영의 눈과 얼굴이 왼쪽 위 전경의 남성을 향한다. 가까이 선 상대를 올려다보는 시선 관계가 명확하며 온화한 미소도 보인다. 남성은 지소영 쪽으로 몸을 돌리고 있지만 눈은 잘려 있어 그의 정확한 시선은 확인되지 않는다.",
        "built_space": "오른쪽에 검은 피아노 한 대와 건반 한 벌, 지소영 뒤에 검은 벤치 한 개가 보인다. 배경에는 곡면 유리 구획, 원형 상부 조명, 실험대와 모니터들, 의자 한 개, 오른쪽 로봇 팔 한 개와 왼쪽 열린 문이 있다. 장소 참조의 연구실 재질과 구조적 특징을 상당 부분 유지한다. 다만 넓은 바닥과 시설, 남성의 어깨까지 포함해 얼굴 클로즈업이 아니며 피아노도 하단의 작은 부분보다 많이 보인다. 보이는 범위에 설비 중복이나 불가능한 반사는 없다.",
        "entities": "지소영은 참조와 유사한 중년 여성의 얼굴, 단정한 짧은 검은 머리, 흰 연구복과 어두운 상의를 갖추었다. 왼쪽에는 짙은 정장과 흰 셔츠를 입은 성인 남성의 턱·목·상체가 크게 보인다. 문맥상 찰리 역할이지만, 가시 인물을 지소영으로 한정한 인물 조건에는 맞지 않는다. 피아노는 검은 광택 마감이며 A보다 오래된 사용감은 덜 뚜렷하다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "가시 인물을 지소영으로 한정한 명시적 조건과 달리, 왼쪽 전경에 성인 남성의 상체를 크게 추가했다."
        ],
        "physics": "지소영은 뒤에 보이는 피아노 벤치 앞부분에 앉은 것으로 읽히며, 골반 접촉점은 화면 아래에 가려져 있다. 몸통을 상대 쪽으로 돌리고 고개를 드는 자세는 가능하다. 남성의 하체는 잘렸지만 서 있는 상체 자세에 부유나 비정상적인 지지 문제는 없다. 벤치와 피아노는 바닥에 놓인 가구로 읽힌다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "지소영의 정체성·연구복·온화한 미소와 올려다보는 동작은 맞지만, 얼굴 클로즈업보다 넓고 피아노가 하단의 작은 일부가 아니라 오른쪽 전경을 크게 차지한다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "찰리를 올려다보는 관계는 명확하지만, 가시 인물을 지소영으로 제한한 조건과 달리 남성을 크게 넣었고 얼굴 클로즈업을 넓은 어깨너머 구도로 바꿨다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "지소영은 턱을 조금 들고 화면 오른쪽 위의 프레임 밖을 바라본다. 다가온 찰리를 올려다보는 방향으로 해석할 수 있으나 상대는 보이지 않아 실제 시선 도착점은 확인할 수 없다. 건반을 내려다보거나 카메라를 응시하지는 않는다.",
        "built_space": "오른쪽에 오래된 목재 업라이트 피아노 한 대와 건반 한 벌이 보인다. 뒤에는 유리 칸막이와 금속 기둥, 실험 장비 및 모니터가 있어 장소 참조의 차가운 연구시설 재질과 이어진다. 지소영은 건반 앞쪽의 연주자 위치에 있으나 좌석은 잘렸다. 피아노 상부까지 오른쪽 화면을 크게 채워, 하단에 건반 측면의 작은 사선 부분만 보여야 한다는 구도와 다르다. 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "보이는 인물은 중년의 한국인 여성으로 제시된 지소영 한 명이며, 얼굴의 연령감과 짧고 단정한 검은 머리가 인물 참조에 대체로 부합한다. 흰 연구복 안에 어두운 상의를 입었고 온화하게 입을 다문 미소를 짓는다. 낡은 목재 피아노는 요청된 소품에 부합한다. 찰리와 분리된 전선, 출입구 상태는 이 구도에서 확인되지 않는다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "목과 어깨, 몸통의 연결과 고개를 들어 올린 자세는 자연스럽고 정지한 순간으로 읽힌다. 손·골반·발과 의자가 프레임 밖이라 착석 접촉은 직접 확인할 수 없지만, 공중에 떠 있다는 증거도 없다. 피아노는 하부가 잘린 정상적인 고정 가구로 보이며 부유하는 물체는 없다."
       },
       {
        "label": "A",
        "direction": "지소영의 눈과 얼굴이 왼쪽 위 전경의 남성을 향한다. 가까이 선 상대를 올려다보는 시선 관계가 명확하며 온화한 미소도 보인다. 남성은 지소영 쪽으로 몸을 돌리고 있지만 눈은 잘려 있어 그의 정확한 시선은 확인되지 않는다.",
        "built_space": "오른쪽에 검은 피아노 한 대와 건반 한 벌, 지소영 뒤에 검은 벤치 한 개가 보인다. 배경에는 곡면 유리 구획, 원형 상부 조명, 실험대와 모니터들, 의자 한 개, 오른쪽 로봇 팔 한 개와 왼쪽 열린 문이 있다. 장소 참조의 연구실 재질과 구조적 특징을 상당 부분 유지한다. 다만 넓은 바닥과 시설, 남성의 어깨까지 포함해 얼굴 클로즈업이 아니며 피아노도 하단의 작은 부분보다 많이 보인다. 보이는 범위에 설비 중복이나 불가능한 반사는 없다.",
        "entities": "지소영은 참조와 유사한 중년 여성의 얼굴, 단정한 짧은 검은 머리, 흰 연구복과 어두운 상의를 갖추었다. 왼쪽에는 짙은 정장과 흰 셔츠를 입은 성인 남성의 턱·목·상체가 크게 보인다. 문맥상 찰리 역할이지만, 가시 인물을 지소영으로 한정한 인물 조건에는 맞지 않는다. 피아노는 검은 광택 마감이며 A보다 오래된 사용감은 덜 뚜렷하다. 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [
         "가시 인물을 지소영으로 한정한 명시적 조건과 달리, 왼쪽 전경에 성인 남성의 상체를 크게 추가했다."
        ],
        "physics": "지소영은 뒤에 보이는 피아노 벤치 앞부분에 앉은 것으로 읽히며, 골반 접촉점은 화면 아래에 가려져 있다. 몸통을 상대 쪽으로 돌리고 고개를 드는 자세는 가능하다. 남성의 하체는 잘렸지만 서 있는 상체 자세에 부유나 비정상적인 지지 문제는 없다. 벤치와 피아노는 바닥에 놓인 가구로 읽힌다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.5,
    "B": 1.571
   },
   "adjusted": {
    "A": 1.25,
    "B": 1.571
   },
   "violations": {
    "A": [
     "[gpt-high] 가시 인물을 지소영으로 한정한 명시적 조건과 달리, 왼쪽 전경에 성인 남성의 상체를 크게 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1250,
   "B": 1571
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1250,
    "verdict_ko": "찰리의 뒷모습을 전경에 배치해 시선을 맞추는 연출을 훌륭히 살렸으며, 화면 하단에 작게 걸친 피아노와 배경 구조물도 지시에 부합합니다.  ★위반: [gpt-high] 가시 인물을 지소영으로 한정한 명시적 조건과 달리, 왼쪽 전경에 성인 남성의 상체를 크게 추가했다."
   },
   {
    "label": "B",
    "score": 1571,
    "verdict_ko": "시선의 대상인 찰리가 등장하지 않아 상호작용이 느껴지지 않으며, 피아노가 지시와 다르게 화면 우측을 과도하게 차지하고 있습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L269B03.png",
    "asset_id": "2d7bcd3a-3137-473f-b939-ac3ee04b6230",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 지소영: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1243508>",
    "asset_id": "c7496f13-cfcf-44a5-976d-96c783d20580",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-f8bc-79ab-ae11-ff1e9f3cf3bc",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S84sh11__bgfirst_bg.png",
   "bg_asset_id": "d36617d7-e206-4f23-93f4-d6161a0fbc32",
   "bg_record_key": "S84sh11::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S84sh11::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T09:14:29.490851+00:00",
  "fingerprint": "e05fbf5f78dad2a39a8c4f4706f5292802ef1865932e63d29bbf8dd13a085071",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S84sh11_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S84sh11_sel.png",
  "source_sha256": "188d28876f252d3e4c57a788f837ea0622191c289af2f3d387b11cd1dd7964c6",
  "file": "S84sh11_cine.png",
  "staged_sha256": "8724e055434b8b67cd33ffd3effe7701cfa70d80ad160894fa9121548b91f6e6",
  "latency_ms": 10053
 },
 "S84sh14::signage": {
  "fp": "9e0bd6fd99ce0dfb",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S84sh14": {
  "input_fingerprint": "f75ed2f8f75276e6",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 지소영의 몸을 양팔로 빈틈없이 감싸 안은 찰리의 밀착된 상체.\n\nLOCATION (lock): At the piano inside the director's spacious modern suite adjoining the research center, in subdued nighttime light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Piano (Old piano, no longer being played) — A partial side and keyboard edge remain behind the joined figures; used as Soft contextual reminder of the music that prompted recognition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient illumination and gentle tonal separation preserve the tenderness of the embrace without introducing a new light source or perceptual effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 지소영 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old piano remains in the director's office. Charlie is free of the laboratory wires, with both arms closed in an embrace. 지소영: She remains in her neat research coat near the piano, her eyes gently closed.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 지소영의 몸을 양팔로 빈틈없이 감싸 안은 찰리의 밀착된 상체.\n\nLOCATION (lock): At the piano inside the director's spacious modern suite adjoining the research center, in subdued nighttime light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Piano (Old piano, no longer being played) — A partial side and keyboard edge remain behind the joined figures; used as Soft contextual reminder of the music that prompted recognition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient illumination and gentle tonal separation preserve the tenderness of the embrace without introducing a new light source or perceptual effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 지소영 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old piano remains in the director's office. Charlie is free of the laboratory wires, with both arms closed in an embrace. 지소영: She remains in her neat research coat near the piano, her eyes gently closed.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 지소영의 몸을 양팔로 빈틈없이 감싸 안은 찰리의 밀착된 상체.\n\nLOCATION (lock): At the piano inside the director's spacious modern suite adjoining the research center, in subdued nighttime light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Piano (Old piano, no longer being played) — A partial side and keyboard edge remain behind the joined figures; used as Soft contextual reminder of the music that prompted recognition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient illumination and gentle tonal separation preserve the tenderness of the embrace without introducing a new light source or perceptual effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 지소영 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old piano remains in the director's office. Charlie is free of the laboratory wires, with both arms closed in an embrace. 지소영: She remains in her neat research coat near the piano, her eyes gently closed.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "찰리는 지소영을 안은 채 고개를 숙이고, 지소영 역시 시선을 아래로 둔 채 눈을 뜨고 있음.",
    "built_space": "피아노가 화면 가로로 길게 배치되어 '일부 측면과 건반 가장자리'만 배경에 보여야 하는 구도 지시를 벗어남.",
    "entities": "찰리의 외형은 일치하나, 지소영은 '눈을 부드럽게 감은' 지시와 다르게 눈을 뜨고 시선을 아래로 향함.",
    "hard_violations": [],
    "physics": "두 인물이 서서 포옹하는 자세는 안정적이며 각 팔의 지지 상태에 물리적인 오류는 보이지 않음."
   },
   {
    "label": "B",
    "direction": "찰리가 지소영을 깊게 안고 얼굴을 맞대며, 지소영은 눈을 감은 채 찰리의 몸에 기대어 있음.",
    "built_space": "피아노의 측면과 건반 일부만이 우측에 배치되어 프롬프트가 지시한 배경 요소의 구도를 완벽하게 따름.",
    "entities": "찰리와 지소영의 외형 및 복장이 레퍼런스와 일치하며, 지소영의 감은 눈 등 세부 지시사항을 정확히 반영함.",
    "hard_violations": [],
    "physics": "서로를 감싸 안은 양팔과 몸의 밀착 등 포옹에 따른 물리적 접촉과 지지가 자연스럽게 묘사됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "이전 샷의 공간 설정(피아노 위치)과 구도 지시를 완벽히 준수하며, 눈을 감은 지소영과 찰리의 포옹을 매우 사실적으로 구현했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "피아노가 화면 전체를 가로지르게 배치되어 구도 지시를 어겼으며, 지소영이 눈을 뜨고 있어 세부 묘사에서 감점을 받았습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "찰리가 지소영을 깊게 안고 얼굴을 맞대며, 지소영은 눈을 감은 채 찰리의 몸에 기대어 있음.",
        "built_space": "피아노의 측면과 건반 일부만이 우측에 배치되어 프롬프트가 지시한 배경 요소의 구도를 완벽하게 따름.",
        "entities": "찰리와 지소영의 외형 및 복장이 레퍼런스와 일치하며, 지소영의 감은 눈 등 세부 지시사항을 정확히 반영함.",
        "hard_violations": [],
        "physics": "서로를 감싸 안은 양팔과 몸의 밀착 등 포옹에 따른 물리적 접촉과 지지가 자연스럽게 묘사됨."
       },
       {
        "label": "A",
        "direction": "찰리는 지소영을 안은 채 고개를 숙이고, 지소영 역시 시선을 아래로 둔 채 눈을 뜨고 있음.",
        "built_space": "피아노가 화면 가로로 길게 배치되어 '일부 측면과 건반 가장자리'만 배경에 보여야 하는 구도 지시를 벗어남.",
        "entities": "찰리의 외형은 일치하나, 지소영은 '눈을 부드럽게 감은' 지시와 다르게 눈을 뜨고 시선을 아래로 향함.",
        "hard_violations": [],
        "physics": "두 인물이 서서 포옹하는 자세는 안정적이며 각 팔의 지지 상태에 물리적인 오류는 보이지 않음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "이전 샷의 공간 설정(피아노 위치)과 구도 지시를 완벽히 준수하며, 눈을 감은 지소영과 찰리의 포옹을 매우 사실적으로 구현했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "피아노가 화면 전체를 가로지르게 배치되어 구도 지시를 어겼으며, 지소영이 눈을 뜨고 있어 세부 묘사에서 감점을 받았습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리가 지소영을 깊게 안고 얼굴을 맞대며, 지소영은 눈을 감은 채 찰리의 몸에 기대어 있음.",
        "built_space": "피아노의 측면과 건반 일부만이 우측에 배치되어 프롬프트가 지시한 배경 요소의 구도를 완벽하게 따름.",
        "entities": "찰리와 지소영의 외형 및 복장이 레퍼런스와 일치하며, 지소영의 감은 눈 등 세부 지시사항을 정확히 반영함.",
        "hard_violations": [],
        "physics": "서로를 감싸 안은 양팔과 몸의 밀착 등 포옹에 따른 물리적 접촉과 지지가 자연스럽게 묘사됨."
       },
       {
        "label": "A",
        "direction": "찰리는 지소영을 안은 채 고개를 숙이고, 지소영 역시 시선을 아래로 둔 채 눈을 뜨고 있음.",
        "built_space": "피아노가 화면 가로로 길게 배치되어 '일부 측면과 건반 가장자리'만 배경에 보여야 하는 구도 지시를 벗어남.",
        "entities": "찰리의 외형은 일치하나, 지소영은 '눈을 부드럽게 감은' 지시와 다르게 눈을 뜨고 시선을 아래로 향함.",
        "hard_violations": [],
        "physics": "두 인물이 서서 포옹하는 자세는 안정적이며 각 팔의 지지 상태에 물리적인 오류는 보이지 않음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "밀착한 상체와 포옹을 중심으로 한 미디엄 숏이며, 피아노의 측면과 건반 일부만 남겨 B보다 지정 구도에 가깝지만 피아노의 화면 비중은 다소 크다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "양팔로 등과 허리를 감싼 동작은 명확하지만, 피아노 정면과 건반 거의 전체를 보여 주어 ‘부분 측면과 건반 가장자리만 남기는’ 배경 지시에서 벗어난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리는 고개와 마스크를 아래로 기울여 지소영의 정수리 쪽을 향한다. 지소영은 눈을 감고 찰리의 가슴에 얼굴을 기댄다. 가까운 팔은 지소영의 몸 앞을 가로질러 반대쪽 옆구리를 감싸고, 다른 손의 일부는 먼 쪽 어깨 부근에 보인다. 두 사람 모두 카메라를 응시하지 않으며 건반을 연주하지 않는다.",
        "built_space": "오른쪽에 낡은 목제 업라이트 피아노 한 대와 그 건반 일부가 있고, 뒤에는 금속 프레임 유리 칸막이, 왼쪽 장비 랙, 중앙 오른쪽 모니터 한 대가 보인다. 두 사람은 피아노 건반 앞쪽의 왼편에서 밀착해 있다. 기준 사진의 차가운 조명과 유리·목재 공간을 유지한다. 피아노는 부분적으로 잘렸지만 오른쪽 전경을 상당히 차지해 부드러운 배경 단서보다는 두드러진다. 불가능한 반사나 중복 피아노는 보이지 않는다.",
        "entities": "찰리와 지소영 두 인물만 있다. 찰리는 육중한 상체, 긴 기계 팔, 마모된 샌드 베이지 장갑판과 흰 마스크를 갖춰 참고 이미지의 정체성에 부합한다. 다만 참고 이미지의 주황색 눈빛은 보이지 않는다. 지소영은 50대 중반 한국인 여성이라는 설정에 맞는 중년 외모와 정돈한 검은 단발, 흰 연구복과 검은 안쪽 옷을 유지하며 눈을 부드럽게 감고 있다. 낡은 목제 피아노가 있으며 외부 실험용 연결선이나 읽을 수 있는 문자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "찰리의 가까운 팔은 어깨와 팔꿈치에서 이어져 굽혀지고 손이 지소영의 몸을 감싼다. 반대쪽 팔은 대부분 몸 뒤에 가려지지만 손 일부가 어깨 옆에 보여 포옹으로 해석 가능하다. 지소영의 손도 찰리 옆구리에 닿아 있고 상체는 그의 가슴에 기대어 있다. 발과 바닥 접점은 프레임 밖이지만 두 몸통은 아래로 이어져 있으며, 공중에 떠 있거나 지지 없이 매달린 자세는 아니다."
       },
       {
        "label": "B",
        "direction": "찰리는 고개를 숙여 지소영의 머리를 향하고, 지소영은 눈을 감은 채 찰리의 가슴 쪽으로 몸을 붙인다. 찰리의 한 손은 지소영의 윗등에, 다른 손은 허리 부근에 놓여 양팔이 몸을 둘러싼다. 지소영의 팔 역시 찰리의 몸통 쪽으로 향한다. 건반에 손을 뻗거나 카메라를 바라보는 인물은 없다.",
        "built_space": "목제 업라이트 피아노 한 대가 두 사람 바로 뒤에 정면으로 놓여 있다. 상판과 전면 패널, 좌우 끝부분 및 건반 대부분이 드러난다. 배경에는 기준 사진과 유사한 유리 칸막이와 금속 프레임, 왼쪽 장비 랙이 있다. 두 사람은 피아노 앞에 서 있어 공간 관계는 성립하지만, 피아노의 부분 측면과 건반 가장자리만 남기라는 구도와는 다르다. 중복 설비나 광학적으로 불가능한 반사는 보이지 않는다.",
        "entities": "찰리와 지소영 두 인물만 보인다. 찰리의 넓은 어깨, 긴 팔, 샌드 베이지 장갑판, 흰 마스크는 참고 이미지와 대체로 일치하나 눈의 주황색 발광은 재현되지 않았다. 지소영은 중년 한국인 여성으로 보이며 검은 단발과 흰 연구복을 유지하고 눈을 감았다. 얼굴은 측면으로 보이고 안쪽 옷은 대부분 가려져 있다. 낡은 목제 피아노가 있으며 실험실 연결선, 추가 인물, 읽을 수 있는 문자는 없다.",
        "hard_violations": [],
        "physics": "양쪽 기계 팔이 각각 어깨·팔꿈치·손목으로 연결되고, 두 손이 지소영의 등과 허리에 직접 닿아 있다. 지소영의 상체와 팔은 찰리의 몸통에 밀착해 자연스러운 상호 포옹을 이룬다. 하체와 발의 지면 접촉은 화면 밖이지만 몸통이 수직으로 이어지며 부유나 비약을 암시하지 않는다. 피아노는 아래로 이어지는 다리 구조가 보이고, 지지 없이 떠 있는 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "밀착한 상체와 포옹을 중심으로 한 미디엄 숏이며, 피아노의 측면과 건반 일부만 남겨 B보다 지정 구도에 가깝지만 피아노의 화면 비중은 다소 크다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "양팔로 등과 허리를 감싼 동작은 명확하지만, 피아노 정면과 건반 거의 전체를 보여 주어 ‘부분 측면과 건반 가장자리만 남기는’ 배경 지시에서 벗어난다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리는 고개와 마스크를 아래로 기울여 지소영의 정수리 쪽을 향한다. 지소영은 눈을 감고 찰리의 가슴에 얼굴을 기댄다. 가까운 팔은 지소영의 몸 앞을 가로질러 반대쪽 옆구리를 감싸고, 다른 손의 일부는 먼 쪽 어깨 부근에 보인다. 두 사람 모두 카메라를 응시하지 않으며 건반을 연주하지 않는다.",
        "built_space": "오른쪽에 낡은 목제 업라이트 피아노 한 대와 그 건반 일부가 있고, 뒤에는 금속 프레임 유리 칸막이, 왼쪽 장비 랙, 중앙 오른쪽 모니터 한 대가 보인다. 두 사람은 피아노 건반 앞쪽의 왼편에서 밀착해 있다. 기준 사진의 차가운 조명과 유리·목재 공간을 유지한다. 피아노는 부분적으로 잘렸지만 오른쪽 전경을 상당히 차지해 부드러운 배경 단서보다는 두드러진다. 불가능한 반사나 중복 피아노는 보이지 않는다.",
        "entities": "찰리와 지소영 두 인물만 있다. 찰리는 육중한 상체, 긴 기계 팔, 마모된 샌드 베이지 장갑판과 흰 마스크를 갖춰 참고 이미지의 정체성에 부합한다. 다만 참고 이미지의 주황색 눈빛은 보이지 않는다. 지소영은 50대 중반 한국인 여성이라는 설정에 맞는 중년 외모와 정돈한 검은 단발, 흰 연구복과 검은 안쪽 옷을 유지하며 눈을 부드럽게 감고 있다. 낡은 목제 피아노가 있으며 외부 실험용 연결선이나 읽을 수 있는 문자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "찰리의 가까운 팔은 어깨와 팔꿈치에서 이어져 굽혀지고 손이 지소영의 몸을 감싼다. 반대쪽 팔은 대부분 몸 뒤에 가려지지만 손 일부가 어깨 옆에 보여 포옹으로 해석 가능하다. 지소영의 손도 찰리 옆구리에 닿아 있고 상체는 그의 가슴에 기대어 있다. 발과 바닥 접점은 프레임 밖이지만 두 몸통은 아래로 이어져 있으며, 공중에 떠 있거나 지지 없이 매달린 자세는 아니다."
       },
       {
        "label": "A",
        "direction": "찰리는 고개를 숙여 지소영의 머리를 향하고, 지소영은 눈을 감은 채 찰리의 가슴 쪽으로 몸을 붙인다. 찰리의 한 손은 지소영의 윗등에, 다른 손은 허리 부근에 놓여 양팔이 몸을 둘러싼다. 지소영의 팔 역시 찰리의 몸통 쪽으로 향한다. 건반에 손을 뻗거나 카메라를 바라보는 인물은 없다.",
        "built_space": "목제 업라이트 피아노 한 대가 두 사람 바로 뒤에 정면으로 놓여 있다. 상판과 전면 패널, 좌우 끝부분 및 건반 대부분이 드러난다. 배경에는 기준 사진과 유사한 유리 칸막이와 금속 프레임, 왼쪽 장비 랙이 있다. 두 사람은 피아노 앞에 서 있어 공간 관계는 성립하지만, 피아노의 부분 측면과 건반 가장자리만 남기라는 구도와는 다르다. 중복 설비나 광학적으로 불가능한 반사는 보이지 않는다.",
        "entities": "찰리와 지소영 두 인물만 보인다. 찰리의 넓은 어깨, 긴 팔, 샌드 베이지 장갑판, 흰 마스크는 참고 이미지와 대체로 일치하나 눈의 주황색 발광은 재현되지 않았다. 지소영은 중년 한국인 여성으로 보이며 검은 단발과 흰 연구복을 유지하고 눈을 감았다. 얼굴은 측면으로 보이고 안쪽 옷은 대부분 가려져 있다. 낡은 목제 피아노가 있으며 실험실 연결선, 추가 인물, 읽을 수 있는 문자는 없다.",
        "hard_violations": [],
        "physics": "양쪽 기계 팔이 각각 어깨·팔꿈치·손목으로 연결되고, 두 손이 지소영의 등과 허리에 직접 닿아 있다. 지소영의 상체와 팔은 찰리의 몸통에 밀착해 자연스러운 상호 포옹을 이룬다. 하체와 발의 지면 접촉은 화면 밖이지만 몸통이 수직으로 이어지며 부유나 비약을 암시하지 않는다. 피아노는 아래로 이어지는 다리 구조가 보이고, 지지 없이 떠 있는 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.375,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.375,
    "B": 2.0
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1375
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "이전 샷의 공간 설정(피아노 위치)과 구도 지시를 완벽히 준수하며, 눈을 감은 지소영과 찰리의 포옹을 매우 사실적으로 구현했습니다."
   },
   {
    "label": "A",
    "score": 1375,
    "verdict_ko": "피아노가 화면 전체를 가로지르게 배치되어 구도 지시를 어겼으며, 지소영이 눈을 뜨고 있어 세부 묘사에서 감점을 받았습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 지소영 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S84sh11_sel.png",
    "asset_id": "2de4be74-5677-43d2-a970-da8e4277b237",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 지소영: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1243508>",
    "asset_id": "c7496f13-cfcf-44a5-976d-96c783d20580",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-fc00-7b99-acf7-d7af303b0dfb",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S84sh11"
  }
 },
 "S84sh14::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T09:15:45.502443+00:00",
  "fingerprint": "ece7bee795a01c454fecf66f7bf93c188a5dd4902390249c02dd67be184b8e09",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S84sh14_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S84sh14_sel.png",
  "source_sha256": "c31503b0d006630a135f006638d0616721dbec325774fd00ce0636ca0abc1629",
  "file": "S84sh14_cine.png",
  "staged_sha256": "b218981e61a1610c3d74e9983ef4eb0e79233ceb574b02b9cb9761cea78ca07d",
  "latency_ms": 9928
 },
 "S85sh11::signage": {
  "fp": "c0fa537ec8cefb43",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S85sh11": {
  "input_fingerprint": "577e5da639c48648",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 회전 중인 한 위치에서 멈춘 가슴 링 틈새에서 짙은 검은 연기가 뿜어져 나오는 순간.\n\nLOCATION (lock): Inside the main research center's glass test enclosure, at the wired test position under laboratory lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Chest ring (Still rotating and slowing as black smoke emerges from its gaps) — The front and one recessed gap are visible obliquely; used as Primary mechanical detail within the larger torso composition; Attached wires (Connected to 찰리's body); used as Peripheral evidence of the ongoing experiment; Black smoke (Emerging from the ring gap); used as Localized visual evidence of failure, leaving the ring's position legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral ambient illumination with controlled contrast distinguishes the dense black smoke from the surrounding torso without adding an unsupported glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie is seated behind the glass wall with multiple wires reattached; black smoke rises from the chest ring as its rotation slows and fails. The experiment display reads “ERROR,” and the previously rising output graphs are falling.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 회전 중인 한 위치에서 멈춘 가슴 링 틈새에서 짙은 검은 연기가 뿜어져 나오는 순간.\n\nLOCATION (lock): Inside the main research center's glass test enclosure, at the wired test position under laboratory lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Chest ring (Still rotating and slowing as black smoke emerges from its gaps) — The front and one recessed gap are visible obliquely; used as Primary mechanical detail within the larger torso composition; Attached wires (Connected to 찰리's body); used as Peripheral evidence of the ongoing experiment; Black smoke (Emerging from the ring gap); used as Localized visual evidence of failure, leaving the ring's position legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral ambient illumination with controlled contrast distinguishes the dense black smoke from the surrounding torso without adding an unsupported glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie is seated behind the glass wall with multiple wires reattached; black smoke rises from the chest ring as its rotation slows and fails. The experiment display reads “ERROR,” and the previously rising output graphs are falling.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 회전 중인 한 위치에서 멈춘 가슴 링 틈새에서 짙은 검은 연기가 뿜어져 나오는 순간.\n\nLOCATION (lock): Inside the main research center's glass test enclosure, at the wired test position under laboratory lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Chest ring (Still rotating and slowing as black smoke emerges from its gaps) — The front and one recessed gap are visible obliquely; used as Primary mechanical detail within the larger torso composition; Attached wires (Connected to 찰리's body); used as Peripheral evidence of the ongoing experiment; Black smoke (Emerging from the ring gap); used as Localized visual evidence of failure, leaving the ring's position legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral ambient illumination with controlled contrast distinguishes the dense black smoke from the surrounding torso without adding an unsupported glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie is seated behind the glass wall with multiple wires reattached; black smoke rises from the chest ring as its rotation slows and fails. The experiment display reads “ERROR,” and the previously rising output graphs are falling.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라가 찰리의 상체를 정면에서 바라보며, 가슴 링에서 검은 연기가 앞쪽으로 뿜어져 나옵니다.",
    "built_space": "배경에 연구소의 유리 벽면이 보이지만, 이전 샷에 명확히 존재했던 금속 테이블과 실험 장비가 사라졌습니다.",
    "entities": "찰리의 몸체와 장갑판, 전선, 검은 연기가 묘사되었으나, 캐릭터가 직립한 상태로 나타납니다.",
    "hard_violations": [
     "[gemini-pro] 이전 샷의 고정된 자세와 환경(금속 테이블 위에 누워있는 상태)을 완전히 무시하고 캐릭터를 서 있는 자세로 변경함"
    ],
    "physics": "캐릭터가 화면 밖의 지면을 딛고 서 있으며, 연기가 공기 중으로 퍼져나갑니다."
   },
   {
    "label": "B",
    "direction": "카메라가 누워 있는 찰리의 상체를 약간 위에서 내려다보며, 가슴 링에서 연기가 위로 솟아오릅니다.",
    "built_space": "이전 샷과 동일한 금속 테이블이 캐릭터를 받치고 있으며, 뒤쪽으로 유리 차폐벽이 올바르게 배치되어 있습니다.",
    "entities": "금속 테이블 위에 누워있는 찰리, 가슴의 원형 링, 연결된 다수의 전선, 뿜어져 나오는 검은 연기가 모두 정확하게 나타납니다.",
    "hard_violations": [],
    "physics": "금속 테이블이 캐릭터의 등과 머리를 안정적으로 지지하고 있으며, 뿜어져 나온 연기가 자연스럽게 위로 피어오릅니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "이전 샷의 자세와 위치를 고정하라는 지시를 무시하고 금속 테이블을 누락한 채 캐릭터를 서 있게 묘사하여 큰 감점을 받았습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "이전 샷에 등장한 금속 테이블과 누워있는 자세를 정확히 유지하면서, 가슴 링에서 피어오르는 검은 연기를 클로즈업으로 훌륭하게 연출했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라가 찰리의 상체를 정면에서 바라보며, 가슴 링에서 검은 연기가 앞쪽으로 뿜어져 나옵니다.",
        "built_space": "배경에 연구소의 유리 벽면이 보이지만, 이전 샷에 명확히 존재했던 금속 테이블과 실험 장비가 사라졌습니다.",
        "entities": "찰리의 몸체와 장갑판, 전선, 검은 연기가 묘사되었으나, 캐릭터가 직립한 상태로 나타납니다.",
        "hard_violations": [
         "이전 샷의 고정된 자세와 환경(금속 테이블 위에 누워있는 상태)을 완전히 무시하고 캐릭터를 서 있는 자세로 변경함"
        ],
        "physics": "캐릭터가 화면 밖의 지면을 딛고 서 있으며, 연기가 공기 중으로 퍼져나갑니다."
       },
       {
        "label": "B",
        "direction": "카메라가 누워 있는 찰리의 상체를 약간 위에서 내려다보며, 가슴 링에서 연기가 위로 솟아오릅니다.",
        "built_space": "이전 샷과 동일한 금속 테이블이 캐릭터를 받치고 있으며, 뒤쪽으로 유리 차폐벽이 올바르게 배치되어 있습니다.",
        "entities": "금속 테이블 위에 누워있는 찰리, 가슴의 원형 링, 연결된 다수의 전선, 뿜어져 나오는 검은 연기가 모두 정확하게 나타납니다.",
        "hard_violations": [],
        "physics": "금속 테이블이 캐릭터의 등과 머리를 안정적으로 지지하고 있으며, 뿜어져 나온 연기가 자연스럽게 위로 피어오릅니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "이전 샷의 자세와 위치를 고정하라는 지시를 무시하고 금속 테이블을 누락한 채 캐릭터를 서 있게 묘사하여 큰 감점을 받았습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "이전 샷에 등장한 금속 테이블과 누워있는 자세를 정확히 유지하면서, 가슴 링에서 피어오르는 검은 연기를 클로즈업으로 훌륭하게 연출했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라가 찰리의 상체를 정면에서 바라보며, 가슴 링에서 검은 연기가 앞쪽으로 뿜어져 나옵니다.",
        "built_space": "배경에 연구소의 유리 벽면이 보이지만, 이전 샷에 명확히 존재했던 금속 테이블과 실험 장비가 사라졌습니다.",
        "entities": "찰리의 몸체와 장갑판, 전선, 검은 연기가 묘사되었으나, 캐릭터가 직립한 상태로 나타납니다.",
        "hard_violations": [
         "이전 샷의 고정된 자세와 환경(금속 테이블 위에 누워있는 상태)을 완전히 무시하고 캐릭터를 서 있는 자세로 변경함"
        ],
        "physics": "캐릭터가 화면 밖의 지면을 딛고 서 있으며, 연기가 공기 중으로 퍼져나갑니다."
       },
       {
        "label": "B",
        "direction": "카메라가 누워 있는 찰리의 상체를 약간 위에서 내려다보며, 가슴 링에서 연기가 위로 솟아오릅니다.",
        "built_space": "이전 샷과 동일한 금속 테이블이 캐릭터를 받치고 있으며, 뒤쪽으로 유리 차폐벽이 올바르게 배치되어 있습니다.",
        "entities": "금속 테이블 위에 누워있는 찰리, 가슴의 원형 링, 연결된 다수의 전선, 뿜어져 나오는 검은 연기가 모두 정확하게 나타납니다.",
        "hard_violations": [],
        "physics": "금속 테이블이 캐릭터의 등과 머리를 안정적으로 지지하고 있으며, 뿜어져 나온 연기가 자연스럽게 위로 피어오릅니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "짙은 검은 연기와 연결 배선은 맞지만, 허벅지까지 들어오는 정면 구도이며 연기가 링의 특정 틈새가 아니라 중앙 개구부 전체에서 솟아 요청한 클로즈업과 분출 위치가 약하다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "가슴 중심의 사선 클로즈업으로 링 전면과 함몰된 측면 틈을 함께 보여 더 충실하지만, 연기의 발생점은 여전히 특정 틈새보다 중앙 개구부로 읽힌다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "가슴 링의 전면은 거의 카메라를 정면으로 향한다. 검은 연기는 링 중앙의 넓은 구멍에서 위로 올라가 목과 얼굴 앞을 덮는다. 지정된 한 틈새에서 분출하는 방향은 구별되지 않는다. 눈은 프레임 상단에서 잘려 시선의 목표를 확인할 수 없다.",
        "built_space": "중앙 가슴 링은 하나이고 양쪽 어깨 장갑이 하나씩 보인다. 몸 뒤에는 좌우로 드러난 경사진 금속 시험대가 있으며, 배경에는 유리 구획의 수직 프레임과 상부 조명이 보인다. 찰리는 시험대에 기대어 있는 배치로 읽혀 이전 장면의 지지 구조와 대체로 이어진다. 화면이나 판독 가능한 글자는 없고, 불가능한 반사는 보이지 않는다.",
        "entities": "찰리 한 명의 장갑 몸체만 보인다. 샌드 베이지 장갑의 긁힘과 마모, 육중한 팔, 부분적으로 보이는 마스크는 참고 이미지의 외형과 대체로 맞는다. 링 하나와 양옆 연결부에 꽂힌 여러 배선, 짙은 검은 연기가 있다. 얼굴 대부분과 다리 길이는 이 구도에서 확인할 수 없으며, 실험 화면도 보이지 않는다.",
        "hard_violations": [],
        "physics": "몸통과 팔 뒤의 경사진 시험대가 몸을 받치는 것으로 보인다. 링은 흉부 하우징에 결합되어 있고 배선은 양쪽 소켓에 연결된 채 아래로 처진다. 연기의 상승 자체는 자연스럽지만 발생 영역이 중앙 구멍 전체로 넓다. 링의 회전이나 감속을 식별할 뚜렷한 운동 흔적은 없다."
       },
       {
        "label": "B",
        "direction": "카메라에 가까운 왼쪽 어깨와 멀어지는 오른쪽 흉부가 사선 시점을 만든다. 링 전면과 오른쪽의 뒤로 파인 둘레 틈이 함께 보인다. 연기는 중앙 개구부 안에서 뭉쳐 위쪽 목 방향으로 올라가며, 측면 틈을 분출 원점으로 특정하기는 어렵다. 눈은 화면 밖이라 시선은 판단할 수 없다.",
        "built_space": "가슴 링 하나, 양쪽 어깨 장갑 두 개, 화면 오른쪽 흉부의 배선 소켓 두 개가 보인다. 배경의 유리벽과 금속 프레임은 흐리게 처리되어 실험실 내부를 유지하고 전경 장치보다 작게 보인다. 좌석과 몸의 접촉점은 크롭 밖이라 정확한 기대기 각도는 확인할 수 없다. 중복된 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "찰리의 가슴과 양팔 일부, 흰 마스크의 턱 부분만 보이며 다른 인물은 없다. 베이지 장갑과 금속 관절, 표면 마모는 참고 외형에 부합한다. 다만 보이는 턱은 이전 장면의 복잡한 마스크보다 단순하고 각져 있다. 링과 연결 배선, 검은 연기는 모두 있으나 연기 상부는 회색으로 옅어진다. 실험 화면과 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "링은 몸통에 고정된 원통형 하우징으로 지지되고, 오른쪽 케이블은 두 소켓에 꽂혀 아래로 휜다. 왼쪽 배선도 몸통을 따라 내려간다. 팔은 어깨와 팔꿈치 관절로 연결되어 있으며 떠 있는 부품은 없다. 하체와 좌석이 화면 밖이라는 이유만으로 몸이 공중에 떠 있다고 볼 수는 없다. 연기는 위로 자연스럽게 퍼지지만 틈새 분출과 회전 감속은 명확하지 않다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "짙은 검은 연기와 연결 배선은 맞지만, 허벅지까지 들어오는 정면 구도이며 연기가 링의 특정 틈새가 아니라 중앙 개구부 전체에서 솟아 요청한 클로즈업과 분출 위치가 약하다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "가슴 중심의 사선 클로즈업으로 링 전면과 함몰된 측면 틈을 함께 보여 더 충실하지만, 연기의 발생점은 여전히 특정 틈새보다 중앙 개구부로 읽힌다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "가슴 링의 전면은 거의 카메라를 정면으로 향한다. 검은 연기는 링 중앙의 넓은 구멍에서 위로 올라가 목과 얼굴 앞을 덮는다. 지정된 한 틈새에서 분출하는 방향은 구별되지 않는다. 눈은 프레임 상단에서 잘려 시선의 목표를 확인할 수 없다.",
        "built_space": "중앙 가슴 링은 하나이고 양쪽 어깨 장갑이 하나씩 보인다. 몸 뒤에는 좌우로 드러난 경사진 금속 시험대가 있으며, 배경에는 유리 구획의 수직 프레임과 상부 조명이 보인다. 찰리는 시험대에 기대어 있는 배치로 읽혀 이전 장면의 지지 구조와 대체로 이어진다. 화면이나 판독 가능한 글자는 없고, 불가능한 반사는 보이지 않는다.",
        "entities": "찰리 한 명의 장갑 몸체만 보인다. 샌드 베이지 장갑의 긁힘과 마모, 육중한 팔, 부분적으로 보이는 마스크는 참고 이미지의 외형과 대체로 맞는다. 링 하나와 양옆 연결부에 꽂힌 여러 배선, 짙은 검은 연기가 있다. 얼굴 대부분과 다리 길이는 이 구도에서 확인할 수 없으며, 실험 화면도 보이지 않는다.",
        "hard_violations": [],
        "physics": "몸통과 팔 뒤의 경사진 시험대가 몸을 받치는 것으로 보인다. 링은 흉부 하우징에 결합되어 있고 배선은 양쪽 소켓에 연결된 채 아래로 처진다. 연기의 상승 자체는 자연스럽지만 발생 영역이 중앙 구멍 전체로 넓다. 링의 회전이나 감속을 식별할 뚜렷한 운동 흔적은 없다."
       },
       {
        "label": "A",
        "direction": "카메라에 가까운 왼쪽 어깨와 멀어지는 오른쪽 흉부가 사선 시점을 만든다. 링 전면과 오른쪽의 뒤로 파인 둘레 틈이 함께 보인다. 연기는 중앙 개구부 안에서 뭉쳐 위쪽 목 방향으로 올라가며, 측면 틈을 분출 원점으로 특정하기는 어렵다. 눈은 화면 밖이라 시선은 판단할 수 없다.",
        "built_space": "가슴 링 하나, 양쪽 어깨 장갑 두 개, 화면 오른쪽 흉부의 배선 소켓 두 개가 보인다. 배경의 유리벽과 금속 프레임은 흐리게 처리되어 실험실 내부를 유지하고 전경 장치보다 작게 보인다. 좌석과 몸의 접촉점은 크롭 밖이라 정확한 기대기 각도는 확인할 수 없다. 중복된 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "찰리의 가슴과 양팔 일부, 흰 마스크의 턱 부분만 보이며 다른 인물은 없다. 베이지 장갑과 금속 관절, 표면 마모는 참고 외형에 부합한다. 다만 보이는 턱은 이전 장면의 복잡한 마스크보다 단순하고 각져 있다. 링과 연결 배선, 검은 연기는 모두 있으나 연기 상부는 회색으로 옅어진다. 실험 화면과 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "링은 몸통에 고정된 원통형 하우징으로 지지되고, 오른쪽 케이블은 두 소켓에 꽂혀 아래로 휜다. 왼쪽 배선도 몸통을 따라 내려간다. 팔은 어깨와 팔꿈치 관절로 연결되어 있으며 떠 있는 부품은 없다. 하체와 좌석이 화면 밖이라는 이유만으로 몸이 공중에 떠 있다고 볼 수는 없다. 연기는 위로 자연스럽게 퍼지지만 틈새 분출과 회전 감속은 명확하지 않다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.429,
    "B": 1.714
   },
   "adjusted": {
    "A": 1.179,
    "B": 1.714
   },
   "violations": {
    "A": [
     "[gemini-pro] 이전 샷의 고정된 자세와 환경(금속 테이블 위에 누워있는 상태)을 완전히 무시하고 캐릭터를 서 있는 자세로 변경함"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "A": 1179,
   "B": 1714
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1179,
    "verdict_ko": "이전 샷의 자세와 위치를 고정하라는 지시를 무시하고 금속 테이블을 누락한 채 캐릭터를 서 있게 묘사하여 큰 감점을 받았습니다.  ★위반: [gemini-pro] 이전 샷의 고정된 자세와 환경(금속 테이블 위에 누워있는 상태)을 완전히 무시하고 캐릭터를 서 있는 자세로 변경함"
   },
   {
    "label": "B",
    "score": 1714,
    "verdict_ko": "이전 샷에 등장한 금속 테이블과 누워있는 자세를 정확히 유지하면서, 가슴 링에서 피어오르는 검은 연기를 클로즈업으로 훌륭하게 연출했습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S84sh1_sel.png",
    "asset_id": "0866ea94-375a-404a-988b-1ddc554b8de1",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-fdb3-7e35-9ecd-118338b45d0c",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S84sh1"
  }
 },
 "S85sh11::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T12:13:37.536963+00:00",
  "fingerprint": "73e839396a4a1cb1833eab1a1f2ab7a2119db13015f49507bb4ad79a8abfe913",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S85sh11_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S85sh11_sel.png",
  "source_sha256": "4bc8c5e11c652d884c4a3b3d808ed410a1dff14bf626fadb27913ef4fc498814",
  "file": "S85sh11_cine.png",
  "staged_sha256": "1818564526c7088cc8e1a5927897a499ff549a16941b682608d647c35ffce54c",
  "latency_ms": 9983
 },
 "S85sh22::signage": {
  "fp": "44894cb3220c2f6f",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::374d1a2346548ec2": {
  "subjects": [],
  "subject_text": "제주도 연구소 메인 연구센터\n발전소처럼 거대한 실내 연구 공간. 각종 컴퓨터와 대형 연구 장치가 배치되고 열대 식물과 꽃이 기계 설비 사이를 채운다.",
  "identity": "canonical",
  "scope_id": "L268",
  "scope_role": "location_interior",
  "scope_sha": "3b89e33f78be0eaf"
 },
 "S85sh22::bgfirst_bg": {
  "input_fingerprint": "ef3385f9c5575580",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 양손으로 자신의 입을 헉 하고 틀어막은 서지민의 당황한 얼굴 클로즈업.\n\nLOCATION (lock): At a dining table inside the research facility's communal cafeteria, in daytime interior light near the meal-service area.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain even ambient illumination and restrained contrast so the exposed eyes remain readable above her hands.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 양손으로 자신의 입을 헉 하고 틀어막은 서지민의 당황한 얼굴 클로즈업.\n\nLOCATION (lock): At a dining table inside the research facility's communal cafeteria, in daytime interior light near the meal-service area.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain even ambient illumination and restrained contrast so the exposed eyes remain readable above her hands.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S85sh22__bgfirst_bg.png",
  "asset_id": "44478bdf-e0e9-48ca-bfd4-2ad6d414a7f9",
  "input_asset_ids": [
   "382f3829-e015-47c7-ab5e-b6be933588ec",
   "a491e625-4c37-4571-b3ba-85a3daf996f9"
  ]
 },
 "S85sh22": {
  "input_fingerprint": "291c3da199433e15",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 양손으로 자신의 입을 헉 하고 틀어막은 서지민의 당황한 얼굴 클로즈업.\n\nLOCATION (lock): At a dining table inside the research facility's communal cafeteria, in daytime interior light near the meal-service area. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain even ambient illumination and restrained contrast so the exposed eyes remain readable above her hands.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Appetizing bread and other food remain on the dining table, still uneaten. 서지민: She is at the dining conversation in her research coat, abruptly checking herself after an unintended disclosure.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 서지민 (한국인 여성, 20세, 앳된 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 양손으로 자신의 입을 헉 하고 틀어막은 서지민의 당황한 얼굴 클로즈업.\n\nLOCATION (lock): At a dining table inside the research facility's communal cafeteria, in daytime interior light near the meal-service area. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain even ambient illumination and restrained contrast so the exposed eyes remain readable above her hands.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Appetizing bread and other food remain on the dining table, still uneaten. 서지민: She is at the dining conversation in her research coat, abruptly checking herself after an unintended disclosure.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 서지민 (한국인 여성, 20세, 앳된 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 양손으로 자신의 입을 헉 하고 틀어막은 서지민의 당황한 얼굴 클로즈업.\n\nLOCATION (lock): At a dining table inside the research facility's communal cafeteria, in daytime interior light near the meal-service area. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain even ambient illumination and restrained contrast so the exposed eyes remain readable above her hands.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Appetizing bread and other food remain on the dining table, still uneaten. 서지민: She is at the dining conversation in her research coat, abruptly checking herself after an unintended disclosure.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 서지민 (한국인 여성, 20세, 앳된 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S85sh22__bgfirst_bg.png",
     "asset_id": "44478bdf-e0e9-48ca-bfd4-2ad6d414a7f9",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S85sh22.png",
     "asset_id": "382f3829-e015-47c7-ab5e-b6be933588ec",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 서지민: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1147284>",
     "asset_id": "a529b479-7055-42c6-be7d-19d14b5e000f",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L268B01.png",
     "asset_id": "a491e625-4c37-4571-b3ba-85a3daf996f9",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 서지민: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1147284>",
     "asset_id": "a529b479-7055-42c6-be7d-19d14b5e000f",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "서지민의 시선이 프레임 좌측 전경에 위치한, 지시되지 않은 인물을 향하고 있습니다.",
    "built_space": "일반적인 구내식당 형태이며, 레퍼런스의 핵심인 원형 유리벽 연구실 구조가 존재하지 않습니다.",
    "entities": "서지민의 외모는 레퍼런스와 일치하나, 프롬프트에 없는 남성의 어깨와 뒷모습이 화면을 크게 가리고 있습니다.",
    "hard_violations": [
     "[gemini-pro] 프롬프트에 명시되지 않은 전경 인물 추가",
     "[gpt-high] 서지민만 등장해야 하는 장면에 전경 대화 상대 1명과 배경 인물 6명을 추가했습니다."
    ],
    "physics": "테이블 앞 의자에 앉아 양손을 들어 입을 가리고 있습니다."
   },
   {
    "label": "B",
    "direction": "서지민의 시선이 카메라 우측 바깥의 특정되지 않은 곳을 당황한 듯 응시하고 있습니다.",
    "built_space": "레퍼런스 이미지와 동일한 원형 유리벽 연구실과 우측의 배식대 시설이 정확히 구현되었습니다.",
    "entities": "서지민의 외모와 복장은 일치하나, 입을 가린 양손의 손가락 개수와 마디가 기형적으로 융합되어 있습니다.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 기형적인 손가락 해부학 구조",
     "[gemini-pro] 프롬프트에 명시되지 않은 배경 인물들 추가",
     "[gpt-high] 서지민만 등장해야 하는 장면에 배경 인물 5명을 추가했습니다."
    ],
    "physics": "테이블 앞 의자에 앉아 양손을 입에 밀착하여 올리고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "프롬프트에 없는 인물이 전경에 크게 추가되었으며, 지정된 레퍼런스 장소와 클로즈업 프레이밍을 명백히 위반했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "지정된 장소와 클로즈업 앵글을 훌륭하게 구현했으나, 손가락의 해부학적 구조가 붕괴되었고 지시되지 않은 인물들이 배경에 나타납니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "서지민의 시선이 프레임 좌측 전경에 위치한, 지시되지 않은 인물을 향하고 있습니다.",
        "built_space": "일반적인 구내식당 형태이며, 레퍼런스의 핵심인 원형 유리벽 연구실 구조가 존재하지 않습니다.",
        "entities": "서지민의 외모는 레퍼런스와 일치하나, 프롬프트에 없는 남성의 어깨와 뒷모습이 화면을 크게 가리고 있습니다.",
        "hard_violations": [
         "프롬프트에 명시되지 않은 전경 인물 추가"
        ],
        "physics": "테이블 앞 의자에 앉아 양손을 들어 입을 가리고 있습니다."
       },
       {
        "label": "B",
        "direction": "서지민의 시선이 카메라 우측 바깥의 특정되지 않은 곳을 당황한 듯 응시하고 있습니다.",
        "built_space": "레퍼런스 이미지와 동일한 원형 유리벽 연구실과 우측의 배식대 시설이 정확히 구현되었습니다.",
        "entities": "서지민의 외모와 복장은 일치하나, 입을 가린 양손의 손가락 개수와 마디가 기형적으로 융합되어 있습니다.",
        "hard_violations": [
         "물리적으로 불가능한 기형적인 손가락 해부학 구조",
         "프롬프트에 명시되지 않은 배경 인물들 추가"
        ],
        "physics": "테이블 앞 의자에 앉아 양손을 입에 밀착하여 올리고 있습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "프롬프트에 없는 인물이 전경에 크게 추가되었으며, 지정된 레퍼런스 장소와 클로즈업 프레이밍을 명백히 위반했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "지정된 장소와 클로즈업 앵글을 훌륭하게 구현했으나, 손가락의 해부학적 구조가 붕괴되었고 지시되지 않은 인물들이 배경에 나타납니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "서지민의 시선이 프레임 좌측 전경에 위치한, 지시되지 않은 인물을 향하고 있습니다.",
        "built_space": "일반적인 구내식당 형태이며, 레퍼런스의 핵심인 원형 유리벽 연구실 구조가 존재하지 않습니다.",
        "entities": "서지민의 외모는 레퍼런스와 일치하나, 프롬프트에 없는 남성의 어깨와 뒷모습이 화면을 크게 가리고 있습니다.",
        "hard_violations": [
         "프롬프트에 명시되지 않은 전경 인물 추가"
        ],
        "physics": "테이블 앞 의자에 앉아 양손을 들어 입을 가리고 있습니다."
       },
       {
        "label": "B",
        "direction": "서지민의 시선이 카메라 우측 바깥의 특정되지 않은 곳을 당황한 듯 응시하고 있습니다.",
        "built_space": "레퍼런스 이미지와 동일한 원형 유리벽 연구실과 우측의 배식대 시설이 정확히 구현되었습니다.",
        "entities": "서지민의 외모와 복장은 일치하나, 입을 가린 양손의 손가락 개수와 마디가 기형적으로 융합되어 있습니다.",
        "hard_violations": [
         "물리적으로 불가능한 기형적인 손가락 해부학 구조",
         "프롬프트에 명시되지 않은 배경 인물들 추가"
        ],
        "physics": "테이블 앞 의자에 앉아 양손을 입에 밀착하여 올리고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "양손으로 입을 막은 얼굴을 더 크게 담고 참조의 곡면 유리 연구공간도 반영했지만, 금지된 배경 인물 5명을 추가해 실격입니다."
       },
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "당황한 표정과 양손 동작은 맞지만, 추가 인물들이 등장하는 넓은 어깨너머 구도로 얼굴 클로즈업을 벗어났으며 장소도 참조와 다릅니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "서지민의 두 눈은 화면 오른쪽의 프레임 밖을 향하며, 특정 상대를 바라보는지는 확인되지 않습니다. 두 손은 자신의 입으로 향해 입술을 가립니다. 배경의 앉은 인물들은 식탁 쪽, 오른쪽 인물들은 배식대와 음식 쪽을 보고 있습니다. 지시문은 서지민의 구체적인 시선 대상을 지정하지 않았습니다.",
        "built_space": "서지민 뒤로 금속 테두리를 두른 곡면 유리 구획 1개, 왼쪽 창과 식재, 오른쪽 배식대 1개가 보입니다. 전경 식탁 1개와 뒤쪽 식탁 2개가 구분됩니다. 곡면 유리와 금속 재료는 장소 참조에 가깝지만 배식 시설은 참조에서 확인되지 않습니다. 서지민은 식탁 앞에 있으며 좌석은 프레임 밖입니다. 얼굴뿐 아니라 가슴과 넓은 배경까지 담아 요구한 얼굴 클로즈업보다는 느슨합니다. 불가능한 반사는 보이지 않습니다.",
        "entities": "중앙에는 검은 중간 길이 머리와 앳된 얼굴의 동아시아계 젊은 여성 1명이 있으며, 보이는 눈매와 머리 형태는 서지민 참조에 비교적 가깝습니다. 연구 가운을 입고 양손으로 입을 가렸으며, 눈과 눈썹으로 당황함을 표현합니다. 왼쪽 식탁에는 먹지 않은 것으로 보이는 빵과 음식이 접시에 놓여 있습니다. 그러나 왼쪽 착석자와 그 뒤의 작은 인물, 중앙 뒤 착석자, 오른쪽 여성과 배식자까지 추가 인물 5명이 보입니다. 읽을 수 있는 글자는 보이지 않습니다.",
        "hard_violations": [
         "서지민만 등장해야 하는 장면에 배경 인물 5명을 추가했습니다."
        ],
        "physics": "두 손은 각각 손목과 팔에 자연스럽게 이어져 입과 얼굴에 닿아 있으며, 손을 급히 올려 입을 막는 동작으로 가능합니다. 서지민의 하체와 좌석은 잘려 있어 지지 상태를 직접 확인할 수 없지만 공중에 떠 있다는 증거는 없습니다. 음식과 식기는 식탁 위에 놓여 있고, 배경 착석자들은 의자에 앉아 있습니다. 명백히 지지되지 않은 물체나 불가능한 신체 자세는 보이지 않습니다."
       },
       {
        "label": "B",
        "direction": "서지민은 화면 왼쪽 전경에 있는 남성의 얼굴 쪽을 바라봅니다. 양손은 자신의 입과 코 아래쪽으로 모여 입을 가립니다. 전경 남성도 서지민 쪽을 향합니다. 뒤쪽 왼편 두 사람은 서로를 향하고, 오른쪽 서 있는 남성은 배식대를 향합니다. 대화 관계는 읽히지만 그 시선 상대 자체가 허용되지 않은 추가 인물입니다.",
        "built_space": "왼쪽에는 직선으로 이어지는 큰 창들, 오른쪽에는 배식대 1개와 작은 펜던트 조명 2개가 보이며, 여러 식탁과 등받이 의자가 줄지어 있습니다. 서지민은 전경 식탁 뒤 긴 등받이 좌석에 앉아 있습니다. 참조의 핵심인 원형 유리 연구 구획과 실험 설비 대신 목재 벽면의 일반 식당을 보여 장소 일치가 약합니다. 전경 상대의 어깨, 서지민의 허리 부근까지, 식탁 음식까지 포함한 넓은 어깨너머 구도로 얼굴 클로즈업이 아닙니다. 불가능한 반사는 보이지 않습니다.",
        "entities": "서지민은 검은 머리의 앳된 동아시아계 젊은 여성으로 표현되었고, 흰 연구 가운과 참조에 가까운 남색 상의를 입었습니다. 양손으로 입을 가리고 정상적인 눈을 크게 떠 당황함을 표현합니다. 전경 식탁에는 온전한 빵 1개와 물이 든 잔 2개가 있습니다. 왼쪽 전경 남성 1명 외에 뒤쪽 식사자 2명, 배식대 앞 남성 1명, 배식자 1명, 오른쪽 착석자 2명으로 추가 인물 총 7명이 보입니다. 읽을 수 있는 글자는 보이지 않습니다.",
        "hard_violations": [
         "서지민만 등장해야 하는 장면에 전경 대화 상대 1명과 배경 인물 6명을 추가했습니다."
        ],
        "physics": "서지민의 상체는 좌석 위에 자연스럽게 놓여 있고 등 뒤로 등받이가 보입니다. 두 팔과 손목이 연결되어 양손을 입에 대는 자세를 지지합니다. 빵과 잔은 식탁과 식판 위에 놓여 있습니다. 배경 인물들도 의자에 앉거나 바닥에 서 있는 것으로 읽히며, 지지 없이 뜬 신체나 물체는 보이지 않습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "양손으로 입을 막은 얼굴을 더 크게 담고 참조의 곡면 유리 연구공간도 반영했지만, 금지된 배경 인물 5명을 추가해 실격입니다."
       },
       {
        "label": "A",
        "score": 1,
        "verdict_ko": "당황한 표정과 양손 동작은 맞지만, 추가 인물들이 등장하는 넓은 어깨너머 구도로 얼굴 클로즈업을 벗어났으며 장소도 참조와 다릅니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "서지민의 두 눈은 화면 오른쪽의 프레임 밖을 향하며, 특정 상대를 바라보는지는 확인되지 않습니다. 두 손은 자신의 입으로 향해 입술을 가립니다. 배경의 앉은 인물들은 식탁 쪽, 오른쪽 인물들은 배식대와 음식 쪽을 보고 있습니다. 지시문은 서지민의 구체적인 시선 대상을 지정하지 않았습니다.",
        "built_space": "서지민 뒤로 금속 테두리를 두른 곡면 유리 구획 1개, 왼쪽 창과 식재, 오른쪽 배식대 1개가 보입니다. 전경 식탁 1개와 뒤쪽 식탁 2개가 구분됩니다. 곡면 유리와 금속 재료는 장소 참조에 가깝지만 배식 시설은 참조에서 확인되지 않습니다. 서지민은 식탁 앞에 있으며 좌석은 프레임 밖입니다. 얼굴뿐 아니라 가슴과 넓은 배경까지 담아 요구한 얼굴 클로즈업보다는 느슨합니다. 불가능한 반사는 보이지 않습니다.",
        "entities": "중앙에는 검은 중간 길이 머리와 앳된 얼굴의 동아시아계 젊은 여성 1명이 있으며, 보이는 눈매와 머리 형태는 서지민 참조에 비교적 가깝습니다. 연구 가운을 입고 양손으로 입을 가렸으며, 눈과 눈썹으로 당황함을 표현합니다. 왼쪽 식탁에는 먹지 않은 것으로 보이는 빵과 음식이 접시에 놓여 있습니다. 그러나 왼쪽 착석자와 그 뒤의 작은 인물, 중앙 뒤 착석자, 오른쪽 여성과 배식자까지 추가 인물 5명이 보입니다. 읽을 수 있는 글자는 보이지 않습니다.",
        "hard_violations": [
         "서지민만 등장해야 하는 장면에 배경 인물 5명을 추가했습니다."
        ],
        "physics": "두 손은 각각 손목과 팔에 자연스럽게 이어져 입과 얼굴에 닿아 있으며, 손을 급히 올려 입을 막는 동작으로 가능합니다. 서지민의 하체와 좌석은 잘려 있어 지지 상태를 직접 확인할 수 없지만 공중에 떠 있다는 증거는 없습니다. 음식과 식기는 식탁 위에 놓여 있고, 배경 착석자들은 의자에 앉아 있습니다. 명백히 지지되지 않은 물체나 불가능한 신체 자세는 보이지 않습니다."
       },
       {
        "label": "A",
        "direction": "서지민은 화면 왼쪽 전경에 있는 남성의 얼굴 쪽을 바라봅니다. 양손은 자신의 입과 코 아래쪽으로 모여 입을 가립니다. 전경 남성도 서지민 쪽을 향합니다. 뒤쪽 왼편 두 사람은 서로를 향하고, 오른쪽 서 있는 남성은 배식대를 향합니다. 대화 관계는 읽히지만 그 시선 상대 자체가 허용되지 않은 추가 인물입니다.",
        "built_space": "왼쪽에는 직선으로 이어지는 큰 창들, 오른쪽에는 배식대 1개와 작은 펜던트 조명 2개가 보이며, 여러 식탁과 등받이 의자가 줄지어 있습니다. 서지민은 전경 식탁 뒤 긴 등받이 좌석에 앉아 있습니다. 참조의 핵심인 원형 유리 연구 구획과 실험 설비 대신 목재 벽면의 일반 식당을 보여 장소 일치가 약합니다. 전경 상대의 어깨, 서지민의 허리 부근까지, 식탁 음식까지 포함한 넓은 어깨너머 구도로 얼굴 클로즈업이 아닙니다. 불가능한 반사는 보이지 않습니다.",
        "entities": "서지민은 검은 머리의 앳된 동아시아계 젊은 여성으로 표현되었고, 흰 연구 가운과 참조에 가까운 남색 상의를 입었습니다. 양손으로 입을 가리고 정상적인 눈을 크게 떠 당황함을 표현합니다. 전경 식탁에는 온전한 빵 1개와 물이 든 잔 2개가 있습니다. 왼쪽 전경 남성 1명 외에 뒤쪽 식사자 2명, 배식대 앞 남성 1명, 배식자 1명, 오른쪽 착석자 2명으로 추가 인물 총 7명이 보입니다. 읽을 수 있는 글자는 보이지 않습니다.",
        "hard_violations": [
         "서지민만 등장해야 하는 장면에 전경 대화 상대 1명과 배경 인물 6명을 추가했습니다."
        ],
        "physics": "서지민의 상체는 좌석 위에 자연스럽게 놓여 있고 등 뒤로 등받이가 보입니다. 두 팔과 손목이 연결되어 양손을 입에 대는 자세를 지지합니다. 빵과 잔은 식탁과 식판 위에 놓여 있습니다. 배경 인물들도 의자에 앉거나 바닥에 서 있는 것으로 읽히며, 지지 없이 뜬 신체나 물체는 보이지 않습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.083,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.833,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 프롬프트에 명시되지 않은 전경 인물 추가",
     "[gpt-high] 서지민만 등장해야 하는 장면에 전경 대화 상대 1명과 배경 인물 6명을 추가했습니다."
    ],
    "B": [
     "[gemini-pro] 물리적으로 불가능한 기형적인 손가락 해부학 구조",
     "[gemini-pro] 프롬프트에 명시되지 않은 배경 인물들 추가",
     "[gpt-high] 서지민만 등장해야 하는 장면에 배경 인물 5명을 추가했습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "A": 833,
   "B": 1750
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 833,
    "verdict_ko": "프롬프트에 없는 인물이 전경에 크게 추가되었으며, 지정된 레퍼런스 장소와 클로즈업 프레이밍을 명백히 위반했습니다.  ★위반: [gemini-pro] 프롬프트에 명시되지 않은 전경 인물 추가 / [gpt-high] 서지민만 등장해야 하는 장면에 전경 대화 상대 1명과 배경 인물 6명을 추가했습니다."
   },
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "지정된 장소와 클로즈업 앵글을 훌륭하게 구현했으나, 손가락의 해부학적 구조가 붕괴되었고 지시되지 않은 인물들이 배경에 나타납니다.  ★위반: [gemini-pro] 물리적으로 불가능한 기형적인 손가락 해부학 구조 / [gemini-pro] 프롬프트에 명시되지 않은 배경 인물들 추가 / [gpt-high] 서지민만 등장해야 하는 장면에 배경 인물 5명을 추가했습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L268B01.png",
    "asset_id": "a491e625-4c37-4571-b3ba-85a3daf996f9",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 서지민: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1147284>",
    "asset_id": "a529b479-7055-42c6-be7d-19d14b5e000f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92c-ff69-7414-8aab-afdda5a40185",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S85sh22__bgfirst_bg.png",
   "bg_asset_id": "44478bdf-e0e9-48ca-bfd4-2ad6d414a7f9",
   "bg_record_key": "S85sh22::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S85sh22::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T09:19:04.941446+00:00",
  "fingerprint": "148de2c82e1c5bf90acc313c8171ea68e3d74cea428d609e106ff65a918491b2",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S85sh22_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S85sh22_sel.png",
  "source_sha256": "1548d64fc9d2cf73c881d4bc01fe2220397ce514f159ab7aec44f6e9dff1b455",
  "file": "S85sh22_cine.png",
  "staged_sha256": "0196372e44e7078b593195fbb0b8bca528726c1197084942658146a601dd298b",
  "latency_ms": 10782
 },
 "S85sh23::signage": {
  "fp": "8760ec37abcff4aa",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S85sh23": {
  "input_fingerprint": "06c8865c7682522a",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 두 주먹을 꽉 쥔 채 분노로 일그러진 표정으로 입을 크게 벌리고 있는 현우의 역동적인 상체.\n\nLOCATION (lock): At the same communal cafeteria table inside the research facility, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Dining table (Bread and food remain in front of 현우) — The near edge runs diagonally across the bottom of the image; used as Maintains the dialogue geography and anchors the fists at a natural scale; Bread and food (Present on the table despite 현우's lack of interest in eating); used as Small peripheral context below the emotional action.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Continue the preceding ambient illumination and controlled contrast without a dramatic relighting for the anger.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The bread and other food remain uneaten on the dining table. 현우: He remains at the dining table, now visibly shocked and angry.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 두 주먹을 꽉 쥔 채 분노로 일그러진 표정으로 입을 크게 벌리고 있는 현우의 역동적인 상체.\n\nLOCATION (lock): At the same communal cafeteria table inside the research facility, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Dining table (Bread and food remain in front of 현우) — The near edge runs diagonally across the bottom of the image; used as Maintains the dialogue geography and anchors the fists at a natural scale; Bread and food (Present on the table despite 현우's lack of interest in eating); used as Small peripheral context below the emotional action.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Continue the preceding ambient illumination and controlled contrast without a dramatic relighting for the anger.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The bread and other food remain uneaten on the dining table. 현우: He remains at the dining table, now visibly shocked and angry.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 두 주먹을 꽉 쥔 채 분노로 일그러진 표정으로 입을 크게 벌리고 있는 현우의 역동적인 상체.\n\nLOCATION (lock): At the same communal cafeteria table inside the research facility, in daytime ambient light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Dining table (Bread and food remain in front of 현우) — The near edge runs diagonally across the bottom of the image; used as Maintains the dialogue geography and anchors the fists at a natural scale; Bread and food (Present on the table despite 현우's lack of interest in eating); used as Small peripheral context below the emotional action.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Continue the preceding ambient illumination and controlled contrast without a dramatic relighting for the anger.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The bread and other food remain uneaten on the dining table. 현우: He remains at the dining table, now visibly shocked and angry.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선은 카메라 정면 약간 아래를 향하고 있으며, 입을 크게 벌려 분노를 표출하고 있습니다.",
    "built_space": "연구소 식당 내부로, 뒤쪽의 원통형 유리 구조물과 오른쪽의 배식대 등 이전 샷의 공간적 특징을 잘 유지하고 있습니다.",
    "entities": "현우의 얼굴과 인상착의는 레퍼런스와 일치하며 테이블 위의 빵과 음식도 존재하지만, 프롬프트에서 금지한 인물(왼쪽 배경의 남성)이 등장합니다.",
    "hard_violations": [
     "[gemini-pro] invented people (프롬프트에서 이전 샷의 인물들을 절대 포함하지 말라고 명시했으나 배경에 인물이 추가됨)",
     "[gpt-high] 현우 외에는 누구도 등장하면 안 되는데 왼쪽 배경에 이전 숏의 인물 한 명이 남아 있습니다."
    ],
    "physics": "현우의 두 주먹은 테이블 가장자리에 닿아 지탱되고 있으며, 식판과 음식들은 테이블 표면 위에 안정적으로 놓여 있습니다."
   },
   {
    "label": "B",
    "direction": "현우의 시선은 프레임 오른쪽 밖을 향하고 있으며, 얼굴을 찌푸리고 입을 크게 벌리고 있습니다.",
    "built_space": "동일한 구내식당 내부로 공간의 구조적 특징과 고정된 요소들은 이전 샷과 일치하게 배치되어 있습니다.",
    "entities": "현우의 외형은 레퍼런스를 잘 반영하고 있으나, 프롬프트에서 명시적으로 제외하라고 지시한 이전 샷의 배경 인물들(앉아있는 남성 2명, 서 있는 남녀 2명)이 모두 그대로 등장합니다.",
    "hard_violations": [
     "[gemini-pro] invented people (프롬프트에서 명시적으로 금지한 이전 샷의 인물 4명이 배경에 모두 포함됨)",
     "[gpt-high] 현우만 등장해야 하는 장면에 이전 숏의 배경 인물 다섯 명을 남겼습니다."
    ],
    "physics": "현우는 허공에 두 주먹을 쥐고 팔 근육으로 지탱하고 있으며, 테이블 위의 음식과 식판은 중력에 맞게 표면에 놓여 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "프롬프트가 명시적으로 금지한 배경 인물이 포함되어 심각한 규칙 위반이 발생했으나, B에 비해 위반 요소가 적고 두 주먹이 식탁에 놓여 구도적 지시를 더 잘 따랐습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "프롬프트에서 절대 포함하지 말라고 강조한 이전 샷의 배경 인물들이 모두 그대로 나타나 심각한 규칙 위반(Hard Violation)이 발생했으며, 주먹의 위치도 지시된 구도와 다릅니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선은 카메라 정면 약간 아래를 향하고 있으며, 입을 크게 벌려 분노를 표출하고 있습니다.",
        "built_space": "연구소 식당 내부로, 뒤쪽의 원통형 유리 구조물과 오른쪽의 배식대 등 이전 샷의 공간적 특징을 잘 유지하고 있습니다.",
        "entities": "현우의 얼굴과 인상착의는 레퍼런스와 일치하며 테이블 위의 빵과 음식도 존재하지만, 프롬프트에서 금지한 인물(왼쪽 배경의 남성)이 등장합니다.",
        "hard_violations": [
         "invented people (프롬프트에서 이전 샷의 인물들을 절대 포함하지 말라고 명시했으나 배경에 인물이 추가됨)"
        ],
        "physics": "현우의 두 주먹은 테이블 가장자리에 닿아 지탱되고 있으며, 식판과 음식들은 테이블 표면 위에 안정적으로 놓여 있습니다."
       },
       {
        "label": "B",
        "direction": "현우의 시선은 프레임 오른쪽 밖을 향하고 있으며, 얼굴을 찌푸리고 입을 크게 벌리고 있습니다.",
        "built_space": "동일한 구내식당 내부로 공간의 구조적 특징과 고정된 요소들은 이전 샷과 일치하게 배치되어 있습니다.",
        "entities": "현우의 외형은 레퍼런스를 잘 반영하고 있으나, 프롬프트에서 명시적으로 제외하라고 지시한 이전 샷의 배경 인물들(앉아있는 남성 2명, 서 있는 남녀 2명)이 모두 그대로 등장합니다.",
        "hard_violations": [
         "invented people (프롬프트에서 명시적으로 금지한 이전 샷의 인물 4명이 배경에 모두 포함됨)"
        ],
        "physics": "현우는 허공에 두 주먹을 쥐고 팔 근육으로 지탱하고 있으며, 테이블 위의 음식과 식판은 중력에 맞게 표면에 놓여 있습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "프롬프트가 명시적으로 금지한 배경 인물이 포함되어 심각한 규칙 위반이 발생했으나, B에 비해 위반 요소가 적고 두 주먹이 식탁에 놓여 구도적 지시를 더 잘 따랐습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "프롬프트에서 절대 포함하지 말라고 강조한 이전 샷의 배경 인물들이 모두 그대로 나타나 심각한 규칙 위반(Hard Violation)이 발생했으며, 주먹의 위치도 지시된 구도와 다릅니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선은 카메라 정면 약간 아래를 향하고 있으며, 입을 크게 벌려 분노를 표출하고 있습니다.",
        "built_space": "연구소 식당 내부로, 뒤쪽의 원통형 유리 구조물과 오른쪽의 배식대 등 이전 샷의 공간적 특징을 잘 유지하고 있습니다.",
        "entities": "현우의 얼굴과 인상착의는 레퍼런스와 일치하며 테이블 위의 빵과 음식도 존재하지만, 프롬프트에서 금지한 인물(왼쪽 배경의 남성)이 등장합니다.",
        "hard_violations": [
         "invented people (프롬프트에서 이전 샷의 인물들을 절대 포함하지 말라고 명시했으나 배경에 인물이 추가됨)"
        ],
        "physics": "현우의 두 주먹은 테이블 가장자리에 닿아 지탱되고 있으며, 식판과 음식들은 테이블 표면 위에 안정적으로 놓여 있습니다."
       },
       {
        "label": "B",
        "direction": "현우의 시선은 프레임 오른쪽 밖을 향하고 있으며, 얼굴을 찌푸리고 입을 크게 벌리고 있습니다.",
        "built_space": "동일한 구내식당 내부로 공간의 구조적 특징과 고정된 요소들은 이전 샷과 일치하게 배치되어 있습니다.",
        "entities": "현우의 외형은 레퍼런스를 잘 반영하고 있으나, 프롬프트에서 명시적으로 제외하라고 지시한 이전 샷의 배경 인물들(앉아있는 남성 2명, 서 있는 남녀 2명)이 모두 그대로 등장합니다.",
        "hard_violations": [
         "invented people (프롬프트에서 명시적으로 금지한 이전 샷의 인물 4명이 배경에 모두 포함됨)"
        ],
        "physics": "현우는 허공에 두 주먹을 쥐고 팔 근육으로 지탱하고 있으며, 테이블 위의 음식과 식판은 중력에 맞게 표면에 놓여 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "분노한 얼굴과 꽉 쥔 두 주먹은 구현했지만, 금지된 배경 인물들을 그대로 남겼고 상체를 지나치게 타이트하게 잘라 지정된 미디엄 숏과 식탁 앞모서리 구도를 놓쳤습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "미디엄 숏의 역동적인 상체와 식탁의 대각선 모서리는 더 충실하지만, 금지된 배경 인물 한 명이 남아 있어 이 후보 역시 탈락 사유가 있습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 얼굴과 시선은 화면 오른쪽의 프레임 밖을 향하며, 입을 크게 벌리고 있습니다. 두 주먹은 턱 아래로 당겨져 있고 특정 상대를 향해 뻗지는 않습니다. 지문은 시선의 구체적인 대상을 지정하지 않았으므로 이 방향 자체는 어긋나지 않습니다.",
        "built_space": "중앙 뒤의 곡면 유리 구획 하나, 오른쪽의 금속 배식대 하나와 유리 가림막, 왼쪽 창과 화분이 이전 장소와 대응합니다. 왼쪽과 중앙 뒤에 식탁과 의자들이 보입니다. 현우 앞 식탁은 화면 왼쪽 아래에 보이지만, 가까운 모서리가 화면 하단을 대각선으로 가로지르는 구도는 드러나지 않습니다. 머리 윗부분과 하부 상체가 잘린 타이트한 구도로, 지정된 미디엄 숏보다 가깝습니다. 불가능한 반사는 보이지 않습니다.",
        "entities": "현우는 앳된 동아시아계 남성으로 보이며 헝클어진 검은 머리와 얼굴 특징은 인물 참고와 대체로 맞습니다. 국적은 외형만으로 확인할 수 없습니다. 남색 티셔츠 위에 참고 인물 사진에는 없는 회색 연구복을 입었습니다. 두 손은 주먹을 쥐었고 얼굴은 분노로 찌푸려져 있습니다. 빵, 다른 음식, 접시와 컵은 식탁에 남아 있습니다. 현우 외에 왼쪽의 앉은 남성, 그 뒤 작은 인물, 중앙 뒤의 앉은 남성, 오른쪽 배식대의 여성과 남성 등 배경 인물 다섯 명이 보입니다. 읽을 수 있는 문구는 보이지 않습니다.",
        "hard_violations": [
         "현우만 등장해야 하는 장면에 이전 숏의 배경 인물 다섯 명을 남겼습니다."
        ],
        "physics": "양 주먹은 손목과 굽힌 팔에 자연스럽게 연결되어 있으며 팔의 힘으로 턱 아래에 유지됩니다. 현우의 좌석과 하체는 프레임 밖이어서 접촉점을 확인할 수 없지만, 몸이 공중에 떠 있다는 징후는 없습니다. 음식과 식기는 식탁 위에 놓여 있고 배경의 앉은 인물들은 의자에 지지됩니다."
       },
       {
        "label": "B",
        "direction": "현우는 몸을 앞으로 기울이고 카메라 왼쪽 가까운 프레임 밖을 바라보며 입을 크게 벌립니다. 두 주먹은 각각 식탁 높이와 허리 앞에 있고 타격 대상을 향해 뻗지는 않습니다. 분노의 상대는 화면에 없으며, 지문에서 특정 표적을 요구하지 않아 방향상 모순은 없습니다.",
        "built_space": "중앙 뒤 곡면 유리 구획 하나, 오른쪽 금속 배식대 하나와 유리 가림막, 왼쪽 창과 화분이 유지되어 같은 구내식당으로 읽힙니다. 전경 식탁 하나와 뒤쪽 식탁들, 여러 빈 의자가 보이며 현우는 전경 식탁 오른쪽 의자에 앉아 있습니다. 식탁의 가까운 모서리가 하단 중앙에서 오른쪽 위로 비스듬히 이어지고, 머리부터 허리 부근까지 포함한 미디엄 숏이 구현됩니다. 유리면에 불가능한 반사는 보이지 않습니다.",
        "entities": "현우의 앳된 동아시아계 남성 외형, 헝클어진 검은 머리와 남색 티셔츠는 참고와 대체로 대응하지만, 회색 연구복이 추가되었습니다. 국적은 외형만으로 판별할 수 없습니다. 양손을 꽉 쥐고 눈썹과 코 주변을 찌푸린 채 입을 벌려 요청한 분노를 표현합니다. 식탁에는 빵과 다른 음식, 접시, 쟁반과 컵이 그대로 놓여 있습니다. 왼쪽 뒤에는 이전 숏의 앉은 남성 한 명이 남아 있습니다. 읽을 수 있는 글자는 보이지 않습니다.",
        "hard_violations": [
         "현우 외에는 누구도 등장하면 안 되는데 왼쪽 배경에 이전 숏의 인물 한 명이 남아 있습니다."
        ],
        "physics": "현우의 골반은 화면 하단의 의자 좌면에 놓인 것으로 읽히고, 의자 등받이와 팔걸이가 몸 뒤와 옆에 보입니다. 화면 왼쪽 주먹의 아래쪽은 식탁에 닿아 있으며 반대편 주먹은 굽힌 팔로 허리 앞에 유지됩니다. 앞으로 기울인 상체는 앉은 자세에서 가능한 동작입니다. 빵과 식기는 식탁이, 배경 남성은 의자가 지지하므로 근거 없이 떠 있는 몸이나 물체는 없습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "분노한 얼굴과 꽉 쥔 두 주먹은 구현했지만, 금지된 배경 인물들을 그대로 남겼고 상체를 지나치게 타이트하게 잘라 지정된 미디엄 숏과 식탁 앞모서리 구도를 놓쳤습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "미디엄 숏의 역동적인 상체와 식탁의 대각선 모서리는 더 충실하지만, 금지된 배경 인물 한 명이 남아 있어 이 후보 역시 탈락 사유가 있습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 얼굴과 시선은 화면 오른쪽의 프레임 밖을 향하며, 입을 크게 벌리고 있습니다. 두 주먹은 턱 아래로 당겨져 있고 특정 상대를 향해 뻗지는 않습니다. 지문은 시선의 구체적인 대상을 지정하지 않았으므로 이 방향 자체는 어긋나지 않습니다.",
        "built_space": "중앙 뒤의 곡면 유리 구획 하나, 오른쪽의 금속 배식대 하나와 유리 가림막, 왼쪽 창과 화분이 이전 장소와 대응합니다. 왼쪽과 중앙 뒤에 식탁과 의자들이 보입니다. 현우 앞 식탁은 화면 왼쪽 아래에 보이지만, 가까운 모서리가 화면 하단을 대각선으로 가로지르는 구도는 드러나지 않습니다. 머리 윗부분과 하부 상체가 잘린 타이트한 구도로, 지정된 미디엄 숏보다 가깝습니다. 불가능한 반사는 보이지 않습니다.",
        "entities": "현우는 앳된 동아시아계 남성으로 보이며 헝클어진 검은 머리와 얼굴 특징은 인물 참고와 대체로 맞습니다. 국적은 외형만으로 확인할 수 없습니다. 남색 티셔츠 위에 참고 인물 사진에는 없는 회색 연구복을 입었습니다. 두 손은 주먹을 쥐었고 얼굴은 분노로 찌푸려져 있습니다. 빵, 다른 음식, 접시와 컵은 식탁에 남아 있습니다. 현우 외에 왼쪽의 앉은 남성, 그 뒤 작은 인물, 중앙 뒤의 앉은 남성, 오른쪽 배식대의 여성과 남성 등 배경 인물 다섯 명이 보입니다. 읽을 수 있는 문구는 보이지 않습니다.",
        "hard_violations": [
         "현우만 등장해야 하는 장면에 이전 숏의 배경 인물 다섯 명을 남겼습니다."
        ],
        "physics": "양 주먹은 손목과 굽힌 팔에 자연스럽게 연결되어 있으며 팔의 힘으로 턱 아래에 유지됩니다. 현우의 좌석과 하체는 프레임 밖이어서 접촉점을 확인할 수 없지만, 몸이 공중에 떠 있다는 징후는 없습니다. 음식과 식기는 식탁 위에 놓여 있고 배경의 앉은 인물들은 의자에 지지됩니다."
       },
       {
        "label": "A",
        "direction": "현우는 몸을 앞으로 기울이고 카메라 왼쪽 가까운 프레임 밖을 바라보며 입을 크게 벌립니다. 두 주먹은 각각 식탁 높이와 허리 앞에 있고 타격 대상을 향해 뻗지는 않습니다. 분노의 상대는 화면에 없으며, 지문에서 특정 표적을 요구하지 않아 방향상 모순은 없습니다.",
        "built_space": "중앙 뒤 곡면 유리 구획 하나, 오른쪽 금속 배식대 하나와 유리 가림막, 왼쪽 창과 화분이 유지되어 같은 구내식당으로 읽힙니다. 전경 식탁 하나와 뒤쪽 식탁들, 여러 빈 의자가 보이며 현우는 전경 식탁 오른쪽 의자에 앉아 있습니다. 식탁의 가까운 모서리가 하단 중앙에서 오른쪽 위로 비스듬히 이어지고, 머리부터 허리 부근까지 포함한 미디엄 숏이 구현됩니다. 유리면에 불가능한 반사는 보이지 않습니다.",
        "entities": "현우의 앳된 동아시아계 남성 외형, 헝클어진 검은 머리와 남색 티셔츠는 참고와 대체로 대응하지만, 회색 연구복이 추가되었습니다. 국적은 외형만으로 판별할 수 없습니다. 양손을 꽉 쥐고 눈썹과 코 주변을 찌푸린 채 입을 벌려 요청한 분노를 표현합니다. 식탁에는 빵과 다른 음식, 접시, 쟁반과 컵이 그대로 놓여 있습니다. 왼쪽 뒤에는 이전 숏의 앉은 남성 한 명이 남아 있습니다. 읽을 수 있는 글자는 보이지 않습니다.",
        "hard_violations": [
         "현우 외에는 누구도 등장하면 안 되는데 왼쪽 배경에 이전 숏의 인물 한 명이 남아 있습니다."
        ],
        "physics": "현우의 골반은 화면 하단의 의자 좌면에 놓인 것으로 읽히고, 의자 등받이와 팔걸이가 몸 뒤와 옆에 보입니다. 화면 왼쪽 주먹의 아래쪽은 식탁에 닿아 있으며 반대편 주먹은 굽힌 팔로 허리 앞에 유지됩니다. 앞으로 기울인 상체는 앉은 자세에서 가능한 동작입니다. 빵과 식기는 식탁이, 배경 남성은 의자가 지지하므로 근거 없이 떠 있는 몸이나 물체는 없습니다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.167
   },
   "adjusted": {
    "A": 1.75,
    "B": 0.917
   },
   "violations": {
    "A": [
     "[gemini-pro] invented people (프롬프트에서 이전 샷의 인물들을 절대 포함하지 말라고 명시했으나 배경에 인물이 추가됨)",
     "[gpt-high] 현우 외에는 누구도 등장하면 안 되는데 왼쪽 배경에 이전 숏의 인물 한 명이 남아 있습니다."
    ],
    "B": [
     "[gemini-pro] invented people (프롬프트에서 명시적으로 금지한 이전 샷의 인물 4명이 배경에 모두 포함됨)",
     "[gpt-high] 현우만 등장해야 하는 장면에 이전 숏의 배경 인물 다섯 명을 남겼습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 917
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "프롬프트가 명시적으로 금지한 배경 인물이 포함되어 심각한 규칙 위반이 발생했으나, B에 비해 위반 요소가 적고 두 주먹이 식탁에 놓여 구도적 지시를 더 잘 따랐습니다.  ★위반: [gemini-pro] invented people (프롬프트에서 이전 샷의 인물들을 절대 포함하지 말라고 명시했으나 배경에 인물이 추가됨) / [gpt-high] 현우 외에는 누구도 등장하면 안 되는데 왼쪽 배경에 이전 숏의 인물 한 명이 남아 있습니다."
   },
   {
    "label": "B",
    "score": 917,
    "verdict_ko": "프롬프트에서 절대 포함하지 말라고 강조한 이전 샷의 배경 인물들이 모두 그대로 나타나 심각한 규칙 위반(Hard Violation)이 발생했으며, 주먹의 위치도 지시된 구도와 다릅니다.  ★위반: [gemini-pro] invented people (프롬프트에서 명시적으로 금지한 이전 샷의 인물 4명이 배경에 모두 포함됨) / [gpt-high] 현우만 등장해야 하는 장면에 이전 숏의 배경 인물 다섯 명을 남겼습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S85sh22_sel.png",
    "asset_id": "0d757fd4-b2af-49e8-a42f-6b7d9f02f5ae",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92d-02b4-74e5-955c-51599e41d8fa",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S85sh22"
  }
 },
 "S85sh23::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T09:20:05.042400+00:00",
  "fingerprint": "d2b5cf76d382e6f68b3d0ca0be50de1605cd9274e68ab1d6d7dbb56b0e8c979a",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S85sh23_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S85sh23_sel.png",
  "source_sha256": "ba20dda2ce3c8e8b8182a08cf40fa38504cd2adfac6db32451564b6ee06f0c80",
  "file": "S85sh23_cine.png",
  "staged_sha256": "2970791fa3481a790bd6cbd09d8c6d9dbf3530fd4df6bdec65b76c2e6ce365e9",
  "latency_ms": 10142
 },
 "S86sh2::signage": {
  "fp": "6611a6f01424ad0c",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S86sh2": {
  "input_fingerprint": "8499f0df03c9243a",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 지소영을 향해 핏발 선 눈으로 삿대질한 채 고함치는 현우의 분노한 상체.\n\nLOCATION (lock): In the open conversation area inside the research director's office, under nighttime office lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office doorway (Open following 현우's entrance) — Seen obliquely behind him on the established entrance side; used as Peripheral spatial anchor connecting his entrance to the confrontation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use setting-appropriate ambient illumination with controlled facial contrast, preserving the bloodshot eyes without inventing a colored source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old piano remains in the modern director's office. 현우: He has entered the director's office and remains standing in an agitated confrontation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 지소영을 향해 핏발 선 눈으로 삿대질한 채 고함치는 현우의 분노한 상체.\n\nLOCATION (lock): In the open conversation area inside the research director's office, under nighttime office lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office doorway (Open following 현우's entrance) — Seen obliquely behind him on the established entrance side; used as Peripheral spatial anchor connecting his entrance to the confrontation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use setting-appropriate ambient illumination with controlled facial contrast, preserving the bloodshot eyes without inventing a colored source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old piano remains in the modern director's office. 현우: He has entered the director's office and remains standing in an agitated confrontation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 지소영을 향해 핏발 선 눈으로 삿대질한 채 고함치는 현우의 분노한 상체.\n\nLOCATION (lock): In the open conversation area inside the research director's office, under nighttime office lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Office doorway (Open following 현우's entrance) — Seen obliquely behind him on the established entrance side; used as Peripheral spatial anchor connecting his entrance to the confrontation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use setting-appropriate ambient illumination with controlled facial contrast, preserving the bloodshot eyes without inventing a colored source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old piano remains in the modern director's office. 현우: He has entered the director's office and remains standing in an agitated confrontation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "시선과 삿대질하는 손가락이 프레임 밖 앞쪽의 대상을 향해 명확하게 일치함.",
    "built_space": "왼쪽 뒤편에 열린 문, 배경의 유리 벽, 오른쪽에 위치한 피아노가 기준 샷과 동일한 위치와 적절한 스케일로 배치됨.",
    "entities": "지시된 현우의 인상착의(얼굴, 헝클어진 머리)와 핏발 선 눈이 잘 구현되었으나, 티셔츠 색상이 기준 이미지(네이비)와 달리 짙은 회색임.",
    "hard_violations": [],
    "physics": "서 있는 몸의 무게 중심과 삿대질을 위해 들어 올린 팔이 자연스럽게 지지되고 있음."
   },
   {
    "label": "B",
    "direction": "시선은 정면을 향하고 있으나 팔은 프레임 왼쪽을 향해 뻗고 있어 지시 대상이 불일치함.",
    "built_space": "왼쪽에 열린 문과 배경 유리가 존재하나, 오른쪽에 있는 피아노의 크기가 기준 샷에 비해 지나치게 축소되어 멀리 밀려남.",
    "entities": "현우의 인상착의와 의상 색상(네이비), 핏발 선 눈이 기준에 부합함.",
    "hard_violations": [],
    "physics": "팔을 뻗은 자세가 신체 중심과 다소 어색하게 연결되며 손가락의 비례가 약간 부자연스러움."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "시선과 삿대질의 방향이 자연스럽게 일치하여 연기의 의도를 명확히 전달하며, 피아노와 배경의 공간 비율을 기준 이미지에 맞게 잘 유지했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "시선과 가리키는 손가락의 방향이 엇갈려 대상이 모호해졌으며, 피아노의 크기와 위치가 기준 샷에 비해 왜곡되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선과 삿대질하는 손가락이 프레임 밖 앞쪽의 대상을 향해 명확하게 일치함.",
        "built_space": "왼쪽 뒤편에 열린 문, 배경의 유리 벽, 오른쪽에 위치한 피아노가 기준 샷과 동일한 위치와 적절한 스케일로 배치됨.",
        "entities": "지시된 현우의 인상착의(얼굴, 헝클어진 머리)와 핏발 선 눈이 잘 구현되었으나, 티셔츠 색상이 기준 이미지(네이비)와 달리 짙은 회색임.",
        "hard_violations": [],
        "physics": "서 있는 몸의 무게 중심과 삿대질을 위해 들어 올린 팔이 자연스럽게 지지되고 있음."
       },
       {
        "label": "B",
        "direction": "시선은 정면을 향하고 있으나 팔은 프레임 왼쪽을 향해 뻗고 있어 지시 대상이 불일치함.",
        "built_space": "왼쪽에 열린 문과 배경 유리가 존재하나, 오른쪽에 있는 피아노의 크기가 기준 샷에 비해 지나치게 축소되어 멀리 밀려남.",
        "entities": "현우의 인상착의와 의상 색상(네이비), 핏발 선 눈이 기준에 부합함.",
        "hard_violations": [],
        "physics": "팔을 뻗은 자세가 신체 중심과 다소 어색하게 연결되며 손가락의 비례가 약간 부자연스러움."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "시선과 삿대질의 방향이 자연스럽게 일치하여 연기의 의도를 명확히 전달하며, 피아노와 배경의 공간 비율을 기준 이미지에 맞게 잘 유지했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "시선과 가리키는 손가락의 방향이 엇갈려 대상이 모호해졌으며, 피아노의 크기와 위치가 기준 샷에 비해 왜곡되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선과 삿대질하는 손가락이 프레임 밖 앞쪽의 대상을 향해 명확하게 일치함.",
        "built_space": "왼쪽 뒤편에 열린 문, 배경의 유리 벽, 오른쪽에 위치한 피아노가 기준 샷과 동일한 위치와 적절한 스케일로 배치됨.",
        "entities": "지시된 현우의 인상착의(얼굴, 헝클어진 머리)와 핏발 선 눈이 잘 구현되었으나, 티셔츠 색상이 기준 이미지(네이비)와 달리 짙은 회색임.",
        "hard_violations": [],
        "physics": "서 있는 몸의 무게 중심과 삿대질을 위해 들어 올린 팔이 자연스럽게 지지되고 있음."
       },
       {
        "label": "B",
        "direction": "시선은 정면을 향하고 있으나 팔은 프레임 왼쪽을 향해 뻗고 있어 지시 대상이 불일치함.",
        "built_space": "왼쪽에 열린 문과 배경 유리가 존재하나, 오른쪽에 있는 피아노의 크기가 기준 샷에 비해 지나치게 축소되어 멀리 밀려남.",
        "entities": "현우의 인상착의와 의상 색상(네이비), 핏발 선 눈이 기준에 부합함.",
        "hard_violations": [],
        "physics": "팔을 뻗은 자세가 신체 중심과 다소 어색하게 연결되며 손가락의 비례가 약간 부자연스러움."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "상체 중심의 미디엄 숏에서 시선과 삿대질이 화면 밖 왼쪽으로 일치하며, 열린 출입문과 남색 티셔츠까지 요구에 더 충실하다."
       },
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "분노한 고함과 충혈된 눈은 잘 구현했지만, 시선과 손끝의 목표 일치가 덜 명료하고 열린 출입문의 공간적 역할과 의상 색상도 A보다 약하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 눈과 얼굴은 화면 밖 왼쪽을 향하고, 뻗은 검지도 같은 왼쪽 전방을 가리킨다. 지소영은 보이지 않으므로 실제 목표 인물은 확인할 수 없지만, 화면 밖 한 상대에게 고함치며 삿대질하는 관계가 명확하다.",
        "built_space": "현우는 유리 칸막이로 둘러싸인 실내에 서 있다. 왼쪽 뒤에 열린 출입문 한 개가 비스듬히 보이고, 오른쪽 뒤에는 목재 업라이트 피아노 한 대와 낮은 의자 한 개가 있다. 유리 너머 장비 캐비닛과 모니터, 천장 조명이 보이며 참고 장소의 유리·금속·목재와 차가운 실내 조명을 유지한다. 피아노의 세부 외형은 참고와 다소 다르지만 배경 크기는 자연스럽다.",
        "entities": "화면에 뚜렷하게 등장하는 사람은 현우 한 명이다. 앳된 동아시아계 남성 외모, 헝클어진 검은 머리, 얼굴 윤곽과 피부 질감, 남색 반소매 티셔츠가 인물 참고와 가깝다. 한국계 미국인이라는 국적·배경은 외모만으로 확인할 수 없다. 눈 주변과 흰자에 붉은 기가 있고 정상적인 홍채와 동공이 유지된다. 입을 벌린 고함과 찌푸린 얼굴이 보이며, 오래된 목재 피아노도 존재한다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "삿대질하는 손은 손목·팔꿈치·어깨를 통해 몸통에 자연스럽게 연결되고, 다른 팔은 아래로 내려간다. 하체와 발은 프레임 밖이지만 상체는 서서 팔을 뻗는 자세로 성립하며 공중에 뜬 징후는 없다. 피아노와 의자는 바닥에 놓이고 출입문은 문틀에 연결되어 있다."
       },
       {
        "label": "B",
        "direction": "얼굴과 시선은 카메라 가까운 전방을 향하고, 검지는 단축된 팔 끝에서 화면 오른쪽 전방을 가리킨다. 같은 화면 밖 상대를 겨냥한 자세로 읽을 여지는 있지만, 시선과 손끝이 하나의 목표에 모이는 관계는 A보다 덜 명확하다. 지소영은 화면에 없어 정확한 목표 위치를 확인할 수 없다.",
        "built_space": "현우 뒤에 유리 칸막이와 장비 캐비닛, 모니터가 있고 오른쪽에는 목재 업라이트 피아노 한 대가 일부 잘려 보인다. 왼쪽 가장자리에는 출입구 한 개가 있으나 문짝과 개방 각도가 충분히 드러나지 않아, 입장 직후 열린 문이라는 공간 표지가 약하다. 참고의 연구실 재질과 천장 조명은 이어지며 불가능한 반사나 중복된 고정 설비는 보이지 않는다.",
        "entities": "현우로 보이는 젊은 동아시아계 남성 한 명만 등장한다. 검은 헝클어진 머리와 앳된 얼굴, 피부 질감은 참고와 가깝지만 티셔츠는 참고의 남색보다 짙은 회색으로 보인다. 국적·배경 자체는 시각적으로 확인할 수 없다. 충혈된 눈과 정상적인 동공·홍채, 크게 벌린 입과 긴장한 얼굴이 분노한 고함을 표현한다. 목재 피아노가 있으며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "오른팔을 들어 팔꿈치를 굽히고 검지를 뻗는 동작이 해부학적으로 자연스럽고, 손은 팔에 정상적으로 연결된다. 몸통은 약간 앞으로 기울어 선 자세이며 발은 미디엄 숏 밖에 있다. 떠 있는 신체나 지지 없는 소품은 보이지 않고, 피아노의 하부는 화면 밖으로 이어진다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "상체 중심의 미디엄 숏에서 시선과 삿대질이 화면 밖 왼쪽으로 일치하며, 열린 출입문과 남색 티셔츠까지 요구에 더 충실하다."
       },
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "분노한 고함과 충혈된 눈은 잘 구현했지만, 시선과 손끝의 목표 일치가 덜 명료하고 열린 출입문의 공간적 역할과 의상 색상도 A보다 약하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "현우의 눈과 얼굴은 화면 밖 왼쪽을 향하고, 뻗은 검지도 같은 왼쪽 전방을 가리킨다. 지소영은 보이지 않으므로 실제 목표 인물은 확인할 수 없지만, 화면 밖 한 상대에게 고함치며 삿대질하는 관계가 명확하다.",
        "built_space": "현우는 유리 칸막이로 둘러싸인 실내에 서 있다. 왼쪽 뒤에 열린 출입문 한 개가 비스듬히 보이고, 오른쪽 뒤에는 목재 업라이트 피아노 한 대와 낮은 의자 한 개가 있다. 유리 너머 장비 캐비닛과 모니터, 천장 조명이 보이며 참고 장소의 유리·금속·목재와 차가운 실내 조명을 유지한다. 피아노의 세부 외형은 참고와 다소 다르지만 배경 크기는 자연스럽다.",
        "entities": "화면에 뚜렷하게 등장하는 사람은 현우 한 명이다. 앳된 동아시아계 남성 외모, 헝클어진 검은 머리, 얼굴 윤곽과 피부 질감, 남색 반소매 티셔츠가 인물 참고와 가깝다. 한국계 미국인이라는 국적·배경은 외모만으로 확인할 수 없다. 눈 주변과 흰자에 붉은 기가 있고 정상적인 홍채와 동공이 유지된다. 입을 벌린 고함과 찌푸린 얼굴이 보이며, 오래된 목재 피아노도 존재한다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "삿대질하는 손은 손목·팔꿈치·어깨를 통해 몸통에 자연스럽게 연결되고, 다른 팔은 아래로 내려간다. 하체와 발은 프레임 밖이지만 상체는 서서 팔을 뻗는 자세로 성립하며 공중에 뜬 징후는 없다. 피아노와 의자는 바닥에 놓이고 출입문은 문틀에 연결되어 있다."
       },
       {
        "label": "A",
        "direction": "얼굴과 시선은 카메라 가까운 전방을 향하고, 검지는 단축된 팔 끝에서 화면 오른쪽 전방을 가리킨다. 같은 화면 밖 상대를 겨냥한 자세로 읽을 여지는 있지만, 시선과 손끝이 하나의 목표에 모이는 관계는 A보다 덜 명확하다. 지소영은 화면에 없어 정확한 목표 위치를 확인할 수 없다.",
        "built_space": "현우 뒤에 유리 칸막이와 장비 캐비닛, 모니터가 있고 오른쪽에는 목재 업라이트 피아노 한 대가 일부 잘려 보인다. 왼쪽 가장자리에는 출입구 한 개가 있으나 문짝과 개방 각도가 충분히 드러나지 않아, 입장 직후 열린 문이라는 공간 표지가 약하다. 참고의 연구실 재질과 천장 조명은 이어지며 불가능한 반사나 중복된 고정 설비는 보이지 않는다.",
        "entities": "현우로 보이는 젊은 동아시아계 남성 한 명만 등장한다. 검은 헝클어진 머리와 앳된 얼굴, 피부 질감은 참고와 가깝지만 티셔츠는 참고의 남색보다 짙은 회색으로 보인다. 국적·배경 자체는 시각적으로 확인할 수 없다. 충혈된 눈과 정상적인 동공·홍채, 크게 벌린 입과 긴장한 얼굴이 분노한 고함을 표현한다. 목재 피아노가 있으며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "오른팔을 들어 팔꿈치를 굽히고 검지를 뻗는 동작이 해부학적으로 자연스럽고, 손은 팔에 정상적으로 연결된다. 몸통은 약간 앞으로 기울어 선 자세이며 발은 미디엄 숏 밖에 있다. 떠 있는 신체나 지지 없는 소품은 보이지 않고, 피아노의 하부는 화면 밖으로 이어진다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.889,
    "B": 1.571
   },
   "adjusted": {
    "A": 1.889,
    "B": 1.571
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1889,
   "B": 1571
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1889,
    "verdict_ko": "시선과 삿대질의 방향이 자연스럽게 일치하여 연기의 의도를 명확히 전달하며, 피아노와 배경의 공간 비율을 기준 이미지에 맞게 잘 유지했습니다."
   },
   {
    "label": "B",
    "score": 1571,
    "verdict_ko": "시선과 가리키는 손가락의 방향이 엇갈려 대상이 모호해졌으며, 피아노의 크기와 위치가 기준 샷에 비해 왜곡되었습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S84sh14_sel.png",
    "asset_id": "e6735990-0fa6-4752-abc5-8848a0be4d3d",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92d-046a-766d-bc76-eb962fad83cb",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S84sh14"
  }
 },
 "S86sh2::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T09:21:11.556879+00:00",
  "fingerprint": "2c8723252a174b3a1263092734ae62e094029696622f2b72d256a437676df8ca",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S86sh2_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S86sh2_sel.png",
  "source_sha256": "9add11156cdeea3076b54a42e32db36506fc20d66e6f8c78daa195f439dc0919",
  "file": "S86sh2_cine.png",
  "staged_sha256": "36b5907f1102c2268ea7c2722b4fa034afd75c163c75a91b032c3910680ff22a",
  "latency_ms": 10827
 },
 "S86sh5::signage": {
  "fp": "740ce7685831109a",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S86sh5": {
  "input_fingerprint": "7b8353b68a1c5147",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 현우를 향해 양손을 펴 보인 채 애써 차분하게 달래는 기색의 지소영 굳은 상체.\n\nLOCATION (lock): Inside the director's office adjoining the research center, facing the visitor under nighttime interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained ambient illumination with sufficient tonal separation to read her open hands and the tension in her face.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old piano remains in the modern director's office. 지소영: She remains in her neat research coat, trying to maintain composure during the confrontation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 현우를 향해 양손을 펴 보인 채 애써 차분하게 달래는 기색의 지소영 굳은 상체.\n\nLOCATION (lock): Inside the director's office adjoining the research center, facing the visitor under nighttime interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained ambient illumination with sufficient tonal separation to read her open hands and the tension in her face.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old piano remains in the modern director's office. 지소영: She remains in her neat research coat, trying to maintain composure during the confrontation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 현우를 향해 양손을 펴 보인 채 애써 차분하게 달래는 기색의 지소영 굳은 상체.\n\nLOCATION (lock): Inside the director's office adjoining the research center, facing the visitor under nighttime interior lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain restrained ambient illumination with sufficient tonal separation to read her open hands and the tension in her face.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The old piano remains in the modern director's office. 지소영: She remains in her neat research coat, trying to maintain composure during the confrontation.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "시선은 정면(현우가 있는 방향)을 향하고 있으며, 양손은 손바닥을 위로 향해 앞으로 뻗어 대상을 달래고 있습니다.",
    "built_space": "레퍼런스와 동일한 연구소장실 내부로, 우측에 피아노, 좌측에 문틀, 배경에 유리벽과 모니터 등 고정 요소들이 정확한 위치에 있습니다.",
    "entities": "지소영(50대 중반 한국인 여성)이 묘사된 외모와 동일하게 흰색 연구 가운을 입고 등장합니다.",
    "hard_violations": [],
    "physics": "화면 밖 바닥에 서서 양팔을 자연스럽게 들어 올린 자세가 안정적으로 유지되고 있습니다."
   },
   {
    "label": "B",
    "direction": "시선은 정면을 향하고 있으며, 양손의 손바닥을 정면으로 펼쳐 대상을 진정시키려는 제스처를 취합니다.",
    "built_space": "레퍼런스와 일치하는 연구소장실로 피아노, 문틀, 유리벽, 배경의 의료 기기들이 올바르게 배치되어 있습니다.",
    "entities": "지소영과 일치하는 50대 중반 여성이 흰색 연구 가운을 입고 있습니다.",
    "hard_violations": [
     "[gemini-pro] 화면 우측에 위치한 인물의 왼손 손가락이 6개로 그려진 물리적 불가능(해부학적 오류)이 있습니다."
    ],
    "physics": "화면 밖 바닥을 지지대로 삼아 서 있으며, 들어 올린 팔의 자세가 작용에 부합합니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지시된 프레이밍과 배경 요소를 완벽하게 유지하며, 양손을 펴고 달래는 인물의 포즈와 표정을 해부학적 오류 없이 자연스럽게 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "배경과 인물의 구도는 지시사항을 잘 따랐으나, 화면 우측 손의 손가락 개수가 잘못되는 해부학적 오류가 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 정면(현우가 있는 방향)을 향하고 있으며, 양손은 손바닥을 위로 향해 앞으로 뻗어 대상을 달래고 있습니다.",
        "built_space": "레퍼런스와 동일한 연구소장실 내부로, 우측에 피아노, 좌측에 문틀, 배경에 유리벽과 모니터 등 고정 요소들이 정확한 위치에 있습니다.",
        "entities": "지소영(50대 중반 한국인 여성)이 묘사된 외모와 동일하게 흰색 연구 가운을 입고 등장합니다.",
        "hard_violations": [],
        "physics": "화면 밖 바닥에 서서 양팔을 자연스럽게 들어 올린 자세가 안정적으로 유지되고 있습니다."
       },
       {
        "label": "B",
        "direction": "시선은 정면을 향하고 있으며, 양손의 손바닥을 정면으로 펼쳐 대상을 진정시키려는 제스처를 취합니다.",
        "built_space": "레퍼런스와 일치하는 연구소장실로 피아노, 문틀, 유리벽, 배경의 의료 기기들이 올바르게 배치되어 있습니다.",
        "entities": "지소영과 일치하는 50대 중반 여성이 흰색 연구 가운을 입고 있습니다.",
        "hard_violations": [
         "화면 우측에 위치한 인물의 왼손 손가락이 6개로 그려진 물리적 불가능(해부학적 오류)이 있습니다."
        ],
        "physics": "화면 밖 바닥을 지지대로 삼아 서 있으며, 들어 올린 팔의 자세가 작용에 부합합니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지시된 프레이밍과 배경 요소를 완벽하게 유지하며, 양손을 펴고 달래는 인물의 포즈와 표정을 해부학적 오류 없이 자연스럽게 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "배경과 인물의 구도는 지시사항을 잘 따랐으나, 화면 우측 손의 손가락 개수가 잘못되는 해부학적 오류가 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 정면(현우가 있는 방향)을 향하고 있으며, 양손은 손바닥을 위로 향해 앞으로 뻗어 대상을 달래고 있습니다.",
        "built_space": "레퍼런스와 동일한 연구소장실 내부로, 우측에 피아노, 좌측에 문틀, 배경에 유리벽과 모니터 등 고정 요소들이 정확한 위치에 있습니다.",
        "entities": "지소영(50대 중반 한국인 여성)이 묘사된 외모와 동일하게 흰색 연구 가운을 입고 등장합니다.",
        "hard_violations": [],
        "physics": "화면 밖 바닥에 서서 양팔을 자연스럽게 들어 올린 자세가 안정적으로 유지되고 있습니다."
       },
       {
        "label": "B",
        "direction": "시선은 정면을 향하고 있으며, 양손의 손바닥을 정면으로 펼쳐 대상을 진정시키려는 제스처를 취합니다.",
        "built_space": "레퍼런스와 일치하는 연구소장실로 피아노, 문틀, 유리벽, 배경의 의료 기기들이 올바르게 배치되어 있습니다.",
        "entities": "지소영과 일치하는 50대 중반 여성이 흰색 연구 가운을 입고 있습니다.",
        "hard_violations": [
         "화면 우측에 위치한 인물의 왼손 손가락이 6개로 그려진 물리적 불가능(해부학적 오류)이 있습니다."
        ],
        "physics": "화면 밖 바닥을 지지대로 삼아 서 있으며, 들어 올린 팔의 자세가 작용에 부합합니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "양손을 펴 상대를 진정시키는 굳은 상체를 미디엄 숏으로 담았고, 단정한 연구복과 기존 공간의 연속성도 잘 지킨다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "애써 침착하려는 표정과 열린 양손은 명확하지만, 가슴 중심의 더 타이트한 구도와 얼룩진 연구복이 지정된 미디엄 숏과 단정한 복장에서 벗어난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 렌즈보다 약간 화면 왼쪽의 화면 밖 상대를 향하고, 앞으로 내민 두 손의 손바닥도 그 상대에게 보이는 방향이다. 현우 자체는 보이지 않지만 현우를 달래는 시선과 제스처의 방향은 일관된다.",
        "built_space": "왼쪽 출입문 틀 하나, 뒤쪽 금속 프레임 유리 칸막이, 연구 장비 랙들, 뒤쪽 모니터 두 대, 오른쪽 낡은 업라이트 피아노 한 대가 보인다. 지소영은 유리 칸막이 앞이자 피아노 왼쪽에 서 있다. 참고 공간의 재료와 차가운 실내 조명이 이어지며, 유리에 비친 천장 조명도 부자연스럽지 않다.",
        "entities": "보이는 인물은 지소영에 해당하는 중년 여성 한 명뿐이다. 한국인 50대 중반이라는 설정에 부합하는 외관이며, 짧게 정돈한 검은 머리와 얼굴 윤곽이 인물 참고와 가깝다. 깨끗한 흰 연구복을 입었고 두 손은 비어 있다. 오래된 목재 피아노가 유지되며 이전 장면의 남성이나 추가 인물은 없다. 읽을 수 있는 문구는 확인되지 않는다.",
        "hard_violations": [],
        "physics": "몸통은 수직으로 이어지고 양팔은 어깨와 굽힌 팔꿈치에서 자연스럽게 지지된다. 손목과 펼친 손가락도 가능한 자세다. 발과 바닥 접촉은 화면 밖이지만 부유를 시사하는 모습은 없다. 피아노와 연구 장비는 정상적으로 설치된 상태로 보인다."
       },
       {
        "label": "B",
        "direction": "시선은 카메라 왼쪽 가까이의 화면 밖 상대를 향한다. 양손을 넓게 벌려 손바닥을 상대 쪽으로 보이고 있어 현우에게 호소하며 진정시키는 행동으로 읽힌다. 직접 렌즈를 응시하는 무표정한 자세는 아니다.",
        "built_space": "왼쪽 출입문 틀 하나와 뒤쪽 유리 칸막이, 장비 랙들, 모니터 두 대, 오른쪽 업라이트 피아노 한 대가 보인다. 지소영은 피아노 왼쪽에서 카메라에 더 가까이 잡혀 배경 일부를 가린다. 참고 장소의 배치와 차가운 조명은 유지되지만, 구도는 허리까지 보여주는 미디엄 숏보다 가슴 중심으로 타이트하다. 불가능한 반사는 보이지 않는다.",
        "entities": "중년 여성 한 명만 등장하며 얼굴의 나이감, 짧은 검은 머리, 체형은 지소영 참고와 대체로 일치한다. 이마와 입 주변의 긴장이 두드러진다. 흰 연구복은 있으나 주머니와 소매에 얼룩이 보여 단정한 상태라는 지시와 차이가 있다. 양손에는 물건이 없고 낡은 목재 피아노는 오른쪽에 남아 있다. 추가 인물이나 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "앞으로 약간 기울인 상체에서 양팔을 뻗고 손바닥을 펼치는 동작은 해부학적으로 가능하다. 전완과 손목의 연결 및 지지가 보이며, 손이 독립적으로 떠 있지 않다. 하체는 구도 밖이므로 발의 접촉은 확인할 수 없지만 몸이 공중에 떠 있다는 근거는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "양손을 펴 상대를 진정시키는 굳은 상체를 미디엄 숏으로 담았고, 단정한 연구복과 기존 공간의 연속성도 잘 지킨다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "애써 침착하려는 표정과 열린 양손은 명확하지만, 가슴 중심의 더 타이트한 구도와 얼룩진 연구복이 지정된 미디엄 숏과 단정한 복장에서 벗어난다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "시선은 렌즈보다 약간 화면 왼쪽의 화면 밖 상대를 향하고, 앞으로 내민 두 손의 손바닥도 그 상대에게 보이는 방향이다. 현우 자체는 보이지 않지만 현우를 달래는 시선과 제스처의 방향은 일관된다.",
        "built_space": "왼쪽 출입문 틀 하나, 뒤쪽 금속 프레임 유리 칸막이, 연구 장비 랙들, 뒤쪽 모니터 두 대, 오른쪽 낡은 업라이트 피아노 한 대가 보인다. 지소영은 유리 칸막이 앞이자 피아노 왼쪽에 서 있다. 참고 공간의 재료와 차가운 실내 조명이 이어지며, 유리에 비친 천장 조명도 부자연스럽지 않다.",
        "entities": "보이는 인물은 지소영에 해당하는 중년 여성 한 명뿐이다. 한국인 50대 중반이라는 설정에 부합하는 외관이며, 짧게 정돈한 검은 머리와 얼굴 윤곽이 인물 참고와 가깝다. 깨끗한 흰 연구복을 입었고 두 손은 비어 있다. 오래된 목재 피아노가 유지되며 이전 장면의 남성이나 추가 인물은 없다. 읽을 수 있는 문구는 확인되지 않는다.",
        "hard_violations": [],
        "physics": "몸통은 수직으로 이어지고 양팔은 어깨와 굽힌 팔꿈치에서 자연스럽게 지지된다. 손목과 펼친 손가락도 가능한 자세다. 발과 바닥 접촉은 화면 밖이지만 부유를 시사하는 모습은 없다. 피아노와 연구 장비는 정상적으로 설치된 상태로 보인다."
       },
       {
        "label": "A",
        "direction": "시선은 카메라 왼쪽 가까이의 화면 밖 상대를 향한다. 양손을 넓게 벌려 손바닥을 상대 쪽으로 보이고 있어 현우에게 호소하며 진정시키는 행동으로 읽힌다. 직접 렌즈를 응시하는 무표정한 자세는 아니다.",
        "built_space": "왼쪽 출입문 틀 하나와 뒤쪽 유리 칸막이, 장비 랙들, 모니터 두 대, 오른쪽 업라이트 피아노 한 대가 보인다. 지소영은 피아노 왼쪽에서 카메라에 더 가까이 잡혀 배경 일부를 가린다. 참고 장소의 배치와 차가운 조명은 유지되지만, 구도는 허리까지 보여주는 미디엄 숏보다 가슴 중심으로 타이트하다. 불가능한 반사는 보이지 않는다.",
        "entities": "중년 여성 한 명만 등장하며 얼굴의 나이감, 짧은 검은 머리, 체형은 지소영 참고와 대체로 일치한다. 이마와 입 주변의 긴장이 두드러진다. 흰 연구복은 있으나 주머니와 소매에 얼룩이 보여 단정한 상태라는 지시와 차이가 있다. 양손에는 물건이 없고 낡은 목재 피아노는 오른쪽에 남아 있다. 추가 인물이나 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "앞으로 약간 기울인 상체에서 양팔을 뻗고 손바닥을 펼치는 동작은 해부학적으로 가능하다. 전완과 손목의 연결 및 지지가 보이며, 손이 독립적으로 떠 있지 않다. 하체는 구도 밖이므로 발의 접촉은 확인할 수 없지만 몸이 공중에 떠 있다는 근거는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.778,
    "B": 1.429
   },
   "adjusted": {
    "A": 1.778,
    "B": 1.179
   },
   "violations": {
    "B": [
     "[gemini-pro] 화면 우측에 위치한 인물의 왼손 손가락이 6개로 그려진 물리적 불가능(해부학적 오류)이 있습니다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1778,
   "B": 1179
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1778,
    "verdict_ko": "지시된 프레이밍과 배경 요소를 완벽하게 유지하며, 양손을 펴고 달래는 인물의 포즈와 표정을 해부학적 오류 없이 자연스럽게 구현했습니다."
   },
   {
    "label": "B",
    "score": 1179,
    "verdict_ko": "배경과 인물의 구도는 지시사항을 잘 따랐으나, 화면 우측 손의 손가락 개수가 잘못되는 해부학적 오류가 발생했습니다.  ★위반: [gemini-pro] 화면 우측에 위치한 인물의 왼손 손가락이 6개로 그려진 물리적 불가능(해부학적 오류)이 있습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S86sh2_sel.png",
    "asset_id": "3119aa13-7a5b-4a9d-b394-0f1405fbbb14",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 지소영: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1243508>",
    "asset_id": "c7496f13-cfcf-44a5-976d-96c783d20580",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92d-0616-7c0d-90f6-39d51f3ebe78",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S86sh2"
  }
 },
 "S86sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T09:22:10.062587+00:00",
  "fingerprint": "da24cbb475548370b2fb4f4439c00dd2fd240bfe7cace57666ec9b268fa7e6cb",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S86sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S86sh5_sel.png",
  "source_sha256": "dca670c01a29494a8faa557fc62a2f0db3a69bfe6268f0a32c2e6e0de49208ee",
  "file": "S86sh5_cine.png",
  "staged_sha256": "ae0e19fb7f77922f47523cc27e1e07fbba18874b43d192525e374bf98cd08dc0",
  "latency_ms": 11175
 },
 "S87sh4::signage": {
  "fp": "4579be01bc0f5441",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::acc0168d60745657": {
  "subjects": [],
  "subject_text": "특임대 이강준의 사무실\n문으로 출입하는 군 부대 사무실. 업무용 책상과 문건을 놓을 수 있는 실무 공간으로 구성된다.",
  "identity": "canonical",
  "scope_id": "L273",
  "scope_role": "location_interior",
  "scope_sha": "d3aa87b72484aa83"
 },
 "groupbg::special_unit_office": {
  "input_fingerprint": "4b5b811772f286dc",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "special_unit_office",
    "tags": [
     "S87sh4"
    ]
   },
   "context_sig": "0a0ddcecc943e9d8"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: Inside the special-operations commander's office, in the briefing area under nighttime office lighting.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n특임대 이강준의 사무실: CCTV 분석 모니터와 보고서가 결재되는 딱딱한 분위기의 군 지휘관 방. (특징: 정돈된 군용 지휘 데스크; 종이 문서가 끼워진 결재판; 드론이나 해안 CCTV 영상이 띄워진 다중 모니터 화면)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 절도 있는 보폭으로 걸어가는 군홧발. 틸업하면 특임대장 사무실\n\nTIME OF DAY (lock): night.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: Inside the special-operations commander's office, in the briefing area under nighttime office lighting.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n특임대 이강준의 사무실: CCTV 분석 모니터와 보고서가 결재되는 딱딱한 분위기의 군 지휘관 방. (특징: 정돈된 군용 지휘 데스크; 종이 문서가 끼워진 결재판; 드론이나 해안 CCTV 영상이 띄워진 다중 모니터 화면)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 절도 있는 보폭으로 걸어가는 군홧발. 틸업하면 특임대장 사무실\n\nTIME OF DAY (lock): night.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_special_unit_office_2f9882.png",
  "asset_id": "4964e681-31dd-4c79-aad2-e3da5f4b52d2",
  "input_asset_ids": [
   "50017786-0851-4a08-8e8f-72323e26d3ef"
  ],
  "origin_tag": "S87sh4",
  "place_text": "Inside the special-operations commander's office, in the briefing area under nighttime office lighting.",
  "origin_inputs": {
   "place_text": "Inside the special-operations commander's office, in the briefing area under nighttime office lighting.",
   "time_of_day_en": "night",
   "conti_asset_id": "50017786-0851-4a08-8e8f-72323e26d3ef"
  }
 },
 "S87sh4::bgfirst_bg": {
  "input_fingerprint": "07f6cefe9afb97ea",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 허공을 향해 주먹을 꽉 쥔 채, 출동을 지시하듯 크게 입을 벌리고 있는 이강준의 핏대 선 상체.\n\nLOCATION (lock): Inside the special-operations commander's office, in the briefing area under nighttime office lighting.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Controlled ambient illumination preserves the strain in his neck and face without introducing an unsupported military-office lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 허공을 향해 주먹을 꽉 쥔 채, 출동을 지시하듯 크게 입을 벌리고 있는 이강준의 핏대 선 상체.\n\nLOCATION (lock): Inside the special-operations commander's office, in the briefing area under nighttime office lighting.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Controlled ambient illumination preserves the strain in his neck and face without introducing an unsupported military-office lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S87sh4__bgfirst_bg.png",
  "asset_id": "4e972bfb-a762-4f2f-9abf-8faa5c7623ed",
  "input_asset_ids": [
   "50017786-0851-4a08-8e8f-72323e26d3ef",
   "4964e681-31dd-4c79-aad2-e3da5f4b52d2"
  ]
 },
 "S87sh4": {
  "input_fingerprint": "ce921913108e96db",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 허공을 향해 주먹을 꽉 쥔 채, 출동을 지시하듯 크게 입을 벌리고 있는 이강준의 핏대 선 상체.\n\nLOCATION (lock): Inside the special-operations commander's office, in the briefing area under nighttime office lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Controlled ambient illumination preserves the strain in his neck and face without introducing an unsupported military-office lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 이강준 (한국인, 성인 남성, 짧은 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 허공을 향해 주먹을 꽉 쥔 채, 출동을 지시하듯 크게 입을 벌리고 있는 이강준의 핏대 선 상체.\n\nLOCATION (lock): Inside the special-operations commander's office, in the briefing area under nighttime office lighting. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Controlled ambient illumination preserves the strain in his neck and face without introducing an unsupported military-office lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 이강준 (한국인, 성인 남성, 짧은 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 허공을 향해 주먹을 꽉 쥔 채, 출동을 지시하듯 크게 입을 벌리고 있는 이강준의 핏대 선 상체.\n\nLOCATION (lock): Inside the special-operations commander's office, in the briefing area under nighttime office lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Controlled ambient illumination preserves the strain in his neck and face without introducing an unsupported military-office lighting effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 이강준 (한국인, 성인 남성, 짧은 머리, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S87sh4__bgfirst_bg.png",
     "asset_id": "4e972bfb-a762-4f2f-9abf-8faa5c7623ed",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S87sh4.png",
     "asset_id": "50017786-0851-4a08-8e8f-72323e26d3ef",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 이강준: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1327185>",
     "asset_id": "6dbc9554-4775-4399-b597-16c2aa160d07",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_special_unit_office_2f9882.png",
     "asset_id": "4964e681-31dd-4c79-aad2-e3da5f4b52d2",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 이강준: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1327185>",
     "asset_id": "6dbc9554-4775-4399-b597-16c2aa160d07",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "주 인물은 화면 오른쪽 전경에 있는 남자를 향해 시선을 두고 소리치고 있음.",
    "built_space": "야간 지휘관실. 데스크, 대형 스크린, 회의 테이블, 깃발 등 레퍼런스의 구조와 일치함.",
    "entities": "이강준(성인 남성, 짧은 검은 머리)이 묘사된 행동을 취하고 있으나, 프롬프트에 없는 남자가 전경에 나타남.",
    "hard_violations": [
     "[gemini-pro] 프롬프트에 명시되지 않은 인물 추가 (전경 오른쪽)",
     "[gpt-high] 화면 오른쪽 전경에 이강준 외의 두 번째 남성을 추가하여, 지정 인물 한 명만 보여야 한다는 조건을 위반했다."
    ],
    "physics": "인물은 바닥에 발을 딛고 자연스럽게 서서 체중을 지탱함."
   },
   {
    "label": "B",
    "direction": "인물은 화면 왼쪽 밖 허공을 향해 시선을 두고 주먹을 들어 올림.",
    "built_space": "야간 지휘관실. 스크린, 깃발, 회의 테이블과 의자 배치가 레퍼런스와 일치함.",
    "entities": "이강준(성인 남성, 짧은 검은 머리)이 레퍼런스의 외형과 일치하며, 지시된 표정과 주먹 쥔 상체를 보여줌.",
    "hard_violations": [],
    "physics": "인물은 안정적인 자세로 서서 행동을 취하고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "프롬프트에 언급되지 않은 인물이 전경에 추가되어 결정적인 위반 사항이 발생했습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "미디엄 샷 구도 내에서 허공을 향한 주먹, 핏대 선 목, 벌린 입 등 명시된 인물의 행동과 표정을 정확하게 재현했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "주 인물은 화면 오른쪽 전경에 있는 남자를 향해 시선을 두고 소리치고 있음.",
        "built_space": "야간 지휘관실. 데스크, 대형 스크린, 회의 테이블, 깃발 등 레퍼런스의 구조와 일치함.",
        "entities": "이강준(성인 남성, 짧은 검은 머리)이 묘사된 행동을 취하고 있으나, 프롬프트에 없는 남자가 전경에 나타남.",
        "hard_violations": [
         "프롬프트에 명시되지 않은 인물 추가 (전경 오른쪽)"
        ],
        "physics": "인물은 바닥에 발을 딛고 자연스럽게 서서 체중을 지탱함."
       },
       {
        "label": "B",
        "direction": "인물은 화면 왼쪽 밖 허공을 향해 시선을 두고 주먹을 들어 올림.",
        "built_space": "야간 지휘관실. 스크린, 깃발, 회의 테이블과 의자 배치가 레퍼런스와 일치함.",
        "entities": "이강준(성인 남성, 짧은 검은 머리)이 레퍼런스의 외형과 일치하며, 지시된 표정과 주먹 쥔 상체를 보여줌.",
        "hard_violations": [],
        "physics": "인물은 안정적인 자세로 서서 행동을 취하고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "프롬프트에 언급되지 않은 인물이 전경에 추가되어 결정적인 위반 사항이 발생했습니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "미디엄 샷 구도 내에서 허공을 향한 주먹, 핏대 선 목, 벌린 입 등 명시된 인물의 행동과 표정을 정확하게 재현했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "주 인물은 화면 오른쪽 전경에 있는 남자를 향해 시선을 두고 소리치고 있음.",
        "built_space": "야간 지휘관실. 데스크, 대형 스크린, 회의 테이블, 깃발 등 레퍼런스의 구조와 일치함.",
        "entities": "이강준(성인 남성, 짧은 검은 머리)이 묘사된 행동을 취하고 있으나, 프롬프트에 없는 남자가 전경에 나타남.",
        "hard_violations": [
         "프롬프트에 명시되지 않은 인물 추가 (전경 오른쪽)"
        ],
        "physics": "인물은 바닥에 발을 딛고 자연스럽게 서서 체중을 지탱함."
       },
       {
        "label": "B",
        "direction": "인물은 화면 왼쪽 밖 허공을 향해 시선을 두고 주먹을 들어 올림.",
        "built_space": "야간 지휘관실. 스크린, 깃발, 회의 테이블과 의자 배치가 레퍼런스와 일치함.",
        "entities": "이강준(성인 남성, 짧은 검은 머리)이 레퍼런스의 외형과 일치하며, 지시된 표정과 주먹 쥔 상체를 보여줌.",
        "hard_violations": [],
        "physics": "인물은 안정적인 자세로 서서 행동을 취하고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "상체 중심의 미디엄 숏에서 허공에 쥔 주먹, 크게 벌린 입과 목의 핏대를 정확히 살렸으며, 기준 인물의 정장과 야간 사무실도 유지했다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "허공의 주먹과 외치는 표정은 맞지만, 금지된 두 번째 인물을 추가했고 허벅지까지 넓힌 구도와 검은 티셔츠도 지정된 상체 숏 및 기준 복장에서 벗어난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "이강준은 화면 오른쪽 바깥을 바라보며 입을 크게 벌리고 있다. 화면 왼쪽의 주먹은 머리보다 높게 들어 빈 공간을 향하며, 사람이나 물체를 가격하는 방향이 아니다. 출동을 지시하는 듯한 시선과 허공에 쥔 주먹이라는 지시가 맞는다.",
        "built_space": "왼쪽 집무 책상 1개, 뒤쪽 벽면의 다중 화면 설비 1세트, 깃발 1개, 지도 부조 1개, 낮은 수납장과 화분 1개, 천장 선형 조명 1개가 보인다. 오른쪽에는 회의 테이블 1개와 명확히 구분되는 의자 4개, 야경이 보이는 창이 있다. 인물은 책상 앞과 회의 구역 사이에 서 있으며 가구와 겹쳐 관통하지 않는다. 회색 벽체와 따뜻한 간접조명, 창의 위치가 장소 기준과 잘 맞고 불가능한 반사는 보이지 않는다.",
        "entities": "성인 남성 1명만 등장하며, 한국인으로 설정된 기준 인물의 짧은 검은 머리, 얼굴 특징과 체격에 대체로 부합한다. 짙은 정장, 흰 셔츠와 어두운 넥타이도 유지된다. 크게 열린 입과 도드라진 목의 힘줄이 보이고 눈은 정상적인 사람의 눈이다. 셔츠 깃과 넥타이는 기준 사진보다 느슨하지만 동작 중 복장으로 자연스럽다. 화면이나 사물에서 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "들어 올린 주먹은 손목과 굽힌 팔, 어깨에 자연스럽게 연결되어 근육으로 지지된다. 앞으로 기운 상체는 화면 아래로 이어지는 몸통에 연결되며 떠 있는 모습이 아니다. 발은 상체 중심 구도 밖이므로 접지는 확인할 수 없지만, 지지 없는 공중 자세를 시사하지 않는다. 책상과 의자는 바닥에 놓이고 화분은 수납장 위에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "중앙 남성은 오른쪽 전경에 있는 다른 남성을 바라보며 외치고, 전경 남성도 그를 향해 몸과 얼굴을 돌리고 있다. 올라간 주먹 자체는 머리 위의 빈 공간을 향해 있어 허공에 쥔 주먹이라는 동작은 맞지만, 시선의 수신자로 허용되지 않은 인물을 화면에 추가했다.",
        "built_space": "왼쪽 집무 책상 1개와 높은 등받이 의자 1개, 깃발 1개, 벽면 화면 설비 1세트, 지도 부조 1개, 수납장과 화분 1개, 천장 선형 조명 1개가 보인다. 오른쪽 회의 구역에는 전경과 후경으로 나뉘어 보이는 탁자 면과 최소 4개의 의자 등받이가 있다. 회색 벽체, 간접조명과 창밖 야경은 기준 장소를 따른다. 주인공은 책상 앞에 서 있지만 오른쪽 전경을 추가 인물의 어깨와 머리가 크게 가려, 지정된 단독 상체 장면이 대화 상대의 어깨 너머 구도로 바뀌었다.",
        "entities": "외치는 성인 남성은 짧은 검은 머리와 얼굴에서 이강준 기준을 대체로 따르지만, 정장·흰 셔츠·넥타이 대신 몸에 붙는 검은 긴팔 티셔츠를 입었다. 오른쪽에는 검은 머리에 정장을 입은 별도의 성인 남성이 부분적으로 등장한다. 허용 인물은 이강준 한 명뿐이므로 이 추가 인물은 명백한 위반이다. 주인공의 열린 입과 목의 핏대는 잘 보이며, 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [
         "화면 오른쪽 전경에 이강준 외의 두 번째 남성을 추가하여, 지정 인물 한 명만 보여야 한다는 조건을 위반했다."
        ],
        "physics": "주인공의 올라간 주먹은 굽힌 팔과 어깨로 지지되고, 내려간 주먹도 팔 끝에 자연스럽게 연결된다. 상체는 벌린 다리 쪽으로 이어져 외치는 서 있는 자세로 가능하다. 발은 화면 밖이지만 몸이 공중에 떠 있는 징후는 없다. 전경 남성 역시 몸통이 화면 아래로 이어지고, 보이는 가구와 소품도 바닥이나 가구 면에 지지되어 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "상체 중심의 미디엄 숏에서 허공에 쥔 주먹, 크게 벌린 입과 목의 핏대를 정확히 살렸으며, 기준 인물의 정장과 야간 사무실도 유지했다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "허공의 주먹과 외치는 표정은 맞지만, 금지된 두 번째 인물을 추가했고 허벅지까지 넓힌 구도와 검은 티셔츠도 지정된 상체 숏 및 기준 복장에서 벗어난다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "이강준은 화면 오른쪽 바깥을 바라보며 입을 크게 벌리고 있다. 화면 왼쪽의 주먹은 머리보다 높게 들어 빈 공간을 향하며, 사람이나 물체를 가격하는 방향이 아니다. 출동을 지시하는 듯한 시선과 허공에 쥔 주먹이라는 지시가 맞는다.",
        "built_space": "왼쪽 집무 책상 1개, 뒤쪽 벽면의 다중 화면 설비 1세트, 깃발 1개, 지도 부조 1개, 낮은 수납장과 화분 1개, 천장 선형 조명 1개가 보인다. 오른쪽에는 회의 테이블 1개와 명확히 구분되는 의자 4개, 야경이 보이는 창이 있다. 인물은 책상 앞과 회의 구역 사이에 서 있으며 가구와 겹쳐 관통하지 않는다. 회색 벽체와 따뜻한 간접조명, 창의 위치가 장소 기준과 잘 맞고 불가능한 반사는 보이지 않는다.",
        "entities": "성인 남성 1명만 등장하며, 한국인으로 설정된 기준 인물의 짧은 검은 머리, 얼굴 특징과 체격에 대체로 부합한다. 짙은 정장, 흰 셔츠와 어두운 넥타이도 유지된다. 크게 열린 입과 도드라진 목의 힘줄이 보이고 눈은 정상적인 사람의 눈이다. 셔츠 깃과 넥타이는 기준 사진보다 느슨하지만 동작 중 복장으로 자연스럽다. 화면이나 사물에서 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "들어 올린 주먹은 손목과 굽힌 팔, 어깨에 자연스럽게 연결되어 근육으로 지지된다. 앞으로 기운 상체는 화면 아래로 이어지는 몸통에 연결되며 떠 있는 모습이 아니다. 발은 상체 중심 구도 밖이므로 접지는 확인할 수 없지만, 지지 없는 공중 자세를 시사하지 않는다. 책상과 의자는 바닥에 놓이고 화분은 수납장 위에 놓여 있다."
       },
       {
        "label": "A",
        "direction": "중앙 남성은 오른쪽 전경에 있는 다른 남성을 바라보며 외치고, 전경 남성도 그를 향해 몸과 얼굴을 돌리고 있다. 올라간 주먹 자체는 머리 위의 빈 공간을 향해 있어 허공에 쥔 주먹이라는 동작은 맞지만, 시선의 수신자로 허용되지 않은 인물을 화면에 추가했다.",
        "built_space": "왼쪽 집무 책상 1개와 높은 등받이 의자 1개, 깃발 1개, 벽면 화면 설비 1세트, 지도 부조 1개, 수납장과 화분 1개, 천장 선형 조명 1개가 보인다. 오른쪽 회의 구역에는 전경과 후경으로 나뉘어 보이는 탁자 면과 최소 4개의 의자 등받이가 있다. 회색 벽체, 간접조명과 창밖 야경은 기준 장소를 따른다. 주인공은 책상 앞에 서 있지만 오른쪽 전경을 추가 인물의 어깨와 머리가 크게 가려, 지정된 단독 상체 장면이 대화 상대의 어깨 너머 구도로 바뀌었다.",
        "entities": "외치는 성인 남성은 짧은 검은 머리와 얼굴에서 이강준 기준을 대체로 따르지만, 정장·흰 셔츠·넥타이 대신 몸에 붙는 검은 긴팔 티셔츠를 입었다. 오른쪽에는 검은 머리에 정장을 입은 별도의 성인 남성이 부분적으로 등장한다. 허용 인물은 이강준 한 명뿐이므로 이 추가 인물은 명백한 위반이다. 주인공의 열린 입과 목의 핏대는 잘 보이며, 판독 가능한 글자는 보이지 않는다.",
        "hard_violations": [
         "화면 오른쪽 전경에 이강준 외의 두 번째 남성을 추가하여, 지정 인물 한 명만 보여야 한다는 조건을 위반했다."
        ],
        "physics": "주인공의 올라간 주먹은 굽힌 팔과 어깨로 지지되고, 내려간 주먹도 팔 끝에 자연스럽게 연결된다. 상체는 벌린 다리 쪽으로 이어져 외치는 서 있는 자세로 가능하다. 발은 화면 밖이지만 몸이 공중에 떠 있는 징후는 없다. 전경 남성 역시 몸통이 화면 아래로 이어지고, 보이는 가구와 소품도 바닥이나 가구 면에 지지되어 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 0.651,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.401,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 프롬프트에 명시되지 않은 인물 추가 (전경 오른쪽)",
     "[gpt-high] 화면 오른쪽 전경에 이강준 외의 두 번째 남성을 추가하여, 지정 인물 한 명만 보여야 한다는 조건을 위반했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "A": 401,
   "B": 2000
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 401,
    "verdict_ko": "프롬프트에 언급되지 않은 인물이 전경에 추가되어 결정적인 위반 사항이 발생했습니다.  ★위반: [gemini-pro] 프롬프트에 명시되지 않은 인물 추가 (전경 오른쪽) / [gpt-high] 화면 오른쪽 전경에 이강준 외의 두 번째 남성을 추가하여, 지정 인물 한 명만 보여야 한다는 조건을 위반했다."
   },
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "미디엄 샷 구도 내에서 허공을 향한 주먹, 핏대 선 목, 벌린 입 등 명시된 인물의 행동과 표정을 정확하게 재현했습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_special_unit_office_2f9882.png",
    "asset_id": "4964e681-31dd-4c79-aad2-e3da5f4b52d2",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 이강준: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1327185>",
    "asset_id": "6dbc9554-4775-4399-b597-16c2aa160d07",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92d-07ce-7d5f-9681-2a9440c6bc37",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S87sh4__bgfirst_bg.png",
   "bg_asset_id": "4e972bfb-a762-4f2f-9abf-8faa5c7623ed",
   "bg_record_key": "S87sh4::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "special_unit_office",
   "groupbg_asset_id": "4964e681-31dd-4c79-aad2-e3da5f4b52d2"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S87sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T13:33:38.921272+00:00",
  "fingerprint": "6af1100358607380c68f50dab733b9cd28acde35d816c6110a5bd59093453ee3",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S87sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S87sh4_sel.png",
  "source_sha256": "2b4ad413de014687fa5dd1fdffcd9b712b6877b8fdc6b6c041cb0c101273f908",
  "file": "S87sh4_cine.png",
  "staged_sha256": "8f357df37bfb071a99c61ad02da5674b5a3b62f8b857eed14a93a51653f8efec",
  "latency_ms": 9562
 },
 "S87sh5::signage": {
  "fp": "9ee41674796092ed",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "groupbg::corporate_chair_office": {
  "input_fingerprint": "748a8191ab997abe",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "corporate_chair_office",
    "tags": [
     "S87sh5"
    ]
   },
   "context_sig": "98a913aacd2b6fe2"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: Inside a corporate chairman's office, at the communications-monitoring position under nighttime office lighting.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n특임대 이강준의 사무실: CCTV 분석 모니터와 보고서가 결재되는 딱딱한 분위기의 군 지휘관 방. (특징: 정돈된 군용 지휘 데스크; 종이 문서가 끼워진 결재판; 드론이나 해안 CCTV 영상이 띄워진 다중 모니터 화면)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- /유빅사 윤성찬 회장실 –N\n\nTIME OF DAY (lock): night.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: Inside a corporate chairman's office, at the communications-monitoring position under nighttime office lighting.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n특임대 이강준의 사무실: CCTV 분석 모니터와 보고서가 결재되는 딱딱한 분위기의 군 지휘관 방. (특징: 정돈된 군용 지휘 데스크; 종이 문서가 끼워진 결재판; 드론이나 해안 CCTV 영상이 띄워진 다중 모니터 화면)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- /유빅사 윤성찬 회장실 –N\n\nTIME OF DAY (lock): night.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_corporate_chair_office_dd85e5.png",
  "asset_id": "0bb05930-fe04-4dc3-b470-8ae77659d308",
  "input_asset_ids": [
   "f85ce6f3-3160-4f21-847d-73d55d234f47"
  ],
  "origin_tag": "S87sh5",
  "place_text": "Inside a corporate chairman's office, at the communications-monitoring position under nighttime office lighting.",
  "origin_inputs": {
   "place_text": "Inside a corporate chairman's office, at the communications-monitoring position under nighttime office lighting.",
   "time_of_day_en": "night",
   "conti_asset_id": "f85ce6f3-3160-4f21-847d-73d55d234f47"
  }
 },
 "S87sh5::bgfirst_bg": {
  "input_fingerprint": "3a302fccdeace35b",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 유빅사 사무실, 귀에 꽂힌 통신용 이어폰에 손을 댄 채 비열하게 입꼬리를 한껏 올린 윤성찬의 얼굴.\n\nLOCATION (lock): Inside a corporate chairman's office, at the communications-monitoring position under nighttime office lighting.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Communication earpiece (Inserted in his ear and touched by his fingers) — Its exposed outer portion is visible beside the near ear; used as Small narrative detail linking his expression to the intercepted command.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient illumination and precise facial contrast keep the hand, earpiece, and calculating smile legible without a device-generated glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 유빅사 사무실, 귀에 꽂힌 통신용 이어폰에 손을 댄 채 비열하게 입꼬리를 한껏 올린 윤성찬의 얼굴.\n\nLOCATION (lock): Inside a corporate chairman's office, at the communications-monitoring position under nighttime office lighting.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Communication earpiece (Inserted in his ear and touched by his fingers) — Its exposed outer portion is visible beside the near ear; used as Small narrative detail linking his expression to the intercepted command.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient illumination and precise facial contrast keep the hand, earpiece, and calculating smile legible without a device-generated glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S87sh5__bgfirst_bg.png",
  "asset_id": "ea1e3ce7-0a30-49e0-a731-7053a0115a46",
  "input_asset_ids": [
   "f85ce6f3-3160-4f21-847d-73d55d234f47",
   "0bb05930-fe04-4dc3-b470-8ae77659d308"
  ]
 },
 "S87sh5": {
  "input_fingerprint": "6d191bcb36c91edc",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 유빅사 사무실, 귀에 꽂힌 통신용 이어폰에 손을 댄 채 비열하게 입꼬리를 한껏 올린 윤성찬의 얼굴.\n\nLOCATION (lock): Inside a corporate chairman's office, at the communications-monitoring position under nighttime office lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Communication earpiece (Inserted in his ear and touched by his fingers) — Its exposed outer portion is visible beside the near ear; used as Small narrative detail linking his expression to the intercepted command.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient illumination and precise facial contrast keep the hand, earpiece, and calculating smile legible without a device-generated glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 윤성찬 right now, so 윤성찬's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 윤성찬: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 유빅사 사무실, 귀에 꽂힌 통신용 이어폰에 손을 댄 채 비열하게 입꼬리를 한껏 올린 윤성찬의 얼굴.\n\nLOCATION (lock): Inside a corporate chairman's office, at the communications-monitoring position under nighttime office lighting. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Communication earpiece (Inserted in his ear and touched by his fingers) — Its exposed outer portion is visible beside the near ear; used as Small narrative detail linking his expression to the intercepted command.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient illumination and precise facial contrast keep the hand, earpiece, and calculating smile legible without a device-generated glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 윤성찬 right now, so 윤성찬's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 윤성찬: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 유빅사 사무실, 귀에 꽂힌 통신용 이어폰에 손을 댄 채 비열하게 입꼬리를 한껏 올린 윤성찬의 얼굴.\n\nLOCATION (lock): Inside a corporate chairman's office, at the communications-monitoring position under nighttime office lighting. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Communication earpiece (Inserted in his ear and touched by his fingers) — Its exposed outer portion is visible beside the near ear; used as Small narrative detail linking his expression to the intercepted command.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained ambient illumination and precise facial contrast keep the hand, earpiece, and calculating smile legible without a device-generated glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 윤성찬 right now, so 윤성찬's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 윤성찬: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S87sh5__bgfirst_bg.png",
     "asset_id": "ea1e3ce7-0a30-49e0-a731-7053a0115a46",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S87sh5.png",
     "asset_id": "f85ce6f3-3160-4f21-847d-73d55d234f47",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 윤성찬: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1453233>",
     "asset_id": "04d34665-3829-49fb-a6d2-25b9fd051d63",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_corporate_chair_office_dd85e5.png",
     "asset_id": "0bb05930-fe04-4dc3-b470-8ae77659d308",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 윤성찬: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1453233>",
     "asset_id": "04d34665-3829-49fb-a6d2-25b9fd051d63",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "시선은 화면 우측 밖을 향하고, 오른손은 귀 부근으로 올라가 있음.",
    "built_space": "참조된 야간 회장실의 책상, 모니터, 창밖 야경이 적절한 원근감으로 배치됨.",
    "entities": "인물의 노화 특징과 의상은 일치하나 표정이 온화함. 얇은 선 형태의 이어폰이 묘사됨.",
    "hard_violations": [
     "[gpt-high] 윤성찬이 지정된 통신 감청 위치가 아니라 앞쪽 회장 책상 앞에 배치되어 있으며, 실제 감청 자리와 빈 의자는 뒤쪽에 별도로 보인다."
    ],
    "physics": "손가락이 귀 부근의 선을 가볍게 쥔 채 허공에 자연스럽게 지지되어 있음."
   },
   {
    "label": "B",
    "direction": "시선은 화면 좌측 밖을 향하고, 오른손 검지가 귀에 꽂힌 기기에 닿아 있음.",
    "built_space": "야간 회장실의 주요 집기와 조명 배치가 참조 이미지와 정확히 호응함.",
    "entities": "인물의 얼굴 특징이 완벽히 일치하며 비열하게 웃는 표정이 명확함. 뚜렷한 형태의 이어폰이 확인됨.",
    "hard_violations": [
     "[gpt-high] 윤성찬이 지정된 통신 감청 위치가 아니라 앞쪽 회장 책상 앞에 배치되어 있으며, 실제 감청 자리와 빈 의자는 뒤쪽에 별도로 보인다."
    ],
    "physics": "손가락이 이어폰 표면을 짚고 있는 물리적 접촉과 무게 중심이 자연스러움."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지문이 강하게 요구한 '비열하게 입꼬리를 올린' 표정과 이어폰을 조작하는 동작을 가장 충실하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "인물 외형과 배경은 잘 일치하나, 표정이 지나치게 온화하여 텍스트의 핵심인 비열한 분위기를 살리지 못했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 화면 우측 밖을 향하고, 오른손은 귀 부근으로 올라가 있음.",
        "built_space": "참조된 야간 회장실의 책상, 모니터, 창밖 야경이 적절한 원근감으로 배치됨.",
        "entities": "인물의 노화 특징과 의상은 일치하나 표정이 온화함. 얇은 선 형태의 이어폰이 묘사됨.",
        "hard_violations": [],
        "physics": "손가락이 귀 부근의 선을 가볍게 쥔 채 허공에 자연스럽게 지지되어 있음."
       },
       {
        "label": "B",
        "direction": "시선은 화면 좌측 밖을 향하고, 오른손 검지가 귀에 꽂힌 기기에 닿아 있음.",
        "built_space": "야간 회장실의 주요 집기와 조명 배치가 참조 이미지와 정확히 호응함.",
        "entities": "인물의 얼굴 특징이 완벽히 일치하며 비열하게 웃는 표정이 명확함. 뚜렷한 형태의 이어폰이 확인됨.",
        "hard_violations": [],
        "physics": "손가락이 이어폰 표면을 짚고 있는 물리적 접촉과 무게 중심이 자연스러움."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지문이 강하게 요구한 '비열하게 입꼬리를 올린' 표정과 이어폰을 조작하는 동작을 가장 충실하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "인물 외형과 배경은 잘 일치하나, 표정이 지나치게 온화하여 텍스트의 핵심인 비열한 분위기를 살리지 못했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 화면 우측 밖을 향하고, 오른손은 귀 부근으로 올라가 있음.",
        "built_space": "참조된 야간 회장실의 책상, 모니터, 창밖 야경이 적절한 원근감으로 배치됨.",
        "entities": "인물의 노화 특징과 의상은 일치하나 표정이 온화함. 얇은 선 형태의 이어폰이 묘사됨.",
        "hard_violations": [],
        "physics": "손가락이 귀 부근의 선을 가볍게 쥔 채 허공에 자연스럽게 지지되어 있음."
       },
       {
        "label": "B",
        "direction": "시선은 화면 좌측 밖을 향하고, 오른손 검지가 귀에 꽂힌 기기에 닿아 있음.",
        "built_space": "야간 회장실의 주요 집기와 조명 배치가 참조 이미지와 정확히 호응함.",
        "entities": "인물의 얼굴 특징이 완벽히 일치하며 비열하게 웃는 표정이 명확함. 뚜렷한 형태의 이어폰이 확인됨.",
        "hard_violations": [],
        "physics": "손가락이 이어폰 표면을 짚고 있는 물리적 접촉과 무게 중심이 자연스러움."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "두 후보 모두 통신 감청 자리에서 벗어나 있지만, A가 얼굴을 더 밀착해 잡고 입꼬리를 크게 올린 비열한 웃음과 이어폰 접촉을 더 분명하게 보여준다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "이어폰을 만지는 동작은 맞지만 감청 자리가 인물 뒤에 비어 있고, 표정도 요구된 한껏 치켜올린 비열한 웃음보다 얌전한 미소에 가깝다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "얼굴은 화면 오른쪽으로 조금 돌아가고 시선은 카메라 부근을 향한다. 지정된 시선 대상은 없다. 화면 왼쪽의 검지가 가까운 귀에 장착된 이어폰 윗부분에 닿아 있어 조작 대상이 명확하다.",
        "built_space": "왼쪽 통창과 야경, 오른쪽 대리석 벽과 깃발 1개, 천장 간접조명이 장소 사진과 대응한다. 앞쪽 회장 책상과 뒤쪽 감청 책상, 각각의 의자 1개씩이 보인다. 감청 모니터는 인물에 가려진 부분을 제외하고 2개가 식별되며, 탁상등도 앞뒤로 1개씩 보인다. 그러나 윤성찬은 앞쪽 책상보다도 카메라 가까이에 있고, 지정된 뒤쪽 통신 감청 자리와 그 의자는 비어 있다. 창의 선형 조명 반사에 명백한 광학적 모순은 없다.",
        "entities": "보이는 사람은 고령의 동아시아계 남성 1명뿐이다. 회색으로 빗어 넘긴 머리, 깊은 이마와 눈가 주름, 검은 재킷과 짙은 줄무늬 셔츠는 윤성찬 참고와 대체로 대응한다. 참고의 안경은 없고 얼굴 윤곽도 완전히 같지는 않다. 손의 주름과 피부 나이는 얼굴과 어울린다. 귀에는 외부 몸체가 드러난 통신 이어폰이 있으며, 치아를 살짝 드러내고 입꼬리를 올린 웃음이 보인다. 읽을 수 있는 글자는 식별되지 않는다.",
        "hard_violations": [
         "윤성찬이 지정된 통신 감청 위치가 아니라 앞쪽 회장 책상 앞에 배치되어 있으며, 실제 감청 자리와 빈 의자는 뒤쪽에 별도로 보인다."
        ],
        "physics": "이어폰은 귀에 걸리고 삽입된 부분으로 지지되며 검지가 그 외측에 접촉한다. 손목과 소매가 이어져 자기 팔을 들어 귀를 만지는 동작으로 성립한다. 하체와 바닥 접촉은 클로즈업 밖이므로 판단할 수 없지만, 공중에 뜬 몸이나 지지 없는 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "얼굴과 두 눈은 화면 오른쪽 위의 프레임 밖을 향하며, 보이는 특정 물체를 응시하지는 않는다. 시선 대상 자체는 지시되지 않았다. 화면 왼쪽 검지는 가까운 귀의 작은 이어폰 외측에 정확히 닿는다.",
        "built_space": "통창, 야간 도시, 천장 간접조명, 오른쪽 대리석 벽과 깃발 1개가 참고 장소와 대응한다. 감청 책상 위에는 모니터 3개가 나란히 있고 그 앞 의자 1개가 비어 있다. 앞쪽 회장 책상과 의자 1개, 앞뒤 탁상등 각 1개도 보인다. 윤성찬은 이 감청 설비에서 떨어진 전경에 배치되어 지정된 감청 위치를 차지하지 않는다. 창 반사에 명백히 불가능한 요소는 없다.",
        "entities": "고령의 동아시아계 남성 1명만 등장한다. 회색 머리, 깊은 주름, 검은 질감의 재킷과 짙은 줄무늬 셔츠는 참고와 유사하지만 참고의 안경은 빠져 있다. 얼굴과 귀를 만지는 손 모두 노년의 피부로 표현된다. 귀에 꽂힌 작은 통신 이어폰과 아래로 이어지는 가는 선이 보인다. 입은 다문 채 살짝 웃고 있어 비열하게 입꼬리를 한껏 올린 표정은 약하다. 읽을 수 있는 문구는 식별되지 않는다.",
        "hard_violations": [
         "윤성찬이 지정된 통신 감청 위치가 아니라 앞쪽 회장 책상 앞에 배치되어 있으며, 실제 감청 자리와 빈 의자는 뒤쪽에 별도로 보인다."
        ],
        "physics": "이어폰은 귀에 고정되어 있고 검지가 외부 부품을 누른다. 손과 손목, 소매의 연결 및 팔을 들어 올린 자세가 물리적으로 가능하다. 하체는 프레임 밖이어서 지지 자세를 확인할 수 없으나, 보이는 신체나 소품에 부유 또는 불가능한 접촉은 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "두 후보 모두 통신 감청 자리에서 벗어나 있지만, A가 얼굴을 더 밀착해 잡고 입꼬리를 크게 올린 비열한 웃음과 이어폰 접촉을 더 분명하게 보여준다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "이어폰을 만지는 동작은 맞지만 감청 자리가 인물 뒤에 비어 있고, 표정도 요구된 한껏 치켜올린 비열한 웃음보다 얌전한 미소에 가깝다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "얼굴은 화면 오른쪽으로 조금 돌아가고 시선은 카메라 부근을 향한다. 지정된 시선 대상은 없다. 화면 왼쪽의 검지가 가까운 귀에 장착된 이어폰 윗부분에 닿아 있어 조작 대상이 명확하다.",
        "built_space": "왼쪽 통창과 야경, 오른쪽 대리석 벽과 깃발 1개, 천장 간접조명이 장소 사진과 대응한다. 앞쪽 회장 책상과 뒤쪽 감청 책상, 각각의 의자 1개씩이 보인다. 감청 모니터는 인물에 가려진 부분을 제외하고 2개가 식별되며, 탁상등도 앞뒤로 1개씩 보인다. 그러나 윤성찬은 앞쪽 책상보다도 카메라 가까이에 있고, 지정된 뒤쪽 통신 감청 자리와 그 의자는 비어 있다. 창의 선형 조명 반사에 명백한 광학적 모순은 없다.",
        "entities": "보이는 사람은 고령의 동아시아계 남성 1명뿐이다. 회색으로 빗어 넘긴 머리, 깊은 이마와 눈가 주름, 검은 재킷과 짙은 줄무늬 셔츠는 윤성찬 참고와 대체로 대응한다. 참고의 안경은 없고 얼굴 윤곽도 완전히 같지는 않다. 손의 주름과 피부 나이는 얼굴과 어울린다. 귀에는 외부 몸체가 드러난 통신 이어폰이 있으며, 치아를 살짝 드러내고 입꼬리를 올린 웃음이 보인다. 읽을 수 있는 글자는 식별되지 않는다.",
        "hard_violations": [
         "윤성찬이 지정된 통신 감청 위치가 아니라 앞쪽 회장 책상 앞에 배치되어 있으며, 실제 감청 자리와 빈 의자는 뒤쪽에 별도로 보인다."
        ],
        "physics": "이어폰은 귀에 걸리고 삽입된 부분으로 지지되며 검지가 그 외측에 접촉한다. 손목과 소매가 이어져 자기 팔을 들어 귀를 만지는 동작으로 성립한다. 하체와 바닥 접촉은 클로즈업 밖이므로 판단할 수 없지만, 공중에 뜬 몸이나 지지 없는 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "얼굴과 두 눈은 화면 오른쪽 위의 프레임 밖을 향하며, 보이는 특정 물체를 응시하지는 않는다. 시선 대상 자체는 지시되지 않았다. 화면 왼쪽 검지는 가까운 귀의 작은 이어폰 외측에 정확히 닿는다.",
        "built_space": "통창, 야간 도시, 천장 간접조명, 오른쪽 대리석 벽과 깃발 1개가 참고 장소와 대응한다. 감청 책상 위에는 모니터 3개가 나란히 있고 그 앞 의자 1개가 비어 있다. 앞쪽 회장 책상과 의자 1개, 앞뒤 탁상등 각 1개도 보인다. 윤성찬은 이 감청 설비에서 떨어진 전경에 배치되어 지정된 감청 위치를 차지하지 않는다. 창 반사에 명백히 불가능한 요소는 없다.",
        "entities": "고령의 동아시아계 남성 1명만 등장한다. 회색 머리, 깊은 주름, 검은 질감의 재킷과 짙은 줄무늬 셔츠는 참고와 유사하지만 참고의 안경은 빠져 있다. 얼굴과 귀를 만지는 손 모두 노년의 피부로 표현된다. 귀에 꽂힌 작은 통신 이어폰과 아래로 이어지는 가는 선이 보인다. 입은 다문 채 살짝 웃고 있어 비열하게 입꼬리를 한껏 올린 표정은 약하다. 읽을 수 있는 문구는 식별되지 않는다.",
        "hard_violations": [
         "윤성찬이 지정된 통신 감청 위치가 아니라 앞쪽 회장 책상 앞에 배치되어 있으며, 실제 감청 자리와 빈 의자는 뒤쪽에 별도로 보인다."
        ],
        "physics": "이어폰은 귀에 고정되어 있고 검지가 외부 부품을 누른다. 손과 손목, 소매의 연결 및 팔을 들어 올린 자세가 물리적으로 가능하다. 하체는 프레임 밖이어서 지지 자세를 확인할 수 없으나, 보이는 신체나 소품에 부유 또는 불가능한 접촉은 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.464,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.214,
    "B": 1.75
   },
   "violations": {
    "B": [
     "[gpt-high] 윤성찬이 지정된 통신 감청 위치가 아니라 앞쪽 회장 책상 앞에 배치되어 있으며, 실제 감청 자리와 빈 의자는 뒤쪽에 별도로 보인다."
    ],
    "A": [
     "[gpt-high] 윤성찬이 지정된 통신 감청 위치가 아니라 앞쪽 회장 책상 앞에 배치되어 있으며, 실제 감청 자리와 빈 의자는 뒤쪽에 별도로 보인다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 1750,
   "A": 1214
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "지문이 강하게 요구한 '비열하게 입꼬리를 올린' 표정과 이어폰을 조작하는 동작을 가장 충실하게 구현했습니다.  ★위반: [gpt-high] 윤성찬이 지정된 통신 감청 위치가 아니라 앞쪽 회장 책상 앞에 배치되어 있으며, 실제 감청 자리와 빈 의자는 뒤쪽에 별도로 보인다."
   },
   {
    "label": "A",
    "score": 1214,
    "verdict_ko": "인물 외형과 배경은 잘 일치하나, 표정이 지나치게 온화하여 텍스트의 핵심인 비열한 분위기를 살리지 못했습니다.  ★위반: [gpt-high] 윤성찬이 지정된 통신 감청 위치가 아니라 앞쪽 회장 책상 앞에 배치되어 있으며, 실제 감청 자리와 빈 의자는 뒤쪽에 별도로 보인다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_corporate_chair_office_dd85e5.png",
    "asset_id": "0bb05930-fe04-4dc3-b470-8ae77659d308",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 윤성찬: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1453233>",
    "asset_id": "04d34665-3829-49fb-a6d2-25b9fd051d63",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92d-0c98-7cb0-8a51-d3cfae6b4768",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S87sh5__bgfirst_bg.png",
   "bg_asset_id": "ea1e3ce7-0a30-49e0-a731-7053a0115a46",
   "bg_record_key": "S87sh5::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "corporate_chair_office",
   "groupbg_asset_id": "0bb05930-fe04-4dc3-b470-8ae77659d308"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S87sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T09:26:12.783989+00:00",
  "fingerprint": "3e6e3d0e1dd64e74947d20b2c32151ad83b5ec387bd506a9ce9809128cccce70",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S87sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S87sh5_sel.png",
  "source_sha256": "0a5cc4e4426dc767880172735e961d25232adb947a6f72f335a1a5d70937b9a0",
  "file": "S87sh5_cine.png",
  "staged_sha256": "0cccd556d8a1d6b0bd39e086e1d87f32699311f2b4d5c1c4e30e7b0e4881fd1e",
  "latency_ms": 10639
 },
 "S88sh7::signage": {
  "fp": "060df420176958a3",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S88sh7": {
  "input_fingerprint": "fad51d7b1100e6fe",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 푸른 조명이 켜진 실험실 안, 차가운 스테인리스 침대 위에 굵은 끈으로 결박당한 채 누워있는 찰리의 낡은 전신.\n\nLOCATION (lock): On the stainless-steel examination bed inside the research laboratory's glass enclosure, under blue laboratory lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Stainless-steel bed, headward end toward upper right in the middle-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Stainless-steel bed (Supporting 찰리's restrained body) — The upper surface, near long edge, and footward end are visible from above; used as Diagonal spatial anchor that exposes the full-body restraint arrangement; Thick restraints (Secured around 찰리, holding him to the bed); used as Readable evidence of confinement across the full-body composition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The specified blue laboratory illumination gives the restrained body a subdued, cold appearance while preserving readable detail in the restraints.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent on the stainless-steel bed with his body restrained against it, including his hands before they are released. The source does not specify the restraint points, his head's direction, or the exact arrangement of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie lies restrained on the stainless-steel laboratory bed; the chest-ring power system remains impaired after the failed test. The laboratory's glass enclosure and monitoring equipment remain in place.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 푸른 조명이 켜진 실험실 안, 차가운 스테인리스 침대 위에 굵은 끈으로 결박당한 채 누워있는 찰리의 낡은 전신.\n\nLOCATION (lock): On the stainless-steel examination bed inside the research laboratory's glass enclosure, under blue laboratory lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Stainless-steel bed, headward end toward upper right in the middle-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Stainless-steel bed (Supporting 찰리's restrained body) — The upper surface, near long edge, and footward end are visible from above; used as Diagonal spatial anchor that exposes the full-body restraint arrangement; Thick restraints (Secured around 찰리, holding him to the bed); used as Readable evidence of confinement across the full-body composition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The specified blue laboratory illumination gives the restrained body a subdued, cold appearance while preserving readable detail in the restraints.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent on the stainless-steel bed with his body restrained against it, including his hands before they are released. The source does not specify the restraint points, his head's direction, or the exact arrangement of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie lies restrained on the stainless-steel laboratory bed; the chest-ring power system remains impaired after the failed test. The laboratory's glass enclosure and monitoring equipment remain in place.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 푸른 조명이 켜진 실험실 안, 차가운 스테인리스 침대 위에 굵은 끈으로 결박당한 채 누워있는 찰리의 낡은 전신.\n\nLOCATION (lock): On the stainless-steel examination bed inside the research laboratory's glass enclosure, under blue laboratory lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Stainless-steel bed, headward end toward upper right in the middle-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Stainless-steel bed (Supporting 찰리's restrained body) — The upper surface, near long edge, and footward end are visible from above; used as Diagonal spatial anchor that exposes the full-body restraint arrangement; Thick restraints (Secured around 찰리, holding him to the bed); used as Readable evidence of confinement across the full-body composition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The specified blue laboratory illumination gives the restrained body a subdued, cold appearance while preserving readable detail in the restraints.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent on the stainless-steel bed with his body restrained against it, including his hands before they are released. The source does not specify the restraint points, his head's direction, or the exact arrangement of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie lies restrained on the stainless-steel laboratory bed; the chest-ring power system remains impaired after the failed test. The laboratory's glass enclosure and monitoring equipment remain in place.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "찰리의 머리는 좌측 상단을 향하며 천장을 바라보고 누워 있음.",
    "built_space": "유리벽으로 둘러싸인 원형 연구실 내부에 로봇 팔 3대와 모니터를 보는 연구원들이 참조 이미지와 동일하게 배치됨.",
    "entities": "고릴라형 로봇 몸체의 찰리가 스테인리스 침대에 두꺼운 끈으로 결박되어 있으나, 프롬프트에 명시되지 않은 다수의 연구원들이 배경에 존재함.",
    "hard_violations": [
     "[gemini-pro] 프롬프트에 없는 인물 등장 (연구원들)",
     "[gemini-pro] 참조 이미지의 카메라 구도 복사 금지 위반",
     "[gemini-pro] 지정된 프레이밍 방향 (머리가 우측 상단) 위반",
     "[gpt-high] 찰리만 등장해야 하는 장면에 이전 장면의 연구원 10명을 그대로 포함했다."
    ],
    "physics": "찰리의 몸은 스테인리스 침대 표면에 밀착되어 중력에 맞게 완전히 지지되고 있음."
   },
   {
    "label": "B",
    "direction": "찰리의 머리는 우측 중앙을 향하며 천장을 바라보고 누워 있음.",
    "built_space": "원형 연구실 내부에 연구원들이 참조 이미지와 동일한 위치에 있으나, 침대의 방향이 방을 기준으로 180도 뒤집혀 있고 로봇 팔이 사라짐.",
    "entities": "찰리가 침대에 결박되어 있으나, 프롬프트에 없는 연구원들이 여전히 배경에 존재함.",
    "hard_violations": [
     "[gemini-pro] 프롬프트에 없는 인물 등장 (연구원들)",
     "[gemini-pro] 이전 샷의 고정된 배경 요소 누락 (로봇 팔)",
     "[gemini-pro] 카메라 시점은 고정된 채 비가동 피사체(침대와 찰리)만 회전한 불가능한 무대 연출",
     "[gpt-high] 찰리만 등장해야 하는 장면에 이전 장면의 연구원 10명을 그대로 포함했다."
    ],
    "physics": "찰리의 몸은 침대 위에 자연스럽게 눕혀져 안정적으로 지지됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지시된 프레이밍 방향(머리가 우측 상단)을 따르지 않고 참조 이미지의 구도를 그대로 복사했으며, 프롬프트에 없는 연구원들을 포함하여 큰 감점을 받았습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "프레이밍 방향을 맞추려 시도했으나, 배경과 카메라는 그대로인 상태에서 침대만 180도 회전시켜 물리적 연속성을 심각하게 파괴했고 필수 배경 요소인 로봇 팔을 누락했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 머리는 좌측 상단을 향하며 천장을 바라보고 누워 있음.",
        "built_space": "유리벽으로 둘러싸인 원형 연구실 내부에 로봇 팔 3대와 모니터를 보는 연구원들이 참조 이미지와 동일하게 배치됨.",
        "entities": "고릴라형 로봇 몸체의 찰리가 스테인리스 침대에 두꺼운 끈으로 결박되어 있으나, 프롬프트에 명시되지 않은 다수의 연구원들이 배경에 존재함.",
        "hard_violations": [
         "프롬프트에 없는 인물 등장 (연구원들)",
         "참조 이미지의 카메라 구도 복사 금지 위반",
         "지정된 프레이밍 방향 (머리가 우측 상단) 위반"
        ],
        "physics": "찰리의 몸은 스테인리스 침대 표면에 밀착되어 중력에 맞게 완전히 지지되고 있음."
       },
       {
        "label": "B",
        "direction": "찰리의 머리는 우측 중앙을 향하며 천장을 바라보고 누워 있음.",
        "built_space": "원형 연구실 내부에 연구원들이 참조 이미지와 동일한 위치에 있으나, 침대의 방향이 방을 기준으로 180도 뒤집혀 있고 로봇 팔이 사라짐.",
        "entities": "찰리가 침대에 결박되어 있으나, 프롬프트에 없는 연구원들이 여전히 배경에 존재함.",
        "hard_violations": [
         "프롬프트에 없는 인물 등장 (연구원들)",
         "이전 샷의 고정된 배경 요소 누락 (로봇 팔)",
         "카메라 시점은 고정된 채 비가동 피사체(침대와 찰리)만 회전한 불가능한 무대 연출"
        ],
        "physics": "찰리의 몸은 침대 위에 자연스럽게 눕혀져 안정적으로 지지됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "지시된 프레이밍 방향(머리가 우측 상단)을 따르지 않고 참조 이미지의 구도를 그대로 복사했으며, 프롬프트에 없는 연구원들을 포함하여 큰 감점을 받았습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "프레이밍 방향을 맞추려 시도했으나, 배경과 카메라는 그대로인 상태에서 침대만 180도 회전시켜 물리적 연속성을 심각하게 파괴했고 필수 배경 요소인 로봇 팔을 누락했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 머리는 좌측 상단을 향하며 천장을 바라보고 누워 있음.",
        "built_space": "유리벽으로 둘러싸인 원형 연구실 내부에 로봇 팔 3대와 모니터를 보는 연구원들이 참조 이미지와 동일하게 배치됨.",
        "entities": "고릴라형 로봇 몸체의 찰리가 스테인리스 침대에 두꺼운 끈으로 결박되어 있으나, 프롬프트에 명시되지 않은 다수의 연구원들이 배경에 존재함.",
        "hard_violations": [
         "프롬프트에 없는 인물 등장 (연구원들)",
         "참조 이미지의 카메라 구도 복사 금지 위반",
         "지정된 프레이밍 방향 (머리가 우측 상단) 위반"
        ],
        "physics": "찰리의 몸은 스테인리스 침대 표면에 밀착되어 중력에 맞게 완전히 지지되고 있음."
       },
       {
        "label": "B",
        "direction": "찰리의 머리는 우측 중앙을 향하며 천장을 바라보고 누워 있음.",
        "built_space": "원형 연구실 내부에 연구원들이 참조 이미지와 동일한 위치에 있으나, 침대의 방향이 방을 기준으로 180도 뒤집혀 있고 로봇 팔이 사라짐.",
        "entities": "찰리가 침대에 결박되어 있으나, 프롬프트에 없는 연구원들이 여전히 배경에 존재함.",
        "hard_violations": [
         "프롬프트에 없는 인물 등장 (연구원들)",
         "이전 샷의 고정된 배경 요소 누락 (로봇 팔)",
         "카메라 시점은 고정된 채 비가동 피사체(침대와 찰리)만 회전한 불가능한 무대 연출"
        ],
        "physics": "찰리의 몸은 침대 위에 자연스럽게 눕혀져 안정적으로 지지됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "머리가 오른쪽 위를 향하고 결박은 선명하지만, 금지된 연구원 10명이 남아 있으며 침대를 전경에 과대하게 배치하고 기존 로봇 팔 3대를 없애 장소와 중경 와이드 구도를 훼손했다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "금지된 연구원 10명이 남아 있어 실격이지만, 중경의 전신 와이드 구도와 침대·유리실·로봇 팔 3대의 공간 관계는 더 충실하다. 다만 머리 쪽이 오른쪽 위가 아니라 왼쪽 위다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 머리는 화면 오른쪽 위, 발은 왼쪽 아래를 향해 지정된 침대 대각선 방향과 맞는다. 마스크 얼굴은 천장을 향하며 특정 대상을 응시하지 않는다. 주변 연구원들은 모니터나 손에 든 자료를 보고 있고 전경 연구원도 화면 쪽을 향한다.",
        "built_space": "원형 유리 격실 1개, 중앙 침대 1개, 둘레 작업대와 여러 모니터가 보인다. 참조의 침대 주변 로봇 팔 3대는 모두 보이지 않는다. 침대는 중경 중앙의 설비라기보다 화면 아래 대부분을 차지하는 거대한 전경 요소로 바뀌었다. 상판과 가까운 긴 변, 발치 끝은 잘 보인다. 참조의 유리문과 원형 바닥 배수로는 유지되지만 침대와 주변 공간의 크기 관계가 크게 달라졌다.",
        "entities": "찰리 1체는 마모된 샌드 베이지 장갑, 육중한 어깨와 팔, 흰색 기계식 마스크를 갖춰 참조의 주요 정체성을 따른다. 스테인리스 침대와 가슴·팔목·발목 부근의 굵은 결박이 보이며 손도 침대에 놓여 있다. 흉부 전원계의 손상 상태는 명확하게 식별되지 않는다. 찰리 외에 둘레 작업대의 연구원 9명과 오른쪽 아래 연구원 1명이 그대로 남아 있다. 푸른 실험실 조명은 구현됐지만 외부가 보이지 않아 아침 여부는 판별할 수 없다.",
        "hard_violations": [
         "찰리만 등장해야 하는 장면에 이전 장면의 연구원 10명을 그대로 포함했다."
        ],
        "physics": "찰리의 몸통과 머리, 팔과 손, 다리는 침대 상판에 놓이고 발뒤꿈치도 상판에 지지된다. 굵은 끈은 몸을 가로질러 침대 가장자리로 이어져 결박으로 읽힌다. 침대는 바닥에 닿는 금속 받침으로 지지된다. 주변 사람들도 의자나 바닥에 지지되어 있으며 명백하게 떠 있는 신체나 물체는 없다."
       },
       {
        "label": "B",
        "direction": "찰리의 머리는 화면 왼쪽 위, 발은 오른쪽 아래를 향한다. 따라서 머리 쪽을 오른쪽 위에 두라는 지시와 반대다. 얼굴은 위쪽을 향하고 특정 표적을 보지 않는다. 로봇 팔 끝은 침대 주변을 향하며 머리 쪽 팔의 도구는 아래로 내려와 있다. 연구원들은 각자의 모니터나 자료를 바라본다.",
        "built_space": "원형 유리 격실 1개와 중앙의 스테인리스 침대 1개가 보이며, 로봇 팔은 참조처럼 머리 뒤쪽·왼쪽 앞·오른쪽에 총 3대가 배치되어 있다. 원형 바닥 배수로, 유리문, 둘레 작업대와 모니터들도 유지된다. 침대는 중경 중앙에 실제적인 크기로 놓이고 상판, 가까운 긴 변, 발치 끝이 내려다보인다. 다만 카메라 구도를 거의 그대로 답습해 요구된 침대 방향으로 전환하지 못했다.",
        "entities": "찰리 1체의 낡은 베이지 장갑과 흰 마스크, 굵은 팔다리는 이전 장면의 외형에 대체로 맞는다. 몸통과 허벅지, 정강이를 가로지르는 넓은 결박이 보이지만 양손을 각각 고정하는 방식은 A보다 덜 명확하다. 흉부 전원계의 손상은 이 크기에서 확정하기 어렵다. 유리 격실과 모니터 설비는 남아 있으나, 제외해야 할 연구원도 둘레에 9명, 오른쪽 아래에 1명 보인다. 푸른 조명은 맞으며 아침을 판별할 외부 단서는 없다.",
        "hard_violations": [
         "찰리만 등장해야 하는 장면에 이전 장면의 연구원 10명을 그대로 포함했다."
        ],
        "physics": "찰리는 등과 머리를 침대에 기대고 팔과 손을 몸 옆 상판에 내려놓고 있다. 다리와 발뒤꿈치 역시 침대가 받친다. 결박은 신체를 가로질러 침대 측면으로 이어진다. 침대와 로봇 팔은 각각 바닥의 받침대에 지지되고, 연구원들은 의자에 앉거나 바닥에 서 있다. 명백하게 지지 없이 떠 있는 물체나 신체는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "머리가 오른쪽 위를 향하고 결박은 선명하지만, 금지된 연구원 10명이 남아 있으며 침대를 전경에 과대하게 배치하고 기존 로봇 팔 3대를 없애 장소와 중경 와이드 구도를 훼손했다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "금지된 연구원 10명이 남아 있어 실격이지만, 중경의 전신 와이드 구도와 침대·유리실·로봇 팔 3대의 공간 관계는 더 충실하다. 다만 머리 쪽이 오른쪽 위가 아니라 왼쪽 위다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 머리는 화면 오른쪽 위, 발은 왼쪽 아래를 향해 지정된 침대 대각선 방향과 맞는다. 마스크 얼굴은 천장을 향하며 특정 대상을 응시하지 않는다. 주변 연구원들은 모니터나 손에 든 자료를 보고 있고 전경 연구원도 화면 쪽을 향한다.",
        "built_space": "원형 유리 격실 1개, 중앙 침대 1개, 둘레 작업대와 여러 모니터가 보인다. 참조의 침대 주변 로봇 팔 3대는 모두 보이지 않는다. 침대는 중경 중앙의 설비라기보다 화면 아래 대부분을 차지하는 거대한 전경 요소로 바뀌었다. 상판과 가까운 긴 변, 발치 끝은 잘 보인다. 참조의 유리문과 원형 바닥 배수로는 유지되지만 침대와 주변 공간의 크기 관계가 크게 달라졌다.",
        "entities": "찰리 1체는 마모된 샌드 베이지 장갑, 육중한 어깨와 팔, 흰색 기계식 마스크를 갖춰 참조의 주요 정체성을 따른다. 스테인리스 침대와 가슴·팔목·발목 부근의 굵은 결박이 보이며 손도 침대에 놓여 있다. 흉부 전원계의 손상 상태는 명확하게 식별되지 않는다. 찰리 외에 둘레 작업대의 연구원 9명과 오른쪽 아래 연구원 1명이 그대로 남아 있다. 푸른 실험실 조명은 구현됐지만 외부가 보이지 않아 아침 여부는 판별할 수 없다.",
        "hard_violations": [
         "찰리만 등장해야 하는 장면에 이전 장면의 연구원 10명을 그대로 포함했다."
        ],
        "physics": "찰리의 몸통과 머리, 팔과 손, 다리는 침대 상판에 놓이고 발뒤꿈치도 상판에 지지된다. 굵은 끈은 몸을 가로질러 침대 가장자리로 이어져 결박으로 읽힌다. 침대는 바닥에 닿는 금속 받침으로 지지된다. 주변 사람들도 의자나 바닥에 지지되어 있으며 명백하게 떠 있는 신체나 물체는 없다."
       },
       {
        "label": "A",
        "direction": "찰리의 머리는 화면 왼쪽 위, 발은 오른쪽 아래를 향한다. 따라서 머리 쪽을 오른쪽 위에 두라는 지시와 반대다. 얼굴은 위쪽을 향하고 특정 표적을 보지 않는다. 로봇 팔 끝은 침대 주변을 향하며 머리 쪽 팔의 도구는 아래로 내려와 있다. 연구원들은 각자의 모니터나 자료를 바라본다.",
        "built_space": "원형 유리 격실 1개와 중앙의 스테인리스 침대 1개가 보이며, 로봇 팔은 참조처럼 머리 뒤쪽·왼쪽 앞·오른쪽에 총 3대가 배치되어 있다. 원형 바닥 배수로, 유리문, 둘레 작업대와 모니터들도 유지된다. 침대는 중경 중앙에 실제적인 크기로 놓이고 상판, 가까운 긴 변, 발치 끝이 내려다보인다. 다만 카메라 구도를 거의 그대로 답습해 요구된 침대 방향으로 전환하지 못했다.",
        "entities": "찰리 1체의 낡은 베이지 장갑과 흰 마스크, 굵은 팔다리는 이전 장면의 외형에 대체로 맞는다. 몸통과 허벅지, 정강이를 가로지르는 넓은 결박이 보이지만 양손을 각각 고정하는 방식은 A보다 덜 명확하다. 흉부 전원계의 손상은 이 크기에서 확정하기 어렵다. 유리 격실과 모니터 설비는 남아 있으나, 제외해야 할 연구원도 둘레에 9명, 오른쪽 아래에 1명 보인다. 푸른 조명은 맞으며 아침을 판별할 외부 단서는 없다.",
        "hard_violations": [
         "찰리만 등장해야 하는 장면에 이전 장면의 연구원 10명을 그대로 포함했다."
        ],
        "physics": "찰리는 등과 머리를 침대에 기대고 팔과 손을 몸 옆 상판에 내려놓고 있다. 다리와 발뒤꿈치 역시 침대가 받친다. 결박은 신체를 가로질러 침대 측면으로 이어진다. 침대와 로봇 팔은 각각 바닥의 받침대에 지지되고, 연구원들은 의자에 앉거나 바닥에 서 있다. 명백하게 지지 없이 떠 있는 물체나 신체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.333
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.083
   },
   "violations": {
    "A": [
     "[gemini-pro] 프롬프트에 없는 인물 등장 (연구원들)",
     "[gemini-pro] 참조 이미지의 카메라 구도 복사 금지 위반",
     "[gemini-pro] 지정된 프레이밍 방향 (머리가 우측 상단) 위반",
     "[gpt-high] 찰리만 등장해야 하는 장면에 이전 장면의 연구원 10명을 그대로 포함했다."
    ],
    "B": [
     "[gemini-pro] 프롬프트에 없는 인물 등장 (연구원들)",
     "[gemini-pro] 이전 샷의 고정된 배경 요소 누락 (로봇 팔)",
     "[gemini-pro] 카메라 시점은 고정된 채 비가동 피사체(침대와 찰리)만 회전한 불가능한 무대 연출",
     "[gpt-high] 찰리만 등장해야 하는 장면에 이전 장면의 연구원 10명을 그대로 포함했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 1083
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "지시된 프레이밍 방향(머리가 우측 상단)을 따르지 않고 참조 이미지의 구도를 그대로 복사했으며, 프롬프트에 없는 연구원들을 포함하여 큰 감점을 받았습니다.  ★위반: [gemini-pro] 프롬프트에 없는 인물 등장 (연구원들) / [gemini-pro] 참조 이미지의 카메라 구도 복사 금지 위반 / [gemini-pro] 지정된 프레이밍 방향 (머리가 우측 상단) 위반 / [gpt-high] 찰리만 등장해야 하는 장면에 이전 장면의 연구원 10명을 그대로 포함했다."
   },
   {
    "label": "B",
    "score": 1083,
    "verdict_ko": "프레이밍 방향을 맞추려 시도했으나, 배경과 카메라는 그대로인 상태에서 침대만 180도 회전시켜 물리적 연속성을 심각하게 파괴했고 필수 배경 요소인 로봇 팔을 누락했습니다.  ★위반: [gemini-pro] 프롬프트에 없는 인물 등장 (연구원들) / [gemini-pro] 이전 샷의 고정된 배경 요소 누락 (로봇 팔) / [gemini-pro] 카메라 시점은 고정된 채 비가동 피사체(침대와 찰리)만 회전한 불가능한 무대 연출 / [gpt-high] 찰리만 등장해야 하는 장면에 이전 장면의 연구원 10명을 그대로 포함했다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S82sh14_sel.png",
    "asset_id": "73257111-705a-49ad-9ee9-a8e062fa4e86",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92d-117e-763e-bb57-c9c34dfadcbe",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S82sh14"
  }
 },
 "S88sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T12:14:46.009892+00:00",
  "fingerprint": "c1209631d4c224358be2e59d030fad1d187034714947141e50e21bd9eac01756",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S88sh7_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S88sh7_sel.png",
  "source_sha256": "7d080a0d4f6e8fe08484f88103c0d363f9d7478cbcf9a2fd6a2e0a5c236a9a73",
  "file": "S88sh7_cine.png",
  "staged_sha256": "aff566e2d8800a79dcb7e95e696aa05001b17f0ea30c95b7d6b0e90f0e842f5d",
  "latency_ms": 10160
 },
 "S88sh24::signage": {
  "fp": "077f014043012f7e",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S88sh24": {
  "input_fingerprint": "ff5632065d4fe4b4",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 슬픈 표정으로 눈물을 흘리는 현우의 뺨에 커다란 금속 손가락을 가만히 얹은 찰리의 다정한 상체.\n\nLOCATION (lock): At the examination bedside inside the glass-walled research laboratory, under its established cool lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Stainless-steel bed (찰리 remains lying on it during the bedside exchange) — A narrow portion of the near long edge and upper surface is visible beneath his shoulder; used as Shared spatial anchor beneath the two figures without obstructing the touch.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain subdued laboratory illumination and gentle tonal separation so the tears and metal finger remain readable without adding a new source or signaling the later alarm.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the stainless-steel bed, glass enclosure, laboratory equipment, and cool base lighting. Exclude the earlier fully secured restraint state; the hand restraints have been released.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie remains recumbent on the stainless-steel bed after his hands are released, extending one freed hand to rest a metal finger against Hyunwoo's cheek. The source does not specify his other hand's position, his legs' arrangement, or which remaining restraints are still fastened.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie remains lying on the stainless-steel bed, but the hand restraints have been released; no release of the remaining restraints has yet been established. The chest-ring system is still impaired following the failed experiment. 현우: He is beside the laboratory bed, crying heavily; the pole stand taken during the confrontation remains in his possession.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 슬픈 표정으로 눈물을 흘리는 현우의 뺨에 커다란 금속 손가락을 가만히 얹은 찰리의 다정한 상체.\n\nLOCATION (lock): At the examination bedside inside the glass-walled research laboratory, under its established cool lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Stainless-steel bed (찰리 remains lying on it during the bedside exchange) — A narrow portion of the near long edge and upper surface is visible beneath his shoulder; used as Shared spatial anchor beneath the two figures without obstructing the touch.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain subdued laboratory illumination and gentle tonal separation so the tears and metal finger remain readable without adding a new source or signaling the later alarm.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the stainless-steel bed, glass enclosure, laboratory equipment, and cool base lighting. Exclude the earlier fully secured restraint state; the hand restraints have been released.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie remains recumbent on the stainless-steel bed after his hands are released, extending one freed hand to rest a metal finger against Hyunwoo's cheek. The source does not specify his other hand's position, his legs' arrangement, or which remaining restraints are still fastened.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie remains lying on the stainless-steel bed, but the hand restraints have been released; no release of the remaining restraints has yet been established. The chest-ring system is still impaired following the failed experiment. 현우: He is beside the laboratory bed, crying heavily; the pole stand taken during the confrontation remains in his possession.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 슬픈 표정으로 눈물을 흘리는 현우의 뺨에 커다란 금속 손가락을 가만히 얹은 찰리의 다정한 상체.\n\nLOCATION (lock): At the examination bedside inside the glass-walled research laboratory, under its established cool lighting. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Stainless-steel bed (찰리 remains lying on it during the bedside exchange) — A narrow portion of the near long edge and upper surface is visible beneath his shoulder; used as Shared spatial anchor beneath the two figures without obstructing the touch.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain subdued laboratory illumination and gentle tonal separation so the tears and metal finger remain readable without adding a new source or signaling the later alarm.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the stainless-steel bed, glass enclosure, laboratory equipment, and cool base lighting. Exclude the earlier fully secured restraint state; the hand restraints have been released.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie remains recumbent on the stainless-steel bed after his hands are released, extending one freed hand to rest a metal finger against Hyunwoo's cheek. The source does not specify his other hand's position, his legs' arrangement, or which remaining restraints are still fastened.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie remains lying on the stainless-steel bed, but the hand restraints have been released; no release of the remaining restraints has yet been established. The chest-ring system is still impaired following the failed experiment. 현우: He is beside the laboratory bed, crying heavily; the pole stand taken during the confrontation remains in his possession.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "전경의 찰리가 현우의 뺨에 손을 대고 있고, 현우는 슬픈 표정으로 시선을 아래로 향함.",
    "built_space": "수술대가 전경과 배경에 2개 존재하며, 배경 유리 너머에 금지된 연구원들이 배치됨.",
    "entities": "현우의 외형은 일치하나 찰리가 2명으로 복제되었고, 제거되어야 할 배경 인물들이 포함됨.",
    "hard_violations": [
     "[gemini-pro] 복제된 신체 (두 명의 찰리)",
     "[gemini-pro] 복제된 구조물 (두 개의 침대)",
     "[gemini-pro] 금지된 배경 인물 (연구원들) 등장",
     "[gemini-pro] 지정된 위치 및 자세 위반 (찰리가 누워있지 않음)",
     "[gpt-high] 찰리를 전경과 배경 침대 위에 각각 한 번씩, 총 두 개체로 중복 생성했다.",
     "[gpt-high] 기존 중앙 검사 침대 외에 전경의 대형 경사 침대를 추가했다.",
     "[gpt-high] 등장이 금지된 이전 장면의 연구원들을 다수 그대로 포함했다."
    ],
    "physics": "현우는 폴대를 잡고 서 있고, 전경의 찰리는 복제된 침대에 체중을 싣고 앉아 있음."
   },
   {
    "label": "B",
    "direction": "찰리가 현우의 뺨에 손가락을 대고 있으며, 현우는 눈물을 흘리며 정면을 응시함.",
    "built_space": "중앙에 수술대가 비어 있고, 배경에 금지된 연구원들이 배치됨.",
    "entities": "현우와 찰리의 외형은 일치하지만, 등장해서는 안 될 배경 인물들이 포함됨.",
    "hard_violations": [
     "[gemini-pro] 지정된 위치 및 자세 위반 (찰리가 침대에 누워있지 않고 서 있음)",
     "[gemini-pro] 금지된 배경 인물 (연구원들) 등장",
     "[gpt-high] 등장이 금지된 이전 장면의 연구원들을 다수 그대로 포함했다.",
     "[gpt-high] 찰리를 침대에 누운 고정 위치가 아니라 침대 밖에 상체를 세운 위치로 옮겼다."
    ],
    "physics": "두 인물 모두 바닥에 서서 체중을 지탱하고 있으며, 현우는 손에 폴대를 쥐고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "찰리와 침대가 복제되었고 금지된 배경 인물이 등장하여 지시를 심각하게 위반했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "찰리가 침대에 누워있지 않고 서 있으며, 등장해서는 안 될 배경 인물이 포함되어 실패했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "전경의 찰리가 현우의 뺨에 손을 대고 있고, 현우는 슬픈 표정으로 시선을 아래로 향함.",
        "built_space": "수술대가 전경과 배경에 2개 존재하며, 배경 유리 너머에 금지된 연구원들이 배치됨.",
        "entities": "현우의 외형은 일치하나 찰리가 2명으로 복제되었고, 제거되어야 할 배경 인물들이 포함됨.",
        "hard_violations": [
         "복제된 신체 (두 명의 찰리)",
         "복제된 구조물 (두 개의 침대)",
         "금지된 배경 인물 (연구원들) 등장",
         "지정된 위치 및 자세 위반 (찰리가 누워있지 않음)"
        ],
        "physics": "현우는 폴대를 잡고 서 있고, 전경의 찰리는 복제된 침대에 체중을 싣고 앉아 있음."
       },
       {
        "label": "B",
        "direction": "찰리가 현우의 뺨에 손가락을 대고 있으며, 현우는 눈물을 흘리며 정면을 응시함.",
        "built_space": "중앙에 수술대가 비어 있고, 배경에 금지된 연구원들이 배치됨.",
        "entities": "현우와 찰리의 외형은 일치하지만, 등장해서는 안 될 배경 인물들이 포함됨.",
        "hard_violations": [
         "지정된 위치 및 자세 위반 (찰리가 침대에 누워있지 않고 서 있음)",
         "금지된 배경 인물 (연구원들) 등장"
        ],
        "physics": "두 인물 모두 바닥에 서서 체중을 지탱하고 있으며, 현우는 손에 폴대를 쥐고 있음."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "찰리와 침대가 복제되었고 금지된 배경 인물이 등장하여 지시를 심각하게 위반했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "찰리가 침대에 누워있지 않고 서 있으며, 등장해서는 안 될 배경 인물이 포함되어 실패했습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "전경의 찰리가 현우의 뺨에 손을 대고 있고, 현우는 슬픈 표정으로 시선을 아래로 향함.",
        "built_space": "수술대가 전경과 배경에 2개 존재하며, 배경 유리 너머에 금지된 연구원들이 배치됨.",
        "entities": "현우의 외형은 일치하나 찰리가 2명으로 복제되었고, 제거되어야 할 배경 인물들이 포함됨.",
        "hard_violations": [
         "복제된 신체 (두 명의 찰리)",
         "복제된 구조물 (두 개의 침대)",
         "금지된 배경 인물 (연구원들) 등장",
         "지정된 위치 및 자세 위반 (찰리가 누워있지 않음)"
        ],
        "physics": "현우는 폴대를 잡고 서 있고, 전경의 찰리는 복제된 침대에 체중을 싣고 앉아 있음."
       },
       {
        "label": "B",
        "direction": "찰리가 현우의 뺨에 손가락을 대고 있으며, 현우는 눈물을 흘리며 정면을 응시함.",
        "built_space": "중앙에 수술대가 비어 있고, 배경에 금지된 연구원들이 배치됨.",
        "entities": "현우와 찰리의 외형은 일치하지만, 등장해서는 안 될 배경 인물들이 포함됨.",
        "hard_violations": [
         "지정된 위치 및 자세 위반 (찰리가 침대에 누워있지 않고 서 있음)",
         "금지된 배경 인물 (연구원들) 등장"
        ],
        "physics": "두 인물 모두 바닥에 서서 체중을 지탱하고 있으며, 현우는 손에 폴대를 쥐고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "손가락의 뺨 접촉과 눈물은 구현했지만 찰리가 침대 밖에 서 있고 금지된 연구원들이 남아 있어 부적합하며, 찰리와 침대까지 복제한 B보다는 위반이 적습니다."
       },
       {
        "label": "B",
        "score": 1,
        "verdict_ko": "침대에 기대어 우는 현우의 뺨을 만지는 관계는 더 가깝지만, 배경에 결박된 찰리와 별도 침대를 그대로 남긴 중복 및 다수의 추가 인물 때문에 부적합합니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 얼굴은 왼쪽의 현우를 향하고, 뻗은 금속 검지 끝은 현우의 코 바로 옆 윗뺨에 닿는다. 접촉 대상은 맞다. 현우는 눈물을 흘리며 정면 아래쪽을 보고 있어 찰리와 시선을 맞추지는 않는다. 현우가 쥔 봉은 수직으로 올라간다.",
        "built_space": "원형 유리 구획, 둘레 작업대와 모니터, 차가운 조명은 이전 장소와 일치한다. 중앙에 빈 검사 침대 한 개와 오른쪽 작업용 로봇팔 한 개가 뚜렷하다. 그러나 침대는 두 인물 뒤에 있고 찰리의 어깨 아래를 받치지 않는다. 현우와 찰리는 침대 곁의 교감 장면이 아니라 유리 구획 앞쪽에 나란히 서 있는 배치다. 배경과 오른쪽 아래에 연구원 약 열 명이 남아 있다.",
        "entities": "현우는 앳된 동아시아계 남성으로, 헝클어진 검은 머리와 남색 반소매 티셔츠가 참조에 가깝고 뺨의 눈물이 보인다. 국적은 외형만으로 확인할 수 없다. 한 손에는 금속 봉을 쥐고 있다. 찰리는 흰 마스크형 얼굴, 모래색 각진 장갑, 육중한 팔을 갖춘 한 개체로 표현되었다. 손목은 풀려 있지만 다른 구속구까지 침대와 바닥에 벗겨져 있어 손만 해제된 연속성이 깨진다. 읽을 수 있는 문구는 뚜렷하지 않다.",
        "hard_violations": [
         "등장이 금지된 이전 장면의 연구원들을 다수 그대로 포함했다.",
         "찰리를 침대에 누운 고정 위치가 아니라 침대 밖에 상체를 세운 위치로 옮겼다."
        ],
        "physics": "금속 손가락은 손과 팔에 연결되어 있고, 관절을 굽혀 현우의 뺨까지 뻗은 동작은 가능하다. 봉은 현우의 손이 실제로 감싸 쥔다. 두 인물의 발은 화면 밖이므로 바닥 접촉은 확인되지 않지만 공중에 떠 있다고 볼 증거는 없다. 핵심 문제는 찰리가 누워 있어야 할 침대의 지지를 전혀 받지 않는 배치이며, 바닥의 풀린 구속구는 바닥에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "앞쪽 찰리는 오른쪽 현우의 얼굴을 내려다보고 금속 손가락들을 현우의 뺨에 댄다. 접촉 대상은 정확하나 한 손가락을 가만히 얹기보다는 손으로 뺨을 감싸는 모습이다. 현우는 눈을 감고 얼굴을 찡그리며 울고 있다. 현우가 쥔 봉은 오른쪽에서 수직으로 뻗는다.",
        "built_space": "유리 구획과 둘레 작업대, 모니터, 차가운 조명은 유지되었다. 그러나 왼쪽 전경에 크게 기울어진 금속 침대가 있고 중앙 배경에도 기존 검사 침대가 있어 침대가 두 개다. 전경 침대는 좁은 가장자리만 보이라는 지시와 달리 화면 왼쪽을 넓게 차지한다. 배경 침대에는 또 다른 찰리가 결박되어 있다. 작업용 로봇팔은 배경에서 두 개가 보이며, 연구원도 최소 여덟 명 남아 있다.",
        "entities": "현우는 검은 머리의 앳된 동아시아계 남성으로 눈물과 슬픈 표정이 선명하고 금속 봉을 쥐고 있다. 다만 참조의 반소매 티셔츠와 달리 긴소매 상의를 입었다. 전경 찰리의 흰 마스크와 모래색 장갑, 큰 팔은 참조의 특징을 따른다. 배경에는 같은 찰리의 몸체가 한 번 더 등장한다. 전경 찰리의 손은 풀려 있으나 배경 복제 몸체에는 이전 결박 상태가 남아 있다. 읽을 수 있는 문구는 뚜렷하지 않다.",
        "hard_violations": [
         "찰리를 전경과 배경 침대 위에 각각 한 번씩, 총 두 개체로 중복 생성했다.",
         "기존 중앙 검사 침대 외에 전경의 대형 경사 침대를 추가했다.",
         "등장이 금지된 이전 장면의 연구원들을 다수 그대로 포함했다."
        ],
        "physics": "전경 찰리의 등과 몸통은 기울어진 금속 침대에 기대어 지지되고, 뻗은 손은 연결된 팔 관절이 받친다. 따라서 손의 접촉 자체를 무지지 부유로 볼 이유는 없다. 배경의 복제 찰리도 별도 침대에 누워 지지된다. 현우의 봉에는 쥔 손이 보이고 아래쪽은 화면 밖으로 이어진다. 물체의 부유보다는 몸체와 지지 침대를 함께 복제한 것이 결정적인 문제다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "손가락의 뺨 접촉과 눈물은 구현했지만 찰리가 침대 밖에 서 있고 금지된 연구원들이 남아 있어 부적합하며, 찰리와 침대까지 복제한 B보다는 위반이 적습니다."
       },
       {
        "label": "A",
        "score": 1,
        "verdict_ko": "침대에 기대어 우는 현우의 뺨을 만지는 관계는 더 가깝지만, 배경에 결박된 찰리와 별도 침대를 그대로 남긴 중복 및 다수의 추가 인물 때문에 부적합합니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 얼굴은 왼쪽의 현우를 향하고, 뻗은 금속 검지 끝은 현우의 코 바로 옆 윗뺨에 닿는다. 접촉 대상은 맞다. 현우는 눈물을 흘리며 정면 아래쪽을 보고 있어 찰리와 시선을 맞추지는 않는다. 현우가 쥔 봉은 수직으로 올라간다.",
        "built_space": "원형 유리 구획, 둘레 작업대와 모니터, 차가운 조명은 이전 장소와 일치한다. 중앙에 빈 검사 침대 한 개와 오른쪽 작업용 로봇팔 한 개가 뚜렷하다. 그러나 침대는 두 인물 뒤에 있고 찰리의 어깨 아래를 받치지 않는다. 현우와 찰리는 침대 곁의 교감 장면이 아니라 유리 구획 앞쪽에 나란히 서 있는 배치다. 배경과 오른쪽 아래에 연구원 약 열 명이 남아 있다.",
        "entities": "현우는 앳된 동아시아계 남성으로, 헝클어진 검은 머리와 남색 반소매 티셔츠가 참조에 가깝고 뺨의 눈물이 보인다. 국적은 외형만으로 확인할 수 없다. 한 손에는 금속 봉을 쥐고 있다. 찰리는 흰 마스크형 얼굴, 모래색 각진 장갑, 육중한 팔을 갖춘 한 개체로 표현되었다. 손목은 풀려 있지만 다른 구속구까지 침대와 바닥에 벗겨져 있어 손만 해제된 연속성이 깨진다. 읽을 수 있는 문구는 뚜렷하지 않다.",
        "hard_violations": [
         "등장이 금지된 이전 장면의 연구원들을 다수 그대로 포함했다.",
         "찰리를 침대에 누운 고정 위치가 아니라 침대 밖에 상체를 세운 위치로 옮겼다."
        ],
        "physics": "금속 손가락은 손과 팔에 연결되어 있고, 관절을 굽혀 현우의 뺨까지 뻗은 동작은 가능하다. 봉은 현우의 손이 실제로 감싸 쥔다. 두 인물의 발은 화면 밖이므로 바닥 접촉은 확인되지 않지만 공중에 떠 있다고 볼 증거는 없다. 핵심 문제는 찰리가 누워 있어야 할 침대의 지지를 전혀 받지 않는 배치이며, 바닥의 풀린 구속구는 바닥에 놓여 있다."
       },
       {
        "label": "A",
        "direction": "앞쪽 찰리는 오른쪽 현우의 얼굴을 내려다보고 금속 손가락들을 현우의 뺨에 댄다. 접촉 대상은 정확하나 한 손가락을 가만히 얹기보다는 손으로 뺨을 감싸는 모습이다. 현우는 눈을 감고 얼굴을 찡그리며 울고 있다. 현우가 쥔 봉은 오른쪽에서 수직으로 뻗는다.",
        "built_space": "유리 구획과 둘레 작업대, 모니터, 차가운 조명은 유지되었다. 그러나 왼쪽 전경에 크게 기울어진 금속 침대가 있고 중앙 배경에도 기존 검사 침대가 있어 침대가 두 개다. 전경 침대는 좁은 가장자리만 보이라는 지시와 달리 화면 왼쪽을 넓게 차지한다. 배경 침대에는 또 다른 찰리가 결박되어 있다. 작업용 로봇팔은 배경에서 두 개가 보이며, 연구원도 최소 여덟 명 남아 있다.",
        "entities": "현우는 검은 머리의 앳된 동아시아계 남성으로 눈물과 슬픈 표정이 선명하고 금속 봉을 쥐고 있다. 다만 참조의 반소매 티셔츠와 달리 긴소매 상의를 입었다. 전경 찰리의 흰 마스크와 모래색 장갑, 큰 팔은 참조의 특징을 따른다. 배경에는 같은 찰리의 몸체가 한 번 더 등장한다. 전경 찰리의 손은 풀려 있으나 배경 복제 몸체에는 이전 결박 상태가 남아 있다. 읽을 수 있는 문구는 뚜렷하지 않다.",
        "hard_violations": [
         "찰리를 전경과 배경 침대 위에 각각 한 번씩, 총 두 개체로 중복 생성했다.",
         "기존 중앙 검사 침대 외에 전경의 대형 경사 침대를 추가했다.",
         "등장이 금지된 이전 장면의 연구원들을 다수 그대로 포함했다."
        ],
        "physics": "전경 찰리의 등과 몸통은 기울어진 금속 침대에 기대어 지지되고, 뻗은 손은 연결된 팔 관절이 받친다. 따라서 손의 접촉 자체를 무지지 부유로 볼 이유는 없다. 배경의 복제 찰리도 별도 침대에 누워 지지된다. 현우의 봉에는 쥔 손이 보이고 아래쪽은 화면 밖으로 이어진다. 물체의 부유보다는 몸체와 지지 침대를 함께 복제한 것이 결정적인 문제다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.25,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.0,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 복제된 신체 (두 명의 찰리)",
     "[gemini-pro] 복제된 구조물 (두 개의 침대)",
     "[gemini-pro] 금지된 배경 인물 (연구원들) 등장",
     "[gemini-pro] 지정된 위치 및 자세 위반 (찰리가 누워있지 않음)",
     "[gpt-high] 찰리를 전경과 배경 침대 위에 각각 한 번씩, 총 두 개체로 중복 생성했다.",
     "[gpt-high] 기존 중앙 검사 침대 외에 전경의 대형 경사 침대를 추가했다.",
     "[gpt-high] 등장이 금지된 이전 장면의 연구원들을 다수 그대로 포함했다."
    ],
    "B": [
     "[gemini-pro] 지정된 위치 및 자세 위반 (찰리가 침대에 누워있지 않고 서 있음)",
     "[gemini-pro] 금지된 배경 인물 (연구원들) 등장",
     "[gpt-high] 등장이 금지된 이전 장면의 연구원들을 다수 그대로 포함했다.",
     "[gpt-high] 찰리를 침대에 누운 고정 위치가 아니라 침대 밖에 상체를 세운 위치로 옮겼다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "A": 1000,
   "B": 1750
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1000,
    "verdict_ko": "찰리와 침대가 복제되었고 금지된 배경 인물이 등장하여 지시를 심각하게 위반했습니다.  ★위반: [gemini-pro] 복제된 신체 (두 명의 찰리) / [gemini-pro] 복제된 구조물 (두 개의 침대) / [gemini-pro] 금지된 배경 인물 (연구원들) 등장 / [gemini-pro] 지정된 위치 및 자세 위반 (찰리가 누워있지 않음) / [gpt-high] 찰리를 전경과 배경 침대 위에 각각 한 번씩, 총 두 개체로 중복 생성했다. / [gpt-high] 기존 중앙 검사 침대 외에 전경의 대형 경사 침대를 추가했다. / [gpt-high] 등장이 금지된 이전 장면의 연구원들을 다수 그대로 포함했다."
   },
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "찰리가 침대에 누워있지 않고 서 있으며, 등장해서는 안 될 배경 인물이 포함되어 실패했습니다.  ★위반: [gemini-pro] 지정된 위치 및 자세 위반 (찰리가 침대에 누워있지 않고 서 있음) / [gemini-pro] 금지된 배경 인물 (연구원들) 등장 / [gpt-high] 등장이 금지된 이전 장면의 연구원들을 다수 그대로 포함했다. / [gpt-high] 찰리를 침대에 누운 고정 위치가 아니라 침대 밖에 상체를 세운 위치로 옮겼다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 찰리 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S88sh7_sel.png",
    "asset_id": "b5b5d2a2-966e-4303-bbcf-f6b602aa2e9a",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92d-1338-77df-b764-713e6d39e350",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S88sh7"
  }
 },
 "S88sh24::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T12:16:11.669406+00:00",
  "fingerprint": "08612094e655a5862499c0d8d8edfce3121fb2475cf52b09aaa1411015cb7f8a",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S88sh24_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S88sh24_sel.png",
  "source_sha256": "da6692e6ae88397a49e23c8c10011710b17e5befe65242fb6b15134f1879cae2",
  "file": "S88sh24_cine.png",
  "staged_sha256": "a804b50e8fb1b9dc66be7551c71779f1cea98a9d6b753ec015cbb708c67c65e7",
  "latency_ms": 9493
 },
 "S88sh29::signage": {
  "fp": "5dcae9eff65f5a79",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S88sh29": {
  "input_fingerprint": "5bbe1e1ac52243ae",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 갑작스러운 붉은 사이렌 불빛 아래서 당황하여 주변의 한 곳으로 시선을 홱 돌린 채 굳어버린 지소영과 현우의 상체.\n\nLOCATION (lock): Inside the glass-walled laboratory's bedside and observation area, suddenly washed in red emergency warning light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Laboratory bed (Present beside the interrupted exchange) — A narrow portion of the stainless-steel side runs diagonally along the lower edge; used as Maintains the established bedside axis without including Charlie in this reaction frame; 지소영's chair (Occupied) — Only the portion supporting her seated body is visible behind her; used as Makes her lower position and interrupted seated movement legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Sudden red alarm light interrupts the existing laboratory illumination while preserving readable expressions and restrained tonal contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the stainless-steel bed, glass enclosure, and fixed laboratory equipment. Exclude the earlier unalarmed lighting state and fully secured hand restraints; apply the newly activated red emergency illumination.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The laboratory is now under a red alert. Charlie remains on the stainless-steel bed with the hands released but the remaining restraints not yet removed. 현우: He remains tearful beside the laboratory bed, still in possession of the stand taken during the confrontation. 지소영: She remains by the chair in her research coat, abruptly turning her attention to the alarm.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 갑작스러운 붉은 사이렌 불빛 아래서 당황하여 주변의 한 곳으로 시선을 홱 돌린 채 굳어버린 지소영과 현우의 상체.\n\nLOCATION (lock): Inside the glass-walled laboratory's bedside and observation area, suddenly washed in red emergency warning light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Laboratory bed (Present beside the interrupted exchange) — A narrow portion of the stainless-steel side runs diagonally along the lower edge; used as Maintains the established bedside axis without including Charlie in this reaction frame; 지소영's chair (Occupied) — Only the portion supporting her seated body is visible behind her; used as Makes her lower position and interrupted seated movement legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Sudden red alarm light interrupts the existing laboratory illumination while preserving readable expressions and restrained tonal contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the stainless-steel bed, glass enclosure, and fixed laboratory equipment. Exclude the earlier unalarmed lighting state and fully secured hand restraints; apply the newly activated red emergency illumination.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The laboratory is now under a red alert. Charlie remains on the stainless-steel bed with the hands released but the remaining restraints not yet removed. 현우: He remains tearful beside the laboratory bed, still in possession of the stand taken during the confrontation. 지소영: She remains by the chair in her research coat, abruptly turning her attention to the alarm.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): morning.\n\nSHOT TEXT (authoritative, Korean): 갑작스러운 붉은 사이렌 불빛 아래서 당황하여 주변의 한 곳으로 시선을 홱 돌린 채 굳어버린 지소영과 현우의 상체.\n\nLOCATION (lock): Inside the glass-walled laboratory's bedside and observation area, suddenly washed in red emergency warning light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Laboratory bed (Present beside the interrupted exchange) — A narrow portion of the stainless-steel side runs diagonally along the lower edge; used as Maintains the established bedside axis without including Charlie in this reaction frame; 지소영's chair (Occupied) — Only the portion supporting her seated body is visible behind her; used as Makes her lower position and interrupted seated movement legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Sudden red alarm light interrupts the existing laboratory illumination while preserving readable expressions and restrained tonal contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Retain the stainless-steel bed, glass enclosure, and fixed laboratory equipment. Exclude the earlier unalarmed lighting state and fully secured hand restraints; apply the newly activated red emergency illumination.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The laboratory is now under a red alert. Charlie remains on the stainless-steel bed with the hands released but the remaining restraints not yet removed. 현우: He remains tearful beside the laboratory bed, still in possession of the stand taken during the confrontation. 지소영: She remains by the chair in her research coat, abruptly turning her attention to the alarm.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "두 인물 모두 화면 왼쪽 밖을 향해 시선을 홱 돌리고 있습니다.",
    "built_space": "유리 벽과 모니터가 있는 연구실 내부입니다. 우측 하단에 스테인리스 침대 모서리가 대각선으로 걸쳐 있고, 좌측에는 회색 책상이 크게 배치되어 있습니다.",
    "entities": "현우는 눈물 자국과 외양(머리, 옷)이 레퍼런스와 일치합니다. 지소영도 나이대와 가운 착용 등 지정된 레퍼런스와 일치합니다.",
    "hard_violations": [
     "[gemini-pro] 발명된 사물: 위치 고정(Location lock) 레퍼런스(이전 샷)의 침대 주변에는 존재하지 않는 대형 회색 책상이 화면 좌측 전경에 임의로 추가됨",
     "[gpt-high] 왼쪽 뒤 콘솔 앞에 앉은 추가 인물이 보여, 지소영과 현우 외에는 누구도 등장시키지 말라는 제한을 위반한다."
    ],
    "physics": "지소영은 의자에 앉아 양손으로 팔걸이를 짚고 체중을 지탱하며 일어서려는 듯한 자세를 취하고 있고, 현우는 서서 프레임 밖 바닥에 지탱되어 있습니다."
   },
   {
    "label": "B",
    "direction": "두 인물 모두 놀란 표정으로 화면 오른쪽 밖의 한 곳을 향해 시선을 고정하고 있습니다.",
    "built_space": "유리 벽으로 둘러싸인 연구실 내부에 붉은 비상조명이 켜져 있습니다. 화면 하단 전체를 스테인리스 침대의 모서리가 대각선으로 가로지르며, 지소영의 뒤로는 의자의 등받이 부분만 보입니다.",
    "entities": "현우는 눈물을 흘리는 모습, 헝클어진 머리, 남색 반팔 티셔츠 등 레퍼런스와 정확히 일치합니다. 지소영 역시 50대 여성의 외양과 깔끔한 검은 머리, 흰색 연구복 착용 상태가 레퍼런스와 일치합니다.",
    "hard_violations": [
     "[gpt-high] 화면 맨 오른쪽 배경에 두 주인공 이외의 사람 얼굴이 보여, 이 반응 숏에는 지소영과 현우만 등장해야 한다는 인물 제한을 위반한다."
    ],
    "physics": "지소영은 화면에 보이는 의자에 앉아 체중을 지탱하고 있으며, 현우는 서 있는 자세로 프레임 밖 바닥에 안정적으로 지탱되어 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "지정된 구도와 붉은 조명, 캐릭터의 외양 및 시선 처리를 지침에 맞게 충실히 구현했으며, 의자의 노출 제한 지시도 정확히 따랐습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "프롬프트가 지시한 의자 노출 제한(등받이만 뒤로 보여야 함)을 어기고 팔걸이를 전면에 노출했으며, 이전 장면에 없던 책상을 추가해 공간의 일관성을 해쳤습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "두 인물 모두 놀란 표정으로 화면 오른쪽 밖의 한 곳을 향해 시선을 고정하고 있습니다.",
        "built_space": "유리 벽으로 둘러싸인 연구실 내부에 붉은 비상조명이 켜져 있습니다. 화면 하단 전체를 스테인리스 침대의 모서리가 대각선으로 가로지르며, 지소영의 뒤로는 의자의 등받이 부분만 보입니다.",
        "entities": "현우는 눈물을 흘리는 모습, 헝클어진 머리, 남색 반팔 티셔츠 등 레퍼런스와 정확히 일치합니다. 지소영 역시 50대 여성의 외양과 깔끔한 검은 머리, 흰색 연구복 착용 상태가 레퍼런스와 일치합니다.",
        "hard_violations": [],
        "physics": "지소영은 화면에 보이는 의자에 앉아 체중을 지탱하고 있으며, 현우는 서 있는 자세로 프레임 밖 바닥에 안정적으로 지탱되어 있습니다."
       },
       {
        "label": "A",
        "direction": "두 인물 모두 화면 왼쪽 밖을 향해 시선을 홱 돌리고 있습니다.",
        "built_space": "유리 벽과 모니터가 있는 연구실 내부입니다. 우측 하단에 스테인리스 침대 모서리가 대각선으로 걸쳐 있고, 좌측에는 회색 책상이 크게 배치되어 있습니다.",
        "entities": "현우는 눈물 자국과 외양(머리, 옷)이 레퍼런스와 일치합니다. 지소영도 나이대와 가운 착용 등 지정된 레퍼런스와 일치합니다.",
        "hard_violations": [
         "발명된 사물: 위치 고정(Location lock) 레퍼런스(이전 샷)의 침대 주변에는 존재하지 않는 대형 회색 책상이 화면 좌측 전경에 임의로 추가됨"
        ],
        "physics": "지소영은 의자에 앉아 양손으로 팔걸이를 짚고 체중을 지탱하며 일어서려는 듯한 자세를 취하고 있고, 현우는 서서 프레임 밖 바닥에 지탱되어 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "지정된 구도와 붉은 조명, 캐릭터의 외양 및 시선 처리를 지침에 맞게 충실히 구현했으며, 의자의 노출 제한 지시도 정확히 따랐습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "프롬프트가 지시한 의자 노출 제한(등받이만 뒤로 보여야 함)을 어기고 팔걸이를 전면에 노출했으며, 이전 장면에 없던 책상을 추가해 공간의 일관성을 해쳤습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "두 인물 모두 놀란 표정으로 화면 오른쪽 밖의 한 곳을 향해 시선을 고정하고 있습니다.",
        "built_space": "유리 벽으로 둘러싸인 연구실 내부에 붉은 비상조명이 켜져 있습니다. 화면 하단 전체를 스테인리스 침대의 모서리가 대각선으로 가로지르며, 지소영의 뒤로는 의자의 등받이 부분만 보입니다.",
        "entities": "현우는 눈물을 흘리는 모습, 헝클어진 머리, 남색 반팔 티셔츠 등 레퍼런스와 정확히 일치합니다. 지소영 역시 50대 여성의 외양과 깔끔한 검은 머리, 흰색 연구복 착용 상태가 레퍼런스와 일치합니다.",
        "hard_violations": [],
        "physics": "지소영은 화면에 보이는 의자에 앉아 체중을 지탱하고 있으며, 현우는 서 있는 자세로 프레임 밖 바닥에 안정적으로 지탱되어 있습니다."
       },
       {
        "label": "A",
        "direction": "두 인물 모두 화면 왼쪽 밖을 향해 시선을 홱 돌리고 있습니다.",
        "built_space": "유리 벽과 모니터가 있는 연구실 내부입니다. 우측 하단에 스테인리스 침대 모서리가 대각선으로 걸쳐 있고, 좌측에는 회색 책상이 크게 배치되어 있습니다.",
        "entities": "현우는 눈물 자국과 외양(머리, 옷)이 레퍼런스와 일치합니다. 지소영도 나이대와 가운 착용 등 지정된 레퍼런스와 일치합니다.",
        "hard_violations": [
         "발명된 사물: 위치 고정(Location lock) 레퍼런스(이전 샷)의 침대 주변에는 존재하지 않는 대형 회색 책상이 화면 좌측 전경에 임의로 추가됨"
        ],
        "physics": "지소영은 의자에 앉아 양손으로 팔걸이를 짚고 체중을 지탱하며 일어서려는 듯한 자세를 취하고 있고, 현우는 서서 프레임 밖 바닥에 지탱되어 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "추가 인물 때문에 실격이지만, 두 사람의 상체 중심 구도와 화면 오른쪽으로 함께 돌린 시선은 B보다 지시된 반응 숏에 가깝다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "추가 인물이 보이고, 두 사람의 시선 방향이 갈리며 지소영의 허벅지와 의자까지 넓게 담아 상체 중심 반응 숏에서 더 벗어난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 고개와 눈을 화면 오른쪽 바깥으로 돌렸고, 지소영도 오른쪽 바깥을 바라본다. 구체적인 대상은 프레임 밖이지만 두 사람이 같은 쪽의 이상 징후에 반응하는 것으로 읽힌다. 현우의 벌어진 입과 지소영의 멈춘 몸에서 당황한 반응이 보인다.",
        "built_space": "금속 프레임으로 나뉜 유리벽, 왼쪽의 장비군, 중앙 위의 경광등 한 개와 유리 반사상이 보인다. 지소영은 왼쪽 의자 한 개에 낮게 앉고 현우는 오른쪽 침대 곁에 서 있다. 침대 한 개의 금속 측면이 하단을 대각선으로 지나지만 오른쪽에서는 상당히 넓어져 '좁은 부분만'이라는 지시보다 두드러진다. 의자는 등받이 일부만 보여 요구에 비교적 가깝다. 이전 장면의 유리와 금속 재질은 유지되지만 고정 장비의 동일한 배치는 확인하기 어렵다.",
        "entities": "주인공 두 사람은 각각 검은 단발의 중년 동아시아계 여성과 헝클어진 검은 머리의 앳된 동아시아계 남성으로 보인다. 지소영의 연구복, 현우의 남색 티셔츠와 눈물 자국은 지시에 맞고 얼굴도 참조와 대체로 유사하다. 다만 오른쪽 뒤 유리 너머에 별도의 사람 얼굴이 보이며 흰 가운 형상들도 있다. 찰리는 보이지 않는다. 현우의 스탠드 소지 여부는 손과 하체가 가려져 확인할 수 없다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "화면 맨 오른쪽 배경에 두 주인공 이외의 사람 얼굴이 보여, 이 반응 숏에는 지소영과 현우만 등장해야 한다는 인물 제한을 위반한다."
        ],
        "physics": "지소영의 낮은 몸은 뒤의 의자 좌판 위치와 연결되고 등은 등받이 쪽에 있어 착석 자세가 성립한다. 현우는 몸통이 수직으로 선 자세이며 발은 프레임 밖이다. 침대 금속판 아래에는 지지 구조가 보이고, 경광등과 장비는 구조물에 부착되거나 받침 위에 놓여 있다. 지지 없이 떠 있는 몸이나 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "현우는 몸통을 비튼 채 카메라 가까운 화면 왼쪽 바깥을 바라본다. 지소영은 화면 오른쪽의 현우 쪽을 보고 있어 두 시선이 같은 경보 발생 지점으로 모이지 않는다. 두 사람 모두 놀란 표정은 있으나 한곳으로 갑자기 시선을 돌린 공동 반응은 A보다 불분명하다.",
        "built_space": "금속 기둥으로 분할된 유리벽과 뒤쪽 관찰용 콘솔·모니터들이 보인다. 지소영은 왼쪽 의자 한 개에 앉아 양쪽 팔걸이를 잡고, 현우는 오른쪽 침대 옆에 있다. 침대 한 개의 금속 측면이 하단을 대각선으로 가로지르지만 넓은 면적으로 드러난다. 지소영의 허벅지, 좌판 주변과 양쪽 팔걸이까지 보여 지시된 상체 중심 구도보다 넓다. 관찰실 콘솔은 이전 장소의 성격과 맞지만 동일한 고정 배치인지는 확인하기 어렵다.",
        "entities": "현우의 앳된 얼굴, 헝클어진 검은 머리, 남색 티셔츠와 뺨의 눈물은 참조에 가깝다. 지소영은 정돈된 검은 단발의 중년 동아시아계 여성으로 연구복과 남색 상의를 입었다. 왼쪽 뒤 콘솔 앞에 등을 보인 별도의 사람이 명확히 있다. 찰리는 프레임에 없다. 현우의 손과 스탠드는 가려져 소지 여부를 판단할 수 없다. 가슴 명찰과 화면들은 있으나 읽을 수 있는 문구는 식별되지 않는다.",
        "hard_violations": [
         "왼쪽 뒤 콘솔 앞에 앉은 추가 인물이 보여, 지소영과 현우 외에는 누구도 등장시키지 말라는 제한을 위반한다."
        ],
        "physics": "지소영의 골반은 의자 좌판에 놓이고 양손은 각각 팔걸이를 잡아, 일어나려다 멈춘 듯한 체중 지지가 자연스럽다. 현우의 비튼 상체도 가능한 자세이며 하체는 침대에 가려져 있다. 침대는 오른쪽 아래 지지부와 연결되고 콘솔 장비는 작업대 위에 놓여 있다. 지지 없이 공중에 뜬 물체나 불가능한 관절 자세는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "추가 인물 때문에 실격이지만, 두 사람의 상체 중심 구도와 화면 오른쪽으로 함께 돌린 시선은 B보다 지시된 반응 숏에 가깝다."
       },
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "추가 인물이 보이고, 두 사람의 시선 방향이 갈리며 지소영의 허벅지와 의자까지 넓게 담아 상체 중심 반응 숏에서 더 벗어난다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 고개와 눈을 화면 오른쪽 바깥으로 돌렸고, 지소영도 오른쪽 바깥을 바라본다. 구체적인 대상은 프레임 밖이지만 두 사람이 같은 쪽의 이상 징후에 반응하는 것으로 읽힌다. 현우의 벌어진 입과 지소영의 멈춘 몸에서 당황한 반응이 보인다.",
        "built_space": "금속 프레임으로 나뉜 유리벽, 왼쪽의 장비군, 중앙 위의 경광등 한 개와 유리 반사상이 보인다. 지소영은 왼쪽 의자 한 개에 낮게 앉고 현우는 오른쪽 침대 곁에 서 있다. 침대 한 개의 금속 측면이 하단을 대각선으로 지나지만 오른쪽에서는 상당히 넓어져 '좁은 부분만'이라는 지시보다 두드러진다. 의자는 등받이 일부만 보여 요구에 비교적 가깝다. 이전 장면의 유리와 금속 재질은 유지되지만 고정 장비의 동일한 배치는 확인하기 어렵다.",
        "entities": "주인공 두 사람은 각각 검은 단발의 중년 동아시아계 여성과 헝클어진 검은 머리의 앳된 동아시아계 남성으로 보인다. 지소영의 연구복, 현우의 남색 티셔츠와 눈물 자국은 지시에 맞고 얼굴도 참조와 대체로 유사하다. 다만 오른쪽 뒤 유리 너머에 별도의 사람 얼굴이 보이며 흰 가운 형상들도 있다. 찰리는 보이지 않는다. 현우의 스탠드 소지 여부는 손과 하체가 가려져 확인할 수 없다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "화면 맨 오른쪽 배경에 두 주인공 이외의 사람 얼굴이 보여, 이 반응 숏에는 지소영과 현우만 등장해야 한다는 인물 제한을 위반한다."
        ],
        "physics": "지소영의 낮은 몸은 뒤의 의자 좌판 위치와 연결되고 등은 등받이 쪽에 있어 착석 자세가 성립한다. 현우는 몸통이 수직으로 선 자세이며 발은 프레임 밖이다. 침대 금속판 아래에는 지지 구조가 보이고, 경광등과 장비는 구조물에 부착되거나 받침 위에 놓여 있다. 지지 없이 떠 있는 몸이나 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "현우는 몸통을 비튼 채 카메라 가까운 화면 왼쪽 바깥을 바라본다. 지소영은 화면 오른쪽의 현우 쪽을 보고 있어 두 시선이 같은 경보 발생 지점으로 모이지 않는다. 두 사람 모두 놀란 표정은 있으나 한곳으로 갑자기 시선을 돌린 공동 반응은 A보다 불분명하다.",
        "built_space": "금속 기둥으로 분할된 유리벽과 뒤쪽 관찰용 콘솔·모니터들이 보인다. 지소영은 왼쪽 의자 한 개에 앉아 양쪽 팔걸이를 잡고, 현우는 오른쪽 침대 옆에 있다. 침대 한 개의 금속 측면이 하단을 대각선으로 가로지르지만 넓은 면적으로 드러난다. 지소영의 허벅지, 좌판 주변과 양쪽 팔걸이까지 보여 지시된 상체 중심 구도보다 넓다. 관찰실 콘솔은 이전 장소의 성격과 맞지만 동일한 고정 배치인지는 확인하기 어렵다.",
        "entities": "현우의 앳된 얼굴, 헝클어진 검은 머리, 남색 티셔츠와 뺨의 눈물은 참조에 가깝다. 지소영은 정돈된 검은 단발의 중년 동아시아계 여성으로 연구복과 남색 상의를 입었다. 왼쪽 뒤 콘솔 앞에 등을 보인 별도의 사람이 명확히 있다. 찰리는 프레임에 없다. 현우의 손과 스탠드는 가려져 소지 여부를 판단할 수 없다. 가슴 명찰과 화면들은 있으나 읽을 수 있는 문구는 식별되지 않는다.",
        "hard_violations": [
         "왼쪽 뒤 콘솔 앞에 앉은 추가 인물이 보여, 지소영과 현우 외에는 누구도 등장시키지 말라는 제한을 위반한다."
        ],
        "physics": "지소영의 골반은 의자 좌판에 놓이고 양손은 각각 팔걸이를 잡아, 일어나려다 멈춘 듯한 체중 지지가 자연스럽다. 현우의 비튼 상체도 가능한 자세이며 하체는 침대에 가려져 있다. 침대는 오른쪽 아래 지지부와 연결되고 콘솔 장비는 작업대 위에 놓여 있다. 지지 없이 공중에 뜬 물체나 불가능한 관절 자세는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.167,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.917,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 발명된 사물: 위치 고정(Location lock) 레퍼런스(이전 샷)의 침대 주변에는 존재하지 않는 대형 회색 책상이 화면 좌측 전경에 임의로 추가됨",
     "[gpt-high] 왼쪽 뒤 콘솔 앞에 앉은 추가 인물이 보여, 지소영과 현우 외에는 누구도 등장시키지 말라는 제한을 위반한다."
    ],
    "B": [
     "[gpt-high] 화면 맨 오른쪽 배경에 두 주인공 이외의 사람 얼굴이 보여, 이 반응 숏에는 지소영과 현우만 등장해야 한다는 인물 제한을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 1750,
   "A": 917
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "지정된 구도와 붉은 조명, 캐릭터의 외양 및 시선 처리를 지침에 맞게 충실히 구현했으며, 의자의 노출 제한 지시도 정확히 따랐습니다.  ★위반: [gpt-high] 화면 맨 오른쪽 배경에 두 주인공 이외의 사람 얼굴이 보여, 이 반응 숏에는 지소영과 현우만 등장해야 한다는 인물 제한을 위반한다."
   },
   {
    "label": "A",
    "score": 917,
    "verdict_ko": "프롬프트가 지시한 의자 노출 제한(등받이만 뒤로 보여야 함)을 어기고 팔걸이를 전면에 노출했으며, 이전 장면에 없던 책상을 추가해 공간의 일관성을 해쳤습니다.  ★위반: [gemini-pro] 발명된 사물: 위치 고정(Location lock) 레퍼런스(이전 샷)의 침대 주변에는 존재하지 않는 대형 회색 책상이 화면 좌측 전경에 임의로 추가됨 / [gpt-high] 왼쪽 뒤 콘솔 앞에 앉은 추가 인물이 보여, 지소영과 현우 외에는 누구도 등장시키지 말라는 제한을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 현우 — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S88sh24_sel.png",
    "asset_id": "733f3193-7aa9-4077-a2e6-25c50ee809e0",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 지소영: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1243508>",
    "asset_id": "c7496f13-cfcf-44a5-976d-96c783d20580",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92d-14f0-764e-8b87-3144e10a68c5",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S88sh24"
  }
 },
 "S88sh29::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T12:17:36.658856+00:00",
  "fingerprint": "584acff5489c5113f907854e113d9b6a4c15d6d181eecfc8bb80c620eaca9c3f",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S88sh29_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S88sh29_sel.png",
  "source_sha256": "846f4740f71df504a42bf79d1398189f9f3a20ed472eff15bbcca7375be6ce5a",
  "file": "S88sh29_cine.png",
  "staged_sha256": "f47d6a8c96b4d7c138d4151d2cea3dae2e8e95f3678bff697a18845e28969f00",
  "latency_ms": 10425
 },
 "S89sh42::signage": {
  "fp": "6f15d9bcb8905787",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::9b5ca709d4a5ecc0": {
  "subjects": [],
  "subject_text": "제주도 연구소 옥상과 안테나 구역\n잔디밭을 갖춘 옥상 공원. 거대한 안테나와 사방을 향하는 스피커가 설치된 개방형 공간이다.",
  "identity": "canonical",
  "scope_id": "L278",
  "scope_role": "location_exterior",
  "scope_sha": "0f3607503a5f57e9"
 },
 "S89sh42::bgfirst_bg": {
  "input_fingerprint": "5101d39a2219e53a",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 헬기 스텝에 한 발을 딛고 선 채 피 흘리는 현우를 향해 입꼬리를 비스듬히 올리고 내려다보는 윤성찬의 거만한 전신.\n\nLOCATION (lock): At the step of a helicopter landed on the island research facility's rooftop, beside the rooftop lawn and antenna area.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 윤성찬's helicopter (Landed on the rooftop) — The step-bearing side is seen obliquely, with most of the aircraft outside the right edge; used as Provides the practical elevation beneath 윤성찬 without overwhelming his full-body silhouette; Research facility rooftop (The confrontation takes place here after the helicopter lands) — The ground plane recedes from 현우 toward the helicopter step; used as Connects the injured foreground figure and elevated antagonist within one continuous space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light keeps the blood, downward smile, and unequal elevations readable without adding theatrical lighting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.\n\nThe THIRD attached image (STRUCTURE LOOK) is the identity source of the fixed structure at this location: its faces, openings, levels, materials and signage are truth. Where it conflicts with the LOCATION PHOTOGRAPH about the structure itself, the STRUCTURE LOOK wins; the photograph still governs the surroundings, time of day and lighting.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 헬기 스텝에 한 발을 딛고 선 채 피 흘리는 현우를 향해 입꼬리를 비스듬히 올리고 내려다보는 윤성찬의 거만한 전신.\n\nLOCATION (lock): At the step of a helicopter landed on the island research facility's rooftop, beside the rooftop lawn and antenna area.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 윤성찬's helicopter (Landed on the rooftop) — The step-bearing side is seen obliquely, with most of the aircraft outside the right edge; used as Provides the practical elevation beneath 윤성찬 without overwhelming his full-body silhouette; Research facility rooftop (The confrontation takes place here after the helicopter lands) — The ground plane recedes from 현우 toward the helicopter step; used as Connects the injured foreground figure and elevated antagonist within one continuous space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light keeps the blood, downward smile, and unequal elevations readable without adding theatrical lighting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.\n\nThe THIRD attached image (STRUCTURE LOOK) is the identity source of the fixed structure at this location: its faces, openings, levels, materials and signage are truth. Where it conflicts with the LOCATION PHOTOGRAPH about the structure itself, the STRUCTURE LOOK wins; the photograph still governs the surroundings, time of day and lighting.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S89sh42__bgfirst_bg.png",
  "asset_id": "99219a2e-c866-4ccd-9da5-2b903d22263a",
  "input_asset_ids": [
   "2a41fdeb-7ce7-4a06-960e-c2b1ec36de8b",
   "6eab84a7-fe74-47c0-bd42-78a5d724b6a9",
   "ae92432a-4956-46ca-9845-8c2d71d637c5"
  ]
 },
 "S89sh42": {
  "input_fingerprint": "f046d53b7fd3dca7",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 헬기 스텝에 한 발을 딛고 선 채 피 흘리는 현우를 향해 입꼬리를 비스듬히 올리고 내려다보는 윤성찬의 거만한 전신.\n\nLOCATION (lock): At the step of a helicopter landed on the island research facility's rooftop, beside the rooftop lawn and antenna area. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nSTRUCTURE LOOK AUTHORITY: the attached STRUCTURE LOOK photograph is the identity of the fixed structure at this location — wherever that structure appears in the frame, its shape, proportions, openings, materials and colors are LOCKED to it. The LOCATION PHOTOGRAPH remains the authority for this shot's sub-space, surroundings, time of day and lighting. If the two conflict on the structure itself, the STRUCTURE LOOK photo wins; for everything else, the LOCATION PHOTOGRAPH wins.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 윤성찬's helicopter (Landed on the rooftop) — The step-bearing side is seen obliquely, with most of the aircraft outside the right edge; used as Provides the practical elevation beneath 윤성찬 without overwhelming his full-body silhouette; Research facility rooftop (The confrontation takes place here after the helicopter lands) — The ground plane recedes from 현우 toward the helicopter step; used as Connects the injured foreground figure and elevated antagonist within one continuous space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light keeps the blood, downward smile, and unequal elevations readable without adding theatrical lighting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rooftop garden retains its large antenna and all-direction speakers, now amid bombardment damage and collapsed combat robots. A Yubik command helicopter has landed on the roof while other armed aircraft surround the facility; Charlie is free of the bed restraints and still carries the experimental sensors. 윤성찬: He is disembarking from the command helicopter on the rooftop, smiling.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 헬기 스텝에 한 발을 딛고 선 채 피 흘리는 현우를 향해 입꼬리를 비스듬히 올리고 내려다보는 윤성찬의 거만한 전신.\n\nLOCATION (lock): At the step of a helicopter landed on the island research facility's rooftop, beside the rooftop lawn and antenna area. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 윤성찬's helicopter (Landed on the rooftop) — The step-bearing side is seen obliquely, with most of the aircraft outside the right edge; used as Provides the practical elevation beneath 윤성찬 without overwhelming his full-body silhouette; Research facility rooftop (The confrontation takes place here after the helicopter lands) — The ground plane recedes from 현우 toward the helicopter step; used as Connects the injured foreground figure and elevated antagonist within one continuous space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light keeps the blood, downward smile, and unequal elevations readable without adding theatrical lighting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rooftop garden retains its large antenna and all-direction speakers, now amid bombardment damage and collapsed combat robots. A Yubik command helicopter has landed on the roof while other armed aircraft surround the facility; Charlie is free of the bed restraints and still carries the experimental sensors. 윤성찬: He is disembarking from the command helicopter on the rooftop, smiling.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 헬기 스텝에 한 발을 딛고 선 채 피 흘리는 현우를 향해 입꼬리를 비스듬히 올리고 내려다보는 윤성찬의 거만한 전신.\n\nLOCATION (lock): At the step of a helicopter landed on the island research facility's rooftop, beside the rooftop lawn and antenna area. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nSTRUCTURE LOOK AUTHORITY: the attached STRUCTURE LOOK photograph is the identity of the fixed structure at this location — wherever that structure appears in the frame, its shape, proportions, openings, materials and colors are LOCKED to it. The LOCATION PHOTOGRAPH remains the authority for this shot's sub-space, surroundings, time of day and lighting. If the two conflict on the structure itself, the STRUCTURE LOOK photo wins; for everything else, the LOCATION PHOTOGRAPH wins.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 윤성찬's helicopter (Landed on the rooftop) — The step-bearing side is seen obliquely, with most of the aircraft outside the right edge; used as Provides the practical elevation beneath 윤성찬 without overwhelming his full-body silhouette; Research facility rooftop (The confrontation takes place here after the helicopter lands) — The ground plane recedes from 현우 toward the helicopter step; used as Connects the injured foreground figure and elevated antagonist within one continuous space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light keeps the blood, downward smile, and unequal elevations readable without adding theatrical lighting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rooftop garden retains its large antenna and all-direction speakers, now amid bombardment damage and collapsed combat robots. A Yubik command helicopter has landed on the roof while other armed aircraft surround the facility; Charlie is free of the bed restraints and still carries the experimental sensors. 윤성찬: He is disembarking from the command helicopter on the rooftop, smiling.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윤성찬 (한국인 남성, 77세, 고령의 얼굴, 깊은 이마 주름, 눈가 주름) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S89sh42__bgfirst_bg.png",
     "asset_id": "99219a2e-c866-4ccd-9da5-2b903d22263a",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S89sh42.png",
     "asset_id": "2a41fdeb-7ce7-4a06-960e-c2b1ec36de8b",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 윤성찬: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1453233>",
     "asset_id": "04d34665-3829-49fb-a6d2-25b9fd051d63",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its spatial layout, surroundings, fixed features, time of day and lighting mood are spatial truth; stage the moment inside this place. If a STRUCTURE LOOK photograph is also attached, that photo wins for the fixed structure itself — this photograph wins for everything around it. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L278B01.png",
     "asset_id": "6eab84a7-fe74-47c0-bd42-78a5d724b6a9",
     "role": "location_plate"
    },
    {
     "label": "STRUCTURE LOOK — the confirmed photograph of the fixed structure at this location: wherever the structure appears in the frame, its shape, proportions, materials, colors and openings are LOCKED to this photo. Never copy its camera framing, time of day or lighting — the shot text and the LOCATION PHOTOGRAPH are the authorities for those.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_island_research_complex_sel.png",
     "asset_id": "ae92432a-4956-46ca-9845-8c2d71d637c5",
     "role": "structure_seed_look"
    },
    {
     "label": "CHARACTER REFERENCE — 윤성찬: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1453233>",
     "asset_id": "04d34665-3829-49fb-a6d2-25b9fd051d63",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "윤성찬의 시선이 화면 좌측 전경에 몸을 웅크린 현우를 정확히 향하고 있음.",
    "built_space": "좌측의 옥상 구조물과 문, 중앙의 잔디밭 및 안테나, 우측의 헬기가 지시된 위치에 정확한 비율과 공간감으로 배치됨.",
    "entities": "윤성찬은 참조 이미지의 이목구비, 헤어, 안경과 일치하며 거만한 미소를 띰. 피 흘리는 현우와 잔디밭에 쓰러진 전투 로봇들이 명확히 묘사됨.",
    "hard_violations": [],
    "physics": "윤성찬은 한 발을 헬기 스텝에 딛고 한 손으로 기체를 잡아 안정적으로 체중을 지지하고 있으며, 현우는 바닥에 엎드려 지탱됨."
   },
   {
    "label": "B",
    "direction": "윤성찬이 좌측 구조물 벽에 기대어 앉은 현우를 내려다보고 있으며, 시선이 맞닿음.",
    "built_space": "옥상의 주요 시설물(문, 안테나, 헬기)이 존재하나, 현우가 전경이 아닌 중경의 벽면에 붙어 있어 요구된 전후경의 깊이 구도가 형성되지 않음.",
    "entities": "윤성찬의 얼굴과 안경은 참조와 일치하며 미소를 짓고 있음. 현우는 피를 흘린 채 앉아 있으며, 주변에 잔해들이 흩어져 있음.",
    "hard_violations": [],
    "physics": "윤성찬의 두 발이 바닥과 스텝에 닿아 있고, 현우 역시 바닥과 벽면에 체중을 의지하여 물리적 지지 상태가 정상적임."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "부상당한 현우를 전경에 배치하여 헬기 스텝까지 이어지는 깊이 있는 구도를 정확히 구현했으며, 윤성찬의 거만한 미소와 포즈가 프롬프트와 완벽히 일치합니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "윤성찬의 인물 표현과 헬기 탑승 포즈는 무난하나, 현우가 전경이 아닌 좌측 벽면에 배치되어 지시된 카메라 구도와 공간적 깊이감을 살리지 못했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "윤성찬의 시선이 화면 좌측 전경에 몸을 웅크린 현우를 정확히 향하고 있음.",
        "built_space": "좌측의 옥상 구조물과 문, 중앙의 잔디밭 및 안테나, 우측의 헬기가 지시된 위치에 정확한 비율과 공간감으로 배치됨.",
        "entities": "윤성찬은 참조 이미지의 이목구비, 헤어, 안경과 일치하며 거만한 미소를 띰. 피 흘리는 현우와 잔디밭에 쓰러진 전투 로봇들이 명확히 묘사됨.",
        "hard_violations": [],
        "physics": "윤성찬은 한 발을 헬기 스텝에 딛고 한 손으로 기체를 잡아 안정적으로 체중을 지지하고 있으며, 현우는 바닥에 엎드려 지탱됨."
       },
       {
        "label": "B",
        "direction": "윤성찬이 좌측 구조물 벽에 기대어 앉은 현우를 내려다보고 있으며, 시선이 맞닿음.",
        "built_space": "옥상의 주요 시설물(문, 안테나, 헬기)이 존재하나, 현우가 전경이 아닌 중경의 벽면에 붙어 있어 요구된 전후경의 깊이 구도가 형성되지 않음.",
        "entities": "윤성찬의 얼굴과 안경은 참조와 일치하며 미소를 짓고 있음. 현우는 피를 흘린 채 앉아 있으며, 주변에 잔해들이 흩어져 있음.",
        "hard_violations": [],
        "physics": "윤성찬의 두 발이 바닥과 스텝에 닿아 있고, 현우 역시 바닥과 벽면에 체중을 의지하여 물리적 지지 상태가 정상적임."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "부상당한 현우를 전경에 배치하여 헬기 스텝까지 이어지는 깊이 있는 구도를 정확히 구현했으며, 윤성찬의 거만한 미소와 포즈가 프롬프트와 완벽히 일치합니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "윤성찬의 인물 표현과 헬기 탑승 포즈는 무난하나, 현우가 전경이 아닌 좌측 벽면에 배치되어 지시된 카메라 구도와 공간적 깊이감을 살리지 못했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "윤성찬의 시선이 화면 좌측 전경에 몸을 웅크린 현우를 정확히 향하고 있음.",
        "built_space": "좌측의 옥상 구조물과 문, 중앙의 잔디밭 및 안테나, 우측의 헬기가 지시된 위치에 정확한 비율과 공간감으로 배치됨.",
        "entities": "윤성찬은 참조 이미지의 이목구비, 헤어, 안경과 일치하며 거만한 미소를 띰. 피 흘리는 현우와 잔디밭에 쓰러진 전투 로봇들이 명확히 묘사됨.",
        "hard_violations": [],
        "physics": "윤성찬은 한 발을 헬기 스텝에 딛고 한 손으로 기체를 잡아 안정적으로 체중을 지지하고 있으며, 현우는 바닥에 엎드려 지탱됨."
       },
       {
        "label": "B",
        "direction": "윤성찬이 좌측 구조물 벽에 기대어 앉은 현우를 내려다보고 있으며, 시선이 맞닿음.",
        "built_space": "옥상의 주요 시설물(문, 안테나, 헬기)이 존재하나, 현우가 전경이 아닌 중경의 벽면에 붙어 있어 요구된 전후경의 깊이 구도가 형성되지 않음.",
        "entities": "윤성찬의 얼굴과 안경은 참조와 일치하며 미소를 짓고 있음. 현우는 피를 흘린 채 앉아 있으며, 주변에 잔해들이 흩어져 있음.",
        "hard_violations": [],
        "physics": "윤성찬의 두 발이 바닥과 스텝에 닿아 있고, 현우 역시 바닥과 벽면에 체중을 의지하여 물리적 지지 상태가 정상적임."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "한 발을 스텝에 올린 전신과 현우를 향한 시선은 맞지만, 비스듬한 냉소보다 활짝 웃는 표정이며 고정 구조물의 확성기 탑도 구조 참조와 다릅니다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "부상당한 현우 너머로 윤성찬의 전신과 내려다보는 냉소를 연결한 구도가 더 정확하지만, 밝은 셔츠와 확성기 탑의 형태는 참조와 다릅니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "윤성찬은 고개와 눈을 화면 왼쪽 아래의 피 흘리는 현우 쪽으로 돌리고 웃는다. 현우는 벽에 기대어 얼굴을 윤성찬 쪽으로 향한다. 요구된 시선 관계는 성립하지만 윤성찬의 표정은 한쪽 입꼬리만 올린 냉소보다 이를 드러낸 큰 웃음에 가깝다. 배경 헬기들의 구체적인 이동 방향은 정지 화면만으로 확정하기 어렵다.",
        "built_space": "왼쪽에 환기구가 달린 문 하나와 벽등 하나, 중앙 잔디에 원형 접시 안테나 탑 하나와 좌우 나팔형 확성기 두 개, 뒤쪽에 콘크리트 난간이 보인다. 오른쪽 헬기는 출입구와 스텝이 비스듬히 보이고 기체 대부분이 화면 밖에 있다. 현우는 왼쪽 벽 아래, 윤성찬은 헬기와 잔디 경계에 있어 연속된 옥상 공간은 성립한다. 다만 탑은 장소 사진의 접시형을 따랐으며, 우선권이 있는 구조 사진의 다층 사각 확성기 탑과 주변 독립 확성기 배치를 재현하지 않았다.",
        "entities": "사람은 윤성찬과 숏 본문에 명시된 현우 두 명이다. 윤성찬은 회색 머리, 안경, 주름진 얼굴의 고령 동아시아계 남성으로 인물 참조와 대체로 부합하고, 어두운 재킷과 셔츠도 가깝다. 현우는 피 묻은 밝은 셔츠를 입은 젊은 동아시아계 남성이다. 착륙 헬기 한 대와 하늘의 헬기 네 대, 훼손된 잔디와 기계 잔해가 보이나 잔해가 전투 로봇임은 뚜렷하지 않다. 읽을 수 있는 문구는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "윤성찬의 한쪽 구두는 헬기 스텝에, 다른 구두는 옥상 바닥에 닿아 체중을 지지한다. 현우는 엉덩이와 다리를 지면에 두고 등을 벽에 기대며 손도 바닥에 댄다. 착륙 헬기는 바퀴로 지지되고 배경 헬기들은 로터가 있는 비행체다. 기계 잔해와 혈흔은 지면 위에 있으며 지지 없이 떠 있는 신체나 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "윤성찬은 왼쪽 전경의 현우를 향해 고개를 숙이고 눈을 내리깔며 입꼬리를 미세하게 올린다. 현우의 뒷머리와 몸통은 윤성찬 쪽을 향하고 있으나 눈은 보이지 않는다. 높은 위치의 윤성찬이 낮은 현우를 내려다보는 관계와 거만한 표정이 읽힌다. 배경 헬기의 특정 표적이나 이동 방향은 확인되지 않는다.",
        "built_space": "왼쪽 옥상 출입 구조물에 문 하나, 환기구 하나, 벽등 하나가 있고 중앙 뒤쪽에는 접시 안테나 탑 하나와 좌우 확성기 두 개가 있다. 현우가 있는 전경의 포장면이 잔디 옆을 지나 오른쪽 헬기 스텝까지 이어진다. 헬기 출입구는 비스듬히 보이며 기체 대부분은 오른쪽 밖으로 잘려 윤성찬의 전신을 압도하지 않는다. 탑과 콘크리트 난간은 장소 사진에 가깝지만 구조 참조의 다층 사각 확성기 탑과 금속 난간 형태는 따르지 않았다.",
        "entities": "윤성찬은 참조와 유사한 회색 머리, 안경, 주름진 얼굴을 가진 고령 동아시아계 남성이다. 다만 회색 정장과 밝은 회청색 셔츠는 참조의 어두운 재킷과 짙은 줄무늬 셔츠에서 벗어난다. 현우는 전경에 머리와 피 묻은 상체가 보이는 남성으로, 얼굴과 정확한 나이는 확인하기 어렵다. 착륙 헬기 한 대와 배경 헬기 세 대, 잔디에 쓰러진 관절형 전투 로봇 두 무더기, 폭격 흔적이 보인다. 읽을 수 있는 문구는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "윤성찬은 한 발을 바닥에, 다른 발을 헬기 스텝에 놓고 한 손으로 출입구의 수직 손잡이를 잡아 안정적으로 지지된다. 현우의 하체는 화면 밖이지만 상체를 낮춘 자세에 공중 부양을 시사하는 부분은 없다. 헬기는 바퀴로 옥상에 착륙해 있고 로봇과 상자들은 지면에 놓여 있다. 배경 헬기에는 비행을 지지하는 로터가 보인다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "한 발을 스텝에 올린 전신과 현우를 향한 시선은 맞지만, 비스듬한 냉소보다 활짝 웃는 표정이며 고정 구조물의 확성기 탑도 구조 참조와 다릅니다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "부상당한 현우 너머로 윤성찬의 전신과 내려다보는 냉소를 연결한 구도가 더 정확하지만, 밝은 셔츠와 확성기 탑의 형태는 참조와 다릅니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "윤성찬은 고개와 눈을 화면 왼쪽 아래의 피 흘리는 현우 쪽으로 돌리고 웃는다. 현우는 벽에 기대어 얼굴을 윤성찬 쪽으로 향한다. 요구된 시선 관계는 성립하지만 윤성찬의 표정은 한쪽 입꼬리만 올린 냉소보다 이를 드러낸 큰 웃음에 가깝다. 배경 헬기들의 구체적인 이동 방향은 정지 화면만으로 확정하기 어렵다.",
        "built_space": "왼쪽에 환기구가 달린 문 하나와 벽등 하나, 중앙 잔디에 원형 접시 안테나 탑 하나와 좌우 나팔형 확성기 두 개, 뒤쪽에 콘크리트 난간이 보인다. 오른쪽 헬기는 출입구와 스텝이 비스듬히 보이고 기체 대부분이 화면 밖에 있다. 현우는 왼쪽 벽 아래, 윤성찬은 헬기와 잔디 경계에 있어 연속된 옥상 공간은 성립한다. 다만 탑은 장소 사진의 접시형을 따랐으며, 우선권이 있는 구조 사진의 다층 사각 확성기 탑과 주변 독립 확성기 배치를 재현하지 않았다.",
        "entities": "사람은 윤성찬과 숏 본문에 명시된 현우 두 명이다. 윤성찬은 회색 머리, 안경, 주름진 얼굴의 고령 동아시아계 남성으로 인물 참조와 대체로 부합하고, 어두운 재킷과 셔츠도 가깝다. 현우는 피 묻은 밝은 셔츠를 입은 젊은 동아시아계 남성이다. 착륙 헬기 한 대와 하늘의 헬기 네 대, 훼손된 잔디와 기계 잔해가 보이나 잔해가 전투 로봇임은 뚜렷하지 않다. 읽을 수 있는 문구는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "윤성찬의 한쪽 구두는 헬기 스텝에, 다른 구두는 옥상 바닥에 닿아 체중을 지지한다. 현우는 엉덩이와 다리를 지면에 두고 등을 벽에 기대며 손도 바닥에 댄다. 착륙 헬기는 바퀴로 지지되고 배경 헬기들은 로터가 있는 비행체다. 기계 잔해와 혈흔은 지면 위에 있으며 지지 없이 떠 있는 신체나 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "윤성찬은 왼쪽 전경의 현우를 향해 고개를 숙이고 눈을 내리깔며 입꼬리를 미세하게 올린다. 현우의 뒷머리와 몸통은 윤성찬 쪽을 향하고 있으나 눈은 보이지 않는다. 높은 위치의 윤성찬이 낮은 현우를 내려다보는 관계와 거만한 표정이 읽힌다. 배경 헬기의 특정 표적이나 이동 방향은 확인되지 않는다.",
        "built_space": "왼쪽 옥상 출입 구조물에 문 하나, 환기구 하나, 벽등 하나가 있고 중앙 뒤쪽에는 접시 안테나 탑 하나와 좌우 확성기 두 개가 있다. 현우가 있는 전경의 포장면이 잔디 옆을 지나 오른쪽 헬기 스텝까지 이어진다. 헬기 출입구는 비스듬히 보이며 기체 대부분은 오른쪽 밖으로 잘려 윤성찬의 전신을 압도하지 않는다. 탑과 콘크리트 난간은 장소 사진에 가깝지만 구조 참조의 다층 사각 확성기 탑과 금속 난간 형태는 따르지 않았다.",
        "entities": "윤성찬은 참조와 유사한 회색 머리, 안경, 주름진 얼굴을 가진 고령 동아시아계 남성이다. 다만 회색 정장과 밝은 회청색 셔츠는 참조의 어두운 재킷과 짙은 줄무늬 셔츠에서 벗어난다. 현우는 전경에 머리와 피 묻은 상체가 보이는 남성으로, 얼굴과 정확한 나이는 확인하기 어렵다. 착륙 헬기 한 대와 배경 헬기 세 대, 잔디에 쓰러진 관절형 전투 로봇 두 무더기, 폭격 흔적이 보인다. 읽을 수 있는 문구는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "윤성찬은 한 발을 바닥에, 다른 발을 헬기 스텝에 놓고 한 손으로 출입구의 수직 손잡이를 잡아 안정적으로 지지된다. 현우의 하체는 화면 밖이지만 상체를 낮춘 자세에 공중 부양을 시사하는 부분은 없다. 헬기는 바퀴로 옥상에 착륙해 있고 로봇과 상자들은 지면에 놓여 있다. 배경 헬기에는 비행을 지지하는 로터가 보인다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.571
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.571
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1571
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "부상당한 현우를 전경에 배치하여 헬기 스텝까지 이어지는 깊이 있는 구도를 정확히 구현했으며, 윤성찬의 거만한 미소와 포즈가 프롬프트와 완벽히 일치합니다."
   },
   {
    "label": "B",
    "score": 1571,
    "verdict_ko": "윤성찬의 인물 표현과 헬기 탑승 포즈는 무난하나, 현우가 전경이 아닌 좌측 벽면에 배치되어 지시된 카메라 구도와 공간적 깊이감을 살리지 못했습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its spatial layout, surroundings, fixed features, time of day and lighting mood are spatial truth; stage the moment inside this place. If a STRUCTURE LOOK photograph is also attached, that photo wins for the fixed structure itself — this photograph wins for everything around it. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L278B01.png",
    "asset_id": "6eab84a7-fe74-47c0-bd42-78a5d724b6a9",
    "role": "location_plate"
   },
   {
    "label": "STRUCTURE LOOK — the confirmed photograph of the fixed structure at this location: wherever the structure appears in the frame, its shape, proportions, materials, colors and openings are LOCKED to this photo. Never copy its camera framing, time of day or lighting — the shot text and the LOCATION PHOTOGRAPH are the authorities for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_island_research_complex_sel.png",
    "asset_id": "ae92432a-4956-46ca-9845-8c2d71d637c5",
    "role": "structure_seed_look"
   },
   {
    "label": "CHARACTER REFERENCE — 윤성찬: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1453233>",
    "asset_id": "04d34665-3829-49fb-a6d2-25b9fd051d63",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92d-169f-78b1-8953-be79ebe6d1d3",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S89sh42__bgfirst_bg.png",
   "bg_asset_id": "99219a2e-c866-4ccd-9da5-2b903d22263a",
   "bg_record_key": "S89sh42::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate",
   "seed_attached": true
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  },
  "lane_policy": "ab_select_ready"
 },
 "S89sh42::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T09:31:18.003093+00:00",
  "fingerprint": "a251d212b98acb61ece15651ac53c8e56972dee53fb72580f0f37987b4a8b749",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S89sh42_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S89sh42_sel.png",
  "source_sha256": "1764d67bd380feb06a33e7cf0572ffb17da13dd2a2394cef8f9317a6c5ea0a6d",
  "file": "S89sh42_cine.png",
  "staged_sha256": "3103947d9be0f26830e963d786ba9d4a07b796a1fb10ba7611f325fdc24298a0",
  "latency_ms": 9798
 },
 "S89sh53::signage": {
  "fp": "4cc60383d8a181c7",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S89sh53": {
  "input_fingerprint": "6e4e27971291368c",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 가슴 링에서 하늘을 향해 거대하고 눈부신 새하얀 광선 기둥이 수직으로 팽팽하게 뻗어 나간 압도적인 폭발 찰나.\n\nLOCATION (lock): On the research facility's exposed rooftop lawn near its large antenna and all-direction speakers, beneath the vertical energy beam. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Vertical white beam continuing directly upward from Charlie's chest ring in the upper-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Rooftop park lawn (Beneath the figures during the beam's eruption) — A shallow strip remains visible across the bottom of the composition; used as Anchors the extraordinary vertical event to the established rooftop location; Sky above the rooftop (Visible around the ascending beam); used as Provides lateral negative space that makes the beam's vertical extension legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The brilliant white chest beam overwhelms the daylight locally, with controlled highlight bloom retaining its origin and the figures at its base.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie's chest ring is spinning at extreme speed and projecting a brilliant, nearly star-white beam vertically into the sky; the experimental sensors remain attached. The battered rooftop, landed command helicopter and surrounding aircraft remain in place as lights surge and dim and the control-room energy graph shoots upward.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 가슴 링에서 하늘을 향해 거대하고 눈부신 새하얀 광선 기둥이 수직으로 팽팽하게 뻗어 나간 압도적인 폭발 찰나.\n\nLOCATION (lock): On the research facility's exposed rooftop lawn near its large antenna and all-direction speakers, beneath the vertical energy beam. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Vertical white beam continuing directly upward from Charlie's chest ring in the upper-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Rooftop park lawn (Beneath the figures during the beam's eruption) — A shallow strip remains visible across the bottom of the composition; used as Anchors the extraordinary vertical event to the established rooftop location; Sky above the rooftop (Visible around the ascending beam); used as Provides lateral negative space that makes the beam's vertical extension legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The brilliant white chest beam overwhelms the daylight locally, with controlled highlight bloom retaining its origin and the figures at its base.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie's chest ring is spinning at extreme speed and projecting a brilliant, nearly star-white beam vertically into the sky; the experimental sensors remain attached. The battered rooftop, landed command helicopter and surrounding aircraft remain in place as lights surge and dim and the control-room energy graph shoots upward.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 가슴 링에서 하늘을 향해 거대하고 눈부신 새하얀 광선 기둥이 수직으로 팽팽하게 뻗어 나간 압도적인 폭발 찰나.\n\nLOCATION (lock): On the research facility's exposed rooftop lawn near its large antenna and all-direction speakers, beneath the vertical energy beam. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Vertical white beam continuing directly upward from Charlie's chest ring in the upper-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Rooftop park lawn (Beneath the figures during the beam's eruption) — A shallow strip remains visible across the bottom of the composition; used as Anchors the extraordinary vertical event to the established rooftop location; Sky above the rooftop (Visible around the ascending beam); used as Provides lateral negative space that makes the beam's vertical extension legible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The brilliant white chest beam overwhelms the daylight locally, with controlled highlight bloom retaining its origin and the figures at its base.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie's chest ring is spinning at extreme speed and projecting a brilliant, nearly star-white beam vertically into the sky; the experimental sensors remain attached. The battered rooftop, landed command helicopter and surrounding aircraft remain in place as lights surge and dim and the control-room energy graph shoots upward.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "광선 기둥이 가슴이 아닌 목 부위에서 위로 뻗어 나감.",
    "built_space": "옥상 잔디밭과 안테나 구조물, 배경 헬기 등이 참조와 일치함.",
    "entities": "찰리의 머리가 누락되었으며 가슴 링이 묘사되지 않음.",
    "hard_violations": [
     "[gemini-pro] physically impossible anatomy (찰리의 머리 누락)"
    ],
    "physics": "찰리가 잔디밭에 양팔을 짚고 안정적으로 서 있음."
   },
   {
    "label": "B",
    "direction": "가슴 링에서 광선 기둥이 수직으로 정확히 뻗어 나감.",
    "built_space": "옥상 잔디밭과 안테나 등 주변 구조물이 참조 공간과 잘 일치함.",
    "entities": "찰리의 외형이 정확하게 묘사되었으나, 샷 텍스트에 없는 두 인물이 바닥에 추가됨.",
    "hard_violations": [
     "[gemini-pro] a person placed where the staging does not put them (프롬프트에 없는 인물들을 찰리 발밑에 임의 배치)",
     "[gpt-high] 찰리만 허용된 장면에 피 묻은 상의의 인물과 군장 차림 인물을 추가했다."
    ],
    "physics": "찰리가 지면을 딛고 서 있으며, 추가된 인물들은 바닥에 엎드려 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "가슴 링에서 뿜어지는 광선과 찰리의 묘사는 훌륭하나, 지시되지 않은 이전 장면의 인물들을 바닥에 임의로 추가한 위반 사항이 있음."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "찰리의 머리가 완전히 누락되었고 광선이 가슴 링이 아닌 목 부위에서 발사되어 핵심 프롬프트를 심각하게 위반함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "광선 기둥이 가슴이 아닌 목 부위에서 위로 뻗어 나감.",
        "built_space": "옥상 잔디밭과 안테나 구조물, 배경 헬기 등이 참조와 일치함.",
        "entities": "찰리의 머리가 누락되었으며 가슴 링이 묘사되지 않음.",
        "hard_violations": [
         "physically impossible anatomy (찰리의 머리 누락)"
        ],
        "physics": "찰리가 잔디밭에 양팔을 짚고 안정적으로 서 있음."
       },
       {
        "label": "B",
        "direction": "가슴 링에서 광선 기둥이 수직으로 정확히 뻗어 나감.",
        "built_space": "옥상 잔디밭과 안테나 등 주변 구조물이 참조 공간과 잘 일치함.",
        "entities": "찰리의 외형이 정확하게 묘사되었으나, 샷 텍스트에 없는 두 인물이 바닥에 추가됨.",
        "hard_violations": [
         "a person placed where the staging does not put them (프롬프트에 없는 인물들을 찰리 발밑에 임의 배치)"
        ],
        "physics": "찰리가 지면을 딛고 서 있으며, 추가된 인물들은 바닥에 엎드려 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "가슴 링에서 뿜어지는 광선과 찰리의 묘사는 훌륭하나, 지시되지 않은 이전 장면의 인물들을 바닥에 임의로 추가한 위반 사항이 있음."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "찰리의 머리가 완전히 누락되었고 광선이 가슴 링이 아닌 목 부위에서 발사되어 핵심 프롬프트를 심각하게 위반함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "광선 기둥이 가슴이 아닌 목 부위에서 위로 뻗어 나감.",
        "built_space": "옥상 잔디밭과 안테나 구조물, 배경 헬기 등이 참조와 일치함.",
        "entities": "찰리의 머리가 누락되었으며 가슴 링이 묘사되지 않음.",
        "hard_violations": [
         "physically impossible anatomy (찰리의 머리 누락)"
        ],
        "physics": "찰리가 잔디밭에 양팔을 짚고 안정적으로 서 있음."
       },
       {
        "label": "B",
        "direction": "가슴 링에서 광선 기둥이 수직으로 정확히 뻗어 나감.",
        "built_space": "옥상 잔디밭과 안테나 등 주변 구조물이 참조 공간과 잘 일치함.",
        "entities": "찰리의 외형이 정확하게 묘사되었으나, 샷 텍스트에 없는 두 인물이 바닥에 추가됨.",
        "hard_violations": [
         "a person placed where the staging does not put them (프롬프트에 없는 인물들을 찰리 발밑에 임의 배치)"
        ],
        "physics": "찰리가 지면을 딛고 서 있으며, 추가된 인물들은 바닥에 엎드려 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "가슴 링에서 하늘로 솟는 수직 광선은 맞지만, 금지된 추가 인물들을 잔디에 배치했고 안테나를 전경에서 크게 보여 실격이다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "추가 인물 없이 얕은 잔디 띠와 넓은 하늘을 살려 상대적으로 낫지만, 광선이 가슴 링이 아니라 목·머리 자리에서 시작해 핵심 발사 지점을 틀렸다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "밝은 원형 가슴 링에서 백색 광선이 화면 위쪽 하늘로 수직 연결된다. 찰리의 얼굴은 약간 화면 오른쪽을 향하며, 특정 대상을 바라보는지는 불명확하다. 늘어진 두 손은 잔디 쪽을 향한다.",
        "built_space": "왼쪽 전경에 접시 안테나 한 개, 좌우로 향한 나팔형 스피커 두 개와 이를 지탱하는 철탑 하나가 보인다. 뒤에는 낮은 난간, 오른쪽 옥탑 구조물 하나, 착륙 헬리콥터 한 대와 멀리 있는 항공기들이 있다. 잔디는 하단 띠로 남지만 안테나가 화면 높이 대부분을 차지해 작은 배경 설비로 두라는 지시와 다르다. 찰리는 잔디 중앙에 서 있고 그 발 앞에 추가 인물들이 놓여 있다.",
        "entities": "찰리의 샌드 베이지 각진 장갑, 긴 육중한 팔, 짧은 다리와 흰 기계식 마스크 얼굴은 참조와 대체로 맞는다. 가슴의 원형 발광 링은 분명하며 회전처럼 보이는 동심원 잔상이 있다. 실험 센서는 별도 장치로 명확히 식별되지 않는다. 그러나 피 묻은 밝은 상의를 입은 검은 머리 인물 한 명과 녹색 군장 차림 인물이 최소 한 명 더 보인다. 이들은 허용된 찰리가 아니며, 얼굴이 가려져 정확한 나이와 민족성은 확인할 수 없다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "찰리만 허용된 장면에 피 묻은 상의의 인물과 군장 차림 인물을 추가했다."
        ],
        "physics": "찰리의 양발은 잔디에 닿고 무릎이 굽혀져 체중을 받는다. 길게 내려온 팔과 손은 몸에 정상적으로 연결되어 있다. 누운 추가 인물들도 잔디에 지지된다. 착륙 헬리콥터는 바퀴로 지면을 딛고, 공중 헬리콥터에는 회전익이 보인다. 광선은 허용된 에너지 현상으로 가슴 링에 연결되어 있으며 지지 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "백색 기둥은 화면 중앙에서 하늘로 수직 상승하지만, 보이는 시작점은 가슴보다 높은 목·머리 자리다. 가슴 장갑은 그 아래에 그대로 보여 가슴 링에서 직접 발사되는 연결이 성립하지 않는다. 얼굴은 광량에 묻혀 시선 방향을 확인할 수 없고 두 손은 아래쪽 잔디를 향한다.",
        "built_space": "왼쪽 중경에 접시 안테나 한 개, 좌우로 향한 스피커 두 개, 철탑 하나와 직사각형 설비함 하나가 있다. 낮은 난간 너머로 산과 도시가 보이고 양쪽 가장자리에 헬리콥터가 부분적으로 나타난다. 잔디는 화면 맨 아래의 얕은 띠이며 하늘이 대부분을 차지해 요청된 공간 배분에 더 가깝다. 찰리는 중앙 잔디 위에 있지만 광선의 시작점은 지정된 상부 중앙보다 낮다.",
        "entities": "보이는 등장 주체는 찰리 하나다. 베이지 장갑판, 큰 어깨, 긴 팔과 짧은 다리는 참조의 체형에 가깝다. 흰 마스크 얼굴은 빛에 가려 확인할 수 없으며, 핵심 소품인 가슴 링도 식별되지 않는다. 가슴에는 참조와 유사한 장갑판이 남아 있다. 부착된 실험 센서는 확실히 구별되지 않는다. 안테나와 스피커, 헬리콥터, 훼손된 옥상 재질은 보이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "찰리의 하체 끝은 잔디에 일부 가려지지만 양다리가 지면까지 이어져 서 있는 체중 지지가 성립한다. 팔과 손은 몸에 연결되어 자연스럽게 내려와 있다. 철탑과 설비함은 옥상에 고정되어 있다. 가장자리 헬리콥터의 하부 접점은 난간에 가려져 있어 부유한다고 판단할 근거는 없다. 광선의 문제는 무지지 부유가 아니라 가슴 링과의 발사 연결이 보이지 않는 점이다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "가슴 링에서 하늘로 솟는 수직 광선은 맞지만, 금지된 추가 인물들을 잔디에 배치했고 안테나를 전경에서 크게 보여 실격이다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "추가 인물 없이 얕은 잔디 띠와 넓은 하늘을 살려 상대적으로 낫지만, 광선이 가슴 링이 아니라 목·머리 자리에서 시작해 핵심 발사 지점을 틀렸다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "밝은 원형 가슴 링에서 백색 광선이 화면 위쪽 하늘로 수직 연결된다. 찰리의 얼굴은 약간 화면 오른쪽을 향하며, 특정 대상을 바라보는지는 불명확하다. 늘어진 두 손은 잔디 쪽을 향한다.",
        "built_space": "왼쪽 전경에 접시 안테나 한 개, 좌우로 향한 나팔형 스피커 두 개와 이를 지탱하는 철탑 하나가 보인다. 뒤에는 낮은 난간, 오른쪽 옥탑 구조물 하나, 착륙 헬리콥터 한 대와 멀리 있는 항공기들이 있다. 잔디는 하단 띠로 남지만 안테나가 화면 높이 대부분을 차지해 작은 배경 설비로 두라는 지시와 다르다. 찰리는 잔디 중앙에 서 있고 그 발 앞에 추가 인물들이 놓여 있다.",
        "entities": "찰리의 샌드 베이지 각진 장갑, 긴 육중한 팔, 짧은 다리와 흰 기계식 마스크 얼굴은 참조와 대체로 맞는다. 가슴의 원형 발광 링은 분명하며 회전처럼 보이는 동심원 잔상이 있다. 실험 센서는 별도 장치로 명확히 식별되지 않는다. 그러나 피 묻은 밝은 상의를 입은 검은 머리 인물 한 명과 녹색 군장 차림 인물이 최소 한 명 더 보인다. 이들은 허용된 찰리가 아니며, 얼굴이 가려져 정확한 나이와 민족성은 확인할 수 없다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "찰리만 허용된 장면에 피 묻은 상의의 인물과 군장 차림 인물을 추가했다."
        ],
        "physics": "찰리의 양발은 잔디에 닿고 무릎이 굽혀져 체중을 받는다. 길게 내려온 팔과 손은 몸에 정상적으로 연결되어 있다. 누운 추가 인물들도 잔디에 지지된다. 착륙 헬리콥터는 바퀴로 지면을 딛고, 공중 헬리콥터에는 회전익이 보인다. 광선은 허용된 에너지 현상으로 가슴 링에 연결되어 있으며 지지 없이 떠 있는 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "백색 기둥은 화면 중앙에서 하늘로 수직 상승하지만, 보이는 시작점은 가슴보다 높은 목·머리 자리다. 가슴 장갑은 그 아래에 그대로 보여 가슴 링에서 직접 발사되는 연결이 성립하지 않는다. 얼굴은 광량에 묻혀 시선 방향을 확인할 수 없고 두 손은 아래쪽 잔디를 향한다.",
        "built_space": "왼쪽 중경에 접시 안테나 한 개, 좌우로 향한 스피커 두 개, 철탑 하나와 직사각형 설비함 하나가 있다. 낮은 난간 너머로 산과 도시가 보이고 양쪽 가장자리에 헬리콥터가 부분적으로 나타난다. 잔디는 화면 맨 아래의 얕은 띠이며 하늘이 대부분을 차지해 요청된 공간 배분에 더 가깝다. 찰리는 중앙 잔디 위에 있지만 광선의 시작점은 지정된 상부 중앙보다 낮다.",
        "entities": "보이는 등장 주체는 찰리 하나다. 베이지 장갑판, 큰 어깨, 긴 팔과 짧은 다리는 참조의 체형에 가깝다. 흰 마스크 얼굴은 빛에 가려 확인할 수 없으며, 핵심 소품인 가슴 링도 식별되지 않는다. 가슴에는 참조와 유사한 장갑판이 남아 있다. 부착된 실험 센서는 확실히 구별되지 않는다. 안테나와 스피커, 헬리콥터, 훼손된 옥상 재질은 보이며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "찰리의 하체 끝은 잔디에 일부 가려지지만 양다리가 지면까지 이어져 서 있는 체중 지지가 성립한다. 팔과 손은 몸에 연결되어 자연스럽게 내려와 있다. 철탑과 설비함은 옥상에 고정되어 있다. 가장자리 헬리콥터의 하부 접점은 난간에 가려져 있어 부유한다고 판단할 근거는 없다. 광선의 문제는 무지지 부유가 아니라 가슴 링과의 발사 연결이 보이지 않는 점이다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.5,
    "B": 1.4
   },
   "adjusted": {
    "A": 1.25,
    "B": 1.15
   },
   "violations": {
    "A": [
     "[gemini-pro] physically impossible anatomy (찰리의 머리 누락)"
    ],
    "B": [
     "[gemini-pro] a person placed where the staging does not put them (프롬프트에 없는 인물들을 찰리 발밑에 임의 배치)",
     "[gpt-high] 찰리만 허용된 장면에 피 묻은 상의의 인물과 군장 차림 인물을 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "A"
   },
   "agreed": false
  },
  "totals": {
   "B": 1150,
   "A": 1250
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1150,
    "verdict_ko": "가슴 링에서 뿜어지는 광선과 찰리의 묘사는 훌륭하나, 지시되지 않은 이전 장면의 인물들을 바닥에 임의로 추가한 위반 사항이 있음.  ★위반: [gemini-pro] a person placed where the staging does not put them (프롬프트에 없는 인물들을 찰리 발밑에 임의 배치) / [gpt-high] 찰리만 허용된 장면에 피 묻은 상의의 인물과 군장 차림 인물을 추가했다."
   },
   {
    "label": "A",
    "score": 1250,
    "verdict_ko": "찰리의 머리가 완전히 누락되었고 광선이 가슴 링이 아닌 목 부위에서 발사되어 핵심 프롬프트를 심각하게 위반함.  ★위반: [gemini-pro] physically impossible anatomy (찰리의 머리 누락)"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S89sh42_sel.png",
    "asset_id": "ddf46900-5244-4a0c-ac47-40f55c7b29d4",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92d-1a08-79c1-b779-b3269ee1cf5a",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S89sh42"
  }
 },
 "S89sh53::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T12:18:47.630248+00:00",
  "fingerprint": "0a494cb580235719c16cd58c82a53862596803fae9548c78e4800c23f4015458",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S89sh53_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S89sh53_sel.png",
  "source_sha256": "ad341255154e9739c87d3bf5a5bad35f482974935017ac28bfb5f7fd987c1c9d",
  "file": "S89sh53_cine.png",
  "staged_sha256": "048d85848c50f30b3abade8db2b0e126e057ff74353267cbd611b5fc3a409f12",
  "latency_ms": 11066
 },
 "S89sh66::signage": {
  "fp": "fa4dd617f446ced4",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S89sh66::bgfirst_bg": {
  "input_fingerprint": "902b8a5a7deb272f",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리의 차가운 잔해를 양팔로 빈틈없이 끌어안은 채 하늘을 향해 고개를 젖히고 입을 크게 벌린 현우의 오열하는 굳은 전신.\n\nLOCATION (lock): In the devastated outdoor area of the island research facility at night, beside the robot's remaining wreckage after the blast.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Ground outside the research facility (The nighttime location of Charlie's surviving remains); used as Surrounds the full-body embrace with unoccupied space as the camera withdraws.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued nighttime ambient illumination preserves the embrace and open-mouthed profile without introducing a visible or specifically colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리의 차가운 잔해를 양팔로 빈틈없이 끌어안은 채 하늘을 향해 고개를 젖히고 입을 크게 벌린 현우의 오열하는 굳은 전신.\n\nLOCATION (lock): In the devastated outdoor area of the island research facility at night, beside the robot's remaining wreckage after the blast.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Ground outside the research facility (The nighttime location of Charlie's surviving remains); used as Surrounds the full-body embrace with unoccupied space as the camera withdraws.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued nighttime ambient illumination preserves the embrace and open-mouthed profile without introducing a visible or specifically colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S89sh66__bgfirst_bg.png",
  "asset_id": "bf0ba719-2d49-4122-ad48-c1530b3f79c4",
  "input_asset_ids": [
   "7273645e-cb70-462d-a06c-54beb3d21fa8",
   "76fc6513-8b50-4f33-b05a-840ac5bb52b7"
  ]
 },
 "S89sh66": {
  "input_fingerprint": "bf58bb91c89f15a0",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 차가운 잔해를 양팔로 빈틈없이 끌어안은 채 하늘을 향해 고개를 젖히고 입을 크게 벌린 현우의 오열하는 굳은 전신.\n\nLOCATION (lock): In the devastated outdoor area of the island research facility at night, beside the robot's remaining wreckage after the blast. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Ground outside the research facility (The nighttime location of Charlie's surviving remains); used as Surrounds the full-body embrace with unoccupied space as the camera withdraws.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued nighttime ambient illumination preserves the embrace and open-mouthed profile without introducing a visible or specifically colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie's powered-off remains are held tightly against Hyunwoo's body in both of his arms. Only part of Charlie's destroyed body remains, and the source does not specify its orientation within the embrace or the arrangement of any surviving appendages.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The beam has vanished, leaving only a partially surviving torso and charred remnants of Charlie outside the research facility at night. Charlie's remaining power has shut down completely. 현우: He remains wounded from the attack and is sobbing with both arms tightly closed around the remains in his grasp.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 차가운 잔해를 양팔로 빈틈없이 끌어안은 채 하늘을 향해 고개를 젖히고 입을 크게 벌린 현우의 오열하는 굳은 전신.\n\nLOCATION (lock): In the devastated outdoor area of the island research facility at night, beside the robot's remaining wreckage after the blast. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Ground outside the research facility (The nighttime location of Charlie's surviving remains); used as Surrounds the full-body embrace with unoccupied space as the camera withdraws.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued nighttime ambient illumination preserves the embrace and open-mouthed profile without introducing a visible or specifically colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie's powered-off remains are held tightly against Hyunwoo's body in both of his arms. Only part of Charlie's destroyed body remains, and the source does not specify its orientation within the embrace or the arrangement of any surviving appendages.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The beam has vanished, leaving only a partially surviving torso and charred remnants of Charlie outside the research facility at night. Charlie's remaining power has shut down completely. 현우: He remains wounded from the attack and is sobbing with both arms tightly closed around the remains in his grasp.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 차가운 잔해를 양팔로 빈틈없이 끌어안은 채 하늘을 향해 고개를 젖히고 입을 크게 벌린 현우의 오열하는 굳은 전신.\n\nLOCATION (lock): In the devastated outdoor area of the island research facility at night, beside the robot's remaining wreckage after the blast. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Ground outside the research facility (The nighttime location of Charlie's surviving remains); used as Surrounds the full-body embrace with unoccupied space as the camera withdraws.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued nighttime ambient illumination preserves the embrace and open-mouthed profile without introducing a visible or specifically colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie's powered-off remains are held tightly against Hyunwoo's body in both of his arms. Only part of Charlie's destroyed body remains, and the source does not specify its orientation within the embrace or the arrangement of any surviving appendages.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The beam has vanished, leaving only a partially surviving torso and charred remnants of Charlie outside the research facility at night. Charlie's remaining power has shut down completely. 현우: He remains wounded from the attack and is sobbing with both arms tightly closed around the remains in his grasp.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 현우 right now, so 현우's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 현우: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S89sh66__bgfirst_bg.png",
     "asset_id": "bf0ba719-2d49-4122-ad48-c1530b3f79c4",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S89sh66.png",
     "asset_id": "7273645e-cb70-462d-a06c-54beb3d21fa8",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L278B02.png",
     "asset_id": "76fc6513-8b50-4f33-b05a-840ac5bb52b7",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1315284>",
     "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "현우의 고개와 시선이 밤하늘을 향해 완전히 젖혀져 있으며, 입을 크게 벌리고 위를 향해 부르짖는 방향성이 명확하게 보입니다.",
    "built_space": "참조 이미지에 주어진 연구소 야외 옥상의 모습(좌측의 건물 구조물, 중앙의 통신 안테나, 우측의 헬리콥터, 바닥의 타일과 잡초, 파편들)이 와이드 샷의 스케일에 맞춰 정확하고 자연스럽게 배치되었습니다.",
    "entities": "현우의 외모(헝클어진 검은 머리, 앳된 얼굴)가 참조와 일치하며, 껴안고 있는 대상이 샌드 베이지색 장갑판과 내부 전선이 드러난 찰리의 '파괴된 잔해(일부)'로 정확히 묘사되었습니다.",
    "hard_violations": [],
    "physics": "현우가 바닥에 무릎을 꿇고 지면의 지지를 받고 있으며, 그의 두 팔과 무릎이 무거운 로봇 잔해의 하중을 자연스럽게 지탱하여 중력의 법칙을 충실히 따르고 있습니다."
   },
   {
    "label": "B",
    "direction": "현우가 하늘을 향해 고개를 젖히고 입을 벌리고 있습니다.",
    "built_space": "배경의 구조물, 안테나, 헬리콥터 등 지형지물은 참조 이미지의 장소와 일치하게 배치되었습니다.",
    "entities": "현우가 등장하지만, 안고 있는 대상이 '일부만 남은 잔해'라는 프롬프트의 지시와 달리 팔, 다리, 머리가 모두 붙어 있는 크기가 축소된 온전한 찰리의 형태입니다.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 해부학적 구조 및 스테이징: 현우가 로봇 뒤에 서서 껴안고 있으나, 현우의 손이 로봇의 앞가슴을 짚고 있는데 팔이 로봇을 어떻게 감싸고 있는지 구조적으로 성립하지 않음 (마치 팔이 로봇의 몸통을 통과한 것처럼 보임).",
     "[gemini-pro] 지탱되지 않는 물체: 현우 앞의 로봇 찰리가 지면에 닿아 있지도 않고 현우의 팔에 의해 제대로 지지되지도 않은 채 허공에 엉성하게 떠 있음."
    ],
    "physics": "로봇이 지면에 발을 딛고 있지 않으며, 현우의 손이 가슴에 얹혀 있을 뿐 로봇 전체의 하중을 지탱하는 물리적인 구조(지지대)가 전혀 보이지 않아 마치 공중에 붙어 있는 것처럼 보입니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 10,
        "verdict_ko": "프롬프트에 명시된 하늘을 향해 오열하는 현우의 자세와 찰리의 파괴된 '잔해'를 두 팔로 단단히 끌어안은 모습을 완벽하고 사실적으로 구현했습니다."
       },
       {
        "label": "B",
        "score": 0,
        "verdict_ko": "로봇이 파괴된 잔해가 아닌 축소된 온전한 형태로 등장했으며, 인물이 로봇을 껴안은 방식이 물리적으로 불가능한 심각한 구조적 오류가 있습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 고개와 시선이 밤하늘을 향해 완전히 젖혀져 있으며, 입을 크게 벌리고 위를 향해 부르짖는 방향성이 명확하게 보입니다.",
        "built_space": "참조 이미지에 주어진 연구소 야외 옥상의 모습(좌측의 건물 구조물, 중앙의 통신 안테나, 우측의 헬리콥터, 바닥의 타일과 잡초, 파편들)이 와이드 샷의 스케일에 맞춰 정확하고 자연스럽게 배치되었습니다.",
        "entities": "현우의 외모(헝클어진 검은 머리, 앳된 얼굴)가 참조와 일치하며, 껴안고 있는 대상이 샌드 베이지색 장갑판과 내부 전선이 드러난 찰리의 '파괴된 잔해(일부)'로 정확히 묘사되었습니다.",
        "hard_violations": [],
        "physics": "현우가 바닥에 무릎을 꿇고 지면의 지지를 받고 있으며, 그의 두 팔과 무릎이 무거운 로봇 잔해의 하중을 자연스럽게 지탱하여 중력의 법칙을 충실히 따르고 있습니다."
       },
       {
        "label": "B",
        "direction": "현우가 하늘을 향해 고개를 젖히고 입을 벌리고 있습니다.",
        "built_space": "배경의 구조물, 안테나, 헬리콥터 등 지형지물은 참조 이미지의 장소와 일치하게 배치되었습니다.",
        "entities": "현우가 등장하지만, 안고 있는 대상이 '일부만 남은 잔해'라는 프롬프트의 지시와 달리 팔, 다리, 머리가 모두 붙어 있는 크기가 축소된 온전한 찰리의 형태입니다.",
        "hard_violations": [
         "물리적으로 불가능한 해부학적 구조 및 스테이징: 현우가 로봇 뒤에 서서 껴안고 있으나, 현우의 손이 로봇의 앞가슴을 짚고 있는데 팔이 로봇을 어떻게 감싸고 있는지 구조적으로 성립하지 않음 (마치 팔이 로봇의 몸통을 통과한 것처럼 보임).",
         "지탱되지 않는 물체: 현우 앞의 로봇 찰리가 지면에 닿아 있지도 않고 현우의 팔에 의해 제대로 지지되지도 않은 채 허공에 엉성하게 떠 있음."
        ],
        "physics": "로봇이 지면에 발을 딛고 있지 않으며, 현우의 손이 가슴에 얹혀 있을 뿐 로봇 전체의 하중을 지탱하는 물리적인 구조(지지대)가 전혀 보이지 않아 마치 공중에 붙어 있는 것처럼 보입니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 10,
        "verdict_ko": "프롬프트에 명시된 하늘을 향해 오열하는 현우의 자세와 찰리의 파괴된 '잔해'를 두 팔로 단단히 끌어안은 모습을 완벽하고 사실적으로 구현했습니다."
       },
       {
        "label": "B",
        "score": 0,
        "verdict_ko": "로봇이 파괴된 잔해가 아닌 축소된 온전한 형태로 등장했으며, 인물이 로봇을 껴안은 방식이 물리적으로 불가능한 심각한 구조적 오류가 있습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 고개와 시선이 밤하늘을 향해 완전히 젖혀져 있으며, 입을 크게 벌리고 위를 향해 부르짖는 방향성이 명확하게 보입니다.",
        "built_space": "참조 이미지에 주어진 연구소 야외 옥상의 모습(좌측의 건물 구조물, 중앙의 통신 안테나, 우측의 헬리콥터, 바닥의 타일과 잡초, 파편들)이 와이드 샷의 스케일에 맞춰 정확하고 자연스럽게 배치되었습니다.",
        "entities": "현우의 외모(헝클어진 검은 머리, 앳된 얼굴)가 참조와 일치하며, 껴안고 있는 대상이 샌드 베이지색 장갑판과 내부 전선이 드러난 찰리의 '파괴된 잔해(일부)'로 정확히 묘사되었습니다.",
        "hard_violations": [],
        "physics": "현우가 바닥에 무릎을 꿇고 지면의 지지를 받고 있으며, 그의 두 팔과 무릎이 무거운 로봇 잔해의 하중을 자연스럽게 지탱하여 중력의 법칙을 충실히 따르고 있습니다."
       },
       {
        "label": "B",
        "direction": "현우가 하늘을 향해 고개를 젖히고 입을 벌리고 있습니다.",
        "built_space": "배경의 구조물, 안테나, 헬리콥터 등 지형지물은 참조 이미지의 장소와 일치하게 배치되었습니다.",
        "entities": "현우가 등장하지만, 안고 있는 대상이 '일부만 남은 잔해'라는 프롬프트의 지시와 달리 팔, 다리, 머리가 모두 붙어 있는 크기가 축소된 온전한 찰리의 형태입니다.",
        "hard_violations": [
         "물리적으로 불가능한 해부학적 구조 및 스테이징: 현우가 로봇 뒤에 서서 껴안고 있으나, 현우의 손이 로봇의 앞가슴을 짚고 있는데 팔이 로봇을 어떻게 감싸고 있는지 구조적으로 성립하지 않음 (마치 팔이 로봇의 몸통을 통과한 것처럼 보임).",
         "지탱되지 않는 물체: 현우 앞의 로봇 찰리가 지면에 닿아 있지도 않고 현우의 팔에 의해 제대로 지지되지도 않은 채 허공에 엉성하게 떠 있음."
        ],
        "physics": "로봇이 지면에 발을 딛고 있지 않으며, 현우의 손이 가슴에 얹혀 있을 뿐 로봇 전체의 하중을 지탱하는 물리적인 구조(지지대)가 전혀 보이지 않아 마치 공중에 붙어 있는 것처럼 보입니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "전신 와이드숏과 위를 향한 오열은 구현했지만, 찰리가 파괴된 몸통 잔해가 아니라 머리만 없는 거의 온전한 로봇이며 양팔의 밀착 포옹도 약하고 명시된 낮 시간대를 어겼다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "빈 공간을 둔 전신 와이드숏에서 파괴된 몸통을 양팔로 밀착해 안고 하늘을 향해 오열하는 순간이 정확하지만, 낮 시간대와 참조의 잔디·폭발 지면 및 상의 색은 맞지 않는다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우는 턱을 들고 얼굴을 위쪽 하늘로 향한 채 입을 크게 벌린다. 눈은 찡그려져 정확한 시선은 불분명하지만 얼굴 방향은 요구와 맞는다. 두 손은 앞에 있는 찰리의 흉부 장갑을 붙잡는다. 찰리의 머리는 없고 몸통 정면은 카메라를 향한다.",
        "built_space": "왼쪽에 환기구가 달린 문 하나와 그 위 벽등 하나, 중앙 뒤에 접시 안테나 탑 하나, 오른쪽에 헬리콥터 한 대가 보인다. 탑의 스피커 일부는 인물에 가려진다. 낮은 외곽 벽, 도시와 산, 젖은 포장 바닥, 잔디와 파헤쳐진 지면이 참조 장소와 대체로 일치한다. 현우는 중앙 전경의 포장 바닥에 서 있으며 주변 여백도 있지만 인물 묶음이 화면 높이의 상당 부분을 차지한다. 젖은 바닥의 반사는 가능한 형태다. 야간 하늘과 도시 불빛은 뒤쪽 야간 지시에는 부합하나 최상단의 명시적 낮 시간 잠금에는 맞지 않는다.",
        "entities": "보이는 인간은 현우 한 명으로, 앳된 동아시아계 남성 외형과 헝클어진 검은 머리는 참조에 대체로 부합한다. 상체 의상 대부분은 로봇에 가려져 있고 뚜렷한 상처는 확인하기 어렵다. 찰리는 샌드 베이지의 각진 장갑과 긴 팔을 갖췄지만 머리만 없고 양팔·골반·양다리가 거의 모두 남아 있어 일부 몸통과 탄 잔해라는 조건을 크게 벗어난다. 흰 마스크 얼굴은 남아 있지 않다. 추가 인물이나 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "현우의 양발은 바닥에 닿아 있다. 찰리의 몸은 현우의 가슴 앞에 놓이고 두 손이 흉부 장갑을 잡아 지지하며 팔과 다리는 아래로 늘어진다. 지지 없이 떠 있는 물체는 아니다. 다만 거의 온전한 육중한 로봇을 앞가슴의 손잡기만으로 들어 올린 모습이라 무게 전달이 설득력 약하고, 팔 전체로 잔해를 빈틈없이 감싸는 포옹보다 세워 든 모습에 가깝다."
       },
       {
        "label": "B",
        "direction": "현우는 고개를 뒤로 젖혀 얼굴을 위쪽 하늘로 향하고 입을 크게 벌린다. 두 팔과 손은 몸 앞의 찰리 잔해를 향해 모여 이를 가슴에 밀착한다. 머리가 없는 찰리에게 별도의 시선은 없고, 끊어진 케이블은 아래로 처진다.",
        "built_space": "왼쪽 문 하나와 문 환기구 하나, 벽등 하나, 뒤쪽 접시 안테나 탑 하나와 좌우 스피커 두 개, 오른쪽 헬리콥터 한 대가 보인다. 외곽 벽과 도시·산 배경도 참조와 대응한다. 현우는 전경 포장 바닥에 무릎을 꿇고 있으며 전신 주변에 넓은 빈 공간이 확보된다. 다만 참조의 넓은 잔디 및 폭발로 뒤집힌 흙 구역이 대부분 포장 바닥으로 바뀌었다. 젖은 지면의 반사는 공간상 가능하다. 야간 표현이므로 명시된 낮 시간 잠금은 충족하지 못한다.",
        "entities": "인간은 현우 한 명이며 젊은 동아시아계 남성 외형과 헝클어진 검은 머리가 맞는다. 젖고 손상된 회색 반소매 상의는 참조의 남색 상의와 다르고, 얼굴 동일성은 젖힌 각도 때문에 제한적으로만 확인된다. 찰리는 샌드 베이지 장갑, 검게 탄 내부 구조와 끊어진 배선이 남은 부분 몸통으로 표현되어 파괴·전원 차단 상태에 부합한다. 머리와 흰 마스크는 잔존하지 않으며 온전한 팔다리는 보이지 않는다. 추가 사람이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우는 양 무릎과 접힌 하체로 바닥에 체중을 싣는다. 한 팔은 잔해 아래쪽을 감싸고 다른 팔과 손은 윗몸통을 가슴 쪽으로 누르며, 잔해 하단도 두 무릎 사이 지면 가까이에 놓인다. 로봇 잔해의 무게가 팔과 몸, 하단 접촉으로 지지되는 구성이고 케이블은 중력 방향으로 늘어진다. 지지 없이 뜬 신체나 물체는 보이지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "전신 와이드숏과 위를 향한 오열은 구현했지만, 찰리가 파괴된 몸통 잔해가 아니라 머리만 없는 거의 온전한 로봇이며 양팔의 밀착 포옹도 약하고 명시된 낮 시간대를 어겼다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "빈 공간을 둔 전신 와이드숏에서 파괴된 몸통을 양팔로 밀착해 안고 하늘을 향해 오열하는 순간이 정확하지만, 낮 시간대와 참조의 잔디·폭발 지면 및 상의 색은 맞지 않는다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "현우는 턱을 들고 얼굴을 위쪽 하늘로 향한 채 입을 크게 벌린다. 눈은 찡그려져 정확한 시선은 불분명하지만 얼굴 방향은 요구와 맞는다. 두 손은 앞에 있는 찰리의 흉부 장갑을 붙잡는다. 찰리의 머리는 없고 몸통 정면은 카메라를 향한다.",
        "built_space": "왼쪽에 환기구가 달린 문 하나와 그 위 벽등 하나, 중앙 뒤에 접시 안테나 탑 하나, 오른쪽에 헬리콥터 한 대가 보인다. 탑의 스피커 일부는 인물에 가려진다. 낮은 외곽 벽, 도시와 산, 젖은 포장 바닥, 잔디와 파헤쳐진 지면이 참조 장소와 대체로 일치한다. 현우는 중앙 전경의 포장 바닥에 서 있으며 주변 여백도 있지만 인물 묶음이 화면 높이의 상당 부분을 차지한다. 젖은 바닥의 반사는 가능한 형태다. 야간 하늘과 도시 불빛은 뒤쪽 야간 지시에는 부합하나 최상단의 명시적 낮 시간 잠금에는 맞지 않는다.",
        "entities": "보이는 인간은 현우 한 명으로, 앳된 동아시아계 남성 외형과 헝클어진 검은 머리는 참조에 대체로 부합한다. 상체 의상 대부분은 로봇에 가려져 있고 뚜렷한 상처는 확인하기 어렵다. 찰리는 샌드 베이지의 각진 장갑과 긴 팔을 갖췄지만 머리만 없고 양팔·골반·양다리가 거의 모두 남아 있어 일부 몸통과 탄 잔해라는 조건을 크게 벗어난다. 흰 마스크 얼굴은 남아 있지 않다. 추가 인물이나 읽을 수 있는 문구는 보이지 않는다.",
        "hard_violations": [],
        "physics": "현우의 양발은 바닥에 닿아 있다. 찰리의 몸은 현우의 가슴 앞에 놓이고 두 손이 흉부 장갑을 잡아 지지하며 팔과 다리는 아래로 늘어진다. 지지 없이 떠 있는 물체는 아니다. 다만 거의 온전한 육중한 로봇을 앞가슴의 손잡기만으로 들어 올린 모습이라 무게 전달이 설득력 약하고, 팔 전체로 잔해를 빈틈없이 감싸는 포옹보다 세워 든 모습에 가깝다."
       },
       {
        "label": "A",
        "direction": "현우는 고개를 뒤로 젖혀 얼굴을 위쪽 하늘로 향하고 입을 크게 벌린다. 두 팔과 손은 몸 앞의 찰리 잔해를 향해 모여 이를 가슴에 밀착한다. 머리가 없는 찰리에게 별도의 시선은 없고, 끊어진 케이블은 아래로 처진다.",
        "built_space": "왼쪽 문 하나와 문 환기구 하나, 벽등 하나, 뒤쪽 접시 안테나 탑 하나와 좌우 스피커 두 개, 오른쪽 헬리콥터 한 대가 보인다. 외곽 벽과 도시·산 배경도 참조와 대응한다. 현우는 전경 포장 바닥에 무릎을 꿇고 있으며 전신 주변에 넓은 빈 공간이 확보된다. 다만 참조의 넓은 잔디 및 폭발로 뒤집힌 흙 구역이 대부분 포장 바닥으로 바뀌었다. 젖은 지면의 반사는 공간상 가능하다. 야간 표현이므로 명시된 낮 시간 잠금은 충족하지 못한다.",
        "entities": "인간은 현우 한 명이며 젊은 동아시아계 남성 외형과 헝클어진 검은 머리가 맞는다. 젖고 손상된 회색 반소매 상의는 참조의 남색 상의와 다르고, 얼굴 동일성은 젖힌 각도 때문에 제한적으로만 확인된다. 찰리는 샌드 베이지 장갑, 검게 탄 내부 구조와 끊어진 배선이 남은 부분 몸통으로 표현되어 파괴·전원 차단 상태에 부합한다. 머리와 흰 마스크는 잔존하지 않으며 온전한 팔다리는 보이지 않는다. 추가 사람이나 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "현우는 양 무릎과 접힌 하체로 바닥에 체중을 싣는다. 한 팔은 잔해 아래쪽을 감싸고 다른 팔과 손은 윗몸통을 가슴 쪽으로 누르며, 잔해 하단도 두 무릎 사이 지면 가까이에 놓인다. 로봇 잔해의 무게가 팔과 몸, 하단 접촉으로 지지되는 구성이고 케이블은 중력 방향으로 늘어진다. 지지 없이 뜬 신체나 물체는 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.571
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.321
   },
   "violations": {
    "B": [
     "[gemini-pro] 물리적으로 불가능한 해부학적 구조 및 스테이징: 현우가 로봇 뒤에 서서 껴안고 있으나, 현우의 손이 로봇의 앞가슴을 짚고 있는데 팔이 로봇을 어떻게 감싸고 있는지 구조적으로 성립하지 않음 (마치 팔이 로봇의 몸통을 통과한 것처럼 보임).",
     "[gemini-pro] 지탱되지 않는 물체: 현우 앞의 로봇 찰리가 지면에 닿아 있지도 않고 현우의 팔에 의해 제대로 지지되지도 않은 채 허공에 엉성하게 떠 있음."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 321
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "프롬프트에 명시된 하늘을 향해 오열하는 현우의 자세와 찰리의 파괴된 '잔해'를 두 팔로 단단히 끌어안은 모습을 완벽하고 사실적으로 구현했습니다."
   },
   {
    "label": "B",
    "score": 321,
    "verdict_ko": "로봇이 파괴된 잔해가 아닌 축소된 온전한 형태로 등장했으며, 인물이 로봇을 껴안은 방식이 물리적으로 불가능한 심각한 구조적 오류가 있습니다.  ★위반: [gemini-pro] 물리적으로 불가능한 해부학적 구조 및 스테이징: 현우가 로봇 뒤에 서서 껴안고 있으나, 현우의 손이 로봇의 앞가슴을 짚고 있는데 팔이 로봇을 어떻게 감싸고 있는지 구조적으로 성립하지 않음 (마치 팔이 로봇의 몸통을 통과한 것처럼 보임). / [gemini-pro] 지탱되지 않는 물체: 현우 앞의 로봇 찰리가 지면에 닿아 있지도 않고 현우의 팔에 의해 제대로 지지되지도 않은 채 허공에 엉성하게 떠 있음."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L278B02.png",
    "asset_id": "76fc6513-8b50-4f33-b05a-840ac5bb52b7",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92d-1bb4-7916-9bf0-90841e8ec35c",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S89sh66__bgfirst_bg.png",
   "bg_asset_id": "bf0ba719-2d49-4122-ad48-c1530b3f79c4",
   "bg_record_key": "S89sh66::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S89sh66::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T09:33:50.919239+00:00",
  "fingerprint": "58047324894ae34885a625ab04348bd5d3dee9b0d7281079f0c9928427a9a64e",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S89sh66_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S89sh66_sel.png",
  "source_sha256": "5afce4dd94abc5fd158575771e075cfcdf2dabd83adb6de3e45194d7330329c4",
  "file": "S89sh66_cine.png",
  "staged_sha256": "08cdc732c10e571536c6c10861e137a8b8396ff73c4689fb7aa26c7a23fd2c35",
  "latency_ms": 9865
 },
 "S90sh4::signage": {
  "fp": "f04fc0d37e0e4402",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::fad8ad38683873bf": {
  "subjects": [],
  "subject_text": "찰리의 부품 보관과 복원이 이루어지는 제주도 연구실\n전체적으로 어두운 넓은 연구실. 대형 모니터 아래 제단처럼 생긴 단상 테이블과 수술대가 있고, 책상에는 조작 버튼이 달려 있다.",
  "identity": "canonical",
  "scope_id": "L282",
  "scope_role": "location_interior",
  "scope_sha": "c219c6865b7d60f4"
 },
 "S90sh4::bgfirst_bg": {
  "input_fingerprint": "7dba88376e5c1b9e",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 현우의 어깨를 한 손으로 부드럽게 감싸 쥔 채 앞으로 작은 데이터 칩을 불쑥 내민 자세의 지소영의 상체.\n\nLOCATION (lock): Beside the altar-like parts table inside a spacious, dark research room, lit by a large display showing deployed energy robots.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Data chip (Held out by 지소영 before the transfer) — Seen obliquely between her fingers, with no invented markings or readable face; used as Links her supportive hand and 현우's response without becoming an oversized foreground object; 현우's seat (Occupied) — A partial side is visible beneath and behind his cropped torso; used as Establishes the seated-to-standing relationship underlying her downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The generally dark laboratory retains soft tonal separation across the hands and faces without assigning an unsupported source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 현우의 어깨를 한 손으로 부드럽게 감싸 쥔 채 앞으로 작은 데이터 칩을 불쑥 내민 자세의 지소영의 상체.\n\nLOCATION (lock): Beside the altar-like parts table inside a spacious, dark research room, lit by a large display showing deployed energy robots.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Data chip (Held out by 지소영 before the transfer) — Seen obliquely between her fingers, with no invented markings or readable face; used as Links her supportive hand and 현우's response without becoming an oversized foreground object; 현우's seat (Occupied) — A partial side is visible beneath and behind his cropped torso; used as Establishes the seated-to-standing relationship underlying her downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The generally dark laboratory retains soft tonal separation across the hands and faces without assigning an unsupported source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S90sh4__bgfirst_bg.png",
  "asset_id": "d20cc5fd-c3cf-45a0-bbb5-940c2dade504",
  "input_asset_ids": [
   "9d655ae0-d34a-4dc7-bb3a-f3c76225f96b",
   "020c6829-306f-4816-9550-7878ddde52fa"
  ]
 },
 "S90sh4": {
  "input_fingerprint": "daccc592816d1254",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 현우의 어깨를 한 손으로 부드럽게 감싸 쥔 채 앞으로 작은 데이터 칩을 불쑥 내민 자세의 지소영의 상체.\n\nLOCATION (lock): Beside the altar-like parts table inside a spacious, dark research room, lit by a large display showing deployed energy robots. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Data chip (Held out by 지소영 before the transfer) — Seen obliquely between her fingers, with no invented markings or readable face; used as Links her supportive hand and 현우's response without becoming an oversized foreground object; 현우's seat (Occupied) — A partial side is visible beneath and behind his cropped torso; used as Establishes the seated-to-standing relationship underlying her downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The generally dark laboratory retains soft tonal separation across the hands and faces without assigning an unsupported source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The large room is dim, with disassembled Charlie components arranged on an altar-like table. The large monitor shows Charlie energy units supplied to locations including Dubai, Europe, China and Africa. 현우: He is seated behind the display table, still covered in wounds. 지소영: She is beside the seated area, holding out the small data chip containing Charlie's stored memories.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 지소영 right now, so 지소영's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 지소영: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 현우의 어깨를 한 손으로 부드럽게 감싸 쥔 채 앞으로 작은 데이터 칩을 불쑥 내민 자세의 지소영의 상체.\n\nLOCATION (lock): Beside the altar-like parts table inside a spacious, dark research room, lit by a large display showing deployed energy robots. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Data chip (Held out by 지소영 before the transfer) — Seen obliquely between her fingers, with no invented markings or readable face; used as Links her supportive hand and 현우's response without becoming an oversized foreground object; 현우's seat (Occupied) — A partial side is visible beneath and behind his cropped torso; used as Establishes the seated-to-standing relationship underlying her downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The generally dark laboratory retains soft tonal separation across the hands and faces without assigning an unsupported source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The large room is dim, with disassembled Charlie components arranged on an altar-like table. The large monitor shows Charlie energy units supplied to locations including Dubai, Europe, China and Africa. 현우: He is seated behind the display table, still covered in wounds. 지소영: She is beside the seated area, holding out the small data chip containing Charlie's stored memories.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 지소영 right now, so 지소영's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 지소영: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 현우의 어깨를 한 손으로 부드럽게 감싸 쥔 채 앞으로 작은 데이터 칩을 불쑥 내민 자세의 지소영의 상체.\n\nLOCATION (lock): Beside the altar-like parts table inside a spacious, dark research room, lit by a large display showing deployed energy robots. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Data chip (Held out by 지소영 before the transfer) — Seen obliquely between her fingers, with no invented markings or readable face; used as Links her supportive hand and 현우's response without becoming an oversized foreground object; 현우's seat (Occupied) — A partial side is visible beneath and behind his cropped torso; used as Establishes the seated-to-standing relationship underlying her downward attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The generally dark laboratory retains soft tonal separation across the hands and faces without assigning an unsupported source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The large room is dim, with disassembled Charlie components arranged on an altar-like table. The large monitor shows Charlie energy units supplied to locations including Dubai, Europe, China and Africa. 현우: He is seated behind the display table, still covered in wounds. 지소영: She is beside the seated area, holding out the small data chip containing Charlie's stored memories.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 지소영 right now, so 지소영's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 지소영: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 지소영 (한국인 여성, 50대 중반, 중년의 얼굴, 깔끔하게 정돈한 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S90sh4__bgfirst_bg.png",
     "asset_id": "d20cc5fd-c3cf-45a0-bbb5-940c2dade504",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S90sh4.png",
     "asset_id": "9d655ae0-d34a-4dc7-bb3a-f3c76225f96b",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 지소영: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1243508>",
     "asset_id": "c7496f13-cfcf-44a5-976d-96c783d20580",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L282B02.png",
     "asset_id": "020c6829-306f-4816-9550-7878ddde52fa",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 지소영: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1243508>",
     "asset_id": "c7496f13-cfcf-44a5-976d-96c783d20580",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "지소영의 시선이 아래쪽의 칩과 현우를 향하고 있으며, 왼손으로 작은 데이터 칩을 정확히 내밀고 있음.",
    "built_space": "어두운 연구실 내에 조립 테이블과 대형 디스플레이가 위치함. 현우는 의자에 앉아 있고 지소영은 그 곁에 서서 올바른 공간적 관계를 형성함.",
    "entities": "지소영은 50대 여성의 외양을 잘 반영하며, 현우는 검은 머리와 목 주변의 상처(반창고)가 확인됨. 데이터 칩 역시 식별 가능한 글씨 없이 묘사됨.",
    "hard_violations": [],
    "physics": "지소영의 오른손이 현우의 왼쪽 어깨를 안정적으로 짚고 있으며, 왼손으로 칩을 들고 있는 자세와 현우가 의자에 앉아 있는 모습 모두 물리적으로 자연스러움."
   },
   {
    "label": "B",
    "direction": "지소영이 현우를 내려다보고 있으며, 칩을 내미는 방향은 현우 쪽을 향함.",
    "built_space": "연구실 배경과 대형 화면, 테이블이 적절히 배치되어 있으며 두 인물의 위치는 샷의 요구와 맞음.",
    "entities": "지소영의 얼굴은 레퍼런스와 일치하나, 현우의 몸에는 지시된 상처가 전혀 보이지 않음.",
    "hard_violations": [
     "[gemini-pro] physically impossible anatomy",
     "[gpt-high] 대형 화면 상단에 식별 가능한 영문 제목이 남아 있어, 읽을 수 있는 글자를 전면 금지한 지시를 위반한다."
    ],
    "physics": "현우의 오른쪽 어깨에 올려진 손의 엄지 방향과 칩을 들고 있는 손의 형태가 모두 오른손의 구조를 띠고 있어 해부학적으로 불가능함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지소영이 어깨를 감싸고 칩을 내미는 자세와 현우의 상처 묘사 등 프롬프트의 요구사항을 충실하고 자연스럽게 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "어깨를 짚은 손과 칩을 든 손이 모두 오른손으로 그려지는 치명적인 해부학적 오류가 있으며 현우의 상처가 누락됨."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "지소영의 시선이 아래쪽의 칩과 현우를 향하고 있으며, 왼손으로 작은 데이터 칩을 정확히 내밀고 있음.",
        "built_space": "어두운 연구실 내에 조립 테이블과 대형 디스플레이가 위치함. 현우는 의자에 앉아 있고 지소영은 그 곁에 서서 올바른 공간적 관계를 형성함.",
        "entities": "지소영은 50대 여성의 외양을 잘 반영하며, 현우는 검은 머리와 목 주변의 상처(반창고)가 확인됨. 데이터 칩 역시 식별 가능한 글씨 없이 묘사됨.",
        "hard_violations": [],
        "physics": "지소영의 오른손이 현우의 왼쪽 어깨를 안정적으로 짚고 있으며, 왼손으로 칩을 들고 있는 자세와 현우가 의자에 앉아 있는 모습 모두 물리적으로 자연스러움."
       },
       {
        "label": "B",
        "direction": "지소영이 현우를 내려다보고 있으며, 칩을 내미는 방향은 현우 쪽을 향함.",
        "built_space": "연구실 배경과 대형 화면, 테이블이 적절히 배치되어 있으며 두 인물의 위치는 샷의 요구와 맞음.",
        "entities": "지소영의 얼굴은 레퍼런스와 일치하나, 현우의 몸에는 지시된 상처가 전혀 보이지 않음.",
        "hard_violations": [
         "physically impossible anatomy"
        ],
        "physics": "현우의 오른쪽 어깨에 올려진 손의 엄지 방향과 칩을 들고 있는 손의 형태가 모두 오른손의 구조를 띠고 있어 해부학적으로 불가능함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "지소영이 어깨를 감싸고 칩을 내미는 자세와 현우의 상처 묘사 등 프롬프트의 요구사항을 충실하고 자연스럽게 구현함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "어깨를 짚은 손과 칩을 든 손이 모두 오른손으로 그려지는 치명적인 해부학적 오류가 있으며 현우의 상처가 누락됨."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "지소영의 시선이 아래쪽의 칩과 현우를 향하고 있으며, 왼손으로 작은 데이터 칩을 정확히 내밀고 있음.",
        "built_space": "어두운 연구실 내에 조립 테이블과 대형 디스플레이가 위치함. 현우는 의자에 앉아 있고 지소영은 그 곁에 서서 올바른 공간적 관계를 형성함.",
        "entities": "지소영은 50대 여성의 외양을 잘 반영하며, 현우는 검은 머리와 목 주변의 상처(반창고)가 확인됨. 데이터 칩 역시 식별 가능한 글씨 없이 묘사됨.",
        "hard_violations": [],
        "physics": "지소영의 오른손이 현우의 왼쪽 어깨를 안정적으로 짚고 있으며, 왼손으로 칩을 들고 있는 자세와 현우가 의자에 앉아 있는 모습 모두 물리적으로 자연스러움."
       },
       {
        "label": "B",
        "direction": "지소영이 현우를 내려다보고 있으며, 칩을 내미는 방향은 현우 쪽을 향함.",
        "built_space": "연구실 배경과 대형 화면, 테이블이 적절히 배치되어 있으며 두 인물의 위치는 샷의 요구와 맞음.",
        "entities": "지소영의 얼굴은 레퍼런스와 일치하나, 현우의 몸에는 지시된 상처가 전혀 보이지 않음.",
        "hard_violations": [
         "physically impossible anatomy"
        ],
        "physics": "현우의 오른쪽 어깨에 올려진 손의 엄지 방향과 칩을 들고 있는 손의 형태가 모두 오른손의 구조를 띠고 있어 해부학적으로 불가능함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "어깨를 감싸며 칩을 내미는 중간 쇼트는 맞지만, 칩의 표식 있는 면을 카메라에 드러내고 배경 화면에도 식별 가능한 영문 제목이 남아 있다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "지소영의 상체, 현우의 어깨를 감싼 손, 비스듬히 내민 작은 칩과 착석 관계를 충실하게 담고 상처와 야간 연구실도 유지한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "지소영은 아래쪽의 현우를 바라보며 한 손을 그의 어깨에 얹고 다른 손의 칩을 그쪽으로 내민다. 현우는 지소영 쪽으로 고개를 돌렸지만 눈은 보이지 않는다. 칩은 전달 대상보다 카메라를 향해 넓은 표면을 드러내므로, 손가락 사이에서 비스듬히 보이라는 지시와 차이가 있다.",
        "built_space": "중앙 금속 부품 테이블 하나, 뒤쪽 대형 벽면 화면 하나, 좌우 장비 수납장과 뒤쪽 작업대가 보인다. 전경 왼쪽의 현우가 앉은 의자와 오른쪽의 빈 의자 일부가 보여 참고 장소의 두 의자 배치와 대체로 맞는다. 지소영은 착석한 현우 오른쪽에 서 있다. 천장 배관과 테이블의 조작 패널도 장소를 유지하며, 불가능한 반사는 보이지 않는다.",
        "entities": "등장인물은 두 명뿐이다. 지소영의 중년 한국인 여성 외관, 정돈된 검은 단발과 남색 상의는 참고와 가깝다. 현우는 검은 헝클어진 머리와 남색 상의를 입었고 얼굴 대부분은 가려져 정확한 얼굴과 상처는 확인하기 어렵다. 작은 칩은 손에 있으나 정면에 어두운 표식이 보인다. 테이블에는 분해된 금속 로봇 부품이 있으며 벽면 화면에는 세계 지도와 공급 현황 형태의 그래픽이 있다. 화면 상단에는 식별 가능한 영문 제목이 남아 있다.",
        "hard_violations": [
         "대형 화면 상단에 식별 가능한 영문 제목이 남아 있어, 읽을 수 있는 글자를 전면 금지한 지시를 위반한다."
        ],
        "physics": "현우의 몸은 의자에 지지되어 있고 지소영의 한 손은 그의 어깨에 실제로 닿는다. 다른 손의 손가락은 칩을 잡고 있으며 손목과 팔의 연결도 자연스럽다. 로봇 부품은 작업대 위에 놓여 있다. 하체가 잘린 지소영에게 부유를 시사하는 정황은 없으며, 지지 없이 떠 있는 물체도 없다."
       },
       {
        "label": "B",
        "direction": "지소영은 앉은 현우의 얼굴을 내려다보고, 현우는 고개를 들어 그녀를 바라본다. 한 손은 현우의 어깨를 부드럽게 감싸고 다른 손은 두 사람 사이로 작은 칩을 내민다. 칩의 면은 카메라에 비스듬하며 전달 전의 방향과 크기가 자연스럽다.",
        "built_space": "뒤쪽 대형 화면 하나, 왼쪽 문 하나, 오른쪽 야경 창과 유리 장비 수납장들이 보인다. 금속 작업대는 뒤쪽과 오른쪽으로 이어져 보이고 그 위에 분해 부품이 놓여 있다. 현우의 의자 등받이와 측면 일부가 잘린 몸통 아래와 뒤에 나타나며, 지소영은 그의 옆에 서서 몸을 돌리고 있다. 참고의 두 번째 의자는 화면에서 확인되지 않지만 이 중간 쇼트에서 반드시 보여야 할 요소는 아니다. 작업대와 부품의 배열은 참고와 조금 다르지만 장소의 주요 재질과 설비는 유지된다.",
        "entities": "두 인물만 등장한다. 지소영의 얼굴, 중년 여성의 연령감, 정돈된 검은 머리와 남색 니트는 참고에 가깝다. 현우는 앳된 동아시아계 남성의 옆얼굴, 헝클어진 검은 머리와 남색 상의를 보이며 뺨과 목의 반창고로 부상 상태가 드러난다. 작은 데이터 칩은 지소영의 손가락 사이에 있고 읽을 수 있는 표식은 없다. 금속 로봇 외장과 관절 부품, 세계 지도 기반 공급 화면이 보이며 창밖은 밤이다.",
        "hard_violations": [],
        "physics": "현우는 의자 등받이 앞에 정상적인 착석 자세로 놓여 있다. 지소영의 어깨 접촉 손과 칩을 집은 손 모두 해당 팔에 자연스럽게 이어지고, 칩은 손가락으로 확실히 지지된다. 상체를 약간 기울여 앉은 사람을 내려다보는 동작도 가능하다. 무거운 로봇 부품은 금속 테이블에 얹혀 있고 떠 있는 인체나 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "어깨를 감싸며 칩을 내미는 중간 쇼트는 맞지만, 칩의 표식 있는 면을 카메라에 드러내고 배경 화면에도 식별 가능한 영문 제목이 남아 있다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "지소영의 상체, 현우의 어깨를 감싼 손, 비스듬히 내민 작은 칩과 착석 관계를 충실하게 담고 상처와 야간 연구실도 유지한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "지소영은 아래쪽의 현우를 바라보며 한 손을 그의 어깨에 얹고 다른 손의 칩을 그쪽으로 내민다. 현우는 지소영 쪽으로 고개를 돌렸지만 눈은 보이지 않는다. 칩은 전달 대상보다 카메라를 향해 넓은 표면을 드러내므로, 손가락 사이에서 비스듬히 보이라는 지시와 차이가 있다.",
        "built_space": "중앙 금속 부품 테이블 하나, 뒤쪽 대형 벽면 화면 하나, 좌우 장비 수납장과 뒤쪽 작업대가 보인다. 전경 왼쪽의 현우가 앉은 의자와 오른쪽의 빈 의자 일부가 보여 참고 장소의 두 의자 배치와 대체로 맞는다. 지소영은 착석한 현우 오른쪽에 서 있다. 천장 배관과 테이블의 조작 패널도 장소를 유지하며, 불가능한 반사는 보이지 않는다.",
        "entities": "등장인물은 두 명뿐이다. 지소영의 중년 한국인 여성 외관, 정돈된 검은 단발과 남색 상의는 참고와 가깝다. 현우는 검은 헝클어진 머리와 남색 상의를 입었고 얼굴 대부분은 가려져 정확한 얼굴과 상처는 확인하기 어렵다. 작은 칩은 손에 있으나 정면에 어두운 표식이 보인다. 테이블에는 분해된 금속 로봇 부품이 있으며 벽면 화면에는 세계 지도와 공급 현황 형태의 그래픽이 있다. 화면 상단에는 식별 가능한 영문 제목이 남아 있다.",
        "hard_violations": [
         "대형 화면 상단에 식별 가능한 영문 제목이 남아 있어, 읽을 수 있는 글자를 전면 금지한 지시를 위반한다."
        ],
        "physics": "현우의 몸은 의자에 지지되어 있고 지소영의 한 손은 그의 어깨에 실제로 닿는다. 다른 손의 손가락은 칩을 잡고 있으며 손목과 팔의 연결도 자연스럽다. 로봇 부품은 작업대 위에 놓여 있다. 하체가 잘린 지소영에게 부유를 시사하는 정황은 없으며, 지지 없이 떠 있는 물체도 없다."
       },
       {
        "label": "A",
        "direction": "지소영은 앉은 현우의 얼굴을 내려다보고, 현우는 고개를 들어 그녀를 바라본다. 한 손은 현우의 어깨를 부드럽게 감싸고 다른 손은 두 사람 사이로 작은 칩을 내민다. 칩의 면은 카메라에 비스듬하며 전달 전의 방향과 크기가 자연스럽다.",
        "built_space": "뒤쪽 대형 화면 하나, 왼쪽 문 하나, 오른쪽 야경 창과 유리 장비 수납장들이 보인다. 금속 작업대는 뒤쪽과 오른쪽으로 이어져 보이고 그 위에 분해 부품이 놓여 있다. 현우의 의자 등받이와 측면 일부가 잘린 몸통 아래와 뒤에 나타나며, 지소영은 그의 옆에 서서 몸을 돌리고 있다. 참고의 두 번째 의자는 화면에서 확인되지 않지만 이 중간 쇼트에서 반드시 보여야 할 요소는 아니다. 작업대와 부품의 배열은 참고와 조금 다르지만 장소의 주요 재질과 설비는 유지된다.",
        "entities": "두 인물만 등장한다. 지소영의 얼굴, 중년 여성의 연령감, 정돈된 검은 머리와 남색 니트는 참고에 가깝다. 현우는 앳된 동아시아계 남성의 옆얼굴, 헝클어진 검은 머리와 남색 상의를 보이며 뺨과 목의 반창고로 부상 상태가 드러난다. 작은 데이터 칩은 지소영의 손가락 사이에 있고 읽을 수 있는 표식은 없다. 금속 로봇 외장과 관절 부품, 세계 지도 기반 공급 화면이 보이며 창밖은 밤이다.",
        "hard_violations": [],
        "physics": "현우는 의자 등받이 앞에 정상적인 착석 자세로 놓여 있다. 지소영의 어깨 접촉 손과 칩을 집은 손 모두 해당 팔에 자연스럽게 이어지고, 칩은 손가락으로 확실히 지지된다. 상체를 약간 기울여 앉은 사람을 내려다보는 동작도 가능하다. 무거운 로봇 부품은 금속 테이블에 얹혀 있고 떠 있는 인체나 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.984
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.734
   },
   "violations": {
    "B": [
     "[gemini-pro] physically impossible anatomy",
     "[gpt-high] 대형 화면 상단에 식별 가능한 영문 제목이 남아 있어, 읽을 수 있는 글자를 전면 금지한 지시를 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 734
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지소영이 어깨를 감싸고 칩을 내미는 자세와 현우의 상처 묘사 등 프롬프트의 요구사항을 충실하고 자연스럽게 구현함."
   },
   {
    "label": "B",
    "score": 734,
    "verdict_ko": "어깨를 짚은 손과 칩을 든 손이 모두 오른손으로 그려지는 치명적인 해부학적 오류가 있으며 현우의 상처가 누락됨.  ★위반: [gemini-pro] physically impossible anatomy / [gpt-high] 대형 화면 상단에 식별 가능한 영문 제목이 남아 있어, 읽을 수 있는 글자를 전면 금지한 지시를 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/episodes/a28fb896-b811-463c-8bcc-ba732968bd7c/images/background_chain/L282B02.png",
    "asset_id": "020c6829-306f-4816-9550-7878ddde52fa",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 지소영: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1243508>",
    "asset_id": "c7496f13-cfcf-44a5-976d-96c783d20580",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92d-1f15-7400-b295-08009833cb42",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S90sh4__bgfirst_bg.png",
   "bg_asset_id": "d20cc5fd-c3cf-45a0-bbb5-940c2dade504",
   "bg_record_key": "S90sh4::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S90sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T09:35:37.456270+00:00",
  "fingerprint": "f77f80a2339bd41c3264ea39d86a3b8905ab4f1fac65164860294496449772c1",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S90sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S90sh4_sel.png",
  "source_sha256": "40ae49eb0bcdf1d5ebed7cc5e0b282deef257b19df698c30a002377ecbf15d1b",
  "file": "S90sh4_cine.png",
  "staged_sha256": "c794cb97f96109c71ca9b1ef9a758647b872a8dd77241c1699691f9b596d2299",
  "latency_ms": 11385
 },
 "S90sh9::signage": {
  "fp": "b7733e579bf3cbc0",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S90sh9": {
  "input_fingerprint": "5844a18b5cc0f8a3",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 연구소 허공에 푸른 제주도 풍경과 함께 찰리의 거대한 홀로그램 영상이 띄워진 찰나.\n\nLOCATION (lock): In the open projection area of the dark restoration laboratory, illuminated by a large island-scene display and a robot hologram. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Large laboratory screen (Displaying the blue Jeju landscape) — The image-bearing face is visible obliquely behind the hologram, with its boundaries retained; used as Supplies the island setting as displayed imagery rather than a literal change of location; the screen occupies less than two-fifths of the frame; 현우's seat (Still occupied before he rises) — A small rear portion appears beside his shoulder at the lower edge; used as Keeps the viewer anchored in the physical laboratory while the projected scene appears.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The blue Jeju imagery and luminous holographic appearance remain distinct against the generally dark laboratory, without implying a blue light source elsewhere in the room.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The display has changed to a Jeju landscape and a holographic Charlie has appeared, while the real disassembled components remain on the table in the dim room. The holographic sequence includes Amber and Raul running toward Charlie.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (홀로그램) (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 연구소 허공에 푸른 제주도 풍경과 함께 찰리의 거대한 홀로그램 영상이 띄워진 찰나.\n\nLOCATION (lock): In the open projection area of the dark restoration laboratory, illuminated by a large island-scene display and a robot hologram. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Large laboratory screen (Displaying the blue Jeju landscape) — The image-bearing face is visible obliquely behind the hologram, with its boundaries retained; used as Supplies the island setting as displayed imagery rather than a literal change of location; the screen occupies less than two-fifths of the frame; 현우's seat (Still occupied before he rises) — A small rear portion appears beside his shoulder at the lower edge; used as Keeps the viewer anchored in the physical laboratory while the projected scene appears.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The blue Jeju imagery and luminous holographic appearance remain distinct against the generally dark laboratory, without implying a blue light source elsewhere in the room.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The display has changed to a Jeju landscape and a holographic Charlie has appeared, while the real disassembled components remain on the table in the dim room. The holographic sequence includes Amber and Raul running toward Charlie.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (홀로그램) (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 연구소 허공에 푸른 제주도 풍경과 함께 찰리의 거대한 홀로그램 영상이 띄워진 찰나.\n\nLOCATION (lock): In the open projection area of the dark restoration laboratory, illuminated by a large island-scene display and a robot hologram. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Large laboratory screen (Displaying the blue Jeju landscape) — The image-bearing face is visible obliquely behind the hologram, with its boundaries retained; used as Supplies the island setting as displayed imagery rather than a literal change of location; the screen occupies less than two-fifths of the frame; 현우's seat (Still occupied before he rises) — A small rear portion appears beside his shoulder at the lower edge; used as Keeps the viewer anchored in the physical laboratory while the projected scene appears.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The blue Jeju imagery and luminous holographic appearance remain distinct against the generally dark laboratory, without implying a blue light source elsewhere in the room.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The display has changed to a Jeju landscape and a holographic Charlie has appeared, while the real disassembled components remain on the table in the dim room. The holographic sequence includes Amber and Raul running toward Charlie.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (홀로그램) (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "찰리의 대형 홀로그램이 정면을 향해 서 있고, 두 명의 작은 홀로그램 인물들이 그를 향해 달리고 있음.",
    "built_space": "레퍼런스와 동일한 구조와 조명을 갖춘 연구실. 중앙 작업대에 분해된 로봇 부품들이 그대로 배치되어 있고, 좌측 뒤편의 모니터, 왼쪽 문, 우측 창문 등 고정 요소가 일치함. 좌측 하단 가장자리에 현우의 의자 일부가 지시대로 나타남.",
    "entities": "찰리 홀로그램은 레퍼런스 이미지의 디자인(고릴라형 몸체, 샌드 베이지 장갑판, 흰 마스크)을 정확히 반영함. 뒤편 모니터에는 프롬프트가 요구한 푸른 제주도 풍경이 표시되어 있음.",
    "hard_violations": [],
    "physics": "찰리 홀로그램의 하단이 테이블을 통과하며 투사되고 있고, 작은 인물들도 허공에 자연스럽게 띄워져 있음. 실제 로봇 부품들은 테이블 표면에 안정적으로 지탱되어 놓여 있음."
   },
   {
    "label": "B",
    "direction": "찰리의 홀로그램이 정면을 향해 서 있고, 두 명의 홀로그램 인물이 그를 향해 달려가고 있음.",
    "built_space": "연구실 공간 구조와 비율이 레퍼런스와 크게 다름. 모니터의 크기가 과도하게 커지고 위치도 벽면 전체로 변경되었으며, 우측 창문의 형태와 구조 등 기존 공간의 고정된 특징이 유지되지 않음.",
    "entities": "찰리의 홀로그램이 제시되었으나, 배경 모니터에는 풍경 대신 지도가 띄워져 있음.",
    "hard_violations": [
     "[gemini-pro] 화면 중앙의 지도에 'Jeju Island'라는 명확히 읽을 수 있는 텍스트가 존재함 (읽을 수 있는 글씨 금지 위반).",
     "[gemini-pro] 홀로그램으로 달리고 있는 두 인물이 이전 샷(레퍼런스)에 등장했던 남녀 작업자의 얼굴과 복장을 그대로 묘사함 (이전 샷 인물 재사용 금지 위반).",
     "[gpt-high] 대형 화면 중앙의 영문 제주 지명이 읽혀, 어디에도 읽을 수 있는 글자를 두지 말라는 조건을 위반한다."
    ],
    "physics": "홀로그램들이 허공에 떠서 투사되고 있으며, 물리적인 부품들은 테이블 위에 놓여 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 10,
        "verdict_ko": "이전 샷의 연구실 환경과 구도를 완벽하게 유지하면서, 지시된 제주도 풍경과 찰리의 홀로그램을 매우 사실적이고 정확하게 구현해 냈습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "화면에 읽을 수 있는 영어 텍스트가 포함되었고, 이전 샷의 인물들을 홀로그램으로 재사용하여 치명적인 지침 위반이 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 대형 홀로그램이 정면을 향해 서 있고, 두 명의 작은 홀로그램 인물들이 그를 향해 달리고 있음.",
        "built_space": "레퍼런스와 동일한 구조와 조명을 갖춘 연구실. 중앙 작업대에 분해된 로봇 부품들이 그대로 배치되어 있고, 좌측 뒤편의 모니터, 왼쪽 문, 우측 창문 등 고정 요소가 일치함. 좌측 하단 가장자리에 현우의 의자 일부가 지시대로 나타남.",
        "entities": "찰리 홀로그램은 레퍼런스 이미지의 디자인(고릴라형 몸체, 샌드 베이지 장갑판, 흰 마스크)을 정확히 반영함. 뒤편 모니터에는 프롬프트가 요구한 푸른 제주도 풍경이 표시되어 있음.",
        "hard_violations": [],
        "physics": "찰리 홀로그램의 하단이 테이블을 통과하며 투사되고 있고, 작은 인물들도 허공에 자연스럽게 띄워져 있음. 실제 로봇 부품들은 테이블 표면에 안정적으로 지탱되어 놓여 있음."
       },
       {
        "label": "B",
        "direction": "찰리의 홀로그램이 정면을 향해 서 있고, 두 명의 홀로그램 인물이 그를 향해 달려가고 있음.",
        "built_space": "연구실 공간 구조와 비율이 레퍼런스와 크게 다름. 모니터의 크기가 과도하게 커지고 위치도 벽면 전체로 변경되었으며, 우측 창문의 형태와 구조 등 기존 공간의 고정된 특징이 유지되지 않음.",
        "entities": "찰리의 홀로그램이 제시되었으나, 배경 모니터에는 풍경 대신 지도가 띄워져 있음.",
        "hard_violations": [
         "화면 중앙의 지도에 'Jeju Island'라는 명확히 읽을 수 있는 텍스트가 존재함 (읽을 수 있는 글씨 금지 위반).",
         "홀로그램으로 달리고 있는 두 인물이 이전 샷(레퍼런스)에 등장했던 남녀 작업자의 얼굴과 복장을 그대로 묘사함 (이전 샷 인물 재사용 금지 위반)."
        ],
        "physics": "홀로그램들이 허공에 떠서 투사되고 있으며, 물리적인 부품들은 테이블 위에 놓여 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 10,
        "verdict_ko": "이전 샷의 연구실 환경과 구도를 완벽하게 유지하면서, 지시된 제주도 풍경과 찰리의 홀로그램을 매우 사실적이고 정확하게 구현해 냈습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "화면에 읽을 수 있는 영어 텍스트가 포함되었고, 이전 샷의 인물들을 홀로그램으로 재사용하여 치명적인 지침 위반이 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 대형 홀로그램이 정면을 향해 서 있고, 두 명의 작은 홀로그램 인물들이 그를 향해 달리고 있음.",
        "built_space": "레퍼런스와 동일한 구조와 조명을 갖춘 연구실. 중앙 작업대에 분해된 로봇 부품들이 그대로 배치되어 있고, 좌측 뒤편의 모니터, 왼쪽 문, 우측 창문 등 고정 요소가 일치함. 좌측 하단 가장자리에 현우의 의자 일부가 지시대로 나타남.",
        "entities": "찰리 홀로그램은 레퍼런스 이미지의 디자인(고릴라형 몸체, 샌드 베이지 장갑판, 흰 마스크)을 정확히 반영함. 뒤편 모니터에는 프롬프트가 요구한 푸른 제주도 풍경이 표시되어 있음.",
        "hard_violations": [],
        "physics": "찰리 홀로그램의 하단이 테이블을 통과하며 투사되고 있고, 작은 인물들도 허공에 자연스럽게 띄워져 있음. 실제 로봇 부품들은 테이블 표면에 안정적으로 지탱되어 놓여 있음."
       },
       {
        "label": "B",
        "direction": "찰리의 홀로그램이 정면을 향해 서 있고, 두 명의 홀로그램 인물이 그를 향해 달려가고 있음.",
        "built_space": "연구실 공간 구조와 비율이 레퍼런스와 크게 다름. 모니터의 크기가 과도하게 커지고 위치도 벽면 전체로 변경되었으며, 우측 창문의 형태와 구조 등 기존 공간의 고정된 특징이 유지되지 않음.",
        "entities": "찰리의 홀로그램이 제시되었으나, 배경 모니터에는 풍경 대신 지도가 띄워져 있음.",
        "hard_violations": [
         "화면 중앙의 지도에 'Jeju Island'라는 명확히 읽을 수 있는 텍스트가 존재함 (읽을 수 있는 글씨 금지 위반).",
         "홀로그램으로 달리고 있는 두 인물이 이전 샷(레퍼런스)에 등장했던 남녀 작업자의 얼굴과 복장을 그대로 묘사함 (이전 샷 인물 재사용 금지 위반)."
        ],
        "physics": "홀로그램들이 허공에 떠서 투사되고 있으며, 물리적인 부품들은 테이블 위에 놓여 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "읽을 수 있는 화면 글자가 금지 조건을 위반하며, 제주 풍경을 지도로 바꾸고 달리는 인물들의 방향도 찰리에게 향하지 않는다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "원래 연구소의 구조와 제주 풍경, 거대한 찰리 및 그에게 달려가는 투영을 잘 재현하지만, 하단 의자가 비어 있고 화면의 사선 배치가 약하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 얼굴은 화면 왼쪽 앞을 향하며 특정 상대를 바라보는지는 불명확하다. 왼쪽의 작은 두 인물은 카메라 쪽으로 달리는 자세여서 오른쪽 찰리를 향하는 동선으로 읽히지 않는다. 제주 화면의 영상 면은 카메라에 비스듬히 보인다.",
        "built_space": "왼쪽에 경계가 보이는 대형 화면 하나, 중앙에 긴 금속 작업대 하나, 오른쪽 아래에 빈 의자 하나가 보인다. 뒤쪽에는 어두운 직사각형 면과 장비 수납장이 있다. 화면은 전체 면적의 5분의 2 미만이지만, 참고의 후면 화면과 오른쪽 항구 창문을 중심으로 한 공간 배치는 잘 보존되지 않았다. 의자는 작은 등받이 일부가 아니라 좌판과 팔걸이까지 드러나며 옆에 현우의 어깨가 없다.",
        "entities": "찰리 한 개체는 흰 마스크형 얼굴, 각진 장갑, 긴 팔과 짧은 다리를 갖춰 참고의 로봇 정체성과 대체로 맞지만 베이지색이 푸른 백색으로 씻겨 있다. 화면에는 제주 자연 풍경이 아니라 섬 지도와 읽을 수 있는 영문 지명이 있다. 작은 인간형 투영 둘은 달리기 시퀀스의 인물로 보이나 앰버와 라울의 신원은 확인할 수 없다. 작업대에는 둥근 외장 부품 두 개와 여러 축·관절 부품이 있다. 현우는 보이지 않는다.",
        "hard_violations": [
         "대형 화면 중앙의 영문 제주 지명이 읽혀, 어디에도 읽을 수 있는 글자를 두지 말라는 조건을 위반한다."
        ],
        "physics": "분해 부품들은 작업대 표면에 놓여 있고 의자는 바닥에 지지된다. 찰리의 발은 작업대 높이에 나타나며 몸 가장자리가 발광해 실물 로봇보다는 투영으로 읽힌다. 작은 두 인물도 반투명 투영이므로 바닥 접촉이 없는 것을 실제 인체의 무지지 부유로 판정하지 않는다. 다만 달리는 자세와 찰리의 위치 사이에 목적지로 이어지는 동선이 없다."
       },
       {
        "label": "B",
        "direction": "찰리는 왼쪽 앞을 내려다보는 방향이며 카메라를 정면 응시하지 않는다. 작업대 위 작은 두 투영은 몸과 달리는 동작이 오른쪽의 찰리를 향해 있어 요청한 접근 방향과 맞는다. 제주 화면은 카메라 쪽으로 거의 정면을 향해 사선 노출 조건은 약하다.",
        "built_space": "왼쪽 문 하나, 후면 대형 화면 하나, 오른쪽 항구 야경 창문, 중앙과 오른쪽의 금속 작업대, 벽면 장비장과 천장 배관이 참고와 대응한다. 대형 화면의 테두리가 모두 보이고 면적은 전체의 5분의 2보다 훨씬 작다. 하단 왼쪽에는 의자 등받이 일부가 있으나 앉은 사람과 어깨는 없다. 왼쪽 후방의 작업용 의자는 별도 작업대에 배치되어 있다.",
        "entities": "찰리의 흰 마스크, 안테나, 넓은 어깨, 긴 중량감 있는 팔과 각진 장갑 형상이 참고와 가깝다. 하체 일부는 작업대와 투영 효과에 가려져 확인할 수 없다. 뒤 화면에는 푸른 해안·섬·산의 실제 풍경이 있고, 오른쪽 작업대에는 둥근 외장과 축 부품이 남아 있다. 작은 인간형 투영 둘은 달리는 시퀀스를 구현하지만 개별 신원과 민족성은 식별할 수 없다. 현우는 없다. 찰리 표면에는 미세한 문자형 무늬가 있으나 확실히 읽히는 문구는 확인되지 않는다.",
        "hard_violations": [],
        "physics": "금속 부품들은 작업대 위에 안정적으로 놓여 있고 가구는 바닥에 지지된다. 찰리는 몸을 통해 배경과 작업대가 비치는 발광 투영이므로 하체의 소실이나 물체와의 중첩은 실물 신체의 지지 실패가 아니다. 작은 두 투영에는 굽힌 무릎과 앞뒤로 흔드는 팔이 보여 달리기 동작으로 읽히며, 이동 방향 끝에 찰리가 있다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "읽을 수 있는 화면 글자가 금지 조건을 위반하며, 제주 풍경을 지도로 바꾸고 달리는 인물들의 방향도 찰리에게 향하지 않는다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "원래 연구소의 구조와 제주 풍경, 거대한 찰리 및 그에게 달려가는 투영을 잘 재현하지만, 하단 의자가 비어 있고 화면의 사선 배치가 약하다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 얼굴은 화면 왼쪽 앞을 향하며 특정 상대를 바라보는지는 불명확하다. 왼쪽의 작은 두 인물은 카메라 쪽으로 달리는 자세여서 오른쪽 찰리를 향하는 동선으로 읽히지 않는다. 제주 화면의 영상 면은 카메라에 비스듬히 보인다.",
        "built_space": "왼쪽에 경계가 보이는 대형 화면 하나, 중앙에 긴 금속 작업대 하나, 오른쪽 아래에 빈 의자 하나가 보인다. 뒤쪽에는 어두운 직사각형 면과 장비 수납장이 있다. 화면은 전체 면적의 5분의 2 미만이지만, 참고의 후면 화면과 오른쪽 항구 창문을 중심으로 한 공간 배치는 잘 보존되지 않았다. 의자는 작은 등받이 일부가 아니라 좌판과 팔걸이까지 드러나며 옆에 현우의 어깨가 없다.",
        "entities": "찰리 한 개체는 흰 마스크형 얼굴, 각진 장갑, 긴 팔과 짧은 다리를 갖춰 참고의 로봇 정체성과 대체로 맞지만 베이지색이 푸른 백색으로 씻겨 있다. 화면에는 제주 자연 풍경이 아니라 섬 지도와 읽을 수 있는 영문 지명이 있다. 작은 인간형 투영 둘은 달리기 시퀀스의 인물로 보이나 앰버와 라울의 신원은 확인할 수 없다. 작업대에는 둥근 외장 부품 두 개와 여러 축·관절 부품이 있다. 현우는 보이지 않는다.",
        "hard_violations": [
         "대형 화면 중앙의 영문 제주 지명이 읽혀, 어디에도 읽을 수 있는 글자를 두지 말라는 조건을 위반한다."
        ],
        "physics": "분해 부품들은 작업대 표면에 놓여 있고 의자는 바닥에 지지된다. 찰리의 발은 작업대 높이에 나타나며 몸 가장자리가 발광해 실물 로봇보다는 투영으로 읽힌다. 작은 두 인물도 반투명 투영이므로 바닥 접촉이 없는 것을 실제 인체의 무지지 부유로 판정하지 않는다. 다만 달리는 자세와 찰리의 위치 사이에 목적지로 이어지는 동선이 없다."
       },
       {
        "label": "A",
        "direction": "찰리는 왼쪽 앞을 내려다보는 방향이며 카메라를 정면 응시하지 않는다. 작업대 위 작은 두 투영은 몸과 달리는 동작이 오른쪽의 찰리를 향해 있어 요청한 접근 방향과 맞는다. 제주 화면은 카메라 쪽으로 거의 정면을 향해 사선 노출 조건은 약하다.",
        "built_space": "왼쪽 문 하나, 후면 대형 화면 하나, 오른쪽 항구 야경 창문, 중앙과 오른쪽의 금속 작업대, 벽면 장비장과 천장 배관이 참고와 대응한다. 대형 화면의 테두리가 모두 보이고 면적은 전체의 5분의 2보다 훨씬 작다. 하단 왼쪽에는 의자 등받이 일부가 있으나 앉은 사람과 어깨는 없다. 왼쪽 후방의 작업용 의자는 별도 작업대에 배치되어 있다.",
        "entities": "찰리의 흰 마스크, 안테나, 넓은 어깨, 긴 중량감 있는 팔과 각진 장갑 형상이 참고와 가깝다. 하체 일부는 작업대와 투영 효과에 가려져 확인할 수 없다. 뒤 화면에는 푸른 해안·섬·산의 실제 풍경이 있고, 오른쪽 작업대에는 둥근 외장과 축 부품이 남아 있다. 작은 인간형 투영 둘은 달리는 시퀀스를 구현하지만 개별 신원과 민족성은 식별할 수 없다. 현우는 없다. 찰리 표면에는 미세한 문자형 무늬가 있으나 확실히 읽히는 문구는 확인되지 않는다.",
        "hard_violations": [],
        "physics": "금속 부품들은 작업대 위에 안정적으로 놓여 있고 가구는 바닥에 지지된다. 찰리는 몸을 통해 배경과 작업대가 비치는 발광 투영이므로 하체의 소실이나 물체와의 중첩은 실물 신체의 지지 실패가 아니다. 작은 두 투영에는 굽힌 무릎과 앞뒤로 흔드는 팔이 보여 달리기 동작으로 읽히며, 이동 방향 끝에 찰리가 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.486
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.236
   },
   "violations": {
    "B": [
     "[gemini-pro] 화면 중앙의 지도에 'Jeju Island'라는 명확히 읽을 수 있는 텍스트가 존재함 (읽을 수 있는 글씨 금지 위반).",
     "[gemini-pro] 홀로그램으로 달리고 있는 두 인물이 이전 샷(레퍼런스)에 등장했던 남녀 작업자의 얼굴과 복장을 그대로 묘사함 (이전 샷 인물 재사용 금지 위반).",
     "[gpt-high] 대형 화면 중앙의 영문 제주 지명이 읽혀, 어디에도 읽을 수 있는 글자를 두지 말라는 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 236
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "이전 샷의 연구실 환경과 구도를 완벽하게 유지하면서, 지시된 제주도 풍경과 찰리의 홀로그램을 매우 사실적이고 정확하게 구현해 냈습니다."
   },
   {
    "label": "B",
    "score": 236,
    "verdict_ko": "화면에 읽을 수 있는 영어 텍스트가 포함되었고, 이전 샷의 인물들을 홀로그램으로 재사용하여 치명적인 지침 위반이 발생했습니다.  ★위반: [gemini-pro] 화면 중앙의 지도에 'Jeju Island'라는 명확히 읽을 수 있는 텍스트가 존재함 (읽을 수 있는 글씨 금지 위반). / [gemini-pro] 홀로그램으로 달리고 있는 두 인물이 이전 샷(레퍼런스)에 등장했던 남녀 작업자의 얼굴과 복장을 그대로 묘사함 (이전 샷 인물 재사용 금지 위반). / [gpt-high] 대형 화면 중앙의 영문 제주 지명이 읽혀, 어디에도 읽을 수 있는 글자를 두지 말라는 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S90sh4_sel.png",
    "asset_id": "d95aed5e-d208-4bb9-910c-b8ba820f7acd",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리 (홀로그램): the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1359377>",
    "asset_id": "eb7d341c-f40c-40ce-8f6b-f3b8d60eb7ed",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": true,
  "shot_run_uid": "06aae92d-226a-7433-ad27-75036e5457c0",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S90sh4"
  }
 },
 "S90sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T13:50:01.752903+00:00",
  "fingerprint": "8cd74973d0dffb728df2f857d0d55f5b140a36ba2f1836c4404c9a25904468cf",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S90sh9_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S90sh9_sel.png",
  "source_sha256": "bf65b76a57907d1a378f4bdf8d1f2d8db7b266365e7873de8c900edfca28671f",
  "file": "S90sh9_cine.png",
  "staged_sha256": "773cc223a59cf0ccddb4a6e7d785b57f4863629ec91ac7cd28f934645d748a23",
  "latency_ms": 10959
 },
 "S90sh13::signage": {
  "fp": "3dfd262870e4321d",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S90sh13": {
  "input_fingerprint": "9c919694e5383caf",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 차가운 수술대 위, 부서진 부품들이 서로 빈틈없이 맞물려 온전한 형태를 갖춘 찰리의 낡은 전신.\n\nLOCATION (lock): On the restoration table inside the spacious research room, with the surrounding displays illuminating the reassembled robot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Operating table (Supporting Charlie's fully fitted body) — The upper supporting surface and foot-side corner are visible from the high oblique position; used as Provides a clear body-length reference while keeping the surrounding laboratory visible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued illumination appropriate to the dark laboratory gently separates the fitted body sections without suggesting an activation flash or dream effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent on the operating table with his broken components fitted back together into a complete body and his eyes closed. The source does not specify his head's direction, which side of his torso faces upward, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie's broken components are being fitted back together on the operating table, with the reconstructed eyes closed. The holographic Jeju presentation remains separate from the physical reconstruction; no return to active operation has yet been established.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 차가운 수술대 위, 부서진 부품들이 서로 빈틈없이 맞물려 온전한 형태를 갖춘 찰리의 낡은 전신.\n\nLOCATION (lock): On the restoration table inside the spacious research room, with the surrounding displays illuminating the reassembled robot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Operating table (Supporting Charlie's fully fitted body) — The upper supporting surface and foot-side corner are visible from the high oblique position; used as Provides a clear body-length reference while keeping the surrounding laboratory visible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued illumination appropriate to the dark laboratory gently separates the fitted body sections without suggesting an activation flash or dream effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent on the operating table with his broken components fitted back together into a complete body and his eyes closed. The source does not specify his head's direction, which side of his torso faces upward, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie's broken components are being fitted back together on the operating table, with the reconstructed eyes closed. The holographic Jeju presentation remains separate from the physical reconstruction; no return to active operation has yet been established.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 차가운 수술대 위, 부서진 부품들이 서로 빈틈없이 맞물려 온전한 형태를 갖춘 찰리의 낡은 전신.\n\nLOCATION (lock): On the restoration table inside the spacious research room, with the surrounding displays illuminating the reassembled robot. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Operating table (Supporting Charlie's fully fitted body) — The upper supporting surface and foot-side corner are visible from the high oblique position; used as Provides a clear body-length reference while keeping the surrounding laboratory visible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Subdued illumination appropriate to the dark laboratory gently separates the fitted body sections without suggesting an activation flash or dream effect.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nIMMOBILE CHARACTER POSE — CANONICAL (identical wherever this character appears in ANY panel; on any conflict THIS POSE WINS): Charlie is recumbent on the operating table with his broken components fitted back together into a complete body and his eyes closed. The source does not specify his head's direction, which side of his torso faces upward, or the positions of his arms and legs.\n\nIMMOBILE BODIES OBEY GRAVITY: a person who is dead or unconscious\nexerts NO muscular effort. Every part of their body — head, torso,\narms, hands, fingers, legs — rests fully on whatever supports it\n(floor, wall, furniture, their own lap) and hangs or slumps with\ngravity. NEVER show any part of an immobile person's body lifted,\nraised, held up in the air, or posed as if presenting something:\nan object in their grip stays clenched in a hand that itself lies\nfallen on a support — the hand does not hold the object up. If the\ncanonical pose leaves a body part unspecified, resolve it as the\nmost gravity-compliant, fully-supported position.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie's broken components are being fitted back together on the operating table, with the reconstructed eyes closed. The holographic Jeju presentation remains separate from the physical reconstruction; no return to active operation has yet been established.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라는 높은 사각(high oblique)에서 수술대 위의 찰리를 내려다보고 있음.",
    "built_space": "이전 샷의 연구실(문, 모니터, 창문)이 배경에 유지됨. 다만 수술대 상판의 투시가 배경과 맞지 않아 평평한 테이블이 가파른 경사면처럼 왜곡되어 묘사됨.",
    "entities": "지시된 외형(샌드 베이지 장갑, 흰 마스크형 얼굴)을 갖춘 찰리가 부품이 결합된 채 누워 있음. 눈은 감겨 있으며 다른 인물은 없음.",
    "hard_violations": [],
    "physics": "로봇의 무거운 몸체와 팔다리가 수술대 표면에 닿아 하중을 지지받고 있음. 가장 가까운 왼손도 손가락이 금속 바닥에 닿아 있어 중력 지침에 순응함."
   },
   {
    "label": "B",
    "direction": "카메라는 높은 사각에서 수술대 위의 찰리를 내려다보고 있음.",
    "built_space": "연구실 배경이 보이며 수술대 상판이 왜곡된 투시로 그려짐. 허공에 임의의 홀로그램 스크린과 데이터 패널들이 떠 있음.",
    "entities": "찰리가 전신이 조립된 상태로 수술대에 누워 있음. 눈은 감겨 있음.",
    "hard_violations": [
     "[gemini-pro] 의식 없는 로봇의 왼손 손목이 꺾인 채 손 전체가 허공에 들려 있음 (중력 및 무생물 지지 지침 위반)"
    ],
    "physics": "화면 앞쪽의 왼손이 바닥에 닿지 않고 들려 있어 그 아래로 그림자가 생길 만큼 허공에 떠 있음. 근육의 힘이 없는 상태에서 중력에 따라 늘어져야 한다는 지침을 명백히 위반함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "중력에 따라 모든 신체 부위가 수술대 위에 온전히 닿아 있는 찰리의 모습을 잘 구현했으며, 불필요한 홀로그램 UI 없이 차분한 분위기를 정확히 표현했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "의식이 없는 상태임에도 왼손이 허공에 들려 있어 중력 지침을 위반했고, 지시되지 않은 홀로그램 UI 패널들을 임의로 추가하여 감점되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 높은 사각(high oblique)에서 수술대 위의 찰리를 내려다보고 있음.",
        "built_space": "이전 샷의 연구실(문, 모니터, 창문)이 배경에 유지됨. 다만 수술대 상판의 투시가 배경과 맞지 않아 평평한 테이블이 가파른 경사면처럼 왜곡되어 묘사됨.",
        "entities": "지시된 외형(샌드 베이지 장갑, 흰 마스크형 얼굴)을 갖춘 찰리가 부품이 결합된 채 누워 있음. 눈은 감겨 있으며 다른 인물은 없음.",
        "hard_violations": [],
        "physics": "로봇의 무거운 몸체와 팔다리가 수술대 표면에 닿아 하중을 지지받고 있음. 가장 가까운 왼손도 손가락이 금속 바닥에 닿아 있어 중력 지침에 순응함."
       },
       {
        "label": "B",
        "direction": "카메라는 높은 사각에서 수술대 위의 찰리를 내려다보고 있음.",
        "built_space": "연구실 배경이 보이며 수술대 상판이 왜곡된 투시로 그려짐. 허공에 임의의 홀로그램 스크린과 데이터 패널들이 떠 있음.",
        "entities": "찰리가 전신이 조립된 상태로 수술대에 누워 있음. 눈은 감겨 있음.",
        "hard_violations": [
         "의식 없는 로봇의 왼손 손목이 꺾인 채 손 전체가 허공에 들려 있음 (중력 및 무생물 지지 지침 위반)"
        ],
        "physics": "화면 앞쪽의 왼손이 바닥에 닿지 않고 들려 있어 그 아래로 그림자가 생길 만큼 허공에 떠 있음. 근육의 힘이 없는 상태에서 중력에 따라 늘어져야 한다는 지침을 명백히 위반함."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "중력에 따라 모든 신체 부위가 수술대 위에 온전히 닿아 있는 찰리의 모습을 잘 구현했으며, 불필요한 홀로그램 UI 없이 차분한 분위기를 정확히 표현했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "의식이 없는 상태임에도 왼손이 허공에 들려 있어 중력 지침을 위반했고, 지시되지 않은 홀로그램 UI 패널들을 임의로 추가하여 감점되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 높은 사각(high oblique)에서 수술대 위의 찰리를 내려다보고 있음.",
        "built_space": "이전 샷의 연구실(문, 모니터, 창문)이 배경에 유지됨. 다만 수술대 상판의 투시가 배경과 맞지 않아 평평한 테이블이 가파른 경사면처럼 왜곡되어 묘사됨.",
        "entities": "지시된 외형(샌드 베이지 장갑, 흰 마스크형 얼굴)을 갖춘 찰리가 부품이 결합된 채 누워 있음. 눈은 감겨 있으며 다른 인물은 없음.",
        "hard_violations": [],
        "physics": "로봇의 무거운 몸체와 팔다리가 수술대 표면에 닿아 하중을 지지받고 있음. 가장 가까운 왼손도 손가락이 금속 바닥에 닿아 있어 중력 지침에 순응함."
       },
       {
        "label": "B",
        "direction": "카메라는 높은 사각에서 수술대 위의 찰리를 내려다보고 있음.",
        "built_space": "연구실 배경이 보이며 수술대 상판이 왜곡된 투시로 그려짐. 허공에 임의의 홀로그램 스크린과 데이터 패널들이 떠 있음.",
        "entities": "찰리가 전신이 조립된 상태로 수술대에 누워 있음. 눈은 감겨 있음.",
        "hard_violations": [
         "의식 없는 로봇의 왼손 손목이 꺾인 채 손 전체가 허공에 들려 있음 (중력 및 무생물 지지 지침 위반)"
        ],
        "physics": "화면 앞쪽의 왼손이 바닥에 닿지 않고 들려 있어 그 아래로 그림자가 생길 만큼 허공에 떠 있음. 근육의 힘이 없는 상태에서 중력에 따라 늘어져야 한다는 지침을 명백히 위반함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "B보다 높은 사선 시점에서 재조립된 전신과 수술대 상판을 보여주고 원래 연구실의 후면 화면도 유지하지만, 추가된 홀로그램 패널은 절제된 조명 요구에서 벗어난다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "눈을 감고 수술대에 누운 찰리는 충실하지만, 더 밀착된 구도와 후면 대형 화면을 블라인드로 바꾼 공간 변경이 와이드 숏 및 장소 고정 지시를 약화한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 머리는 화면 오른쪽 위, 두 발은 왼쪽 아래를 향하고 얼굴과 가슴은 위로 향한다. 눈은 닫힌 어두운 틈으로 보이며 응시 대상이나 작동 중인 눈빛은 없다. 무기나 겨누는 물체는 없다. 뒤쪽 풍경 화면과 주변 반투명 패널의 표시 면은 실내 수술대 쪽을 향한다.",
        "built_space": "중앙에 찰리를 받치는 금속 수술대 한 개가 있으며 상판과 발치 쪽 가장자리, 한쪽 모서리가 보인다. 왼쪽 문 한 개, 작업등 한 개와 장비 작업대, 후면의 대형 풍경 화면 한 개, 오른쪽 야경 창과 부품 작업대가 보여 참고 공간의 주요 배치를 유지한다. 다만 원래 없던 반투명 표시 패널 여러 면이 수술대 주변에 추가되어 있다. 카메라는 B보다 상판을 더 내려다보며 연구실도 함께 담는다.",
        "entities": "인물은 비인간 로봇 찰리 한 몸뿐이다. 흰 각진 마스크형 얼굴, 샌드 베이지 장갑, 둥근 어깨 관절, 육중하고 긴 팔, 상대적으로 짧은 다리와 표면 마모가 인물 참고와 대체로 맞는다. 머리·몸통·팔다리가 연결된 완성된 몸으로 보이며 눈의 주황색 발광은 없다. 오른쪽 작업대의 원통형 부품은 이전 장소 참고에도 있다. 패널에는 도표와 작은 문자 모양이 있으나 확실히 읽을 수 있는 문구는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "등과 골반, 다리 뒤쪽이 금속 상판에 놓여 있고, 가까운 팔과 손도 상판 위에 내려져 있다. 발끝은 위를 향하지만 발뒤꿈치와 하퇴가 받쳐져 있어 몸이 떠 있는 자세는 아니다. 머리 뒤쪽은 목·등의 기계 구조와 이어져 있으며 공중에 독립적으로 매달린 부분은 보이지 않는다. 반투명 패널은 물리적 판재가 떠 있는 모습보다는 광학 투사 표시로 표현되어 있다."
       },
       {
        "label": "B",
        "direction": "머리는 오른쪽 위, 발은 왼쪽 아래로 향하고 얼굴은 천장 쪽으로 들려 보인다. 눈은 닫힌 검은 틈으로 표현되어 응시나 활성화는 보이지 않는다. 두 팔은 몸 옆으로 내려가며 손이 무엇을 겨누거나 내보이는 행동은 없다. 오른쪽 뒤의 풍경 화면은 실내를 향한다.",
        "built_space": "금속 수술대 한 개의 상판과 발치 가장자리가 보이지만 찰리가 화면을 더 크게 차지하고 수술대 하부와 주변 공간은 더 적게 나온다. 왼쪽 문 한 개와 작업등 한 개, 장비 작업대, 오른쪽 야경 창과 부품 작업대는 유지된다. 그러나 참고의 후면 중앙 대형 풍경 화면 자리에 블라인드가 생겼고 풍경 화면은 머리 뒤 오른쪽에 나타나, 고정된 연구실의 배치가 달라졌다.",
        "entities": "찰리 한 몸만 보이며 추가 사람은 없다. 베이지색의 낡고 각진 장갑, 흰 마스크형 얼굴, 큰 어깨와 긴 팔, 짧고 두꺼운 다리가 참고의 정체성과 맞는다. 전신 부품은 연결되어 있고 눈은 빛나지 않는다. 수술대와 오른쪽의 기존 기계 부품도 보인다. 주변 표시 장치가 몸에 가려져 재조립된 몸을 둘러싼 디스플레이의 존재감은 약하다. 읽을 수 있는 문구는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "몸통과 골반은 상판에 기대어 있고 다리 뒤쪽과 발뒤꿈치가 상판에 놓여 있다. 가까운 전완과 손도 상판의 지지를 받는다. 머리가 다소 높지만 뒤쪽 목·등 구조와 연결되어 있어 지지 없이 떠 있다고 볼 근거는 없다. 보이는 자세는 기계 몸체가 수술대에 누워 있는 상태로 가능하며, 도약이나 운동을 암시하는 부위는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "B보다 높은 사선 시점에서 재조립된 전신과 수술대 상판을 보여주고 원래 연구실의 후면 화면도 유지하지만, 추가된 홀로그램 패널은 절제된 조명 요구에서 벗어난다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "눈을 감고 수술대에 누운 찰리는 충실하지만, 더 밀착된 구도와 후면 대형 화면을 블라인드로 바꾼 공간 변경이 와이드 숏 및 장소 고정 지시를 약화한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 머리는 화면 오른쪽 위, 두 발은 왼쪽 아래를 향하고 얼굴과 가슴은 위로 향한다. 눈은 닫힌 어두운 틈으로 보이며 응시 대상이나 작동 중인 눈빛은 없다. 무기나 겨누는 물체는 없다. 뒤쪽 풍경 화면과 주변 반투명 패널의 표시 면은 실내 수술대 쪽을 향한다.",
        "built_space": "중앙에 찰리를 받치는 금속 수술대 한 개가 있으며 상판과 발치 쪽 가장자리, 한쪽 모서리가 보인다. 왼쪽 문 한 개, 작업등 한 개와 장비 작업대, 후면의 대형 풍경 화면 한 개, 오른쪽 야경 창과 부품 작업대가 보여 참고 공간의 주요 배치를 유지한다. 다만 원래 없던 반투명 표시 패널 여러 면이 수술대 주변에 추가되어 있다. 카메라는 B보다 상판을 더 내려다보며 연구실도 함께 담는다.",
        "entities": "인물은 비인간 로봇 찰리 한 몸뿐이다. 흰 각진 마스크형 얼굴, 샌드 베이지 장갑, 둥근 어깨 관절, 육중하고 긴 팔, 상대적으로 짧은 다리와 표면 마모가 인물 참고와 대체로 맞는다. 머리·몸통·팔다리가 연결된 완성된 몸으로 보이며 눈의 주황색 발광은 없다. 오른쪽 작업대의 원통형 부품은 이전 장소 참고에도 있다. 패널에는 도표와 작은 문자 모양이 있으나 확실히 읽을 수 있는 문구는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "등과 골반, 다리 뒤쪽이 금속 상판에 놓여 있고, 가까운 팔과 손도 상판 위에 내려져 있다. 발끝은 위를 향하지만 발뒤꿈치와 하퇴가 받쳐져 있어 몸이 떠 있는 자세는 아니다. 머리 뒤쪽은 목·등의 기계 구조와 이어져 있으며 공중에 독립적으로 매달린 부분은 보이지 않는다. 반투명 패널은 물리적 판재가 떠 있는 모습보다는 광학 투사 표시로 표현되어 있다."
       },
       {
        "label": "A",
        "direction": "머리는 오른쪽 위, 발은 왼쪽 아래로 향하고 얼굴은 천장 쪽으로 들려 보인다. 눈은 닫힌 검은 틈으로 표현되어 응시나 활성화는 보이지 않는다. 두 팔은 몸 옆으로 내려가며 손이 무엇을 겨누거나 내보이는 행동은 없다. 오른쪽 뒤의 풍경 화면은 실내를 향한다.",
        "built_space": "금속 수술대 한 개의 상판과 발치 가장자리가 보이지만 찰리가 화면을 더 크게 차지하고 수술대 하부와 주변 공간은 더 적게 나온다. 왼쪽 문 한 개와 작업등 한 개, 장비 작업대, 오른쪽 야경 창과 부품 작업대는 유지된다. 그러나 참고의 후면 중앙 대형 풍경 화면 자리에 블라인드가 생겼고 풍경 화면은 머리 뒤 오른쪽에 나타나, 고정된 연구실의 배치가 달라졌다.",
        "entities": "찰리 한 몸만 보이며 추가 사람은 없다. 베이지색의 낡고 각진 장갑, 흰 마스크형 얼굴, 큰 어깨와 긴 팔, 짧고 두꺼운 다리가 참고의 정체성과 맞는다. 전신 부품은 연결되어 있고 눈은 빛나지 않는다. 수술대와 오른쪽의 기존 기계 부품도 보인다. 주변 표시 장치가 몸에 가려져 재조립된 몸을 둘러싼 디스플레이의 존재감은 약하다. 읽을 수 있는 문구는 식별되지 않는다.",
        "hard_violations": [],
        "physics": "몸통과 골반은 상판에 기대어 있고 다리 뒤쪽과 발뒤꿈치가 상판에 놓여 있다. 가까운 전완과 손도 상판의 지지를 받는다. 머리가 다소 높지만 뒤쪽 목·등 구조와 연결되어 있어 지지 없이 떠 있다고 볼 근거는 없다. 보이는 자세는 기계 몸체가 수술대에 누워 있는 상태로 가능하며, 도약이나 운동을 암시하는 부위는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.75,
    "B": 1.5
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.25
   },
   "violations": {
    "B": [
     "[gemini-pro] 의식 없는 로봇의 왼손 손목이 꺾인 채 손 전체가 허공에 들려 있음 (중력 및 무생물 지지 지침 위반)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1750,
   "B": 1250
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "중력에 따라 모든 신체 부위가 수술대 위에 온전히 닿아 있는 찰리의 모습을 잘 구현했으며, 불필요한 홀로그램 UI 없이 차분한 분위기를 정확히 표현했습니다."
   },
   {
    "label": "B",
    "score": 1250,
    "verdict_ko": "의식이 없는 상태임에도 왼손이 허공에 들려 있어 중력 지침을 위반했고, 지시되지 않은 홀로그램 UI 패널들을 임의로 추가하여 감점되었습니다.  ★위반: [gemini-pro] 의식 없는 로봇의 왼손 손목이 꺾인 채 손 전체가 허공에 들려 있음 (중력 및 무생물 지지 지침 위반)"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S90sh9_sel.png",
    "asset_id": "4732321f-dfed-41a1-9a36-b9d8ed88233f",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": true,
  "shot_run_uid": "06aae931-5137-7260-ae76-356080fad261",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S90sh9"
  }
 },
 "S90sh13::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T13:51:18.624551+00:00",
  "fingerprint": "6200b11aa1ad11c5e1c2ab540d1b41e36bd0612341b3188f36fbd48d834647c9",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S90sh13_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S90sh13_sel.png",
  "source_sha256": "30352bcc49832fd7f7eb11bef6815d1c51da236cd5bf398a65d0443e25150972",
  "file": "S90sh13_cine.png",
  "staged_sha256": "6a41a55130adc2decc8491d368ddfb69d765e5ea132f1995d8eec1013203a2d9",
  "latency_ms": 9959
 },
 "S91sh4::signage": {
  "fp": "db4e4da97b9b5332",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::986d272d2d1a6fd3": {
  "subjects": [],
  "subject_text": "에필로그 바위섬과 임시 식탁\n바다 위로 드러난 바위섬의 야외 공간. 임시 식탁 위에 여러 음식이 차려져 있으며 밝은 낮빛이 비친다.",
  "identity": "canonical",
  "scope_id": "L284",
  "scope_role": "location_exterior",
  "scope_sha": "27853240e279399c"
 },
 "groupbg::islet_picnic_shore": {
  "input_fingerprint": "e0eac67a9dac1ad9",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "islet_picnic_shore",
    "tags": [
     "S91sh4",
     "S91sh5"
    ]
   },
   "context_sig": "638e73bc2f63dae3"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At a temporary dining table set on the exposed rock surface of a small sea islet in daylight.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n에필로그 바위섬과 임시 식탁: 푸른 바다를 배경으로 거친 암석 위에 하얀 식탁이 근사하게 차려진 피크닉 풍경. (특징: 바다와 맞닿은 거칠고 넓은 잿빛 섬바위 표면; 바위 위에 놓인 깨끗한 테이블과 접시, 풍성한 음식들; 사람 어깨 위에 사뿐히 내려앉은 작은 새; 따뜻하고 밝게 빛나는 주간 야외 채광)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 91. 에필로그 / 어느 바다 섬바위 – D 자막: 한 달 후\n- 현우는 지민과 식탁에 근사하게 음식을 차리고 있고.\n- 어디선가 날아 온 아기 새, 찰리의 어깨 위로 내려앉는다.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At a temporary dining table set on the exposed rock surface of a small sea islet in daylight.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n에필로그 바위섬과 임시 식탁: 푸른 바다를 배경으로 거친 암석 위에 하얀 식탁이 근사하게 차려진 피크닉 풍경. (특징: 바다와 맞닿은 거칠고 넓은 잿빛 섬바위 표면; 바위 위에 놓인 깨끗한 테이블과 접시, 풍성한 음식들; 사람 어깨 위에 사뿐히 내려앉은 작은 새; 따뜻하고 밝게 빛나는 주간 야외 채광)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 91. 에필로그 / 어느 바다 섬바위 – D 자막: 한 달 후\n- 현우는 지민과 식탁에 근사하게 음식을 차리고 있고.\n- 어디선가 날아 온 아기 새, 찰리의 어깨 위로 내려앉는다.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_islet_picnic_shore_1b9adf.png",
  "asset_id": "080c8b58-a1b3-40f1-8b38-c1b52f00ffae",
  "input_asset_ids": [
   "8048621a-1847-4b10-b7c7-7f2985a29c65"
  ],
  "origin_tag": "S91sh4",
  "place_text": "At a temporary dining table set on the exposed rock surface of a small sea islet in daylight.",
  "origin_inputs": {
   "place_text": "At a temporary dining table set on the exposed rock surface of a small sea islet in daylight.",
   "time_of_day_en": "day",
   "conti_asset_id": "8048621a-1847-4b10-b7c7-7f2985a29c65"
  }
 },
 "S91sh4::bgfirst_bg": {
  "input_fingerprint": "92cfe0c383842d43",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 바위섬 위 임시 식탁에 예쁜 접시를 올려둔 채 서로를 향해 활짝 웃고 있는 현우와 서지민의 전신.\n\nLOCATION (lock): At a temporary dining table set on the exposed rock surface of a small sea islet in daylight.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Temporary dining table (Being set with food and attractive plates) — Its top and near side are visible diagonally between the two figures; used as Connects their separate gestures while preserving full-body visibility at its adjacent sides; Food and plates (Partly arranged for the meal) — The plate interiors and food are visible from the oblique camera position; used as Provide the shared focus of their activity at natural scale; Island rock (Supporting the table and both figures); used as Keeps the meal grounded in the island setting without adding furnishings.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight renders the meal and smiling faces with gentle contrast and restrained saturation appropriate to the setting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 바위섬 위 임시 식탁에 예쁜 접시를 올려둔 채 서로를 향해 활짝 웃고 있는 현우와 서지민의 전신.\n\nLOCATION (lock): At a temporary dining table set on the exposed rock surface of a small sea islet in daylight.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Temporary dining table (Being set with food and attractive plates) — Its top and near side are visible diagonally between the two figures; used as Connects their separate gestures while preserving full-body visibility at its adjacent sides; Food and plates (Partly arranged for the meal) — The plate interiors and food are visible from the oblique camera position; used as Provide the shared focus of their activity at natural scale; Island rock (Supporting the table and both figures); used as Keeps the meal grounded in the island setting without adding furnishings.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight renders the meal and smiling faces with gentle contrast and restrained saturation appropriate to the setting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S91sh4__bgfirst_bg.png",
  "asset_id": "410404cd-7317-4419-a279-f0a70de5bf97",
  "input_asset_ids": [
   "8048621a-1847-4b10-b7c7-7f2985a29c65",
   "080c8b58-a1b3-40f1-8b38-c1b52f00ffae"
  ]
 },
 "S91sh4": {
  "input_fingerprint": "34525825ba77a662",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바위섬 위 임시 식탁에 예쁜 접시를 올려둔 채 서로를 향해 활짝 웃고 있는 현우와 서지민의 전신.\n\nLOCATION (lock): At a temporary dining table set on the exposed rock surface of a small sea islet in daylight. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Temporary dining table (Being set with food and attractive plates) — Its top and near side are visible diagonally between the two figures; used as Connects their separate gestures while preserving full-body visibility at its adjacent sides; Food and plates (Partly arranged for the meal) — The plate interiors and food are visible from the oblique camera position; used as Provide the shared focus of their activity at natural scale; Island rock (Supporting the table and both figures); used as Keeps the meal grounded in the island setting without adding furnishings.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight renders the meal and smiling faces with gentle contrast and restrained saturation appropriate to the setting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Clear seawater reveals fish beneath the rocky island, and an outdoor dining table is being laid with an ample meal. Charlie is fully reassembled and active again, trying to catch fish. 현우: He is at the dining table arranging food. 서지민: She is at the dining table arranging food.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 서지민 (한국인 여성, 20세, 앳된 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바위섬 위 임시 식탁에 예쁜 접시를 올려둔 채 서로를 향해 활짝 웃고 있는 현우와 서지민의 전신.\n\nLOCATION (lock): At a temporary dining table set on the exposed rock surface of a small sea islet in daylight. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Temporary dining table (Being set with food and attractive plates) — Its top and near side are visible diagonally between the two figures; used as Connects their separate gestures while preserving full-body visibility at its adjacent sides; Food and plates (Partly arranged for the meal) — The plate interiors and food are visible from the oblique camera position; used as Provide the shared focus of their activity at natural scale; Island rock (Supporting the table and both figures); used as Keeps the meal grounded in the island setting without adding furnishings.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight renders the meal and smiling faces with gentle contrast and restrained saturation appropriate to the setting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Clear seawater reveals fish beneath the rocky island, and an outdoor dining table is being laid with an ample meal. Charlie is fully reassembled and active again, trying to catch fish. 현우: He is at the dining table arranging food. 서지민: She is at the dining table arranging food.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 서지민 (한국인 여성, 20세, 앳된 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 바위섬 위 임시 식탁에 예쁜 접시를 올려둔 채 서로를 향해 활짝 웃고 있는 현우와 서지민의 전신.\n\nLOCATION (lock): At a temporary dining table set on the exposed rock surface of a small sea islet in daylight. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Temporary dining table (Being set with food and attractive plates) — Its top and near side are visible diagonally between the two figures; used as Connects their separate gestures while preserving full-body visibility at its adjacent sides; Food and plates (Partly arranged for the meal) — The plate interiors and food are visible from the oblique camera position; used as Provide the shared focus of their activity at natural scale; Island rock (Supporting the table and both figures); used as Keeps the meal grounded in the island setting without adding furnishings.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight renders the meal and smiling faces with gentle contrast and restrained saturation appropriate to the setting.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Clear seawater reveals fish beneath the rocky island, and an outdoor dining table is being laid with an ample meal. Charlie is fully reassembled and active again, trying to catch fish. 현우: He is at the dining table arranging food. 서지민: She is at the dining table arranging food.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리); 서지민 (한국인 여성, 20세, 앳된 얼굴, 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S91sh4__bgfirst_bg.png",
     "asset_id": "410404cd-7317-4419-a279-f0a70de5bf97",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S91sh4.png",
     "asset_id": "8048621a-1847-4b10-b7c7-7f2985a29c65",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 서지민: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1147284>",
     "asset_id": "a529b479-7055-42c6-be7d-19d14b5e000f",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_islet_picnic_shore_1b9adf.png",
     "asset_id": "080c8b58-a1b3-40f1-8b38-c1b52f00ffae",
     "role": "bgfirst_group_bg"
    },
    {
     "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1443881>",
     "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
     "role": "character_ref"
    },
    {
     "label": "CHARACTER REFERENCE — 서지민: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1147284>",
     "asset_id": "a529b479-7055-42c6-be7d-19d14b5e000f",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "두 인물이 서로를 향해 웃고 있으며, 시선이 자연스럽게 교차합니다.",
    "built_space": "바위섬 배경은 일치하나, 남성 앞의 메인 식탁 뒤로 또 다른 식탁 상판이 겹쳐서 생성되었습니다.",
    "entities": "현우, 서지민의 외모와 의상은 참조와 일치하며 지정된 소품들이 존재합니다.",
    "hard_violations": [
     "[gemini-pro] 중복된 가구 (메인 식탁 외에 불필요한 식탁 상판이 추가로 생성됨)",
     "[gemini-pro] 물리적으로 불가능한 스테이징 (남성의 손이 접시와 융합됨)"
    ],
    "physics": "불가능한 형태의 식탁 상판이 겹쳐 있으며, 남성의 왼손이 접시 안으로 파고들어 형체가 무너졌습니다."
   },
   {
    "label": "B",
    "direction": "두 인물이 서로 시선을 맞추며 활짝 웃고 있습니다.",
    "built_space": "바위섬 배경과 단일 임시 식탁이 참조 이미지와 동일하게 정확히 배치되었습니다.",
    "entities": "현우와 서지민의 외모, 단일 식탁, 예쁜 접시 등 모든 요소가 프롬프트와 일치합니다.",
    "hard_violations": [],
    "physics": "두 사람의 발이 바위 위에 단단히 지지되어 있으며, 남성이 접시를 안정적으로 들고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지정된 전신 구도와 인물들의 상호작용, 단일 식탁의 구조를 물리적 결함 없이 완벽하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "인물의 표정과 구도는 좋으나, 식탁 상판이 중복 생성되는 치명적인 구조적 오류가 있습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "두 인물이 서로 시선을 맞추며 활짝 웃고 있습니다.",
        "built_space": "바위섬 배경과 단일 임시 식탁이 참조 이미지와 동일하게 정확히 배치되었습니다.",
        "entities": "현우와 서지민의 외모, 단일 식탁, 예쁜 접시 등 모든 요소가 프롬프트와 일치합니다.",
        "hard_violations": [],
        "physics": "두 사람의 발이 바위 위에 단단히 지지되어 있으며, 남성이 접시를 안정적으로 들고 있습니다."
       },
       {
        "label": "A",
        "direction": "두 인물이 서로를 향해 웃고 있으며, 시선이 자연스럽게 교차합니다.",
        "built_space": "바위섬 배경은 일치하나, 남성 앞의 메인 식탁 뒤로 또 다른 식탁 상판이 겹쳐서 생성되었습니다.",
        "entities": "현우, 서지민의 외모와 의상은 참조와 일치하며 지정된 소품들이 존재합니다.",
        "hard_violations": [
         "중복된 가구 (메인 식탁 외에 불필요한 식탁 상판이 추가로 생성됨)",
         "물리적으로 불가능한 스테이징 (남성의 손이 접시와 융합됨)"
        ],
        "physics": "불가능한 형태의 식탁 상판이 겹쳐 있으며, 남성의 왼손이 접시 안으로 파고들어 형체가 무너졌습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "지정된 전신 구도와 인물들의 상호작용, 단일 식탁의 구조를 물리적 결함 없이 완벽하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "인물의 표정과 구도는 좋으나, 식탁 상판이 중복 생성되는 치명적인 구조적 오류가 있습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "두 인물이 서로 시선을 맞추며 활짝 웃고 있습니다.",
        "built_space": "바위섬 배경과 단일 임시 식탁이 참조 이미지와 동일하게 정확히 배치되었습니다.",
        "entities": "현우와 서지민의 외모, 단일 식탁, 예쁜 접시 등 모든 요소가 프롬프트와 일치합니다.",
        "hard_violations": [],
        "physics": "두 사람의 발이 바위 위에 단단히 지지되어 있으며, 남성이 접시를 안정적으로 들고 있습니다."
       },
       {
        "label": "A",
        "direction": "두 인물이 서로를 향해 웃고 있으며, 시선이 자연스럽게 교차합니다.",
        "built_space": "바위섬 배경은 일치하나, 남성 앞의 메인 식탁 뒤로 또 다른 식탁 상판이 겹쳐서 생성되었습니다.",
        "entities": "현우, 서지민의 외모와 의상은 참조와 일치하며 지정된 소품들이 존재합니다.",
        "hard_violations": [
         "중복된 가구 (메인 식탁 외에 불필요한 식탁 상판이 추가로 생성됨)",
         "물리적으로 불가능한 스테이징 (남성의 손이 접시와 융합됨)"
        ],
        "physics": "불가능한 형태의 식탁 상판이 겹쳐 있으며, 남성의 왼손이 접시 안으로 파고들어 형체가 무너졌습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "두 사람의 전신과 서로의 얼굴을 향한 환한 웃음, 상차림 동작을 충실히 구현했으나 식탁의 대각선 배치는 약하다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "전신과 풍성한 상차림은 충실하지만 두 사람의 시선이 상대 얼굴보다 음식 쪽으로 내려가 있어 핵심인 ‘서로를 향해 활짝 웃는’ 순간이 약하다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "왼쪽 현우는 오른쪽 서지민의 얼굴을 바라보며 입을 벌려 웃고, 서지민도 현우의 얼굴을 향해 웃는다. 현우가 받친 접시와 서지민의 두 손은 식탁 위 상차림을 향한다.",
        "built_space": "노출된 암반 위에 흰 접이식 식탁 한 개가 있고, 두 사람은 그 좌우에 서 있다. 상판 아래 양쪽 접이식 지지대와 네 다리가 보이며 추가 의자는 없다. 상판과 가까운 테두리, 접시 내부와 음식이 보이고 두 사람의 신발까지 프레임에 들어온다. 다만 식탁 앞 모서리는 화면에 거의 수평이어서 요구한 대각선 구도는 약하다. 뒤쪽 바다와 오른쪽 수목이 있는 바위섬, 주변 암반은 장소 참조와 잘 대응한다.",
        "entities": "앳된 동아시아계 남녀 두 명만 보인다. 현우의 흐트러진 짧은 검은 머리, 서지민의 어깨 길이 검은 머리와 두 사람의 남색 반팔은 참조와 대체로 맞는다. 정확한 나이와 국적은 외관만으로 확정할 수 없다. 흰 식탁 위에 밝은색 접시 여러 개, 생선과 채소 등 음식, 물병 한 개, 잔 두 개, 작은 꽃병 한 개가 보인다. 읽을 수 있는 글자는 없다. 찰리와 물속 물고기는 이 구도에 나타나지 않는다.",
        "hard_violations": [],
        "physics": "두 사람 모두 신발을 암반에 딛고 서서 상체를 식탁 쪽으로 조금 기울인다. 현우의 접시는 양손이 받치며, 서지민의 손은 상판 위 그릇에 닿아 있다. 다른 음식과 식기는 상판에 놓이고 식탁 다리는 암반으로 이어진다. 지지 없이 떠 있는 몸이나 물체는 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "현우는 고개와 시선을 자신이 옮기는 음식 접시 쪽으로 내리며 웃는다. 서지민도 왼쪽 아래의 식탁과 음식 쪽을 향해 웃어, 상대 얼굴을 바라보는 관계는 A보다 불명확하다. 두 사람의 손은 각각 접시와 샐러드 그릇을 식탁에 놓는 방향이다.",
        "built_space": "암반 위 흰 접이식 식탁 한 개의 좌우에 두 사람이 서 있고 전신과 신발이 모두 보인다. 식탁의 네 다리와 양쪽 접이식 지지 구조가 보이며 추가 가구는 없다. 상판과 앞 테두리, 음식과 접시 내부가 드러나지만 앞 테두리가 거의 수평이라 대각선 배치 요구는 약하다. 바다, 주변 암반과 오른쪽 먼 바위섬은 장소 참조와 대체로 일치한다.",
        "entities": "젊은 동아시아계 남성 한 명과 여성 한 명만 보인다. 남성의 짧고 흐트러진 검은 머리, 여성의 어깨 길이 검은 머리, 남색 반팔은 인물 참조와 대체로 맞으며 정확한 나이와 국적은 외관만으로 확정할 수 없다. 두 사람은 검은 바지와 흰 운동화를 착용한다. 식탁에는 밝은 접시 여러 개와 생선, 채소 등 풍성한 음식, 물병 한 개와 잔 두 개가 있다. 읽을 수 있는 글자는 없고 찰리나 물속 물고기는 보이지 않는다.",
        "hard_violations": [],
        "physics": "두 사람의 발은 암반에 닿아 체중을 지탱한다. 현우는 음식 접시를 두 손으로 받치고, 서지민은 큰 샐러드 그릇의 양쪽을 잡고 있다. 나머지 식기는 식탁 상판이 받치며 식탁 다리도 암반에 닿는다. 상차림 중 앞으로 기울인 자세는 가능한 동작이고, 지지 없이 떠 있는 물체나 인물은 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "두 사람의 전신과 서로의 얼굴을 향한 환한 웃음, 상차림 동작을 충실히 구현했으나 식탁의 대각선 배치는 약하다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "전신과 풍성한 상차림은 충실하지만 두 사람의 시선이 상대 얼굴보다 음식 쪽으로 내려가 있어 핵심인 ‘서로를 향해 활짝 웃는’ 순간이 약하다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "왼쪽 현우는 오른쪽 서지민의 얼굴을 바라보며 입을 벌려 웃고, 서지민도 현우의 얼굴을 향해 웃는다. 현우가 받친 접시와 서지민의 두 손은 식탁 위 상차림을 향한다.",
        "built_space": "노출된 암반 위에 흰 접이식 식탁 한 개가 있고, 두 사람은 그 좌우에 서 있다. 상판 아래 양쪽 접이식 지지대와 네 다리가 보이며 추가 의자는 없다. 상판과 가까운 테두리, 접시 내부와 음식이 보이고 두 사람의 신발까지 프레임에 들어온다. 다만 식탁 앞 모서리는 화면에 거의 수평이어서 요구한 대각선 구도는 약하다. 뒤쪽 바다와 오른쪽 수목이 있는 바위섬, 주변 암반은 장소 참조와 잘 대응한다.",
        "entities": "앳된 동아시아계 남녀 두 명만 보인다. 현우의 흐트러진 짧은 검은 머리, 서지민의 어깨 길이 검은 머리와 두 사람의 남색 반팔은 참조와 대체로 맞는다. 정확한 나이와 국적은 외관만으로 확정할 수 없다. 흰 식탁 위에 밝은색 접시 여러 개, 생선과 채소 등 음식, 물병 한 개, 잔 두 개, 작은 꽃병 한 개가 보인다. 읽을 수 있는 글자는 없다. 찰리와 물속 물고기는 이 구도에 나타나지 않는다.",
        "hard_violations": [],
        "physics": "두 사람 모두 신발을 암반에 딛고 서서 상체를 식탁 쪽으로 조금 기울인다. 현우의 접시는 양손이 받치며, 서지민의 손은 상판 위 그릇에 닿아 있다. 다른 음식과 식기는 상판에 놓이고 식탁 다리는 암반으로 이어진다. 지지 없이 떠 있는 몸이나 물체는 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "현우는 고개와 시선을 자신이 옮기는 음식 접시 쪽으로 내리며 웃는다. 서지민도 왼쪽 아래의 식탁과 음식 쪽을 향해 웃어, 상대 얼굴을 바라보는 관계는 A보다 불명확하다. 두 사람의 손은 각각 접시와 샐러드 그릇을 식탁에 놓는 방향이다.",
        "built_space": "암반 위 흰 접이식 식탁 한 개의 좌우에 두 사람이 서 있고 전신과 신발이 모두 보인다. 식탁의 네 다리와 양쪽 접이식 지지 구조가 보이며 추가 가구는 없다. 상판과 앞 테두리, 음식과 접시 내부가 드러나지만 앞 테두리가 거의 수평이라 대각선 배치 요구는 약하다. 바다, 주변 암반과 오른쪽 먼 바위섬은 장소 참조와 대체로 일치한다.",
        "entities": "젊은 동아시아계 남성 한 명과 여성 한 명만 보인다. 남성의 짧고 흐트러진 검은 머리, 여성의 어깨 길이 검은 머리, 남색 반팔은 인물 참조와 대체로 맞으며 정확한 나이와 국적은 외관만으로 확정할 수 없다. 두 사람은 검은 바지와 흰 운동화를 착용한다. 식탁에는 밝은 접시 여러 개와 생선, 채소 등 풍성한 음식, 물병 한 개와 잔 두 개가 있다. 읽을 수 있는 글자는 없고 찰리나 물속 물고기는 보이지 않는다.",
        "hard_violations": [],
        "physics": "두 사람의 발은 암반에 닿아 체중을 지탱한다. 현우는 음식 접시를 두 손으로 받치고, 서지민은 큰 샐러드 그릇의 양쪽을 잡고 있다. 나머지 식기는 식탁 상판이 받치며 식탁 다리도 암반에 닿는다. 상차림 중 앞으로 기울인 자세는 가능한 동작이고, 지지 없이 떠 있는 물체나 인물은 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 1.304,
    "B": 2.0
   },
   "adjusted": {
    "A": 1.054,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 중복된 가구 (메인 식탁 외에 불필요한 식탁 상판이 추가로 생성됨)",
     "[gemini-pro] 물리적으로 불가능한 스테이징 (남성의 손이 접시와 융합됨)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt-high": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 1054
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "지정된 전신 구도와 인물들의 상호작용, 단일 식탁의 구조를 물리적 결함 없이 완벽하게 구현했습니다."
   },
   {
    "label": "A",
    "score": 1054,
    "verdict_ko": "인물의 표정과 구도는 좋으나, 식탁 상판이 중복 생성되는 치명적인 구조적 오류가 있습니다.  ★위반: [gemini-pro] 중복된 가구 (메인 식탁 외에 불필요한 식탁 상판이 추가로 생성됨) / [gemini-pro] 물리적으로 불가능한 스테이징 (남성의 손이 접시와 융합됨)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_islet_picnic_shore_1b9adf.png",
    "asset_id": "080c8b58-a1b3-40f1-8b38-c1b52f00ffae",
    "role": "bgfirst_group_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 서지민: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1147284>",
    "asset_id": "a529b479-7055-42c6-be7d-19d14b5e000f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae936-0f19-7ddc-9b09-4d733dab2759",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S91sh4__bgfirst_bg.png",
   "bg_asset_id": "410404cd-7317-4419-a279-f0a70de5bf97",
   "bg_record_key": "S91sh4::bgfirst_bg",
   "chain_winner": false,
   "authority": "groupbg",
   "group_key": "islet_picnic_shore",
   "groupbg_asset_id": "080c8b58-a1b3-40f1-8b38-c1b52f00ffae"
  },
  "ref_mode": "그룹 배경+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S91sh4::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T09:39:46.486891+00:00",
  "fingerprint": "7aad97877a3b4e2beb77e9184bab0e20f089bc13412ec0139facdc794b562e2c",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S91sh4_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S91sh4_sel.png",
  "source_sha256": "1fbf1617d55172b649769cb49df49eb3efd7006617edaceb092b6412a7cfd0b7",
  "file": "S91sh4_cine.png",
  "staged_sha256": "42c312a6243eabc2974218d2fbd628b66523d0284313317d7ab8adc3d9a1082b",
  "latency_ms": 9586
 },
 "S91sh5::signage": {
  "fp": "b38697630282dd68",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S91sh5": {
  "input_fingerprint": "d7e3558c31073a94",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 낡은 금속 어깨 위로 깃털이 부스스한 아기 새(B-200이 돌보던 새)가 가만히 내려앉은 클로즈업.\n\nLOCATION (lock): At the water's edge beside the small rocky islet, where the robot is fishing in clear daylight. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Sea below Charlie (Clear water containing fish) — Seen downward beyond the shoulder, with the water below rather than a reflected scene as the background; used as Provides softly resolved spatial context for Charlie's continued fishing attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight preserves fine feather and worn-metal detail with soft contrast, without introducing unsupported reflections or atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie remains fully reassembled and active, with the chest ring intact; a baby bird has settled on the metal shoulder. Fish remain visible in the clear seawater, and the meal is being arranged on the nearby table.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 아기 새(B-200이 돌보던 새) (손가락 위에 앉는 작은 크기, 둥근 몸통, 작은 부리, 짧은 날개, 가느다란 발) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 낡은 금속 어깨 위로 깃털이 부스스한 아기 새(B-200이 돌보던 새)가 가만히 내려앉은 클로즈업.\n\nLOCATION (lock): At the water's edge beside the small rocky islet, where the robot is fishing in clear daylight. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Sea below Charlie (Clear water containing fish) — Seen downward beyond the shoulder, with the water below rather than a reflected scene as the background; used as Provides softly resolved spatial context for Charlie's continued fishing attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight preserves fine feather and worn-metal detail with soft contrast, without introducing unsupported reflections or atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie remains fully reassembled and active, with the chest ring intact; a baby bird has settled on the metal shoulder. Fish remain visible in the clear seawater, and the meal is being arranged on the nearby table.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 아기 새(B-200이 돌보던 새) (손가락 위에 앉는 작은 크기, 둥근 몸통, 작은 부리, 짧은 날개, 가느다란 발) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 찰리의 낡은 금속 어깨 위로 깃털이 부스스한 아기 새(B-200이 돌보던 새)가 가만히 내려앉은 클로즈업.\n\nLOCATION (lock): At the water's edge beside the small rocky islet, where the robot is fishing in clear daylight. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Sea below Charlie (Clear water containing fish) — Seen downward beyond the shoulder, with the water below rather than a reflected scene as the background; used as Provides softly resolved spatial context for Charlie's continued fishing attention.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight preserves fine feather and worn-metal detail with soft contrast, without introducing unsupported reflections or atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Charlie remains fully reassembled and active, with the chest ring intact; a baby bird has settled on the metal shoulder. Fish remain visible in the clear seawater, and the meal is being arranged on the nearby table.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 찰리 (고릴라형 몸체, 육중하고 긴 팔, 짧은 다리, 각진 샌드 베이지 장갑판, 점과 선으로 구성된 흰 마스크형 얼굴); 아기 새(B-200이 돌보던 새) (손가락 위에 앉는 작은 크기, 둥근 몸통, 작은 부리, 짧은 날개, 가느다란 발) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라는 로봇의 어깨 너머 아래쪽의 바다를 향하고 있으며, 아기 새는 우측을 바라보고 있습니다.",
    "built_space": "인공 구조물은 없으며, 로봇의 어깨와 바다 배경으로 구성되어 있습니다.",
    "entities": "아기 새는 레퍼런스와 일치하는 작고 부스스한 깃털을 가졌으며, 로봇은 찰리의 낡은 베이지색 금속 장갑판과 일치합니다. 물속에 흐릿한 물고기들이 보입니다.",
    "hard_violations": [],
    "physics": "아기 새가 로봇의 금속 어깨 위를 두 발로 안정적으로 딛고 서 있습니다."
   },
   {
    "label": "B",
    "direction": "카메라는 로봇의 어깨 너머 아래쪽의 바다를 비추며, 아기 새는 우측 전방을 바라봅니다.",
    "built_space": "인공 구조물은 없으며, 로봇의 어깨 부품과 바다 배경이 존재합니다.",
    "entities": "아기 새의 형태가 레퍼런스와 유사하며, 로봇의 어깨 장갑판이 찰리의 특징을 담고 있습니다. 물속에 물고기들이 비교적 선명하게 보입니다.",
    "hard_violations": [],
    "physics": "아기 새가 금속 표면 위에 서 있으며, 발의 접촉면과 그림자가 물리적으로 자연스럽게 구현되었습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "부스스한 아기 새와 낡은 금속 어깨의 질감을 훌륭하게 구현했으며, 물고기가 있는 바다 배경을 부드럽게 아웃포커싱하여 프롬프트가 요구한 공간적 맥락을 정확히 표현했습니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "요구된 주요 요소들은 잘 배치되었으나, 배경의 물고기가 너무 선명하게 표현되어 '부드럽게 묘사된(softly resolved) 배경'이라는 지시사항의 의도를 다소 벗어났습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 로봇의 어깨 너머 아래쪽의 바다를 향하고 있으며, 아기 새는 우측을 바라보고 있습니다.",
        "built_space": "인공 구조물은 없으며, 로봇의 어깨와 바다 배경으로 구성되어 있습니다.",
        "entities": "아기 새는 레퍼런스와 일치하는 작고 부스스한 깃털을 가졌으며, 로봇은 찰리의 낡은 베이지색 금속 장갑판과 일치합니다. 물속에 흐릿한 물고기들이 보입니다.",
        "hard_violations": [],
        "physics": "아기 새가 로봇의 금속 어깨 위를 두 발로 안정적으로 딛고 서 있습니다."
       },
       {
        "label": "B",
        "direction": "카메라는 로봇의 어깨 너머 아래쪽의 바다를 비추며, 아기 새는 우측 전방을 바라봅니다.",
        "built_space": "인공 구조물은 없으며, 로봇의 어깨 부품과 바다 배경이 존재합니다.",
        "entities": "아기 새의 형태가 레퍼런스와 유사하며, 로봇의 어깨 장갑판이 찰리의 특징을 담고 있습니다. 물속에 물고기들이 비교적 선명하게 보입니다.",
        "hard_violations": [],
        "physics": "아기 새가 금속 표면 위에 서 있으며, 발의 접촉면과 그림자가 물리적으로 자연스럽게 구현되었습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "부스스한 아기 새와 낡은 금속 어깨의 질감을 훌륭하게 구현했으며, 물고기가 있는 바다 배경을 부드럽게 아웃포커싱하여 프롬프트가 요구한 공간적 맥락을 정확히 표현했습니다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "요구된 주요 요소들은 잘 배치되었으나, 배경의 물고기가 너무 선명하게 표현되어 '부드럽게 묘사된(softly resolved) 배경'이라는 지시사항의 의도를 다소 벗어났습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 로봇의 어깨 너머 아래쪽의 바다를 향하고 있으며, 아기 새는 우측을 바라보고 있습니다.",
        "built_space": "인공 구조물은 없으며, 로봇의 어깨와 바다 배경으로 구성되어 있습니다.",
        "entities": "아기 새는 레퍼런스와 일치하는 작고 부스스한 깃털을 가졌으며, 로봇은 찰리의 낡은 베이지색 금속 장갑판과 일치합니다. 물속에 흐릿한 물고기들이 보입니다.",
        "hard_violations": [],
        "physics": "아기 새가 로봇의 금속 어깨 위를 두 발로 안정적으로 딛고 서 있습니다."
       },
       {
        "label": "B",
        "direction": "카메라는 로봇의 어깨 너머 아래쪽의 바다를 비추며, 아기 새는 우측 전방을 바라봅니다.",
        "built_space": "인공 구조물은 없으며, 로봇의 어깨 부품과 바다 배경이 존재합니다.",
        "entities": "아기 새의 형태가 레퍼런스와 유사하며, 로봇의 어깨 장갑판이 찰리의 특징을 담고 있습니다. 물속에 물고기들이 비교적 선명하게 보입니다.",
        "hard_violations": [],
        "physics": "아기 새가 금속 표면 위에 서 있으며, 발의 접촉면과 그림자가 물리적으로 자연스럽게 구현되었습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "낡은 어깨에 앉은 작은 새와 아래쪽 맑은 바다를 클로즈업으로 정확히 담았지만, 새의 깃털이 비교적 가지런해 ‘부스스한 아기 새’라는 핵심 묘사는 B보다 약하다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "부스스한 아기 새가 낡은 금속 어깨에 발을 붙이고 가만히 앉은 순간을 밀착해 보여주며, 찰리의 아래쪽 바다를 향한 주의와 물속 물고기도 유지한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "찰리의 옆얼굴은 화면 왼쪽 아래의 바다를 향한다. 눈 전체가 보이지 않아 정확한 시선 종착점은 단정할 수 없지만 머리 방향은 낚시에 집중하는 설정과 맞는다. 새의 부리와 눈은 화면 오른쪽을 향하며 카메라를 정면으로 응시하지 않는다. 배경에는 수면 아래 여러 물고기의 길쭉한 몸이 보인다.",
        "built_space": "전경의 어깨 장갑 하나가 화면 중앙과 하단을 차지하고, 그 뒤 오른쪽에 목과 머리 일부 및 연결 장갑이 보인다. 배경은 어깨 너머 아래로 내려다본 맑은 바다와 수중 바위이며, 반사된 육상 풍경이 아니다. 인공 구조물이나 중복된 고정 설비는 없다. 식탁과 먼 바위섬은 이 클로즈업 밖에 있어 확인할 수 없다.",
        "entities": "찰리 한 개체와 아기 새 한 마리만 보이며 이전 장면의 사람들은 없다. 찰리의 샌드 베이지 장갑, 마모된 금속, 검은 목 기구와 흰 얼굴 일부는 참조와 부합한다. 새는 작은 부리, 둥근 몸통, 짧게 접힌 날개와 가는 발을 갖췄지만 참조보다 다소 짙은 갈색이고 깃털의 흐트러짐이 약하다. 맑은 물속 물고기가 보이고 읽을 수 있는 글자는 없다. 가슴 고리와 전신 비율은 프레임 밖이다.",
        "hard_violations": [],
        "physics": "새의 몸 아래 발과 발가락이 어깨 장갑 윗면에 닿아 체중을 지지하고, 금속 표면에 접촉 그림자가 생긴다. 날개를 접은 정착 자세로, 공중에 떠 있지 않다. 어깨는 찰리의 몸통과 기계 관절에 연결되어 있다. 찰리의 지면 접촉은 프레임 밖이며, 물고기는 물속에 있다."
       },
       {
        "label": "B",
        "direction": "찰리의 옆얼굴은 화면 오른쪽 아래의 바다 쪽으로 기울어 있다. 정확한 눈동자는 보이지 않지만 머리 방향은 아래쪽 물에 주의를 둔 상태로 읽힌다. 새도 부리와 눈을 오른쪽 아래로 향하고 있으며 날아가는 동작은 없다. 배경 물고기들은 수면 아래에서 서로 다른 방향으로 놓여 있다.",
        "built_space": "새가 앉은 어깨 장갑 하나가 중앙 하단에 크게 보이고, 왼쪽 위에 찰리의 목과 옆얼굴, 왼쪽 가장자리에 연결 장갑이 이어진다. 어깨 뒤와 아래의 바다가 배경을 채우며 수중 바위와 물고기가 부드럽게 풀려 보인다. 수평선이나 반사 풍경으로 배경을 대체하지 않았다. 식탁과 섬의 전체 형상은 밀착 구도 밖이며 중복 설비는 없다.",
        "entities": "찰리 한 개체와 작은 새 한 마리만 있다. 찰리의 각진 베이지 장갑, 긁힌 도장, 목의 케이블과 흰 마스크 옆면은 참조 정체성을 유지한다. 새는 참조처럼 황갈색 솜털, 둥근 몸통, 작은 부리, 짧은 날개와 가는 발을 갖추고 특히 머리와 등의 깃털이 부스스하다. 물속 물고기가 보이며 다른 사람이나 읽을 수 있는 글자는 없다. 가슴 고리와 하체는 구도상 보이지 않는다.",
        "hard_violations": [],
        "physics": "새의 두 발과 구부러진 발가락이 어깨 장갑의 윗면과 경사면에 닿아 몸을 받친다. 다리가 몸 아래에 놓이고 날개는 접혀 있어 가만히 내려앉은 뒤의 자세로 성립한다. 깃털과 발의 그림자도 장갑 표면에 자연스럽게 맺힌다. 어깨와 목은 몸통에 연결되어 있으며 지지 없이 떠 있는 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "낡은 어깨에 앉은 작은 새와 아래쪽 맑은 바다를 클로즈업으로 정확히 담았지만, 새의 깃털이 비교적 가지런해 ‘부스스한 아기 새’라는 핵심 묘사는 B보다 약하다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "부스스한 아기 새가 낡은 금속 어깨에 발을 붙이고 가만히 앉은 순간을 밀착해 보여주며, 찰리의 아래쪽 바다를 향한 주의와 물속 물고기도 유지한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "찰리의 옆얼굴은 화면 왼쪽 아래의 바다를 향한다. 눈 전체가 보이지 않아 정확한 시선 종착점은 단정할 수 없지만 머리 방향은 낚시에 집중하는 설정과 맞는다. 새의 부리와 눈은 화면 오른쪽을 향하며 카메라를 정면으로 응시하지 않는다. 배경에는 수면 아래 여러 물고기의 길쭉한 몸이 보인다.",
        "built_space": "전경의 어깨 장갑 하나가 화면 중앙과 하단을 차지하고, 그 뒤 오른쪽에 목과 머리 일부 및 연결 장갑이 보인다. 배경은 어깨 너머 아래로 내려다본 맑은 바다와 수중 바위이며, 반사된 육상 풍경이 아니다. 인공 구조물이나 중복된 고정 설비는 없다. 식탁과 먼 바위섬은 이 클로즈업 밖에 있어 확인할 수 없다.",
        "entities": "찰리 한 개체와 아기 새 한 마리만 보이며 이전 장면의 사람들은 없다. 찰리의 샌드 베이지 장갑, 마모된 금속, 검은 목 기구와 흰 얼굴 일부는 참조와 부합한다. 새는 작은 부리, 둥근 몸통, 짧게 접힌 날개와 가는 발을 갖췄지만 참조보다 다소 짙은 갈색이고 깃털의 흐트러짐이 약하다. 맑은 물속 물고기가 보이고 읽을 수 있는 글자는 없다. 가슴 고리와 전신 비율은 프레임 밖이다.",
        "hard_violations": [],
        "physics": "새의 몸 아래 발과 발가락이 어깨 장갑 윗면에 닿아 체중을 지지하고, 금속 표면에 접촉 그림자가 생긴다. 날개를 접은 정착 자세로, 공중에 떠 있지 않다. 어깨는 찰리의 몸통과 기계 관절에 연결되어 있다. 찰리의 지면 접촉은 프레임 밖이며, 물고기는 물속에 있다."
       },
       {
        "label": "A",
        "direction": "찰리의 옆얼굴은 화면 오른쪽 아래의 바다 쪽으로 기울어 있다. 정확한 눈동자는 보이지 않지만 머리 방향은 아래쪽 물에 주의를 둔 상태로 읽힌다. 새도 부리와 눈을 오른쪽 아래로 향하고 있으며 날아가는 동작은 없다. 배경 물고기들은 수면 아래에서 서로 다른 방향으로 놓여 있다.",
        "built_space": "새가 앉은 어깨 장갑 하나가 중앙 하단에 크게 보이고, 왼쪽 위에 찰리의 목과 옆얼굴, 왼쪽 가장자리에 연결 장갑이 이어진다. 어깨 뒤와 아래의 바다가 배경을 채우며 수중 바위와 물고기가 부드럽게 풀려 보인다. 수평선이나 반사 풍경으로 배경을 대체하지 않았다. 식탁과 섬의 전체 형상은 밀착 구도 밖이며 중복 설비는 없다.",
        "entities": "찰리 한 개체와 작은 새 한 마리만 있다. 찰리의 각진 베이지 장갑, 긁힌 도장, 목의 케이블과 흰 마스크 옆면은 참조 정체성을 유지한다. 새는 참조처럼 황갈색 솜털, 둥근 몸통, 작은 부리, 짧은 날개와 가는 발을 갖추고 특히 머리와 등의 깃털이 부스스하다. 물속 물고기가 보이며 다른 사람이나 읽을 수 있는 글자는 없다. 가슴 고리와 하체는 구도상 보이지 않는다.",
        "hard_violations": [],
        "physics": "새의 두 발과 구부러진 발가락이 어깨 장갑의 윗면과 경사면에 닿아 몸을 받친다. 다리가 몸 아래에 놓이고 날개는 접혀 있어 가만히 내려앉은 뒤의 자세로 성립한다. 깃털과 발의 그림자도 장갑 표면에 자연스럽게 맺힌다. 어깨와 목은 몸통에 연결되어 있으며 지지 없이 떠 있는 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.639
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.639
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt-high": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1639
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "부스스한 아기 새와 낡은 금속 어깨의 질감을 훌륭하게 구현했으며, 물고기가 있는 바다 배경을 부드럽게 아웃포커싱하여 프롬프트가 요구한 공간적 맥락을 정확히 표현했습니다."
   },
   {
    "label": "B",
    "score": 1639,
    "verdict_ko": "요구된 주요 요소들은 잘 배치되었으나, 배경의 물고기가 너무 선명하게 표현되어 '부드럽게 묘사된(softly resolved) 배경'이라는 지시사항의 의도를 다소 벗어났습니다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S91sh4_sel.png",
    "asset_id": "a4bb55ea-1d08-4624-9e8a-1ea09a57a191",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 찰리: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1315284>",
    "asset_id": "bffcb4b3-236e-4732-825b-982c7970e059",
    "role": "character_ref"
   },
   {
    "label": "CHARACTER REFERENCE — 아기 새(B-200이 돌보던 새): the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:833571>",
    "asset_id": "0eaf7589-8c85-4da6-aebd-27ac974cb20f",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae936-140c-7b9e-adc2-6ab1827e8f6a",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S91sh4"
  }
 },
 "S91sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T12:22:06.809716+00:00",
  "fingerprint": "cd5a9d6a972314d3a2df42ca0d3d28f6169b4079b1215c0416c18164c8b20a4c",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S91sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S91sh5_sel.png",
  "source_sha256": "39274b42c0cd8ec571a59300bbaa7180fe468a7b89aba28a8f5c223fa01ef614",
  "file": "S91sh5_cine.png",
  "staged_sha256": "1249eab33bcfbbdc8901dc5256fb9a5b9a42a3d743b0cb8ef67bbe1321906194",
  "latency_ms": 10095
 },
 "S72sh43::cine": {
  "applied": false,
  "attempted_at": "2026-09-19T09:44:42.231220+00:00",
  "fingerprint": "f495faf3b8caba0485601c28a588c87343d94d35283843603013c1ba4e807d72",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S72sh43_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S72sh43_sel.png",
  "source_sha256": "287445ad11e0685d81c96331bb75878da8d4b02a4f675233399cc1d9693769db",
  "error": "RuntimeError: moderation blocked: {\"code\":\"imagine:content-moderated\",\"error\":\"Generated image rejected by content moderation.\",\"usage\":{\"cost_in_usd_ticks\":220000000}}",
  "moderation_refusals": 1,
  "declined": true,
  "declined_reason": "moderation"
 },
 "S72sh43::cine::fb_grok": {
  "applied": false,
  "attempted_at": "2026-09-19T09:44:46.167518+00:00",
  "fingerprint": "e741983211c20eeecff0bab11f00a2f276baef08c7f93af80b06dbbccc419e60",
  "fingerprint_version": 2,
  "provider": "grok",
  "endpoint": "openrouter/chat-completions",
  "model": "x-ai/grok-imagine-image-quality",
  "multiroll_tag": "still_S72sh43_cine_fb_grok",
  "slot": "fb_grok",
  "pack": "24.202608252115",
  "source_file": "S72sh43_sel.png",
  "source_sha256": "287445ad11e0685d81c96331bb75878da8d4b02a4f675233399cc1d9693769db",
  "error": "RuntimeError: Grok image API error 400: {\"error\":{\"message\":\"Provider returned error\",\"code\":400,\"metadata\":{\"raw\":\"{\\\"code\\\":\\\"imagine:content-moderated\\\",\\\"error\\\":\\\"Generated image rejected by content moderation.\\\",\\\"usage\\\":{\\\"cost_in_usd_ticks\\\":600000000}}\",\"provider_name\":\"xAI\",\"is_byok\":false}},\"user_id\":\"org_3HnYgDbzLu9jaNLuyBrNsWTAFk2\"}",
  "moderation_refusals": 1,
  "declined": true,
  "declined_reason": "moderation"
 },
 "S72sh43::cine::fb_mai": {
  "applied": false,
  "attempted_at": "2026-09-19T09:44:55.879634+00:00",
  "fingerprint": "b52603a921c6bda58173dcc9c3ee887f26e32a268f15f12ebd1d15910ee02fad",
  "fingerprint_version": 2,
  "provider": "mai",
  "endpoint": "openrouter/chat-completions",
  "model": "microsoft/mai-image-2.6",
  "multiroll_tag": "still_S72sh43_cine_fb_mai",
  "slot": "fb_mai",
  "pack": "24.202608252115",
  "source_file": "S72sh43_sel.png",
  "source_sha256": "287445ad11e0685d81c96331bb75878da8d4b02a4f675233399cc1d9693769db",
  "error": "RuntimeError: Grok image API error 400: {\"error\":{\"message\":\"Provider returned error\",\"code\":400,\"metadata\":{\"raw\":\"{\\\"error\\\":{\\\"code\\\":\\\"content_safety_violation\\\",\\\"message\\\":\\\"Response content blocked by label 'MultiSeverity_ViolenceScore'.\\\",\\\"details\\\":\\\"Response content blocked by label 'MultiSeverity_ViolenceScore'.\\\"}}\",\"provider_name\":\"Azure\",\"is_byok\":false,\"provider_error_code\":\"content_safety_violation\"}},\"user_id\":\"org_3HnYgDbzLu9jaNLuyBrNsWTAFk2\"}",
  "moderation_refusals": 1,
  "declined": true,
  "declined_reason": "moderation"
 },
 "S72sh43::cine::fb_seedream": {
  "applied": true,
  "attempted_at": "2026-09-19T10:24:14.237085+00:00",
  "fingerprint": "defd913d6021803b1723637d46e7fcb37dd5d6377d7199f9bd54be5d2900af75",
  "fingerprint_version": 2,
  "provider": "seedream",
  "endpoint": "openrouter/images+aspect_ratio",
  "model": "bytedance-seed/seedream-5-0-pro",
  "multiroll_tag": "still_S72sh43_cine_fb_seedream",
  "slot": "fb_seedream",
  "pack": "24.202608252115",
  "source_file": "S72sh43_sel.png",
  "source_sha256": "287445ad11e0685d81c96331bb75878da8d4b02a4f675233399cc1d9693769db",
  "file": "S72sh43_cine_fb_seedream.png",
  "staged_sha256": "20980c78ad9e693292395e8f43992b59008a81b34d8e467e275d2f3f444f9afe",
  "latency_ms": 112870
 },
 "S11sh1::ab_noconti": {
  "input_fingerprint": "3514e7148b68891b",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 어둠이 내린 좁은 골목길을 향해 닫히는 도중인 육중한 철창문 틈새로 몸을 비틀어 구겨 넣은 자세의 현우의 다급한 전신.\n\nLOCATION (lock): At the narrowing opening of the refugee settlement's heavy barred entrance gate, opening onto a dark alley. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Closing entrance gate in the middle-left of the frame, midground; Alley beyond the entrance in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Heavy barred entrance gate (Closing, leaving a narrowing passage) — Seen obliquely from inside the entrance, with its closing edge beside 현우; used as Creates the constricting left boundary without obscuring his complete body; Entrance threshold and alley (현우 is crossing into the settlement) — The threshold crosses the lower field, and the alley continues toward screen right; used as Establishes an unambiguous inward route for the subsequent pan.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The nighttime entrance remains dim after the settlement's lights go out, with only enough ambient visibility to distinguish his body from the gate.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camp entrance is closing for curfew; shop doors are shutting and the container homes and streetlights are being extinguished, leaving the streets dark. 현우: Hyunwoo is hurrying through the entrance before it closes, still carrying the untreated dog-bite injury and earlier head injury.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 어둠이 내린 좁은 골목길을 향해 닫히는 도중인 육중한 철창문 틈새로 몸을 비틀어 구겨 넣은 자세의 현우의 다급한 전신.\n\nLOCATION (lock): At the narrowing opening of the refugee settlement's heavy barred entrance gate, opening onto a dark alley. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Closing entrance gate in the middle-left of the frame, midground; Alley beyond the entrance in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Heavy barred entrance gate (Closing, leaving a narrowing passage) — Seen obliquely from inside the entrance, with its closing edge beside 현우; used as Creates the constricting left boundary without obscuring his complete body; Entrance threshold and alley (현우 is crossing into the settlement) — The threshold crosses the lower field, and the alley continues toward screen right; used as Establishes an unambiguous inward route for the subsequent pan.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The nighttime entrance remains dim after the settlement's lights go out, with only enough ambient visibility to distinguish his body from the gate.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camp entrance is closing for curfew; shop doors are shutting and the container homes and streetlights are being extinguished, leaving the streets dark. 현우: Hyunwoo is hurrying through the entrance before it closes, still carrying the untreated dog-bite injury and earlier head injury.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "현우의 시선과 얼굴은 프레임 왼쪽 바깥을 다급하게 주시하고 있습니다.",
    "built_space": "화면 왼쪽 중경에 닫히는 철문이 있고 우측 배경으로 어두운 골목길이 보이나, 고정된 로케이션 레퍼런스 사진의 '마름모 패턴이 있는 접이식 철문'이 아닌 수직 창살로만 이루어진 완전히 다른 형태의 슬라이딩 철문이 렌더링되었습니다.",
    "entities": "현우의 외모, 앳된 얼굴과 헝클어진 머리는 캐릭터 레퍼런스와 일치하며, 프롬프트에 명시된 머리와 신체의 부상 역시 붕대로 표현되었습니다.",
    "hard_violations": [
     "절대적으로 고정된 로케이션 레퍼런스의 구조(마름모꼴 패턴의 접이식 철문)를 무시하고 임의로 창작된 형태의 수직 철창문을 렌더링함."
    ],
    "physics": "현우는 양발을 땅에 안정적으로 딛고 양손으로 철문을 잡은 채 서 있으며, 프롬프트가 요구한 '문 틈새로 몸을 비틀어 구겨 넣는' 물리적인 압박감이나 역동적인 포즈는 전혀 형성되지 않았습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "현우의 인물 레퍼런스와 부상 상태(머리 및 팔의 붕대)는 훌륭히 구현되었으나, 로케이션 레퍼런스에 명시된 마름모꼴 접이식 철문의 형태를 완전히 무시하였고 문 틈새로 몸을 비틀어 구겨 넣는다는 핵심 동작이 누락되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선과 얼굴은 프레임 왼쪽 바깥을 다급하게 주시하고 있습니다.",
        "built_space": "화면 왼쪽 중경에 닫히는 철문이 있고 우측 배경으로 어두운 골목길이 보이나, 고정된 로케이션 레퍼런스 사진의 '마름모 패턴이 있는 접이식 철문'이 아닌 수직 창살로만 이루어진 완전히 다른 형태의 슬라이딩 철문이 렌더링되었습니다.",
        "entities": "현우의 외모, 앳된 얼굴과 헝클어진 머리는 캐릭터 레퍼런스와 일치하며, 프롬프트에 명시된 머리와 신체의 부상 역시 붕대로 표현되었습니다.",
        "hard_violations": [
         "절대적으로 고정된 로케이션 레퍼런스의 구조(마름모꼴 패턴의 접이식 철문)를 무시하고 임의로 창작된 형태의 수직 철창문을 렌더링함."
        ],
        "physics": "현우는 양발을 땅에 안정적으로 딛고 양손으로 철문을 잡은 채 서 있으며, 프롬프트가 요구한 '문 틈새로 몸을 비틀어 구겨 넣는' 물리적인 압박감이나 역동적인 포즈는 전혀 형성되지 않았습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "현우의 인물 레퍼런스와 부상 상태(머리 및 팔의 붕대)는 훌륭히 구현되었으나, 로케이션 레퍼런스에 명시된 마름모꼴 접이식 철문의 형태를 완전히 무시하였고 문 틈새로 몸을 비틀어 구겨 넣는다는 핵심 동작이 누락되었습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "현우의 시선과 얼굴은 프레임 왼쪽 바깥을 다급하게 주시하고 있습니다.",
        "built_space": "화면 왼쪽 중경에 닫히는 철문이 있고 우측 배경으로 어두운 골목길이 보이나, 고정된 로케이션 레퍼런스 사진의 '마름모 패턴이 있는 접이식 철문'이 아닌 수직 창살로만 이루어진 완전히 다른 형태의 슬라이딩 철문이 렌더링되었습니다.",
        "entities": "현우의 외모, 앳된 얼굴과 헝클어진 머리는 캐릭터 레퍼런스와 일치하며, 프롬프트에 명시된 머리와 신체의 부상 역시 붕대로 표현되었습니다.",
        "hard_violations": [
         "절대적으로 고정된 로케이션 레퍼런스의 구조(마름모꼴 패턴의 접이식 철문)를 무시하고 임의로 창작된 형태의 수직 철창문을 렌더링함."
        ],
        "physics": "현우는 양발을 땅에 안정적으로 딛고 양손으로 철문을 잡은 채 서 있으며, 프롬프트가 요구한 '문 틈새로 몸을 비틀어 구겨 넣는' 물리적인 압박감이나 역동적인 포즈는 전혀 형성되지 않았습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "밤의 다급한 전신 동작은 보이지만 고정 장소와 의상이 크게 다르고 틈새에 몸을 구겨 넣는 순간도 약하다. B가 없어 비교 우승은 확정할 수 없다."
       },
       {
        "label": "B",
        "score": 0,
        "verdict_ko": "후보 B 이미지가 제공되지 않아 평가 불가하며, 0점은 이미지 품질에 대한 판정이 아니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 얼굴과 시선은 카메라 쪽 앞 공간을 향하고, 두 손은 왼쪽 철문 끝을 붙잡는다. 한쪽 다리를 화면 오른쪽으로 벌렸지만 문턱을 넘어 정착촌 안으로 진입하는 방향은 명확하지 않다. 골목은 오른쪽 배경으로 이어진다.",
        "built_space": "왼쪽에 높은 철문 한 짝, 오른쪽에 철제 문 또는 울타리 면 하나, 양옆 벽체와 하단 문턱이 보인다. 참조의 낮은 마름모 장식 철문과 개방된 출입구 대신 상단이 화면 밖까지 솟은 직선 철창과 불투명 금속판을 갖춘 다른 출입구다. 문 끝과 오른쪽 문설주 사이도 몸을 비틀어야 할 만큼 좁아 보이지 않는다. 문이 왼쪽 다리 일부를 가린다.",
        "entities": "인물은 한 명이며, 앳된 동아시아계 남성의 얼굴과 검은 흐트러진 머리는 현우의 설정에 대체로 맞는다. 다만 참조의 남색 둥근목 티셔츠 대신 회색 계열 겉옷과 안옷을 입었다. 이마 상처와 손목 붕대는 보이지만 치료하지 않은 개 물림 상처는 식별되지 않는다. 무거운 철문과 어두운 골목은 있으나 철문의 정체성은 장소 참조와 다르다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [
         "고정 장소의 핵심 구조인 낮은 마름모 장식 출입문을 높은 직선 철창·금속판 출입구로 대체해, 참조와 다른 장소 구조를 만들었다."
        ],
        "physics": "화면 오른쪽 신발이 바닥에 닿고 반대쪽 발도 문 아래 바닥 부근에 보이며, 두 손은 문 가장자리를 잡는다. 벌린 다리와 기울인 상체에는 지지가 있어 부유하거나 불가능한 자세는 아니다. 문도 바닥의 레일 또는 하부 지지부에 놓여 있다. 다만 자세는 좁아지는 틈을 통과한다기보다 문을 붙들고 버티는 동작에 가깝다."
       },
       {
        "label": "B",
        "direction": "이미지가 제공되지 않아 시선과 이동 방향을 관찰할 수 없다.",
        "built_space": "이미지가 제공되지 않아 구조와 배치를 관찰할 수 없다.",
        "entities": "이미지가 제공되지 않아 인물과 사물을 확인할 수 없다.",
        "hard_violations": [],
        "physics": "이미지가 제공되지 않아 신체와 사물의 지지 관계를 확인할 수 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": null,
     "ok": false
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gpt-high"
   ],
   "route": "single_forward"
  },
  "totals": {
   "A": 2
  },
  "selected": "A",
  "ranking": [
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2,
    "verdict_ko": "현우의 인물 레퍼런스와 부상 상태(머리 및 팔의 붕대)는 훌륭히 구현되었으나, 로케이션 레퍼런스에 명시된 마름모꼴 접이식 철문의 형태를 완전히 무시하였고 문 틈새로 몸을 비틀어 구겨 넣는다는 핵심 동작이 누락되었습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_refugee_gate_sel.png",
    "asset_id": "331ee13b-1a46-4cba-9a37-07f52e1d1493",
    "role": "location_seed_bg"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": true,
  "shot_run_uid": "06aae5a9-44fb-76cb-b780-94b77fadbb95"
 },
 "S11sh1::conti_ab_decision": {
  "fingerprint": "d19074241d68200d",
  "winner": "A",
  "outer": {
   "winner": "A",
   "verdicts": [
    {
     "order": "AB",
     "winner_label": "1",
     "winner_branch": "A",
     "reason_ko": "1번 후보는 인물의 전신을 가리지 않고 철창문 틈새로 몸을 구겨 넣는 다급한 자세와 요구된 화면 구성을 정확하게 구현했습니다."
    },
    {
     "order": "BA",
     "winner_label": "2",
     "winner_branch": "A",
     "reason_ko": "닫히는 철창문 틈새로 몸을 비틀어 통과하는 지정된 동작과 프레이밍을 정확히 구현한 반면, 1은 문 뒤에 서 있는 모습임."
    }
   ]
  }
 },
 "S11sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T09:50:44.063610+00:00",
  "fingerprint": "49d012b273b8fa4c0375f7da27aeb0501f1ac7a1bf72247b190c888ecd750a4d",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S11sh1_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S11sh1_sel.png",
  "source_sha256": "06f20e53d6cec76ee596496c74ea61e986e574b777f028207fba0d7afaf0b8bf",
  "file": "S11sh1_cine.png",
  "staged_sha256": "f1a7303be1c1819043510adec3b043e95a1df1895610165a7ac532db99192df2",
  "latency_ms": 10220
 },
 "S11sh1": {
  "input_fingerprint": "72106a82e86350ac",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 어둠이 내린 좁은 골목길을 향해 닫히는 도중인 육중한 철창문 틈새로 몸을 비틀어 구겨 넣은 자세의 현우의 다급한 전신.\n\nLOCATION (lock): At the narrowing opening of the refugee settlement's heavy barred entrance gate, opening onto a dark alley. The shot takes place here — the attached LOCATION STRUCTURE PHOTOGRAPH is the single authority for this exact place — its fixed structure and permanent site details are LOCKED to it. No separate location photograph exists for this place. Build everything else strictly from the location text above and the shot text; the layout sketch (when attached) governs framing and placement only, and the shot text governs time of day, lighting and action.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Closing entrance gate in the middle-left of the frame, midground; Alley beyond the entrance in the middle-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: Heavy barred entrance gate (Closing, leaving a narrowing passage) — Seen obliquely from inside the entrance, with its closing edge beside 현우; used as Creates the constricting left boundary without obscuring his complete body; Entrance threshold and alley (현우 is crossing into the settlement) — The threshold crosses the lower field, and the alley continues toward screen right; used as Establishes an unambiguous inward route for the subsequent pan.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The nighttime entrance remains dim after the settlement's lights go out, with only enough ambient visibility to distinguish his body from the gate.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The camp entrance is closing for curfew; shop doors are shutting and the container homes and streetlights are being extinguished, leaving the streets dark. 현우: Hyunwoo is hurrying through the entrance before it closes, still carrying the untreated dog-bite injury and earlier head injury.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 현우 (18세 남성, 한국계 미국인, 앳된 얼굴, 헝클어진 검은 머리) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "shot_run_spend_attempt_count": 3,
  "readings": [
   {
    "label": "A",
    "direction": "시선은 골목길 안쪽(화면 우측)을 향하고 있으며, 몸통은 좁아지는 문틈을 통과하는 방향으로 움직이고 있음.",
    "built_space": "마름모 장식의 철제 대문, 바닥 레일, 좌측 컨테이너 구조물과 우측 원형 조명 기둥이 레퍼런스와 일치하며 지정된 야간 환경에 맞게 배치됨.",
    "entities": "현우(젊은 아시아인 남성, 헝클어진 검은 머리)의 외형이 일치하며, 얼굴의 상처와 바지 다리 부분의 핏자국(개에게 물린 상처)이 확인됨.",
    "hard_violations": [],
    "physics": "양발로 지면을 딛고 오른손으로 닫히는 철문을 짚으며 틈새를 비집고 들어가는 체중 이동과 지지가 물리적으로 자연스러움."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "지정된 야간 조명 속에서 닫히는 문틈을 빠져나오는 다급한 동작과 프레이밍, 그리고 머리와 다리의 상처 세부 사항까지 샷 텍스트를 충실히 구현함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선은 골목길 안쪽(화면 우측)을 향하고 있으며, 몸통은 좁아지는 문틈을 통과하는 방향으로 움직이고 있음.",
        "built_space": "마름모 장식의 철제 대문, 바닥 레일, 좌측 컨테이너 구조물과 우측 원형 조명 기둥이 레퍼런스와 일치하며 지정된 야간 환경에 맞게 배치됨.",
        "entities": "현우(젊은 아시아인 남성, 헝클어진 검은 머리)의 외형이 일치하며, 얼굴의 상처와 바지 다리 부분의 핏자국(개에게 물린 상처)이 확인됨.",
        "hard_violations": [],
        "physics": "양발로 지면을 딛고 오른손으로 닫히는 철문을 짚으며 틈새를 비집고 들어가는 체중 이동과 지지가 물리적으로 자연스러움."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "지정된 야간 조명 속에서 닫히는 문틈을 빠져나오는 다급한 동작과 프레이밍, 그리고 머리와 다리의 상처 세부 사항까지 샷 텍스트를 충실히 구현함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "시선은 골목길 안쪽(화면 우측)을 향하고 있으며, 몸통은 좁아지는 문틈을 통과하는 방향으로 움직이고 있음.",
        "built_space": "마름모 장식의 철제 대문, 바닥 레일, 좌측 컨테이너 구조물과 우측 원형 조명 기둥이 레퍼런스와 일치하며 지정된 야간 환경에 맞게 배치됨.",
        "entities": "현우(젊은 아시아인 남성, 헝클어진 검은 머리)의 외형이 일치하며, 얼굴의 상처와 바지 다리 부분의 핏자국(개에게 물린 상처)이 확인됨.",
        "hard_violations": [],
        "physics": "양발로 지면을 딛고 오른손으로 닫히는 철문을 짚으며 틈새를 비집고 들어가는 체중 이동과 지지가 물리적으로 자연스러움."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "전신 와이드 구도와 오른쪽 골목, 문을 붙잡고 비집고 들어오는 동작은 부합하지만 철문이 장소 사진보다 지나치게 높으며, B가 제공되지 않아 비교 승자는 잠정적입니다."
       },
       {
        "label": "B",
        "score": 0,
        "verdict_ko": "이미지가 제공되지 않아 평가할 수 없으며, 0점과 최하위 배치는 실제 충실도 판정이 아닌 미제공 표시입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "현우의 얼굴과 시선은 화면 오른쪽 골목 방향을 향한다. 상체와 앞다리도 오른쪽으로 나가고 뒤로 뻗은 손은 철문 가장자리를 잡는다. 문을 통과하는 동작은 보이지만, 배경에 정착촌 컨테이너들이 보여 정착촌 안으로 들어오는 방향인지는 명확하지 않다.",
        "built_space": "왼쪽의 큰 철문 면과 중앙의 비스듬한 철문 면 사이에 현우가 있다. 하단에는 문턱과 문 바퀴가 있고, 오른쪽 문기둥에는 구형 등 하나와 확성기 하나가 보인다. 왼쪽 가장자리에는 볼록거울 하나와 일부 잘린 구형 등이 있다. 오른쪽 뒤로 컨테이너 사이 통로가 이어진다. 마름모 장식과 재료는 참고 장소와 유사하지만, 철문 높이와 장식 비례는 사진보다 크게 확대되었다.",
        "entities": "인물은 한 명으로, 앳된 동아시아계 남성의 외모와 헝클어진 검은 머리, 짙은 반팔 티셔츠가 현우 참고 이미지에 대체로 맞는다. 국적은 외관만으로 확인할 수 없다. 이마의 상처와 바지 정강이 부근의 혈흔은 보이지만 혈흔만으로 개에 물린 상처인지 확정할 수 없다. 철창문과 어두운 골목이 있으며 읽을 수 있는 글자는 보이지 않는다.",
        "hard_violations": [],
        "physics": "뒤로 뻗은 손이 문 가장자리를 잡고 있어 몸을 비틀 때의 버팀점이 있다. 앞발은 지면 가까이 내딛는 중이고 뒷발의 접지는 문 가장자리에 가려 확인하기 어렵다. 허리와 무릎의 굽힘은 틈을 통과하는 보행 동작으로 가능하며, 명백한 무지지 부유로 단정할 근거는 없다. 철문은 하단 바퀴와 문기둥 쪽 구조로 지지된다."
       },
       {
        "label": "B",
        "direction": "이미지가 제공되지 않아 시선과 이동 방향을 관찰할 수 없다.",
        "built_space": "이미지가 제공되지 않아 문과 통로의 구조 및 인물 배치를 관찰할 수 없다.",
        "entities": "이미지가 제공되지 않아 인물과 사물을 확인할 수 없다.",
        "hard_violations": [],
        "physics": "이미지가 제공되지 않아 접지, 지지점 및 동작의 물리적 타당성을 확인할 수 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": null,
     "ok": false
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gpt-high"
   ],
   "route": "single_forward"
  },
  "totals": {
   "A": 8
  },
  "selected": "A",
  "ranking": [
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 8,
    "verdict_ko": "지정된 야간 조명 속에서 닫히는 문틈을 빠져나오는 다급한 동작과 프레이밍, 그리고 머리와 다리의 상처 세부 사항까지 샷 텍스트를 충실히 구현함."
   }
  ],
  "refs": [
   {
    "label": "LOCATION STRUCTURE PHOTOGRAPH — the confirmed photograph of this exact place and its fixed structure: it is the SINGLE authority for the location, the structure's shape, proportions, materials, colors, openings and every permanent site detail. Never copy its camera framing, time of day or lighting — the shot text is the authority for those.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/background_chain/seed_bg_refugee_gate_sel.png",
    "asset_id": "331ee13b-1a46-4cba-9a37-07f52e1d1493",
    "role": "location_seed_bg"
   },
   {
    "label": "LAYOUT SKETCH — a bare thin-line layout guide, a REFERENCE ONLY: take from it ONLY the camera framing, figure placement, pose and size/depth order. It carries ZERO visual style — every texture, material, light and all realism come from the text and the photographic reference. Never let any line-drawing quality leak into the output.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/conti/conti_S11sh1.png",
    "asset_id": "37dd0d43-2fdf-429b-82e4-045c520df701",
    "role": "conti_light"
   },
   {
    "label": "CHARACTER REFERENCE — 현우: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1443881>",
    "asset_id": "d578d6b8-dec6-4e4e-b744-95ce13a02eaf",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae926-c7d0-79ae-94ce-7c30073bde42",
  "conti_ab": {
   "winner": "A",
   "outer": {
    "winner": "A",
    "verdicts": [
     {
      "order": "AB",
      "winner_label": "1",
      "winner_branch": "A",
      "reason_ko": "1번 후보는 인물의 전신을 가리지 않고 철창문 틈새로 몸을 구겨 넣는 다급한 자세와 요구된 화면 구성을 정확하게 구현했습니다."
     },
     {
      "order": "BA",
      "winner_label": "2",
      "winner_branch": "A",
      "reason_ko": "닫히는 철창문 틈새로 몸을 비틀어 통과하는 지정된 동작과 프레이밍을 정확히 구현한 반면, 1은 문 뒤에 서 있는 모습임."
     }
    ]
   },
   "outer_judged_this_run": false,
   "sel_conti": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S11sh1__ab_conti_sel.png",
   "sel_noconti": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S11sh1__ab_noconti_sel.png"
  },
  "ref_mode": "seed-bg+콘티+엔티티 (복잡구조물 A/B: 콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  },
  "lane_policy": "ab_select_ready"
 },
 "S48sh5::confined_fp": {
  "reads": {
   "controls": "The only primary control drawn is the steering wheel, attached to the front-left driver station immediately behind the windshield. It is not attached to Amber’s front-right passenger station. Both front seats face toward the windshield.",
   "mirrors": "No mirror or other explicitly reflective surface is marked. The windshield is identified as a window, with no indicated reflective face or reflection.",
   "camera": "The camera symbol sits immediately outboard of the front-right passenger seat and points inward, across the vehicle toward Amber. Its dashed viewing cone specifies a close-up of her face, rather than a forward view through the windshield. From this camera, the vehicle’s front is screen-right and its rear is screen-left.",
   "occupants": "Amber occupies the front-right passenger seat. Charlie occupies the rear-right seat and has a blanket over him. The driver seat and rear-left seat are not marked as occupied. Two police figures stand at the checkpoint outside, beyond the windshield; they are outside the indicated close-up."
  },
  "mismatches": [
   "The LOCATION specifies looking through the windshield toward the checkpoint, but the drawn camera points sideways inward at Amber from beside the passenger seat. The windshield and checkpoint are consequently outside this indicated close-up, toward screen-right."
  ],
  "scene_description_en": "The camera looks inward from beside the front-right passenger station, framing Amber’s face close and centrally rather than looking forward through the windshield. Amber occupies that passenger seat, which faces the vehicle’s front toward screen-right, making this a side-on view of her forward-facing position. Only a narrow portion of her passenger seat belongs beside and behind her in the close framing, toward screen-left. The unoccupied driver station lies farther across the cabin, with its steering wheel forward of the seat toward screen-right and outside the tight facial crop. Charlie occupies the rear-right seat under a blanket, facing forward toward screen-right but located off-screen to the left; the rear-left seat is also outside the frame and has no marked occupant. The windshield is forward and off-screen to the right, with the distant road barricades, two police figures, checkpoint booth, police vehicle, dead trees, and ruined houses beyond it, not visible in this sideways close-up. No mirror is drawn, so there is no mirror face or reflected view in the composition.",
  "readback_fallback": {
   "first_model": "gemini-pro",
   "first_error": "LLM returned empty response for step=confined_fp_readback_S48sh5_fix, model=gemini-pro, finish_reason='content_filter'",
   "model": "gpt-high",
   "physical_model": "gpt-6-astra"
  },
  "fixed": true,
  "input_fingerprint": "24c8c21baf352f38"
 },
 "era_assess::a0596258cfca1fe0": {
  "subjects": [],
  "subject_text": "화성의 황량한 도로\n황량한 들판을 가로지르는 도로. 주변에는 메마른 나무와 폐가가 드문드문 남아 있고 멀리 검문 시설이 보인다.",
  "identity": "canonical",
  "scope_id": "L203",
  "scope_role": "location_exterior",
  "scope_sha": "baf8d1a915ba6e43"
 },
 "S48sh5": {
  "input_fingerprint": "f25f1166452f49fe",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 멀리 도로 끝에 세워진 바리케이드와 경찰 검문소를 바라보며 두 눈을 동그랗게 뜬 앰버의 놀란 얼굴.\n\nLOCATION (lock): Inside the camper's front passenger seat, looking through the windshield toward a distant road checkpoint in daylight. The shot takes place here — the LOCATION text above is the only authority for this place — no location photograph is attached. Build the place strictly from that text and the shot text, inventing nothing beyond them.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Front passenger seat (Occupied by 앰버) — A narrow portion of the seat beside her shoulder is visible; used as Grounds the close reaction within the vehicle without revealing the other passengers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light preserves clear eye detail and restrained contrast without exaggerating her surprise.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The aging camper has a leaking roof and window frames, with the gathered supplies aboard; a police checkpoint stands farther along the road through barren fields, dead trees and abandoned houses. Charlie retains his worn metal body and blanket covering and is in the rear seat. 앰버: She occupies the passenger seat with the map and wears the replacement shoes obtained at the unmanned store.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera looks inward from beside the front-right passenger station, framing Amber’s face close and centrally rather than looking forward through the windshield. Amber occupies that passenger seat, which faces the vehicle’s front toward screen-right, making this a side-on view of her forward-facing position. Only a narrow portion of her passenger seat belongs beside and behind her in the close framing, toward screen-left. The unoccupied driver station lies farther across the cabin, with its steering wheel forward of the seat toward screen-right and outside the tight facial crop. Charlie occupies the rear-right seat under a blanket, facing forward toward screen-right but located off-screen to the left; the rear-left seat is also outside the frame and has no marked occupant. The windshield is forward and off-screen to the right, with the distant road barricades, two police figures, checkpoint booth, police vehicle, dead trees, and ruined houses beyond it, not visible in this sideways close-up. No mirror is drawn, so there is no mirror face or reflected view in the composition.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 멀리 도로 끝에 세워진 바리케이드와 경찰 검문소를 바라보며 두 눈을 동그랗게 뜬 앰버의 놀란 얼굴.\n\nLOCATION (lock): Inside the camper's front passenger seat, looking through the windshield toward a distant road checkpoint in daylight. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light preserves clear eye detail and restrained contrast without exaggerating her surprise.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The aging camper has a leaking roof and window frames, with the gathered supplies aboard; a police checkpoint stands farther along the road through barren fields, dead trees and abandoned houses. Charlie retains his worn metal body and blanket covering and is in the rear seat. 앰버: She occupies the passenger seat with the map and wears the replacement shoes obtained at the unmanned store.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE 16:9 photorealistic film still for the brief below.\n\nThe FIRST attached image is a top-down FLOOR PLAN of this interior and\nthe SCENE LAYOUT text below is what a careful reader saw in it.\nTogether they are the ONLY authority for physical arrangement: which\nseat/station each person occupies, which station every primary control\nbelongs to, where any mirror/reflective surface sits and what it can\nphysically reflect, where the camera stands and what appears on which\nside of the screen. If any other sentence seems to contradict them, the\nfloor plan wins. The floor plan is a diagram, not scenery — none of its\nlines, arrows or labels may appear in the photograph. WHO the people\nare and what they do comes from the SHOT TEXT and the attached\nCHARACTER/PROP references — never add a person the SHOT TEXT does not\nplace here. No text, no watermarks.\n\nSCENE LAYOUT (what a careful reader saw in the attached floor plan):\nThe camera looks inward from beside the front-right passenger station, framing Amber’s face close and centrally rather than looking forward through the windshield. Amber occupies that passenger seat, which faces the vehicle’s front toward screen-right, making this a side-on view of her forward-facing position. Only a narrow portion of her passenger seat belongs beside and behind her in the close framing, toward screen-left. The unoccupied driver station lies farther across the cabin, with its steering wheel forward of the seat toward screen-right and outside the tight facial crop. Charlie occupies the rear-right seat under a blanket, facing forward toward screen-right but located off-screen to the left; the rear-left seat is also outside the frame and has no marked occupant. The windshield is forward and off-screen to the right, with the distant road barricades, two police figures, checkpoint booth, police vehicle, dead trees, and ruined houses beyond it, not visible in this sideways close-up. No mirror is drawn, so there is no mirror face or reflected view in the composition.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational. TIME OF DAY (lock): day.\n\nSHOT TEXT (authoritative, Korean): 멀리 도로 끝에 세워진 바리케이드와 경찰 검문소를 바라보며 두 눈을 동그랗게 뜬 앰버의 놀란 얼굴.\n\nLOCATION (lock): Inside the camper's front passenger seat, looking through the windshield toward a distant road checkpoint in daylight. The shot takes place here.\n\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light preserves clear eye detail and restrained contrast without exaggerating her surprise.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The aging camper has a leaking roof and window frames, with the gathered supplies aboard; a police checkpoint stands farther along the road through barren fields, dead trees and abandoned houses. Charlie retains his worn metal body and blanket covering and is in the rear seat. 앰버: She occupies the passenger seat with the map and wears the replacement shoes obtained at the unmanned store.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앰버 (10세 여자아이, 한국계 백인 혼혈, 금발 머리, 커다란 눈, 둥근 얼굴 윤곽) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S48sh5_confinedfp.png",
     "asset_id": null,
     "role": null
    },
    {
     "label": "앰버",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S48sh5_confinedfp.png",
     "asset_id": null,
     "role": null
    },
    {
     "label": "앰버",
     "path": "<bytes:1294604>",
     "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "B",
    "direction": "앰버는 앞유리 너머 검문소를 등진 채 얼굴을 뒤쪽으로 돌리고 화면 오른쪽을 바라본다. 시선은 화면 중앙 앞유리 밖의 바리케이드나 경찰 검문소에 닿지 않는다. 경찰 두 명은 캠핑카 쪽을 향해 서 있다.",
    "built_space": "뒤 좌석 부근에서 앞을 보는 넓은 구도로, 앞좌석 등받이 두 개, 운전대 하나, 대시보드, 앞유리 하나와 선바이저 두 개가 보인다. 앰버는 오른쪽 조수석에 있지만 어깨 옆 좌석 일부만 보여야 하는 요구와 달리 양쪽 좌석과 운전 공간이 크게 노출된다. 앞유리 밖에는 콘크리트 차단물 두 개, 줄무늬 바리케이드 두 개, 초소 하나와 경찰차 한 대가 보이며 검문소가 비교적 가깝고 크게 표현되었다.",
    "entities": "앰버는 금발에 둥근 얼굴을 가진 약 10세 여자아이로, 참조의 얼굴과 남색 상의에 대체로 부합한다. 한국계 백인 혼혈이라는 배경은 외형만으로 확정할 수 없다. 지도는 화면 아래에 일부 보이고, 신발과 찰리 및 보급품은 프레임 밖이라 평가하지 않는다. 황량한 들판, 고사목, 폐가, 바리케이드와 초소는 보이지만, 이 숏에 허용되지 않은 경찰 두 명도 명확히 등장한다. 뚜렷하게 읽히는 글자는 없다.",
    "hard_violations": [
     "앰버만 등장하도록 제한된 숏에 경찰 두 명을 추가했다."
    ],
    "physics": "앰버의 하체는 가려져 있지만 몸통은 조수석 등받이 앞에 위치하며 뒤돌아보는 자세 자체는 가능하다. 지도는 무릎 부근에 놓인 것으로 보이고 접촉 지점은 화면 아래에 가려져 있어 공중 부유라고 볼 근거는 없다. 경찰은 발을 도로에 딛고 있고 차단물과 초소, 차량도 지면에 놓여 있다."
   },
   {
    "label": "A",
    "direction": "앰버의 얼굴과 두 눈은 카메라 가까운 화면 왼쪽 방향을 향한다. 검문소는 얼굴 뒤 화면 오른쪽 창밖에 있어 그녀가 바라보는 대상과 일치하지 않는다. 경찰 두 명은 차량 쪽을 향해 서 있다.",
    "built_space": "얼굴 중심 클로즈업이며 화면 왼쪽에 좌석 등받이와 머리받침 하나가 일부 보인다. 오른쪽에는 문 손잡이와 팔걸이가 달린 측면 문 하나, 측면 창 하나, 사이드미러 하나가 보이고 위에는 선바이저 하나가 걸쳐 있다. 검문소는 이 측면 창 너머에 배치되어 앞유리를 통해 도로 끝을 바라보는 장소 설정과 맞지 않는다. 밖에는 줄무늬 바리케이드 두 개, 초소 하나, 경찰차 한 대가 보인다.",
    "entities": "금발, 큰 눈, 둥근 얼굴과 남색 상의가 참조 앰버에 대체로 맞으며 약 10세 여자아이로 보인다. 혼혈 배경 자체는 외형만으로 검증할 수 없다. 눈을 크게 뜨고 입을 살짝 벌린 놀람이 보이며 눈의 해부학적 형태는 정상이다. 고사목과 폐가, 검문 시설이 보이지만 허용되지 않은 경찰 두 명도 등장한다. 지도, 신발, 찰리와 보급품은 클로즈업 밖이므로 미노출을 감점하지 않는다. 차량 측면 표기는 흐려 명확한 판독이 어렵다.",
    "hard_violations": [
     "앰버만 등장하도록 제한된 숏에 경찰 두 명을 추가했다."
    ],
    "physics": "앰버는 좌석 등받이 앞에 상체를 세우고 있으며 목과 어깨의 연결 및 앉은 자세에 물리적 모순은 보이지 않는다. 골반과 발은 프레임 밖이다. 경찰 두 명과 바리케이드, 초소, 경찰차는 모두 도로 위에 지지되어 있고 떠 있는 인물이나 물체는 없다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": null,
     "normalized": null,
     "ok": false
    },
    {
     "model": "gpt-high",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "얼굴 클로즈업 대신 차량 내부를 넓게 보여주며, 앰버가 앞의 검문소가 아닌 뒤쪽을 돌아보고 허용되지 않은 경찰 두 명까지 등장한다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "얼굴 클로즈업과 동그랗게 뜬 눈은 더 충실하지만, 검문소가 앞유리가 아닌 측면 창밖에 있고 시선도 그곳을 향하지 않으며 경찰 두 명이 추가되었다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앰버는 앞유리 너머 검문소를 등진 채 얼굴을 뒤쪽으로 돌리고 화면 오른쪽을 바라본다. 시선은 화면 중앙 앞유리 밖의 바리케이드나 경찰 검문소에 닿지 않는다. 경찰 두 명은 캠핑카 쪽을 향해 서 있다.",
        "built_space": "뒤 좌석 부근에서 앞을 보는 넓은 구도로, 앞좌석 등받이 두 개, 운전대 하나, 대시보드, 앞유리 하나와 선바이저 두 개가 보인다. 앰버는 오른쪽 조수석에 있지만 어깨 옆 좌석 일부만 보여야 하는 요구와 달리 양쪽 좌석과 운전 공간이 크게 노출된다. 앞유리 밖에는 콘크리트 차단물 두 개, 줄무늬 바리케이드 두 개, 초소 하나와 경찰차 한 대가 보이며 검문소가 비교적 가깝고 크게 표현되었다.",
        "entities": "앰버는 금발에 둥근 얼굴을 가진 약 10세 여자아이로, 참조의 얼굴과 남색 상의에 대체로 부합한다. 한국계 백인 혼혈이라는 배경은 외형만으로 확정할 수 없다. 지도는 화면 아래에 일부 보이고, 신발과 찰리 및 보급품은 프레임 밖이라 평가하지 않는다. 황량한 들판, 고사목, 폐가, 바리케이드와 초소는 보이지만, 이 숏에 허용되지 않은 경찰 두 명도 명확히 등장한다. 뚜렷하게 읽히는 글자는 없다.",
        "hard_violations": [
         "앰버만 등장하도록 제한된 숏에 경찰 두 명을 추가했다."
        ],
        "physics": "앰버의 하체는 가려져 있지만 몸통은 조수석 등받이 앞에 위치하며 뒤돌아보는 자세 자체는 가능하다. 지도는 무릎 부근에 놓인 것으로 보이고 접촉 지점은 화면 아래에 가려져 있어 공중 부유라고 볼 근거는 없다. 경찰은 발을 도로에 딛고 있고 차단물과 초소, 차량도 지면에 놓여 있다."
       },
       {
        "label": "B",
        "direction": "앰버의 얼굴과 두 눈은 카메라 가까운 화면 왼쪽 방향을 향한다. 검문소는 얼굴 뒤 화면 오른쪽 창밖에 있어 그녀가 바라보는 대상과 일치하지 않는다. 경찰 두 명은 차량 쪽을 향해 서 있다.",
        "built_space": "얼굴 중심 클로즈업이며 화면 왼쪽에 좌석 등받이와 머리받침 하나가 일부 보인다. 오른쪽에는 문 손잡이와 팔걸이가 달린 측면 문 하나, 측면 창 하나, 사이드미러 하나가 보이고 위에는 선바이저 하나가 걸쳐 있다. 검문소는 이 측면 창 너머에 배치되어 앞유리를 통해 도로 끝을 바라보는 장소 설정과 맞지 않는다. 밖에는 줄무늬 바리케이드 두 개, 초소 하나, 경찰차 한 대가 보인다.",
        "entities": "금발, 큰 눈, 둥근 얼굴과 남색 상의가 참조 앰버에 대체로 맞으며 약 10세 여자아이로 보인다. 혼혈 배경 자체는 외형만으로 검증할 수 없다. 눈을 크게 뜨고 입을 살짝 벌린 놀람이 보이며 눈의 해부학적 형태는 정상이다. 고사목과 폐가, 검문 시설이 보이지만 허용되지 않은 경찰 두 명도 등장한다. 지도, 신발, 찰리와 보급품은 클로즈업 밖이므로 미노출을 감점하지 않는다. 차량 측면 표기는 흐려 명확한 판독이 어렵다.",
        "hard_violations": [
         "앰버만 등장하도록 제한된 숏에 경찰 두 명을 추가했다."
        ],
        "physics": "앰버는 좌석 등받이 앞에 상체를 세우고 있으며 목과 어깨의 연결 및 앉은 자세에 물리적 모순은 보이지 않는다. 골반과 발은 프레임 밖이다. 경찰 두 명과 바리케이드, 초소, 경찰차는 모두 도로 위에 지지되어 있고 떠 있는 인물이나 물체는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "얼굴 클로즈업 대신 차량 내부를 넓게 보여주며, 앰버가 앞의 검문소가 아닌 뒤쪽을 돌아보고 허용되지 않은 경찰 두 명까지 등장한다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "얼굴 클로즈업과 동그랗게 뜬 눈은 더 충실하지만, 검문소가 앞유리가 아닌 측면 창밖에 있고 시선도 그곳을 향하지 않으며 경찰 두 명이 추가되었다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "앰버는 앞유리 너머 검문소를 등진 채 얼굴을 뒤쪽으로 돌리고 화면 오른쪽을 바라본다. 시선은 화면 중앙 앞유리 밖의 바리케이드나 경찰 검문소에 닿지 않는다. 경찰 두 명은 캠핑카 쪽을 향해 서 있다.",
        "built_space": "뒤 좌석 부근에서 앞을 보는 넓은 구도로, 앞좌석 등받이 두 개, 운전대 하나, 대시보드, 앞유리 하나와 선바이저 두 개가 보인다. 앰버는 오른쪽 조수석에 있지만 어깨 옆 좌석 일부만 보여야 하는 요구와 달리 양쪽 좌석과 운전 공간이 크게 노출된다. 앞유리 밖에는 콘크리트 차단물 두 개, 줄무늬 바리케이드 두 개, 초소 하나와 경찰차 한 대가 보이며 검문소가 비교적 가깝고 크게 표현되었다.",
        "entities": "앰버는 금발에 둥근 얼굴을 가진 약 10세 여자아이로, 참조의 얼굴과 남색 상의에 대체로 부합한다. 한국계 백인 혼혈이라는 배경은 외형만으로 확정할 수 없다. 지도는 화면 아래에 일부 보이고, 신발과 찰리 및 보급품은 프레임 밖이라 평가하지 않는다. 황량한 들판, 고사목, 폐가, 바리케이드와 초소는 보이지만, 이 숏에 허용되지 않은 경찰 두 명도 명확히 등장한다. 뚜렷하게 읽히는 글자는 없다.",
        "hard_violations": [
         "앰버만 등장하도록 제한된 숏에 경찰 두 명을 추가했다."
        ],
        "physics": "앰버의 하체는 가려져 있지만 몸통은 조수석 등받이 앞에 위치하며 뒤돌아보는 자세 자체는 가능하다. 지도는 무릎 부근에 놓인 것으로 보이고 접촉 지점은 화면 아래에 가려져 있어 공중 부유라고 볼 근거는 없다. 경찰은 발을 도로에 딛고 있고 차단물과 초소, 차량도 지면에 놓여 있다."
       },
       {
        "label": "A",
        "direction": "앰버의 얼굴과 두 눈은 카메라 가까운 화면 왼쪽 방향을 향한다. 검문소는 얼굴 뒤 화면 오른쪽 창밖에 있어 그녀가 바라보는 대상과 일치하지 않는다. 경찰 두 명은 차량 쪽을 향해 서 있다.",
        "built_space": "얼굴 중심 클로즈업이며 화면 왼쪽에 좌석 등받이와 머리받침 하나가 일부 보인다. 오른쪽에는 문 손잡이와 팔걸이가 달린 측면 문 하나, 측면 창 하나, 사이드미러 하나가 보이고 위에는 선바이저 하나가 걸쳐 있다. 검문소는 이 측면 창 너머에 배치되어 앞유리를 통해 도로 끝을 바라보는 장소 설정과 맞지 않는다. 밖에는 줄무늬 바리케이드 두 개, 초소 하나, 경찰차 한 대가 보인다.",
        "entities": "금발, 큰 눈, 둥근 얼굴과 남색 상의가 참조 앰버에 대체로 맞으며 약 10세 여자아이로 보인다. 혼혈 배경 자체는 외형만으로 검증할 수 없다. 눈을 크게 뜨고 입을 살짝 벌린 놀람이 보이며 눈의 해부학적 형태는 정상이다. 고사목과 폐가, 검문 시설이 보이지만 허용되지 않은 경찰 두 명도 등장한다. 지도, 신발, 찰리와 보급품은 클로즈업 밖이므로 미노출을 감점하지 않는다. 차량 측면 표기는 흐려 명확한 판독이 어렵다.",
        "hard_violations": [
         "앰버만 등장하도록 제한된 숏에 경찰 두 명을 추가했다."
        ],
        "physics": "앰버는 좌석 등받이 앞에 상체를 세우고 있으며 목과 어깨의 연결 및 앉은 자세에 물리적 모순은 보이지 않는다. 골반과 발은 프레임 밖이다. 경찰 두 명과 바리케이드, 초소, 경찰차는 모두 도로 위에 지지되어 있고 떠 있는 인물이나 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt-high"
   ],
   "failed": [
    "gemini-pro"
   ],
   "route": "single_reverse"
  },
  "totals": {
   "B": 2,
   "A": 3
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2,
    "verdict_ko": "얼굴 클로즈업 대신 차량 내부를 넓게 보여주며, 앰버가 앞의 검문소가 아닌 뒤쪽을 돌아보고 허용되지 않은 경찰 두 명까지 등장한다."
   },
   {
    "label": "A",
    "score": 3,
    "verdict_ko": "얼굴 클로즈업과 동그랗게 뜬 눈은 더 충실하지만, 검문소가 앞유리가 아닌 측면 창밖에 있고 시선도 그곳을 향하지 않으며 경찰 두 명이 추가되었다."
   }
  ],
  "refs": [
   {
    "label": "FLOOR PLAN — layout authority, a diagram, never scenery",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S48sh5_confinedfp.png",
    "asset_id": null,
    "role": null
   },
   {
    "label": "앰버",
    "path": "<bytes:1294604>",
    "asset_id": "8b82464d-b554-42c6-87b1-aa97055d6b16",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": false,
  "shot_run_uid": "06aae92a-25ba-7eee-82c2-788510e59a12",
  "confined_fp": {
   "base_key": "confinedfp::0261ae55cee7",
   "apt_reason": "이 샷은 캠핑카 내부 조수석을 배경으로 하며, 앰버가 조수석에 앉아 앞유리를 통해 전방의 경찰 검문소를 바라보는 상황입니다. 차량 내부의 정확한 좌석 배치와 인물의 시선 방향이 어긋나면 화면의 일관성과 몰입을 해칠 수 있으므로 평면도 형태의 레이아웃 가이드가 필요합니다.",
   "fixed": true,
   "mismatches": [
    "The LOCATION specifies looking through the windshield toward the checkpoint, but the drawn camera points sideways inward at Amber from beside the passenger seat. The windshield and checkpoint are consequently outside this indicated close-up, toward screen-right."
   ]
  },
  "ref_mode": "confined_fp: 도면+장면설명+엔티티",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S48sh5::cine": {
  "applied": true,
  "attempted_at": "2026-09-19T10:10:38.561608+00:00",
  "fingerprint": "78257b5cefd7b4e37e15e3039ca10f14399e692c877c246d977d7324f3db1fd7",
  "fingerprint_version": 2,
  "provider": "xai",
  "endpoint": "xai/images/edits",
  "model": "grok-imagine-image",
  "multiroll_tag": "still_S48sh5_cine",
  "slot": "",
  "pack": "24.202608252115",
  "source_file": "S48sh5_sel.png",
  "source_sha256": "cbf8cc2a36c4ebd52527ca4054a573855c8d7017218d9bb66b28175b8d024b27",
  "file": "S48sh5_cine.png",
  "staged_sha256": "f432779492e0f67637be29bc70c9ef46710ec3a2642a661b2a3c19f7dbba3052",
  "latency_ms": 9817
 },
 "era_assess::01eb3d1e2773b936": {
  "subjects": [],
  "subject_text": "쓰레기 수거선 갑판\n대형 수거선의 넓은 금속 갑판. 수거용 크레인과 그물, 쌓인 해양 폐기물과 고철이 낮빛 아래 드러난다.",
  "identity": "canonical",
  "scope_id": "L147",
  "scope_role": "location_exterior",
  "scope_sha": "007731c4bba546f4"
 },
 "groupbg::collection_deck": {
  "input_fingerprint": "ddda1847e8c8fde8",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "collection_deck",
    "tags": [
     "S2sh3"
    ]
   },
   "context_sig": "681bc2b5ba32f5eb"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On the open deck of a garbage collection ship, among freshly dumped rubbish and tangled fishing net.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n쓰레기 수거선 갑판: 바다에서 건져 올린 쓰레기가 쏟아지는 크고 거친 금속 갑판. (특징: 철제 선박 갑판; 위에서 우수수 떨어지는 쓰레기 무더기; 그물에 엉켜 있는 낡은 고철 로봇)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 쓰레기 수거선의 갑판으로 우수수 떨어지는 쓰레기들.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On the open deck of a garbage collection ship, among freshly dumped rubbish and tangled fishing net.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n쓰레기 수거선 갑판: 바다에서 건져 올린 쓰레기가 쏟아지는 크고 거친 금속 갑판. (특징: 철제 선박 갑판; 위에서 우수수 떨어지는 쓰레기 무더기; 그물에 엉켜 있는 낡은 고철 로봇)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 쓰레기 수거선의 갑판으로 우수수 떨어지는 쓰레기들.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_collection_deck_b9e7b9.png",
  "asset_id": "ad655f79-61dd-4341-a8cb-31e490c3f150",
  "input_asset_ids": [
   "b077dc93-9217-44ee-9e06-e7c3ca0dcd21"
  ],
  "origin_tag": "S2sh3",
  "place_text": "On the open deck of a garbage collection ship, among freshly dumped rubbish and tangled fishing net.",
  "origin_inputs": {
   "place_text": "On the open deck of a garbage collection ship, among freshly dumped rubbish and tangled fishing net.",
   "time_of_day_en": "day",
   "conti_asset_id": "b077dc93-9217-44ee-9e06-e7c3ca0dcd21"
  }
 },
 "S2sh3::bgfirst_bg": {
  "input_fingerprint": "d391387eb9ee00b4",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 쏟아진 쓰레기 더미 사이, 그물에 몸이 감긴 채 널브러져 있는 낡은 고철 로봇 찰리의 전신.\n\nLOCATION (lock): On the open deck of a garbage collection ship, among freshly dumped rubbish and tangled fishing net.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Net around 찰리 (Wrapped around his body among the dumped rubbish); used as Crossing lines that reveal confinement while leaving the full-body silhouette readable; Dumped rubbish (Deposited on the collection ship's deck around 찰리); used as Uneven foreground and background layers surrounding the revealed body; Collection ship deck (Receiving the collected rubbish) — Its upper surface is seen obliquely beneath gaps in the rubbish; used as A spatial base establishing that the body is now aboard the ship.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light with controlled contrast keeps the net and aged robot body legible without romanticizing the discarded surroundings.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 쏟아진 쓰레기 더미 사이, 그물에 몸이 감긴 채 널브러져 있는 낡은 고철 로봇 찰리의 전신.\n\nLOCATION (lock): On the open deck of a garbage collection ship, among freshly dumped rubbish and tangled fishing net.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Net around 찰리 (Wrapped around his body among the dumped rubbish); used as Crossing lines that reveal confinement while leaving the full-body silhouette readable; Dumped rubbish (Deposited on the collection ship's deck around 찰리); used as Uneven foreground and background layers surrounding the revealed body; Collection ship deck (Receiving the collected rubbish) — Its upper surface is seen obliquely beneath gaps in the rubbish; used as A spatial base establishing that the body is now aboard the ship.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light with controlled contrast keeps the net and aged robot body legible without romanticizing the discarded surroundings.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S2sh3__bgfirst_bg.png",
  "asset_id": "22001d0d-2139-4c65-801f-8fe7b7978313",
  "input_asset_ids": [
   "b077dc93-9217-44ee-9e06-e7c3ca0dcd21",
   "ad655f79-61dd-4341-a8cb-31e490c3f150"
  ]
 },
 "groupbg::dump_emergence": {
  "input_fingerprint": "a058d0a8ccb0e7ce",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "dump_emergence",
    "tags": [
     "S6sh15",
     "S6sh17"
    ]
   },
   "context_sig": "f05b52e5623c370f"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: Within an exposed rubbish mound at the refugee settlement's dump, at the spot where a large robot emerges.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n인천 난민촌 쓰레기장: 산처럼 거대하게 쌓인 폐기물과 고철 더미로 이루어진 구역. (특징: 하늘을 가릴 듯 높은 거대한 쓰레기 산; 우그러진 로봇 팔, 기어 등 각종 고철 부품; 부서진 낡은 주크박스 몸체와 나팔형 스피커; 바닥에 깔린 모포와 널브러진 공구, LP판)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 쓰레기 더미에서 몸을 서서히 드러내는 건.... 고릴라 모양의 로봇이다.\n- 앰버를 내려다보다 방긋 웃더니, 이내 와락 껴안는다!\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: Within an exposed rubbish mound at the refugee settlement's dump, at the spot where a large robot emerges.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n인천 난민촌 쓰레기장: 산처럼 거대하게 쌓인 폐기물과 고철 더미로 이루어진 구역. (특징: 하늘을 가릴 듯 높은 거대한 쓰레기 산; 우그러진 로봇 팔, 기어 등 각종 고철 부품; 부서진 낡은 주크박스 몸체와 나팔형 스피커; 바닥에 깔린 모포와 널브러진 공구, LP판)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 쓰레기 더미에서 몸을 서서히 드러내는 건.... 고릴라 모양의 로봇이다.\n- 앰버를 내려다보다 방긋 웃더니, 이내 와락 껴안는다!\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_dump_emergence_3bffab.png",
  "asset_id": "27c70e61-a2a2-4e06-b8da-d8a7b9325903",
  "input_asset_ids": [
   "c554fa18-115e-4e6b-9462-3d5f65b3a88e"
  ],
  "origin_tag": "S6sh15",
  "place_text": "Within an exposed rubbish mound at the refugee settlement's dump, at the spot where a large robot emerges.",
  "origin_inputs": {
   "place_text": "Within an exposed rubbish mound at the refugee settlement's dump, at the spot where a large robot emerges.",
   "time_of_day_en": "day",
   "conti_asset_id": "c554fa18-115e-4e6b-9462-3d5f65b3a88e"
  }
 },
 "S6sh15::bgfirst_bg": {
  "input_fingerprint": "f2e37e8b99b592df",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 쓰레기를 머리에 뒤집어쓴 채 상체를 반쯤 일으킨 지탱 자세로 고릴라 형태의 로봇 찰리의 두 눈에 파란 불빛이 켜져 있는 순간.\n\nLOCATION (lock): Within an exposed rubbish mound at the refugee settlement's dump, at the spot where a large robot emerges.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Refuse surrounding 찰리 (Disturbed by his emergence, with rubbish still resting on his head and surrounding his lower body); used as Frames the supporting arms and makes the effort of emergence physically readable; Larger rubbish heap (Extending behind the emergence point); used as Provides layered background context without obscuring the head outline.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light remains steady while the blue illumination in 찰리's eyes becomes the restrained focal accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 쓰레기를 머리에 뒤집어쓴 채 상체를 반쯤 일으킨 지탱 자세로 고릴라 형태의 로봇 찰리의 두 눈에 파란 불빛이 켜져 있는 순간.\n\nLOCATION (lock): Within an exposed rubbish mound at the refugee settlement's dump, at the spot where a large robot emerges.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Refuse surrounding 찰리 (Disturbed by his emergence, with rubbish still resting on his head and surrounding his lower body); used as Frames the supporting arms and makes the effort of emergence physically readable; Larger rubbish heap (Extending behind the emergence point); used as Provides layered background context without obscuring the head outline.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light remains steady while the blue illumination in 찰리's eyes becomes the restrained focal accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S6sh15__bgfirst_bg.png",
  "asset_id": "8d75d911-3d18-4135-b6ab-fc26e495c901",
  "input_asset_ids": [
   "c554fa18-115e-4e6b-9462-3d5f65b3a88e",
   "27c70e61-a2a2-4e06-b8da-d8a7b9325903"
  ]
 },
 "era_assess::f2fe65929099e9dd": {
  "subjects": [],
  "subject_text": "라울의 컨테이너 앞 공터\n낡은 철제 컨테이너 출입문 앞에 마련된 작은 공터. 주변에 컨테이너 하우스가 늘어서고 바닥에는 물 호스가 놓여 있다.",
  "identity": "canonical",
  "scope_id": "L169",
  "scope_role": "location_exterior",
  "scope_sha": "d5e7f7fd70a3c7c2"
 },
 "S16sh3::bgfirst_bg": {
  "input_fingerprint": "c93785fe45234431",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 흙먼지가 씻겨나간 젖은 몸으로 양어깨를 치켜올린 채 즐거워하는 찰리의 상체.\n\nLOCATION (lock): In the open washing area directly in front of a container home in the refugee settlement.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Hose water stream (Striking Charlie and washing away the dirt) — Enters laterally from the off-screen hose position toward his torso; used as Connect the raised shoulders to the physical sensation without obscuring his face; Front of 라울's container home (Visible behind the washing action) — A partial exterior backdrop is retained beside the upper body; used as Locate the intimate action without widening to the observers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daylight with controlled highlights on the explicitly wet body and water stream, preserving a gentle rather than harsh mood.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 흙먼지가 씻겨나간 젖은 몸으로 양어깨를 치켜올린 채 즐거워하는 찰리의 상체.\n\nLOCATION (lock): In the open washing area directly in front of a container home in the refugee settlement.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Hose water stream (Striking Charlie and washing away the dirt) — Enters laterally from the off-screen hose position toward his torso; used as Connect the raised shoulders to the physical sensation without obscuring his face; Front of 라울's container home (Visible behind the washing action) — A partial exterior backdrop is retained beside the upper body; used as Locate the intimate action without widening to the observers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daylight with controlled highlights on the explicitly wet body and water stream, preserving a gentle rather than harsh mood.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S16sh3__bgfirst_bg.png",
  "asset_id": "58f097e3-cb89-4739-96b1-20ba4de0ccc6",
  "input_asset_ids": [
   "083f7cfb-ee18-4327-afb0-899eddfe6874",
   "a2e9a0db-ecb3-47e7-96eb-291c9f873ffb"
  ]
 },
 "groupbg::banana_market_stall": {
  "input_fingerprint": "d286953e23f5d0a1",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "banana_market_stall",
    "tags": [
     "S22sh3",
     "S22sh7",
     "S22sh9"
    ]
   },
   "context_sig": "6d6114ecb3b8f4ca"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At the street-facing banana display of an outdoor market stall in the refugee settlement.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n인천 난민촌 시장과 상점 골목: 물건을 파는 매대와 천막들이 복잡하게 얽혀 있는 좁고 혼잡한 시장통. (특징: 조악하게 지어진 상점과 노점상들; 바나나 등 식료품이 진열된 매대; 통로에 쌓여 있는 종이 상자와 물건들)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 그러다 바나나를 파는 가게 앞에 멈추는 찰리.\n- 인파 틈으로 사라지는 찰리. 그 자리로 뛰어오는 앰버와 라울.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At the street-facing banana display of an outdoor market stall in the refugee settlement.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n인천 난민촌 시장과 상점 골목: 물건을 파는 매대와 천막들이 복잡하게 얽혀 있는 좁고 혼잡한 시장통. (특징: 조악하게 지어진 상점과 노점상들; 바나나 등 식료품이 진열된 매대; 통로에 쌓여 있는 종이 상자와 물건들)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 그러다 바나나를 파는 가게 앞에 멈추는 찰리.\n- 인파 틈으로 사라지는 찰리. 그 자리로 뛰어오는 앰버와 라울.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_banana_market_stall_8677ef.png",
  "asset_id": "0333daa2-9ee6-4388-bf59-c510fb37f958",
  "input_asset_ids": [
   "774388e6-e167-4f76-a648-7a8dce68ade1"
  ],
  "origin_tag": "S22sh3",
  "place_text": "At the street-facing banana display of an outdoor market stall in the refugee settlement.",
  "origin_inputs": {
   "place_text": "At the street-facing banana display of an outdoor market stall in the refugee settlement.",
   "time_of_day_en": "day",
   "conti_asset_id": "774388e6-e167-4f76-a648-7a8dce68ade1"
  }
 },
 "S22sh3::bgfirst_bg": {
  "input_fingerprint": "16b23c64cd826776",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 바나나를 향해 커다란 금속 손가락을 뻗은 찰리의 손 클로즈업.\n\nLOCATION (lock): At the street-facing banana display of an outdoor market stall in the refugee settlement.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Bananas (Offered for sale and not yet touched by 찰리); used as Provide the clearly visible destination of the fingers while occupying less than a third of the image; Banana stall (Visible in a limited area around the offered fruit) — The camera views the selling area obliquely from the side of 찰리's reaching arm; used as Anchor the hand and fruit within the market rather than isolating them against an undefined background.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daytime ambient light with restrained contrast, allowing the metal hand and the bananas' own color to remain distinct without adding a colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 바나나를 향해 커다란 금속 손가락을 뻗은 찰리의 손 클로즈업.\n\nLOCATION (lock): At the street-facing banana display of an outdoor market stall in the refugee settlement.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Bananas (Offered for sale and not yet touched by 찰리); used as Provide the clearly visible destination of the fingers while occupying less than a third of the image; Banana stall (Visible in a limited area around the offered fruit) — The camera views the selling area obliquely from the side of 찰리's reaching arm; used as Anchor the hand and fruit within the market rather than isolating them against an undefined background.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use neutral daytime ambient light with restrained contrast, allowing the metal hand and the bananas' own color to remain distinct without adding a colored light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S22sh3__bgfirst_bg.png",
  "asset_id": "e6ccf295-a01a-4188-a4ca-5ce36e1609e9",
  "input_asset_ids": [
   "774388e6-e167-4f76-a648-7a8dce68ade1",
   "0333daa2-9ee6-4388-bf59-c510fb37f958"
  ]
 },
 "groupbg::reception_clearing": {
  "input_fingerprint": "4459c25d668f8883",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "reception_clearing",
    "tags": [
     "S25sh12",
     "S25sh19",
     "S26sh7"
    ]
   },
   "context_sig": "4d84b3e4221710d9"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: In the refugee settlement's open wedding-reception lot, beneath makeshift canopies and newly illuminated streetlights.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n인천 난민촌 지하 하수도: 지상에서 맨홀과 사다리를 통해 내려오는 습하고 좁은 지하 콘크리트 통로. (특징: 수직으로 이어진 낡은 철제 사다리; 갈라지는 두 갈래의 콘크리트 길; 바닥의 물기와 어두운 조명) / 인천 난민촌 피로연장: 빈 공터에 천막을 치고 음악을 틀어놓은 조악한 파티장. (특징: 공터 위로 드리워진 허름한 천막; 빛나는 주크박스; 순간적으로 꺼졌다가 과부하로 밝게 터져나가는 가로등과 전구들)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- (*피로연장이라고 해봤자 그냥 빈 공터에 천막 치는 정도)\n- /피로연장 -N\n- 하객들은 한가운데 무리 지어있고, 그들을 총과 몽둥이로 위협하고 있는 상황\n\nTIME OF DAY (lock): sunset.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: In the refugee settlement's open wedding-reception lot, beneath makeshift canopies and newly illuminated streetlights.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n인천 난민촌 지하 하수도: 지상에서 맨홀과 사다리를 통해 내려오는 습하고 좁은 지하 콘크리트 통로. (특징: 수직으로 이어진 낡은 철제 사다리; 갈라지는 두 갈래의 콘크리트 길; 바닥의 물기와 어두운 조명) / 인천 난민촌 피로연장: 빈 공터에 천막을 치고 음악을 틀어놓은 조악한 파티장. (특징: 공터 위로 드리워진 허름한 천막; 빛나는 주크박스; 순간적으로 꺼졌다가 과부하로 밝게 터져나가는 가로등과 전구들)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- (*피로연장이라고 해봤자 그냥 빈 공터에 천막 치는 정도)\n- /피로연장 -N\n- 하객들은 한가운데 무리 지어있고, 그들을 총과 몽둥이로 위협하고 있는 상황\n\nTIME OF DAY (lock): sunset.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_reception_clearing_3488b1.png",
  "asset_id": "7f28b2fd-054b-41f4-b119-0c0d0c391ef1",
  "input_asset_ids": [
   "b30b0e2b-b930-4260-9eaf-f19f33633ca4"
  ],
  "origin_tag": "S25sh12",
  "place_text": "In the refugee settlement's open wedding-reception lot, beneath makeshift canopies and newly illuminated streetlights.",
  "origin_inputs": {
   "place_text": "In the refugee settlement's open wedding-reception lot, beneath makeshift canopies and newly illuminated streetlights.",
   "time_of_day_en": "sunset",
   "conti_asset_id": "b30b0e2b-b930-4260-9eaf-f19f33633ca4"
  }
 },
 "S25sh12::bgfirst_bg": {
  "input_fingerprint": "d9dcc640cf231108",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리의 가슴에서 뿜어진 빛과 함께 주변 가로등들이 일제히 탁 켜져 빛을 발하는 눈부신 순간의 광경.\n\nLOCATION (lock): In the refugee settlement's open wedding-reception lot, beneath makeshift canopies and newly illuminated streetlights.\n\nTIME OF DAY (lock): sunset.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Streetlights behind and left of the unobstructed chest in the upper-left of the frame, background; Streetlights behind and right of the unobstructed chest in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: 주변 가로등 (Arranged around the reception area and continuing into the settlement) — Seen from above and obliquely, distributed behind and to both sides of 찰리; used as Their physical distribution gives the expanding event depth and scale; 피로연장 공터 (An open clearing used for the reception); used as Provides visible separation between 찰리 and the surrounding streetlights; 피로연장 천막 (Set up in the clearing) — An oblique upper and side view appears along a background edge; used as Anchors the spectacle to the modest reception setting without blocking the chest; 난민촌 (Extending beyond the reception clearing); used as Supplies the wider environmental scale revealed by the retreat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: At dusk, the ring-shaped chest emission and rapidly relighting streetlights create a dazzling expansion of illumination while controlled exposure preserves 찰리's outline.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리의 가슴에서 뿜어진 빛과 함께 주변 가로등들이 일제히 탁 켜져 빛을 발하는 눈부신 순간의 광경.\n\nLOCATION (lock): In the refugee settlement's open wedding-reception lot, beneath makeshift canopies and newly illuminated streetlights.\n\nTIME OF DAY (lock): sunset.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Streetlights behind and left of the unobstructed chest in the upper-left of the frame, background; Streetlights behind and right of the unobstructed chest in the upper-right of the frame, background.\n- KEY BACKGROUND ELEMENTS: 주변 가로등 (Arranged around the reception area and continuing into the settlement) — Seen from above and obliquely, distributed behind and to both sides of 찰리; used as Their physical distribution gives the expanding event depth and scale; 피로연장 공터 (An open clearing used for the reception); used as Provides visible separation between 찰리 and the surrounding streetlights; 피로연장 천막 (Set up in the clearing) — An oblique upper and side view appears along a background edge; used as Anchors the spectacle to the modest reception setting without blocking the chest; 난민촌 (Extending beyond the reception clearing); used as Supplies the wider environmental scale revealed by the retreat.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: At dusk, the ring-shaped chest emission and rapidly relighting streetlights create a dazzling expansion of illumination while controlled exposure preserves 찰리's outline.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S25sh12__bgfirst_bg.png",
  "asset_id": "44dbc893-09f5-4451-985f-95d270aac79f",
  "input_asset_ids": [
   "b30b0e2b-b930-4260-9eaf-f19f33633ca4",
   "7f28b2fd-054b-41f4-b119-0c0d0c391ef1"
  ]
 },
 "era_assess::f3c165a73d599006": {
  "subjects": [],
  "subject_text": "인천 난민촌 도로와 맨홀 주변\n컨테이너 주거지 사이를 지나는 도로. 노면에 둥근 맨홀 뚜껑이 있고 길을 따라 가로등이 이어진다.",
  "identity": "canonical",
  "scope_id": "L177",
  "scope_role": "location_exterior",
  "scope_sha": "50531cbe908dd43e"
 },
 "S27sh12::bgfirst_bg": {
  "input_fingerprint": "9bfccaadca8a231e",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 도로 위를 달리며 한 발이 공중에 뜬 찰리를 향해 빠른 속도로 돌진하는 거대한 자동차의 전경.\n\nLOCATION (lock): On the open roadway in the refugee settlement, in the path of an approaching vehicle's headlights.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Approaching large vehicle in the middle-left of the frame, foreground, moves toward 찰리's running path in the right midground.\n- KEY BACKGROUND ELEMENTS: Approaching large vehicle (Moving rapidly toward 찰리 before the emergency stop) — Its front and one side are visible diagonally from the road edge; used as Forms the left foreground threat while remaining fully within the frame; Road (찰리 and the approaching vehicle occupy converging paths) — The road extends diagonally from the foreground toward the distance; used as Keeps the collision geometry and remaining separation visible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the approaching headlights and the streetlights' established switching behavior to shape the nighttime threat, keeping the road and 찰리 readable without added atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.\n\nThe THIRD attached image (STRUCTURE LOOK) is the identity source of the fixed structure at this location: its faces, openings, levels, materials and signage are truth. Where it conflicts with the LOCATION PHOTOGRAPH about the structure itself, the STRUCTURE LOOK wins; the photograph still governs the surroundings, time of day and lighting.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 도로 위를 달리며 한 발이 공중에 뜬 찰리를 향해 빠른 속도로 돌진하는 거대한 자동차의 전경.\n\nLOCATION (lock): On the open roadway in the refugee settlement, in the path of an approaching vehicle's headlights.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: Approaching large vehicle in the middle-left of the frame, foreground, moves toward 찰리's running path in the right midground.\n- KEY BACKGROUND ELEMENTS: Approaching large vehicle (Moving rapidly toward 찰리 before the emergency stop) — Its front and one side are visible diagonally from the road edge; used as Forms the left foreground threat while remaining fully within the frame; Road (찰리 and the approaching vehicle occupy converging paths) — The road extends diagonally from the foreground toward the distance; used as Keeps the collision geometry and remaining separation visible.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the approaching headlights and the streetlights' established switching behavior to shape the nighttime threat, keeping the road and 찰리 readable without added atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.\n\nThe THIRD attached image (STRUCTURE LOOK) is the identity source of the fixed structure at this location: its faces, openings, levels, materials and signage are truth. Where it conflicts with the LOCATION PHOTOGRAPH about the structure itself, the STRUCTURE LOOK wins; the photograph still governs the surroundings, time of day and lighting.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S27sh12__bgfirst_bg.png",
  "asset_id": "164bd776-d0dc-43e4-bf32-6ef9f6356b78",
  "input_asset_ids": [
   "29e8ef08-4d4f-4022-ac9f-847c4599b99d",
   "785d4e92-6642-43a7-a3f4-70e3c4f7a9b3",
   "2b4b407b-73bc-405e-8265-11c6b8d30ef7"
  ]
 },
 "era_assess::678a9ff44c021387": {
  "subjects": [],
  "subject_text": "무인점포 내부\n신발과 가방, 식품, 생활용품이 진열된 무인 매장. 여러 진열대 사이로 통로가 나 있고 의약품 코너와 현금인출기가 설치돼 있다.",
  "identity": "canonical",
  "scope_id": "L195",
  "scope_role": "location_interior",
  "scope_sha": "ef0c39aef8ac337b"
 },
 "S41sh14::bgfirst_bg": {
  "input_fingerprint": "6eba475ba1d90762",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리의 거대한 금속 주먹이 현금인출기 앞면을 완전히 뚫고 들어간 파괴적인 찰나.\n\nLOCATION (lock): At the cash machine inside the unattended shop's sales floor, in daytime shop light near stocked aisles.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 현금인출기 (Front punctured by 찰리's fist; the fist has not yet withdrawn) — The front and a narrow adjacent side are visible obliquely, exposing the penetration rather than presenting the machine square-on; used as Impact surface and immediate spatial context for the extended arm; 점포 진열대 (Visible only as a narrow background fragment) — Seen obliquely beyond the cash machine; used as Preserves store context and prevents the impact from becoming an isolated effects image.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained ambient illumination appropriate to the store, with controlled contrast that keeps the fist and broken opening legible without added impact effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리의 거대한 금속 주먹이 현금인출기 앞면을 완전히 뚫고 들어간 파괴적인 찰나.\n\nLOCATION (lock): At the cash machine inside the unattended shop's sales floor, in daytime shop light near stocked aisles.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 현금인출기 (Front punctured by 찰리's fist; the fist has not yet withdrawn) — The front and a narrow adjacent side are visible obliquely, exposing the penetration rather than presenting the machine square-on; used as Impact surface and immediate spatial context for the extended arm; 점포 진열대 (Visible only as a narrow background fragment) — Seen obliquely beyond the cash machine; used as Preserves store context and prevents the impact from becoming an isolated effects image.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use restrained ambient illumination appropriate to the store, with controlled contrast that keeps the fist and broken opening legible without added impact effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S41sh14__bgfirst_bg.png",
  "asset_id": "6398fb6a-6f79-4ad1-90ee-719e7320869a",
  "input_asset_ids": [
   "2b2818ac-406a-4c9d-805c-3eb8df283def",
   "830dae63-3f71-4c94-9045-5c77a027e48e"
  ]
 },
 "groupbg::camp_gate": {
  "input_fingerprint": "9acb2b35bb0ae94e",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "camp_gate",
    "tags": [
     "S11sh1",
     "S42sh11",
     "S42sh19",
     "S42sh2"
    ]
   },
   "context_sig": "d84d14ae7a17fa66"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At ground level near the refugee settlement entrance at night, looking into an opened body bag.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n인천 난민촌 입구와 철창문 앞: 통금을 알리는 방송 장비가 설치된 거대한 수용소 출입구 구역. (특징: 난민 거주 구역을 분리하는 거대한 철문; 곳곳에 설치된 낡은 확성기 스피커; 가로등이 소등되어 컴컴해지는 거리)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 출입문 닫기 전에 헐레벌떡 뛰어오는 난민들, 그 틈에 현우도 끼여서 간신히 들어온다.\n- 42. 난민촌 입구 – N\n\nTIME OF DAY (lock): night.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: At ground level near the refugee settlement entrance at night, looking into an opened body bag.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n인천 난민촌 입구와 철창문 앞: 통금을 알리는 방송 장비가 설치된 거대한 수용소 출입구 구역. (특징: 난민 거주 구역을 분리하는 거대한 철문; 곳곳에 설치된 낡은 확성기 스피커; 가로등이 소등되어 컴컴해지는 거리)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 출입문 닫기 전에 헐레벌떡 뛰어오는 난민들, 그 틈에 현우도 끼여서 간신히 들어온다.\n- 42. 난민촌 입구 – N\n\nTIME OF DAY (lock): night.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_camp_gate_b39179.png",
  "asset_id": "83651a7a-615d-46c7-ba87-2459a6bc8bf2",
  "input_asset_ids": [
   "ed925a67-f2a3-41f5-9413-65abf6516887"
  ],
  "origin_tag": "S42sh2",
  "place_text": "At ground level near the refugee settlement entrance at night, looking into an opened body bag.",
  "origin_inputs": {
   "place_text": "At ground level near the refugee settlement entrance at night, looking into an opened body bag.",
   "time_of_day_en": "night",
   "conti_asset_id": "ed925a67-f2a3-41f5-9413-65abf6516887"
  }
 },
 "S42sh2::bgfirst_bg": {
  "input_fingerprint": "d9da0cfaff19a919",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 보디백 안, 핏기 없이 죽은 구도환의 창백한 얼굴을 내려다보는 시점 쇼트.\n\nLOCATION (lock): At ground level near the refugee settlement entrance at night, looking into an opened body bag.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 보디백 (Unzipped around 구도환's exposed face) — The opened interior and near rim are seen from above; used as Frames the face with evidence of death while remaining subordinate to the human subject.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued nighttime ambient illumination sufficient to reveal 구도환's bloodless pallor, with no dreamlike distortion or invented visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 보디백 안, 핏기 없이 죽은 구도환의 창백한 얼굴을 내려다보는 시점 쇼트.\n\nLOCATION (lock): At ground level near the refugee settlement entrance at night, looking into an opened body bag.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: 보디백 (Unzipped around 구도환's exposed face) — The opened interior and near rim are seen from above; used as Frames the face with evidence of death while remaining subordinate to the human subject.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use subdued nighttime ambient illumination sufficient to reveal 구도환's bloodless pallor, with no dreamlike distortion or invented visible light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S42sh2__bgfirst_bg.png",
  "asset_id": "0428d63b-de1f-42c5-ad4e-e5abb4f95632",
  "input_asset_ids": [
   "ed925a67-f2a3-41f5-9413-65abf6516887",
   "83651a7a-615d-46c7-ba87-2459a6bc8bf2"
  ]
 },
 "era_assess::912540b6f9cc78bf": {
  "subjects": [],
  "subject_text": "찰리가 홀로 걷는 지방도로\n도시 밖으로 길게 이어지는 지방도로. 양옆으로 길가 지면이 드러나고 인근 숲과 산길로 이어지는 갈림길이 있다.",
  "identity": "canonical",
  "scope_id": "L212",
  "scope_role": "location_exterior",
  "scope_sha": "53aa184621808354"
 },
 "groupbg::rural_walking_road": {
  "input_fingerprint": "3d470ba3998b1a53",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "rural_walking_road",
    "tags": [
     "S52sh1"
    ]
   },
   "context_sig": "f7a1a96b56ac9a0d"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On an empty provincial asphalt road, with the disguised robot walking alone along the exposed roadway.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n찰리가 홀로 걷는 지방도로: 비 온 뒤 젖어 있는 외곽 차도를 독특한 복장을 한 인물이 걷는 풍경. (특징: 넓은 챙의 밀짚모자; 커다란 고무장화; 알록달록한 비닐 우비; 적막한 시골 도로)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 52. 지방도로 – D\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On an empty provincial asphalt road, with the disguised robot walking alone along the exposed roadway.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n찰리가 홀로 걷는 지방도로: 비 온 뒤 젖어 있는 외곽 차도를 독특한 복장을 한 인물이 걷는 풍경. (특징: 넓은 챙의 밀짚모자; 커다란 고무장화; 알록달록한 비닐 우비; 적막한 시골 도로)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 52. 지방도로 – D\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_rural_walking_road_5c3d2f.png",
  "asset_id": "2c211078-5c11-4ab2-abaa-988d6405f566",
  "input_asset_ids": [
   "364391b6-7633-4178-a784-c9f3499899a6"
  ],
  "origin_tag": "S52sh1",
  "place_text": "On an empty provincial asphalt road, with the disguised robot walking alone along the exposed roadway.",
  "origin_inputs": {
   "place_text": "On an empty provincial asphalt road, with the disguised robot walking alone along the exposed roadway.",
   "time_of_day_en": "day",
   "conti_asset_id": "364391b6-7633-4178-a784-c9f3499899a6"
  }
 },
 "S52sh1::bgfirst_bg": {
  "input_fingerprint": "84b0daefeff7cd87",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 커다란 밀짚모자를 쓰고 알록달록한 우비를 걸친 채 텅 빈 잿빛 도로 위를 걷는 도중 뒷발로 바닥을 밀어내고 앞발을 든 mid-stride 자세의 찰리의 낡은 금속 전신.\n\nLOCATION (lock): On an empty provincial asphalt road, with the disguised robot walking alone along the exposed roadway.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Provincial road (Empty and gray around the walking figure) — The route extends diagonally ahead of 찰리 toward frame right; used as Open space around the full figure establishes isolation and makes the lifted-foot phase readable; Oversized straw hat, raincoat, and boots (Worn by 찰리; the raincoat is multicolored) — Seen obliquely with the hat brim and separated boots clearly readable; used as Costume silhouettes establish the comic discrepancy between his appearance and determined walk.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light maintains restrained tonal contrast while allowing the explicitly multicolored raincoat to retain its stronger color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 커다란 밀짚모자를 쓰고 알록달록한 우비를 걸친 채 텅 빈 잿빛 도로 위를 걷는 도중 뒷발로 바닥을 밀어내고 앞발을 든 mid-stride 자세의 찰리의 낡은 금속 전신.\n\nLOCATION (lock): On an empty provincial asphalt road, with the disguised robot walking alone along the exposed roadway.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Provincial road (Empty and gray around the walking figure) — The route extends diagonally ahead of 찰리 toward frame right; used as Open space around the full figure establishes isolation and makes the lifted-foot phase readable; Oversized straw hat, raincoat, and boots (Worn by 찰리; the raincoat is multicolored) — Seen obliquely with the hat brim and separated boots clearly readable; used as Costume silhouettes establish the comic discrepancy between his appearance and determined walk.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daytime ambient light maintains restrained tonal contrast while allowing the explicitly multicolored raincoat to retain its stronger color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S52sh1__bgfirst_bg.png",
  "asset_id": "4bc8f5c0-751d-4e2e-8589-aaf2e188a7ae",
  "input_asset_ids": [
   "364391b6-7633-4178-a784-c9f3499899a6",
   "2c211078-5c11-4ab2-abaa-988d6405f566"
  ]
 },
 "groupbg::forest_trap_site": {
  "input_fingerprint": "490cf693a7eb6de1",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "forest_trap_site",
    "tags": [
     "S56sh14",
     "S56sh8"
    ]
   },
   "context_sig": "fe15f055e3293da6"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On a flower-lined forest path near the marsh, at the clearing where the robot is caught in a falling net.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n익산 늪지대의 캠핑카 고립 지점: 차량 바퀴가 푹 빠지는 질척한 뻘과 낡은 경고판이 꽂혀 있는 지대. (특징: 차체 하부가 빠져 움직이지 못하는 진흙 구덩이; 주변의 메마른 늪지대 환경; 녹이 슬어 글자가 지워진 구형 방사능 표지판과 출입 금지 팻말)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- /숲 길 -D\n- 현우가 도착한 곳은 찰리가 있던 장소, 그러나 아무도 없고...\n- 순간 트랩에 빠지는 현우와 앰버\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On a flower-lined forest path near the marsh, at the clearing where the robot is caught in a falling net.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n익산 늪지대의 캠핑카 고립 지점: 차량 바퀴가 푹 빠지는 질척한 뻘과 낡은 경고판이 꽂혀 있는 지대. (특징: 차체 하부가 빠져 움직이지 못하는 진흙 구덩이; 주변의 메마른 늪지대 환경; 녹이 슬어 글자가 지워진 구형 방사능 표지판과 출입 금지 팻말)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- /숲 길 -D\n- 현우가 도착한 곳은 찰리가 있던 장소, 그러나 아무도 없고...\n- 순간 트랩에 빠지는 현우와 앰버\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_forest_trap_site_08ebbc.png",
  "asset_id": "0948ec1a-4562-45ba-965f-9a74745e6847",
  "input_asset_ids": [
   "b891d386-6990-4b27-8c39-13f03b4127da"
  ],
  "origin_tag": "S56sh8",
  "place_text": "On a flower-lined forest path near the marsh, at the clearing where the robot is caught in a falling net.",
  "origin_inputs": {
   "place_text": "On a flower-lined forest path near the marsh, at the clearing where the robot is caught in a falling net.",
   "time_of_day_en": "day",
   "conti_asset_id": "b891d386-6990-4b27-8c39-13f03b4127da"
  }
 },
 "S56sh8::bgfirst_bg": {
  "input_fingerprint": "a24eece6f5f15e38",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 허공에서 거대한 그물망이 찰리의 육중한 금속 몸통을 덮치는 찰나.\n\nLOCATION (lock): On a flower-lined forest path near the marsh, at the clearing where the robot is caught in a falling net.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Descending net (Just contacting 찰리's torso) — Its falling edge enters from above and folds across the near side of his body; used as Interrupts the open space above his reaching gesture; Forest trees (Surrounding the capture location); used as Peripheral depth and body-scale reference behind the net.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight and controlled highlights preserve the distinction between 찰리's metal body, the net, and the forest without added atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 허공에서 거대한 그물망이 찰리의 육중한 금속 몸통을 덮치는 찰나.\n\nLOCATION (lock): On a flower-lined forest path near the marsh, at the clearing where the robot is caught in a falling net.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: Descending net (Just contacting 찰리's torso) — Its falling edge enters from above and folds across the near side of his body; used as Interrupts the open space above his reaching gesture; Forest trees (Surrounding the capture location); used as Peripheral depth and body-scale reference behind the net.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight and controlled highlights preserve the distinction between 찰리's metal body, the net, and the forest without added atmospheric effects.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S56sh8__bgfirst_bg.png",
  "asset_id": "7e75322c-4c67-4f9e-95d0-69b421b60a39",
  "input_asset_ids": [
   "b891d386-6990-4b27-8c39-13f03b4127da",
   "0948ec1a-4562-45ba-965f-9a74745e6847"
  ]
 },
 "S59sh36::bgfirst_bg": {
  "input_fingerprint": "5f87be3da579985f",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): (회상/판화) 뚱뚱해진 체형에 번쩍이는 금빛 하회탈을 쓴 채 군중을 굽어보는 백산의 위압적인 목판화 질감 그림.\n\nLOCATION (lock): A stylized woodcut flashback of the village ruler elevated above a crowd; no specific architectural setting is established.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: foreground crowd within the woodcut in the lower-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: 금빛 하회탈 (Intact and worn by the heavyset 백산) — The carved facial front is seen obliquely from below, inclined toward the foreground crowd; used as Provides a compact emblem of authority within the larger figure; 판화 속 전경 군중 (Gathered beneath 백산 within the recollection) — Unevenly overlapping backs and partial profiles face inward toward 백산, with varied head tilts and shoulder levels; used as Preserves the human scale and the upward relationship that gives the print its oppressive force.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Express the remembered scene through controlled woodcut light-and-dark blocks, reserving selective gold brilliance for the mask rather than introducing a realistic light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): (회상/판화) 뚱뚱해진 체형에 번쩍이는 금빛 하회탈을 쓴 채 군중을 굽어보는 백산의 위압적인 목판화 질감 그림.\n\nLOCATION (lock): A stylized woodcut flashback of the village ruler elevated above a crowd; no specific architectural setting is established.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: foreground crowd within the woodcut in the lower-center of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: 금빛 하회탈 (Intact and worn by the heavyset 백산) — The carved facial front is seen obliquely from below, inclined toward the foreground crowd; used as Provides a compact emblem of authority within the larger figure; 판화 속 전경 군중 (Gathered beneath 백산 within the recollection) — Unevenly overlapping backs and partial profiles face inward toward 백산, with varied head tilts and shoulder levels; used as Preserves the human scale and the upward relationship that gives the print its oppressive force.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Express the remembered scene through controlled woodcut light-and-dark blocks, reserving selective gold brilliance for the mask rather than introducing a realistic light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S59sh36__bgfirst_bg.png",
  "asset_id": "c373b774-4997-4590-9fd6-e9ba287393af",
  "input_asset_ids": [
   "f5ca6eee-4fd8-46c4-909f-e7354ed6208f",
   "dfd97579-913b-4622-86a5-0d6cd32f2ea6"
  ]
 },
 "S60sh4::bgfirst_bg": {
  "input_fingerprint": "0749be9b35ce6836",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리의 반대편에 우뚝 선 거대하고 육중한 전투 병기 B-200의 위협적인 실루엣.\n\nLOCATION (lock): On the open-air fighting ground in front of the village church, opposite the arriving robot and surrounded by spectators and torches.\n\nTIME OF DAY (lock): sunset.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 열린 케이지 출입구 (Open after 찰리's arrival at ground level) — Only the near side of the opening appears at the far left edge behind 찰리; used as Anchors the camera's departure point without obstructing either robot; 격투장 바닥 (An open interval separates 찰리 and B-200 before their fight); used as Supplies shared ground and a credible distance comparison without foreground scale exaggeration.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the established sunset illumination and arena torchlight to articulate B-200's threatening outline while retaining enough tonal detail to read both bodies.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.\n\nThe THIRD attached image (STRUCTURE LOOK) is the identity source of the fixed structure at this location: its faces, openings, levels, materials and signage are truth. Where it conflicts with the LOCATION PHOTOGRAPH about the structure itself, the STRUCTURE LOOK wins; the photograph still governs the surroundings, time of day and lighting.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리의 반대편에 우뚝 선 거대하고 육중한 전투 병기 B-200의 위협적인 실루엣.\n\nLOCATION (lock): On the open-air fighting ground in front of the village church, opposite the arriving robot and surrounded by spectators and torches.\n\nTIME OF DAY (lock): sunset.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: 열린 케이지 출입구 (Open after 찰리's arrival at ground level) — Only the near side of the opening appears at the far left edge behind 찰리; used as Anchors the camera's departure point without obstructing either robot; 격투장 바닥 (An open interval separates 찰리 and B-200 before their fight); used as Supplies shared ground and a credible distance comparison without foreground scale exaggeration.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Use the established sunset illumination and arena torchlight to articulate B-200's threatening outline while retaining enough tonal detail to read both bodies.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.\n\nThe THIRD attached image (STRUCTURE LOOK) is the identity source of the fixed structure at this location: its faces, openings, levels, materials and signage are truth. Where it conflicts with the LOCATION PHOTOGRAPH about the structure itself, the STRUCTURE LOOK wins; the photograph still governs the surroundings, time of day and lighting.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S60sh4__bgfirst_bg.png",
  "asset_id": "ab433a49-41f6-42f3-9c2c-d68dd6bf3b32",
  "input_asset_ids": [
   "6a6601d1-910a-4eba-a636-eb4b443cdfb9",
   "239110df-f90f-4ca2-b3ef-4a8f15c13fe9",
   "7cc6d446-071f-4801-886a-a504b66fe88c"
  ]
 },
 "era_assess::61cd8d32cafd0b75": {
  "subjects": [],
  "subject_text": "익산 마을의 고장 난 로봇 보관 폐창고\n기계 부품과 고장 난 로봇 몸체가 산더미처럼 쌓인 대형 폐창고. 반쯤 열린 문으로 빛이 들어오고 구석에는 상자들이 놓여 있다.",
  "identity": "canonical",
  "scope_id": "L238",
  "scope_role": "location_interior",
  "scope_sha": "bbf27bbaa44eebab"
 },
 "S64sh7::bgfirst_bg": {
  "input_fingerprint": "297b9287270f504b",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리를 향해 두꺼운 팔뚝의 고사포 총구를 매섭게 정조준한 B-200의 위협적인 자세.\n\nLOCATION (lock): In a shadowed corner inside the village's large machine-parts warehouse, with daylight entering through its half-open door.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: B-200's forearm gun (Raised and aimed at 찰리) — The barrel is viewed obliquely from the side, with its muzzle directed toward the left foreground figure, not the camera; used as Creates a readable threat line between the characters while remaining proportionate to B-200's body; Pile of broken robots (Heaped in a corner of the warehouse) — Irregular portions of the piled bodies remain visible beyond the confrontation; used as Provides subdued warehouse context and a background against which B-200's living movement registers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the warehouse corner's established darkness while retaining enough restrained tonal separation to read the gun, crouched body, and intervening space.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 찰리를 향해 두꺼운 팔뚝의 고사포 총구를 매섭게 정조준한 B-200의 위협적인 자세.\n\nLOCATION (lock): In a shadowed corner inside the village's large machine-parts warehouse, with daylight entering through its half-open door.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- KEY BACKGROUND ELEMENTS: B-200's forearm gun (Raised and aimed at 찰리) — The barrel is viewed obliquely from the side, with its muzzle directed toward the left foreground figure, not the camera; used as Creates a readable threat line between the characters while remaining proportionate to B-200's body; Pile of broken robots (Heaped in a corner of the warehouse) — Irregular portions of the piled bodies remain visible beyond the confrontation; used as Provides subdued warehouse context and a background against which B-200's living movement registers.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Preserve the warehouse corner's established darkness while retaining enough restrained tonal separation to read the gun, crouched body, and intervening space.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S64sh7__bgfirst_bg.png",
  "asset_id": "2c542339-d519-40a3-aee9-40b352e75792",
  "input_asset_ids": [
   "cbafc64b-e6e6-4acf-8c80-6fb23ad1a5d9",
   "e1a9f48b-d9e5-4c90-99a8-fadcfb66b2d8"
  ]
 },
 "groupbg::village_truck_ambush": {
  "input_fingerprint": "063be8a200e52513",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "village_truck_ambush",
    "tags": [
     "S67sh49",
     "S67sh70"
    ]
   },
   "context_sig": "89cb53bfef34571e"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On the road leading out of the village, directly ahead of the departing militia truck convoy at night.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n익산 한옥마을 거리와 공터: 황폐화된 전통 건축물 사이로 불길이 일고 낡은 공을 차는 흙바닥 넓은 터. (특징: 부서진 기와와 낡은 목조 한옥 잔해들; 밤을 밝히는 드럼통 모닥불과 횃불; 창, 도끼, 몽둥이를 든 하회탈 무리; 흙먼지 날리는 공터 바닥과 낡은 축구공)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 빠르게 멀어지는 민병대 트럭!\n- 트럭 앞에서 대기하고 있는...B-200.\n- 찰리, 다친 몸을 이끌고 B-200의 폭파지점으로 향한다.\n\nTIME OF DAY (lock): night, bright moonlight.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: On the road leading out of the village, directly ahead of the departing militia truck convoy at night.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n익산 한옥마을 거리와 공터: 황폐화된 전통 건축물 사이로 불길이 일고 낡은 공을 차는 흙바닥 넓은 터. (특징: 부서진 기와와 낡은 목조 한옥 잔해들; 밤을 밝히는 드럼통 모닥불과 횃불; 창, 도끼, 몽둥이를 든 하회탈 무리; 흙먼지 날리는 공터 바닥과 낡은 축구공)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 빠르게 멀어지는 민병대 트럭!\n- 트럭 앞에서 대기하고 있는...B-200.\n- 찰리, 다친 몸을 이끌고 B-200의 폭파지점으로 향한다.\n\nTIME OF DAY (lock): night, bright moonlight.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_village_truck_ambush_bde2df.png",
  "asset_id": "bb89f6c6-a1e1-4ea1-b9c4-7696c210aebc",
  "input_asset_ids": [
   "f64f2986-6e3d-4497-9926-988ba8db7762"
  ],
  "origin_tag": "S67sh49",
  "place_text": "On the road leading out of the village, directly ahead of the departing militia truck convoy at night.",
  "origin_inputs": {
   "place_text": "On the road leading out of the village, directly ahead of the departing militia truck convoy at night.",
   "time_of_day_en": "night, bright moonlight",
   "conti_asset_id": "f64f2986-6e3d-4497-9926-988ba8db7762"
  }
 },
 "S67sh49::bgfirst_bg": {
  "input_fingerprint": "fe84cf31c82330ef",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 달리는 트럭들 앞을 우뚝 가로막고 선 채 거대한 고사포를 정조준한 B-200의 위협적인 전신.\n\nLOCATION (lock): On the road leading out of the village, directly ahead of the departing militia truck convoy at night.\n\nTIME OF DAY (lock): night, bright moonlight.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Trucks' forward route (Blocked by B-200 before the convoy overturns) — The approach runs from the lower-left foreground toward B-200 in the middle distance; used as Preserve the opposing movement and aiming directions without placing the camera on the firing line.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established bright moonlit night with restrained tonal separation, withholding muzzle flash because B-200 has not yet fired.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 달리는 트럭들 앞을 우뚝 가로막고 선 채 거대한 고사포를 정조준한 B-200의 위협적인 전신.\n\nLOCATION (lock): On the road leading out of the village, directly ahead of the departing militia truck convoy at night.\n\nTIME OF DAY (lock): night, bright moonlight.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Trucks' forward route (Blocked by B-200 before the convoy overturns) — The approach runs from the lower-left foreground toward B-200 in the middle distance; used as Preserve the opposing movement and aiming directions without placing the camera on the firing line.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Maintain the established bright moonlit night with restrained tonal separation, withholding muzzle flash because B-200 has not yet fired.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S67sh49__bgfirst_bg.png",
  "asset_id": "5b89bd25-951b-4371-9526-4d76774415ae",
  "input_asset_ids": [
   "f64f2986-6e3d-4497-9926-988ba8db7762",
   "bb89f6c6-a1e1-4ea1-b9c4-7696c210aebc"
  ]
 },
 "S67sh70::bgfirst_bg": {
  "input_fingerprint": "fe7986440477c893",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): B-200의 부품을 가슴에 깊이 품은 채 두 눈의 디지털 불빛이 완전히 꺼진 찰리의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): At the blast site on the village departure road, among wrecked trucks and robot debris after the nighttime battle.\n\nTIME OF DAY (lock): night, bright moonlight.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: B-200's remaining chest component (Retrieved after the explosion and held against 찰리's chest); used as Remain partly visible at the lower edge as the tangible object of grief; no intact B-200 figure appears.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The established moonlit ambience softly separates 찰리's damaged face from the subdued background, while both eyes remain completely unlit.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): B-200의 부품을 가슴에 깊이 품은 채 두 눈의 디지털 불빛이 완전히 꺼진 찰리의 굳은 얼굴 클로즈업.\n\nLOCATION (lock): At the blast site on the village departure road, among wrecked trucks and robot debris after the nighttime battle.\n\nTIME OF DAY (lock): night, bright moonlight.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: B-200's remaining chest component (Retrieved after the explosion and held against 찰리's chest); used as Remain partly visible at the lower edge as the tangible object of grief; no intact B-200 figure appears.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The established moonlit ambience softly separates 찰리's damaged face from the subdued background, while both eyes remain completely unlit.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S67sh70__bgfirst_bg.png",
  "asset_id": "24c611fd-00da-42ad-afe0-c4324cb61919",
  "input_asset_ids": [
   "0909ba75-540e-4c7a-b173-2342ee20e201",
   "bb89f6c6-a1e1-4ea1-b9c4-7696c210aebc"
  ]
 },
 "era_assess::78e004d7320278aa": {
  "subjects": [],
  "subject_text": "익산과 서남부 내륙의 지방도로\n완만한 능선과 들판 사이로 이어지는 지방도로. 마을 밖 갈림길에서 도로가 두 방향으로 나뉘며 멀리 낮은 산자락이 보인다.",
  "identity": "canonical",
  "scope_id": "L236",
  "scope_role": "location_exterior",
  "scope_sha": "ed8bfc6afb2c3029"
 },
 "groupbg::open_truck_bed": {
  "input_fingerprint": "f5dd248f6c584f8e",
  "meta": {
   "model": "gpt-image-2.5-sunburst",
   "size": "1536x864",
   "pack": "11.202607220237",
   "contract": "bgfirst_full_v3",
   "group_sig": {
    "key": "open_truck_bed",
    "tags": [
     "S68sh7"
    ]
   },
   "context_sig": "b534b76a0a5277eb"
  },
  "prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: In the open rear cargo bed of a moving old truck on the rural road, with unobstructed daylight above.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n익산과 서남부 내륙의 지방도로: 차량 여러 대가 열을 지어 달리다가 길이 양옆으로 갈라지는 확 트인 농로 및 국도. (특징: 나란히 달리는 지붕 없는 구형 트럭들; Y자 형태로 갈라지는 교차로/갈림길 노면; 흙먼지를 일으키며 멀어지는 타이어 궤적)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 짐칸에 앉은 찰리 B-200의 가슴 부품을 꺼내본다.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create ONE empty live-action location background photograph — NO PEOPLE, no figures, no body parts, no silhouettes, no shadows or reflections of people anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods and produce in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch of a shot that happens at this location: use it ONLY as spatial evidence — what this place contains, how its ground, structures and landmarks are arranged and proportioned. Ignore the sketched people and arrows entirely, and do NOT copy its line style: render a fully photographic, physically plausible real place that fits THE LOCATION text below.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nTHE LOCATION — South Korea, 2069, primarily Korean-speaking; non-refugee locals are Korean unless otherwise stated, while refugee communities are multiethnic and multinational: In the open rear cargo bed of a moving old truck on the rural road, with unobstructed daylight above.\n\nLOCATION DETAIL (descriptive evidence about this place from the production's location records — evidence, not a staging order): use it to understand what kind of place this is and what it permanently contains. Carry over ONLY its enduring physical features — layout, structures, ground and materials, fixed equipment and installations. Do NOT treat any lighting, weather or time-of-day wording, momentary object states, depicted events or subjective impressions in it as fixed facts of the place; the TIME OF DAY lock below is the sole authority for time and lighting.\n익산과 서남부 내륙의 지방도로: 차량 여러 대가 열을 지어 달리다가 길이 양옆으로 갈라지는 확 트인 농로 및 국도. (특징: 나란히 달리는 지붕 없는 구형 트럭들; Y자 형태로 갈라지는 교차로/갈림길 노면; 흙먼지를 일으키며 멀어지는 타이어 궤적)\n\nSCENE EVIDENCE (verbatim quotes from the screenplay about this place — treat them as evidence of what the location physically contains and looks like; stage the PLACE those moments happen in, but do NOT depict the momentary actions, people or staged props themselves):\n- 짐칸에 앉은 찰리 B-200의 가슴 부품을 꺼내본다.\n\nTIME OF DAY (lock): day.\n\nRender ONE photorealistic empty location photograph, 16:9, neutral enough that every shot of this place can be staged from it later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/groupbg_open_truck_bed_9f9871.png",
  "asset_id": "f8675ac4-ad27-42ab-a598-758f4a775a49",
  "input_asset_ids": [
   "aa8d3a60-5597-4fad-886f-90bc9f48a604"
  ],
  "origin_tag": "S68sh7",
  "place_text": "In the open rear cargo bed of a moving old truck on the rural road, with unobstructed daylight above.",
  "origin_inputs": {
   "place_text": "In the open rear cargo bed of a moving old truck on the rural road, with unobstructed daylight above.",
   "time_of_day_en": "day",
   "conti_asset_id": "aa8d3a60-5597-4fad-886f-90bc9f48a604"
  }
 },
 "S68sh7::bgfirst_bg": {
  "input_fingerprint": "d508cb28332e1ada",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 손가락 위의 아기 새(B-200이 돌보던 새)를 가만히 내려다보며 눈을 지그시 감은 찰리의 낡은 금속 얼굴 클로즈업.\n\nLOCATION (lock): In the open rear cargo bed of a moving old truck on the rural road, with unobstructed daylight above.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Truck cargo bed (Carrying the seated 찰리 during the journey) — Only a narrow oblique portion remains visible behind his lower shoulder; used as Anchor the farewell in the moving truck before the dream transition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight reveals the worn metal face and the small bird with gentle contrast, without introducing dream coloration before the dissolve.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 손가락 위의 아기 새(B-200이 돌보던 새)를 가만히 내려다보며 눈을 지그시 감은 찰리의 낡은 금속 얼굴 클로즈업.\n\nLOCATION (lock): In the open rear cargo bed of a moving old truck on the rural road, with unobstructed daylight above.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Truck cargo bed (Carrying the seated 찰리 during the journey) — Only a narrow oblique portion remains visible behind his lower shoulder; used as Anchor the farewell in the moving truck before the dream transition.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Daylight reveals the worn metal face and the small bird with gentle contrast, without introducing dream coloration before the dissolve.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S68sh7__bgfirst_bg.png",
  "asset_id": "16f6020b-ef17-4edc-b91e-0f09780d6056",
  "input_asset_ids": [
   "aa8d3a60-5597-4fad-886f-90bc9f48a604",
   "f8675ac4-ad27-42ab-a598-758f4a775a49"
  ]
 },
 "era_assess::8b9cd93cb333094d": {
  "subjects": [],
  "subject_text": "크리스의 선박 선원실\n두 명 정도가 누울 수 있는 작은 선원실. 벽 가까이 소파가 놓여 있고 미닫이 출입문으로 바깥 통로와 연결된다.",
  "identity": "canonical",
  "scope_id": "L263",
  "scope_role": "location_interior",
  "scope_sha": "963e62a141b2a059"
 },
 "S79sh10::bgfirst_bg": {
  "input_fingerprint": "2d74ae5216f638dc",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 현우의 물음에 복잡한 표정의 디지털 눈 이모티콘을 띄운 찰리의 낡은 얼굴 클로즈업.\n\nLOCATION (lock): At the sofa inside the boat's small two-person crew cabin, under modest nighttime cabin lighting.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Sofa (Supporting both reclining figures) — A narrow portion behind 찰리 and beneath 현우's shoulder remains visible; used as Maintains the shared resting position without competing with 찰리's face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient illumination and gentle tonal separation keep the digital expression readable without adding an unsupported light source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 현우의 물음에 복잡한 표정의 디지털 눈 이모티콘을 띄운 찰리의 낡은 얼굴 클로즈업.\n\nLOCATION (lock): At the sofa inside the boat's small two-person crew cabin, under modest nighttime cabin lighting.\n\nTIME OF DAY (lock): night.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Sofa (Supporting both reclining figures) — A narrow portion behind 찰리 and beneath 현우's shoulder remains visible; used as Maintains the shared resting position without competing with 찰리's face.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Setting-appropriate ambient illumination and gentle tonal separation keep the digital expression readable without adding an unsupported light source or color.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S79sh10__bgfirst_bg.png",
  "asset_id": "9ca43f94-b62c-4185-bbd5-d874b7069c57",
  "input_asset_ids": [
   "471ec738-cab9-44f7-b95c-4082e4ab2709",
   "a3a7948a-4ed6-4034-bdbc-66720a85c15b"
  ]
 },
 "S82sh14::bgfirst_bg": {
  "input_fingerprint": "629fe2bea12b2c63",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 실험실 중앙의 차가운 스테인리스 침대 위로 미동 없이 축 늘어져 누운 찰리의 거대한 전신.\n\nLOCATION (lock): Inside the circular glass enclosure of the main center's adjoining laboratory, on a stainless-steel examination bed under cool laboratory lighting.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Circular glass enclosure (Separates the central bed from the surrounding researchers) — 찰리, the bed, and the scanning arms are visible through the near section of glass; used as Establishes the physical barrier before the camera approaches 현우; Stainless-steel bed (Supports 찰리's motionless body) — Its long side runs diagonally across the elevated view; used as Central support and scale reference; Artificial-intelligence scanning arms (Scanning different areas of 찰리's body) — Articulated sections approach the body from different positions around the bed; used as Surround the still figure with purposeful mechanical activity; Peripheral computer workstations (In use by roughly ten researchers, each engaged at a separate computer) — Seen obliquely around the outer laboratory, without emphasis on screen contents; used as Peripheral scale and asynchronous background activity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral laboratory ambient illumination and controlled tonal contrast reveal the stainless-steel bed and scanning equipment without theatrical highlights.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 실험실 중앙의 차가운 스테인리스 침대 위로 미동 없이 축 늘어져 누운 찰리의 거대한 전신.\n\nLOCATION (lock): Inside the circular glass enclosure of the main center's adjoining laboratory, on a stainless-steel examination bed under cool laboratory lighting.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- KEY BACKGROUND ELEMENTS: Circular glass enclosure (Separates the central bed from the surrounding researchers) — 찰리, the bed, and the scanning arms are visible through the near section of glass; used as Establishes the physical barrier before the camera approaches 현우; Stainless-steel bed (Supports 찰리's motionless body) — Its long side runs diagonally across the elevated view; used as Central support and scale reference; Artificial-intelligence scanning arms (Scanning different areas of 찰리's body) — Articulated sections approach the body from different positions around the bed; used as Surround the still figure with purposeful mechanical activity; Peripheral computer workstations (In use by roughly ten researchers, each engaged at a separate computer) — Seen obliquely around the outer laboratory, without emphasis on screen contents; used as Peripheral scale and asynchronous background activity.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral laboratory ambient illumination and controlled tonal contrast reveal the stainless-steel bed and scanning equipment without theatrical highlights.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S82sh14__bgfirst_bg.png",
  "asset_id": "09c1e32d-ad58-4bd4-a51a-d8b7432ecb22",
  "input_asset_ids": [
   "90735080-f327-4c19-9101-6d98f10c3031",
   "a491e625-4c37-4571-b3ba-85a3daf996f9"
  ]
 },
 "S83sh5::bgfirst_bg": {
  "input_fingerprint": "8df7f491a6de2e97",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 모니터 화면 속, 활짝 웃고 있는 앰버와 라울의 얼굴 클로즈업.\n\nLOCATION (lock): On a video-call monitor inside the research facility's temporary-care room, showing two remote callers without an identifiable background location.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Video-call monitor (Displaying 앰버 and 라울 smiling during the live call) — The image-bearing front is seen obliquely, with its boundary retained around the displayed faces; used as Mediates the close view and distinguishes remote people from local space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the tonal separation between the electronic call image and the local ambient surroundings without adding an unsupported colored glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 모니터 화면 속, 활짝 웃고 있는 앰버와 라울의 얼굴 클로즈업.\n\nLOCATION (lock): On a video-call monitor inside the research facility's temporary-care room, showing two remote callers without an identifiable background location.\n\nTIME OF DAY (lock): day.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Video-call monitor (Displaying 앰버 and 라울 smiling during the live call) — The image-bearing front is seen obliquely, with its boundary retained around the displayed faces; used as Mediates the close view and distinguishes remote people from local space.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Retain the tonal separation between the electronic call image and the local ambient surroundings without adding an unsupported colored glow.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/8f8a58f6-cc7d-4143-940a-1c9c6ab707f7/images/a28fb896-b811-463c-8bcc-ba732968bd7c/scene/recipe/S83sh5__bgfirst_bg.png",
  "asset_id": "3bc12182-b8a5-4b6b-bbfb-1fc47da3fb74",
  "input_asset_ids": [
   "481275cf-ea10-437e-8452-befc99e03a7f",
   "78924c8c-f3a0-4a9f-b31d-34a3fce75e4a"
  ]
 }
}